NTH

Interpretable Discriminative Text Representations via Agreement and Label Disentanglement

AuthorsTong Wang, Yiqing Xu, Leo Yang Yang

May 26, 2026 3 min read
Watch on YouTube
The one-line take

This paper turns text-model interpretability into an auditability test: a feature should be both understandable to humans and genuinely separate from the prediction label.

Key results

0.708 vs. 0.707
Mean balanced accuracy

LFD matches the strongest interpretable baseline, TBM, on average balanced accuracy across the benchmark.

232 raters
Human audit size

The human audit used 232 raters to evaluate reproducibility and label leakage of LFD features versus TBM concepts.

0.68–0.69 vs. 0.27–0.45
Human-human Cohen’s κ

Independent human raters agreed much more on LFD features than on TBM concepts, indicating clearer definitions.

0.99 on iSarcasm and 0.84 on AG-News Sci/Tech
Human-LLM agreement

LLM-produced labels for LFD features aligned closely with human-majority labels on the audited tasks.

κ ≥ 0.70
Cross-LLM κ threshold

LFD admits candidate features only if two independent LLMs applying the same definition achieve substantial agreement.

What the paper found

This paper proposes a concrete auditability standard for interpretable text representations: each coordinate must satisfy conceptual clarity, measured by chance-adjusted inter-rater agreement, and label disentanglement, meaning it should not merely paraphrase the target label. The authors instantiate this in LFD, an LLM-assisted Feature Discovery pipeline that mines contrastive outcome-opposed text pairs, has one LLM propose lexical or semantic features, and requires a separate examiner LLM to reproduce the same feature labels before admission; candidates must clear Cohen’s κ ≥ 0.70 and add residual held-out predictive gain under a LightGBM head. A stylized analysis links the κ screen to a per-feature annotation-noise bound, giving an operating-point interpretation for reliability rather than only post-hoc plausibility. Across ten classification tasks on seven corpora, including AdParaphrase, Loan Application, CFPB Complaints, Gaslighting, iSarcasm, Dec. Reviews, and AG News sub-tasks, LFD matches the predictive performance of the strongest interpretable baseline, TBM, with mean balanced accuracy 0.708 versus 0.707, while producing materially clearer features. Human audits with 232 raters find LFD features much more reproducible, with human-human κ around 0.68–0.69 versus 0.27–0.45 for TBM concepts, and far less label leakage; human-LLM agreement is also higher, reaching 0.99 on iSarcasm and 0.84 on AG-News Sci/Tech. Importantly, LFD achieves this without a post-hoc disentanglement gate, suggesting that cross-LLM agreement plus residual contrastive selection is enough to generate named coordinates that are both predictive and independently auditable.

Original abstract

Interpretable text representations should expose coordinates that are not only predictive, but also meaningful enough for independent auditors to apply. Existing discriminative representations often use anonymous embedding directions, while concept-bottleneck and LLM-assisted methods attach natural-language names to features without ensuring that those definitions are reproducible or distinct from the target label. We propose an operational criterion for interpretable discriminative text representations: each coordinate should satisfy conceptual clarity, measured by chance-adjusted agreement between independent annotators applying the feature definition, and label disentanglement, meaning the feature should not merely paraphrase the prediction target. We instantiate this criterion in LLM-assisted Feature Discovery (LFD), an iterative method that proposes lexical and semantic features from contrastive outcome-opposed text pairs, screens candidates using cross-LLM Cohen's $κ$, and selects features by residual held-out predictive gain. A stylized analysis connects the $κ$ screen to a per-feature annotation-noise bound, formalizing agreement as a reliability check. Across ten text-classification tasks spanning seven corpora, LFD matches the predictive performance of a strong text bottleneck baseline while producing substantially clearer and less label-entangled features. Human audits with 232 raters show that LFD features achieve higher human--human and human--LLM agreement than baseline concepts, and raters consistently judge them as less label-leaking. These results suggest that agreement-tested, label-disentangled coordinates provide a practical auditability standard for interpretable text classification.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →