LLM Explainability with Counterfactual Chains and Causal Graphs
AuthorsNirit Nussbaum-Hoffer, Nitay Calderon, Liat Ein-Dor, Roi Reichart
Resources
This paper turns LLM decision-making into a causal graph of human-interpretable concepts, using counterfactual chains to make the model’s reasoning more transparent.
Key results
Disease diagnosis dataset size
Sentiment analysis dataset size
LLM-as-a-judge dataset size
Counterfactual expansion depth per seed example
Recursive refinement attempts per rejected counterfactual
Allowed non-target concept changes during counterfactual acceptance
What the paper found
This paper, from Technion and IBM Research, reframes LLM explainability as a concept-level causal discovery problem over the model’s own inference process, not the external world. Using Gemini-2-Flash, Qwen3-14B, and OpenAI’s gpt-OSS-20b, it builds a four-stage pipeline: replace gold labels with LLM predictions, extract human-interpretable discriminative concepts, generate counterfactual text chains with an MCMC-inspired acceptance test, and learn a causal graph with σ-CG. The method is evaluated on three classification settings: LIBERTY disease diagnosis with 1,448 examples, IMDB sentiment analysis with 2,096 examples, and an LLM-as-a-judge Reddit preference task with 395 examples. Across 10-fold cross-validation, graph-based parent sets consistently outperform alternative concept subsets; for example, on Gemini-2-Flash disease diagnosis, prediction accuracy rises from 0.61 for competing subsets to 0.67 for the learned causal parents, while concept-node prediction reaches 0.54 versus 0.51. The MCMC augmentation uses 11 counterfactual steps per seed, up to 5 refinements per proposal, and a drift tolerance of 1 to 2 non-target concepts, and its KL-divergence traces converge toward the perfect-overlap bound with structural Hamming distance dropping to 0 for Gemini and Qwen on the synthetic and IMDB settings. The resulting graphs expose model-specific heuristics: shared clinical concepts on LIBERTY, but divergent latent features on natural data, showing that faithful LLM explanations can be made both causal and interpretable at the concept level.
Original abstract
Causal graphs provide a high-level language for making mechanisms transparent. Recent work uses Large Language Models (LLMs) to recover causal graphs of external-world processes. Instead, in this paper, we use causal graphs to model LLM inference itself, providing stakeholders with a transparent view of how the model perceives and organizes high-level concepts to produce a prediction. We propose a four-phase method for constructing such graphs. Given a target LLM and a set of textual examples, our method discovers class-discriminative, human-interpretable concepts and maps each input to LLM-perceived concept states. We then introduce an MCMC-inspired counterfactual augmentation procedure that expands the sparse observational data through chains of counterfactuals. This enables stable causal discovery with $σ$-CG, yielding informative, interpretable graphs. We apply our method to three LLMs across disease diagnosis, sentiment analysis, and LLM-as-a-judge classification tasks. We evaluate the learned graphs for predictive fidelity and structural stability, and the MCMC-inspired augmentation for convergence and downstream utility. Our results show that the discovered causal graphs capture meaningful dependencies consistent with LLMs' reasoning. Together, this paper provides a foundation for concept-level explainability of LLMs.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.