NTH

LLM Explainability with Counterfactual Chains and Causal Graphs

AuthorsNirit Nussbaum-Hoffer, Nitay Calderon, Liat Ein-Dor, Roi Reichart

June 24, 2026 2 min read
Watch on YouTube
The one-line take

This paper turns LLM decision-making into a causal graph of human-interpretable concepts, using counterfactual chains to make the model’s reasoning more transparent.

Key results

1448
LIBERTY

Disease diagnosis dataset size

2096
IMDB

Sentiment analysis dataset size

395
LAJ

LLM-as-a-judge dataset size

11
MCMC steps

Counterfactual expansion depth per seed example

5
Max retries

Recursive refinement attempts per rejected counterfactual

1-2
Drift tolerance

Allowed non-target concept changes during counterfactual acceptance

What the paper found

This paper, from Technion and IBM Research, reframes LLM explainability as a concept-level causal discovery problem over the model’s own inference process, not the external world. Using Gemini-2-Flash, Qwen3-14B, and OpenAI’s gpt-OSS-20b, it builds a four-stage pipeline: replace gold labels with LLM predictions, extract human-interpretable discriminative concepts, generate counterfactual text chains with an MCMC-inspired acceptance test, and learn a causal graph with σ-CG. The method is evaluated on three classification settings: LIBERTY disease diagnosis with 1,448 examples, IMDB sentiment analysis with 2,096 examples, and an LLM-as-a-judge Reddit preference task with 395 examples. Across 10-fold cross-validation, graph-based parent sets consistently outperform alternative concept subsets; for example, on Gemini-2-Flash disease diagnosis, prediction accuracy rises from 0.61 for competing subsets to 0.67 for the learned causal parents, while concept-node prediction reaches 0.54 versus 0.51. The MCMC augmentation uses 11 counterfactual steps per seed, up to 5 refinements per proposal, and a drift tolerance of 1 to 2 non-target concepts, and its KL-divergence traces converge toward the perfect-overlap bound with structural Hamming distance dropping to 0 for Gemini and Qwen on the synthetic and IMDB settings. The resulting graphs expose model-specific heuristics: shared clinical concepts on LIBERTY, but divergent latent features on natural data, showing that faithful LLM explanations can be made both causal and interpretable at the concept level.

Original abstract

Causal graphs provide a high-level language for making mechanisms transparent. Recent work uses Large Language Models (LLMs) to recover causal graphs of external-world processes. Instead, in this paper, we use causal graphs to model LLM inference itself, providing stakeholders with a transparent view of how the model perceives and organizes high-level concepts to produce a prediction. We propose a four-phase method for constructing such graphs. Given a target LLM and a set of textual examples, our method discovers class-discriminative, human-interpretable concepts and maps each input to LLM-perceived concept states. We then introduce an MCMC-inspired counterfactual augmentation procedure that expands the sparse observational data through chains of counterfactuals. This enables stable causal discovery with $σ$-CG, yielding informative, interpretable graphs. We apply our method to three LLMs across disease diagnosis, sentiment analysis, and LLM-as-a-judge classification tasks. We evaluate the learned graphs for predictive fidelity and structural stability, and the MCMC-inspired augmentation for convergence and downstream utility. Our results show that the discovered causal graphs capture meaningful dependencies consistent with LLMs' reasoning. Together, this paper provides a foundation for concept-level explainability of LLMs.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis