NTH

The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study

AuthorsVictoria Lin, Taedong Yun, Maja Matarić, John Canny, Arthur Gretton, Alexander D'Amour

May 25, 2026 2 min read
Watch on YouTube
The one-line take

This paper argues that many LLM simulation experiments are really observational studies in disguise, and it proposes causal tools to detect and reduce the hidden bias from shifting synthetic users.

Key results

six models
Models evaluated

The study evaluates user-drift behavior across six LLMs, including Qwen3-30B-A3B, GPT-OSS-20B, Gemma-4-31B-it, Gemma-3-4B-it, and Gemini 3 Flash.

OpinionQA, Book Opinions, and MovieLens
Primary datasets

These are the three benchmark datasets used for survey-style and agent-evaluation simulations.

30 independent trials per user
Trials per user

Each synthetic user is evaluated in 30 independent trials with both intervention conditions run in parallel.

1 - Very unlikely to 4 - Very likely
Outcome scale in agent evaluation

In the book/movie recommendation setting, the primary outcome asks how likely the respondent is to read or watch the item on a four-point scale.

What the paper found

This paper shows that LLM-based “synthetic user” experiments often do not implement randomized trials; they behave like observational studies with intervention-dependent selection bias. The authors formalize “user drift” as a causal confounding problem: when a persona L is held fixed, the latent user state X \ L can still shift differently under treatment A = 0 versus A = 1, so the observed effect τ_obs estimates a mixture of treatment effect and selection bias rather than the average treatment effect. They diagnose drift with negative control outcomes Z—attributes that should be invariant to the intervention—and quantify distributional shift with total variation distance between P(Z | A = 1, L) and P(Z | A = 0, L). Across six models, including Qwen3-30B-A3B, GPT-OSS-20B, Gemma-4-31B-it, Gemma-3-4B-it, and Gemini 3 Flash, they find substantial baseline TVD in OpinionQA, Book Opinions, and MovieLens, with the smallest Gemma-3-4B showing especially large drift. They then mitigate bias by iteratively eliciting additional confounders L′, especially task-relevant variables such as political ideology, media habits, or book/movie selection behavior. This adjustment often reduces TVD and stabilizes effect estimates, but generic demographics can initially worsen bias before targeted confounders bring it down. The key novelty is methodological: the paper argues that persona prompting plus intervention is not enough for causal identification, and that internal validity in LLM simulations requires negative controls and confounder adjustment, not just better mimicry of human responses.

Original abstract

Large language models (LLMs) show potential as simulators of human behavior, offering a scalable way to study responses to interventions. However, because LLMs are trained largely on observational data, interventions in experiments with LLM-simulated synthetic users can induce unintended shifts in latent user attributes, causing user drift where the implicit simulated population differs across treatment conditions, potentially distorting effect estimates. We formalize the confounding or selection bias that can arise due to user drift and show how intervention-dependent shifts can inflate or attenuate observed differences in user responses under intervention. To diagnose confounding, we propose using negative control outcomes--attributes that should remain invariant under intervention--to identify distribution shifts across intervention conditions, providing evidence of user drift. To mitigate drift, we study adjusting the persona specification by eliciting additional confounders, finding that targeted, setting-relevant confounders can substantially reduce bias across survey-style and multi-turn agent evaluations.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis