The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World
AuthorsAshritha Gonuguntla
Resources
Replay-based tests may judge LLM routers in a world that agents never actually encounter, so this paper replaces static logs with live branching rollouts.
Key results
Approximate number of rollouts used across six paired experiments
Normalized edit-distance increase for a model swap over matched controls
Lower bound of early swaps diverging at the first post-fork action
Share of replayed states remaining valid for early swaps
Lower bound of similarity between replay-predicted and actually produced patches
What the paper found
This paper argues that static replay benchmarks evaluate the wrong world when routing models inside multi-step LLM agents. Using branching rollouts on SWE-bench Verified, the researchers forked live mini-SWE-agent trajectories, reconstructed each environment, and continued them with either Qwen3-4B-Instruct served in FP8 or Qwen3-14B served in AWQ, producing about 900 rollouts across six paired experiments. Model swaps rewrote 61–94% of post-fork actions, exceeding same-model control noise by as much as 0.66 normalized edit distance; 74–77% of early swaps diverged on the very first new action, leaving only 3% of replayed states valid. The failure is causal: the agent’s action changes the environment, so later observations and tool calls no longer match the logged trajectory. All 5 observed outcome flips occurred in swap branches, while 0 occurred across 359 control forks. A log-stitching evaluator mispredicted every success-relevant outcome and generated patches with 0.00–0.11 similarity to the patches actually produced. The study also shows that temperature-0 determinism depends on serving configuration: FP8 controls diverged on over 90% of forks, whereas AWQ controls were near-identical. The authors recommend live or branched evaluation, with off-policy estimators such as importance sampling validated against branch-level ground truth, and caution that stronger models can exhaust a 50-step budget more often because they explore and verify more.
Original abstract
LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged trajectories and substituting another model's recorded outputs, assuming the rest of the trajectory is unaffected. We test this assumption with branching rollouts: we fork live SWE-bench agent trajectories at controlled points, rebuild the environment, continue each fork with a different model, and compare against same-model control forks that isolate sampling and replay noise. Across six paired runs (~900 rollouts), swaps exceed their matched control floors by +0.25 to +0.66 normalized edit distance (multiplicity-corrected CIs exclude zero), rewriting 61-94% of post-fork actions; 74-77% of early swaps diverge at the first post-fork action, versus 6-35% of controls, leaving only 3% of replayed states valid. Divergence decreases with fork depth in both directions. All five outcome flips we observe occur in swap arms, upgrades rescuing unsolved instances and a downgrade losing the sole solve, and zero occur across 359 control forks. Scoring these same swaps with a log-stitching replay evaluator, replay mispredicts every success-relevant outcome call and predicts patches with 0.00-0.11 similarity to reality. Auditing the noise floor, temperature-0 "determinism" is configuration-dependent: FP8-served controls diverge on over 90% of forks while AWQ-served ones remain near-identical; and under tight budgets the stronger model more often exhausts its steps without submitting. Replay-based benchmarks score the wrong world for agentic routing; we release our harness and all trajectories.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.