AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
AuthorsCheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke
Resources
A powerful model can effectively teach a weaker one new problem-solving habits at test time by building a smart software harness around it.
Key results
Total evaluation items across BigToM, Hi-ToM, MMToM-QA, and MuMA-ToM.
Portion of benchmark data exposed to builders during harness construction.
Macro-average accuracy without a scaffold.
GPT-5.5-built harness evaluated with GPT-5.4-mini through GPT Codex.
Pearson correlation between deterministic workload fraction and final accuracy.
What the paper found
This paper introduces strong-to-weak scaffolding, a form of test-time capability transfer in which a powerful builder model designs an inference-time harness for a weaker target without changing its parameters. Across 4 Theory-of-Mind benchmarks—BigToM, Hi-ToM, MMToM-QA, and MuMA-ToM—the system used 5% validation data to iteratively create routing rules, prompt templates, deterministic solvers, state extraction, verification, and strict output formatting, then evaluated on a 3900-item hidden test set. For GPT-5.4-mini, direct prompting achieved only 0.488 macro-average accuracy, while the best harness, built by GPT-5.5 through GPT Codex, reached 0.912. The strongest gains came not from longer chain-of-thought or broader sampling, but from externalizing unstable reasoning into code, routing each benchmark by subtype, enforcing answer formats, and reducing the target model’s cognitive load. Deterministic offloading correlated with final accuracy at 0.72, and builder reasoning effort improved harness quality monotonically for Opus-4.7. Platform choice mattered less than builder capability, with Cursor and Claude Code showing only modest, conditional differences. Weaker targets benefited most: GPT-5.4-mini gained substantially more than Gemini-3.5-flash, whose already strong performance sometimes declined under over-scaffolding. The remaining errors concentrated in higher-order Hi-ToM recursion, deception, and Bayesian goal inference, showing that harnesses can compile regular structure but do not eliminate the need for genuine model reasoning.
Original abstract
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such transfer can instead occur at test time. We study strong-to-weak scaffolding: whether a stronger builder model can construct inference-time harnesses that help a weaker target model solve tasks more reliably without any parameter updates. Using four representative Theory-of-Mind benchmarks, each builder model uses 5% of the data as a validation set to iteratively refine its harness over multiple rounds, after which the finalized harness is evaluated on the full test set. Empirically, this form of test-time capability transfer is highly effective, nearly doubling average target-model performance from 0.49 to 0.91. Our analysis shows that the gains come primarily from offloading unstable model reasoning into deterministic code, benchmark-specific routing, and strict answer-format enforcement, rather than from encouraging the target model to reason more extensively or sample more broadly. We further find that builder-model reasoning effort improves harness quality monotonically, platform effects are modest relative to the builder model's own capability, and weaker target models receive the largest gains. These results suggest that inference-time harness design is an important complement to conventional training-time distillation, enabling strong models to transfer cognitive structure to weaker models without retraining.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.