NTH

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

AuthorsCheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke

August 20, 2026 2 min read
Watch on YouTube
The one-line take

A powerful model can effectively teach a weaker one new problem-solving habits at test time by building a smart software harness around it.

Key results

3900
Hidden test set

Total evaluation items across BigToM, Hi-ToM, MMToM-QA, and MuMA-ToM.

5%
Validation-data share

Portion of benchmark data exposed to builders during harness construction.

0.488
GPT-5.4-mini vanilla accuracy

Macro-average accuracy without a scaffold.

0.912
Best scaffold accuracy

GPT-5.5-built harness evaluated with GPT-5.4-mini through GPT Codex.

0.72
Deterministic-offloading correlation

Pearson correlation between deterministic workload fraction and final accuracy.

What the paper found

This paper introduces strong-to-weak scaffolding, a form of test-time capability transfer in which a powerful builder model designs an inference-time harness for a weaker target without changing its parameters. Across 4 Theory-of-Mind benchmarks—BigToM, Hi-ToM, MMToM-QA, and MuMA-ToM—the system used 5% validation data to iteratively create routing rules, prompt templates, deterministic solvers, state extraction, verification, and strict output formatting, then evaluated on a 3900-item hidden test set. For GPT-5.4-mini, direct prompting achieved only 0.488 macro-average accuracy, while the best harness, built by GPT-5.5 through GPT Codex, reached 0.912. The strongest gains came not from longer chain-of-thought or broader sampling, but from externalizing unstable reasoning into code, routing each benchmark by subtype, enforcing answer formats, and reducing the target model’s cognitive load. Deterministic offloading correlated with final accuracy at 0.72, and builder reasoning effort improved harness quality monotonically for Opus-4.7. Platform choice mattered less than builder capability, with Cursor and Claude Code showing only modest, conditional differences. Weaker targets benefited most: GPT-5.4-mini gained substantially more than Gemini-3.5-flash, whose already strong performance sometimes declined under over-scaffolding. The remaining errors concentrated in higher-order Hi-ToM recursion, deception, and Bayesian goal inference, showing that harnesses can compile regular structure but do not eliminate the need for genuine model reasoning.

Original abstract

Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such transfer can instead occur at test time. We study strong-to-weak scaffolding: whether a stronger builder model can construct inference-time harnesses that help a weaker target model solve tasks more reliably without any parameter updates. Using four representative Theory-of-Mind benchmarks, each builder model uses 5% of the data as a validation set to iteratively refine its harness over multiple rounds, after which the finalized harness is evaluated on the full test set. Empirically, this form of test-time capability transfer is highly effective, nearly doubling average target-model performance from 0.49 to 0.91. Our analysis shows that the gains come primarily from offloading unstable model reasoning into deterministic code, benchmark-specific routing, and strict answer-format enforcement, rather than from encouraging the target model to reason more extensively or sample more broadly. We further find that builder-model reasoning effort improves harness quality monotonically, platform effects are modest relative to the builder model's own capability, and weaker target models receive the largest gains. These results suggest that inference-time harness design is an important complement to conventional training-time distillation, enabling strong models to transfer cognitive structure to weaker models without retraining.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis