Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
AuthorsChanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
AffiliationsKAIST · DeepAuto.ai(†\dagger: Equal advising)
Resources
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
Key results
Hard problems across six benchmarks used for FlyBy evaluation
Average pass@8 across ArXivMath, GPQA-Diamond, SuperGPQA, ChemBench, MedXpertQA, and MMLU-Pro
FlyBy-4B serves at 2.7 times lower cost than Qwen3-14B
Hard-problem pass@1, exceeding Qwen3-8B’s 15.31%
Average pass@8 after scaling the FlyBy cognitive core to 8B parameters
What the paper found
This paper argues that small reasoning models should not automatically think longer: their failures split into execution bottlenecks, where self-refinement can recover an already reachable solution, and knowledge bottlenecks, where missing external information is required. Interventions across Qwen3 and Gemma models show that epistemic verbalizations such as “wait” mostly concentrate probability on existing solution paths rather than creating new ones. The proposed FlyBy framework trains Qwen3-based 4B and 8B models to reason first, diagnose the bottleneck, and selectively query stronger backends without revealing the original problem. Supervised fine-tuning teaches query and answer-integration actions, while cost-aware reinforcement learning learns when to query, what to ask, and whether to use shallow or deep assistance from models including DeepSeek-V4-Pro. On 1,158 hard problems spanning ArXivMath, GPQA-Diamond, SuperGPQA, ChemBench, MedXpertQA, and MMLU-Pro, FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B at 2.7 times lower serving cost; it also reaches 16.85% pass@1, above Qwen3-8B’s 15.31%. Scaling to FlyBy-8B raises pass@8 to 51.81%. The learned policy performs targeted, sometimes repeated queries rather than outsourcing full solutions, and transfers to alternative backends such as OpenAI’s GPT-5.6 Luna, indicating that the key capability is diagnosing when internal reasoning is insufficient.
Original abstract
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.
Read the original paperMore in AI Reasoning
Browse all 39 papers →On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.
Neuro-Formal Verification: Agentic Language-Agnostic Formal Program Reasoning
Shuvendu K. Lahiri
NFV uses AI agents to translate ordinary code into machine-checkable formal proofs, making software verification more accessible while revealing the limits of end-to-end soundness.