NTH

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

AuthorsChanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang

AffiliationsKAIST · DeepAuto.ai(†\dagger: Equal advising)

October 5, 2026 2 min read
Watch on YouTube
The one-line take

FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.

Key results

1,158
Evaluation problems

Hard problems across six benchmarks used for FlyBy evaluation

45.96%
FlyBy-4B pass@8

Average pass@8 across ArXivMath, GPQA-Diamond, SuperGPQA, ChemBench, MedXpertQA, and MMLU-Pro

2.7
Cost advantage over Qwen3-14B

FlyBy-4B serves at 2.7 times lower cost than Qwen3-14B

16.85%
FlyBy-4B pass@1

Hard-problem pass@1, exceeding Qwen3-8B’s 15.31%

51.81%
FlyBy-8B pass@8

Average pass@8 after scaling the FlyBy cognitive core to 8B parameters

What the paper found

This paper argues that small reasoning models should not automatically think longer: their failures split into execution bottlenecks, where self-refinement can recover an already reachable solution, and knowledge bottlenecks, where missing external information is required. Interventions across Qwen3 and Gemma models show that epistemic verbalizations such as “wait” mostly concentrate probability on existing solution paths rather than creating new ones. The proposed FlyBy framework trains Qwen3-based 4B and 8B models to reason first, diagnose the bottleneck, and selectively query stronger backends without revealing the original problem. Supervised fine-tuning teaches query and answer-integration actions, while cost-aware reinforcement learning learns when to query, what to ask, and whether to use shallow or deep assistance from models including DeepSeek-V4-Pro. On 1,158 hard problems spanning ArXivMath, GPQA-Diamond, SuperGPQA, ChemBench, MedXpertQA, and MMLU-Pro, FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B at 2.7 times lower serving cost; it also reaches 16.85% pass@1, above Qwen3-8B’s 15.31%. Scaling to FlyBy-8B raises pass@8 to 51.81%. The learned policy performs targeted, sometimes repeated queries rather than outsourcing full solutions, and transfers to alternative backends such as OpenAI’s GPT-5.6 Luna, indicating that the key capability is diagnosing when internal reasoning is insufficient.

Original abstract

Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
01Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis