NTH

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

AuthorsDingwei Chen, Zefang Zong, Zhipeng Ma, Leo Luo, Yang Li, Chengming Li, Peng Chen, Jie Jiang

June 30, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches LLM agents to know when they really need a tool—and when they should trust their own knowledge—cutting unnecessary tool calls while improving accuracy.

Key results

1.85
avg EM gain

Average Exact Match improvement across seven QA benchmarks

18%
tool call reduction

Average reduction in tool calls over standard agentic RL

25%
tool productivity gain

Reported increase in tool productivity

3.16
Qwen3-4B Multi-Hop TC

GRPO baseline average tool calls on Qwen3-4B Multi-Hop

2.60
Qwen3-4B Multi-Hop TC

AKBE average tool calls on Qwen3-4B Multi-Hop

15%
training speedup

Average per-step training time reduction versus GRPO

What the paper found

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement, proposed by Tencent researchers, introduces AKBE, an on-policy module for LLM agents that detects whether a question lies inside the model’s intrinsic knowledge boundary by running parallel with-tool and no-tool rollouts, then supervising the policy with boundary-guided trajectories instead of reward shaping. The method separates three failure modes: Tool-dependent cases, where tools are necessary and the minimum correct tool-call trajectory is reinforced; Efficiency cases, where tools are unnecessary and no-tool correct trajectories suppress redundant calls; and Hallucination cases, where tools actively mislead the model and no-tool trajectories are preferred. Across seven QA benchmarks on Qwen3-4B and Qwen2.5-7B, AKBE raises average Exact Match by 1.85 while cutting average tool calls by 18%, which the paper translates into about 25% higher tool productivity. On Qwen3-4B Multi-Hop, the average tool-call count drops from 3.16 to 2.60, and in plug-and-play experiments AKBE also improves GRPO, DAPO, GSPO, and AEPO. The authors report that the method is 15% faster per training step despite adding no-tool rollouts, because no-tool inference is cheaper and later rollouts shorten as tool use declines. An ablation shows that removing Tool-dependent supervision sharply hurts accuracy, confirming that AKBE’s gains come from preserving necessary tool use while eliminating redundant calls.

Original abstract

Agentic reinforcement learning (RL) has proven effective for training LLM-based agents with external tool-use capabilities. However, we identify that agentic RL training induces increasing redundant tool calls and blurs the model's intrinsic knowledge boundary, where the model fails to distinguish when tools are needed versus when parametric knowledge suffices. Existing solutions based on reward shaping create coarse-grained optimization targets that tend to incentivize indiscriminate tool-call suppression, leading to reward hacking. In this paper, we propose AKBE (Agentic Knowledge Boundary Enhancement), an on-policy method that dynamically probes the model's intrinsic knowledge boundary through dual-path (with-tool and no-tool) rollouts during training. We define the knowledge boundary as the per-instance determination of whether tools are required and the minimum tool calls necessary. By comparing correctness across paths, AKBE categorizes trajectories and constructs targeted supervisory signals that guide efficient tool-use patterns for each question. These signals are integrated seamlessly into the agentic RL training loop. Experiments on seven QA benchmarks demonstrate that AKBE improves task accuracy by +1.85 on average and reduces tool calls by 18% over standard agentic RL, yielding 25% higher tool productivity without any accuracy-efficiency trade-off. Further analysis suggests its plug-and-play compatibility across different RL algorithms and the mechanism of each signal category. Our code is available at https://github.com/CuSO4-Chen/AKBE.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis