Tunable Tool-Call Rates in LLM Agents via Representation Steering
AuthorsYuqi Chen, Vincent Siu, Yang Liu, Dawn Song, Chenguang Wang
A simple inference-time steering vector lets LLM agents smoothly dial tool use up or down, improving the tradeoff between answer quality, cost, and unwanted actions.
Key results
Accuracy with no searches before representation steering.
Accuracy reached with live search under positive steering.
Average searches per question at 0.56 PopQA accuracy.
Models tested across dense, MoE, and multimodal architectures.
Unseen tools used to test tool-generalization.
What the paper found
This paper introduces a training-free way to control when an LLM agent uses external tools by steering one linear direction in its residual stream. The direction is extracted with a difference-of-means contrast over the model’s own tool-call opener probability, requiring no gradients, labels, retraining, or prompt changes; adjusting its coefficient α suppresses or induces calls while generally preserving valid tool formatting and model-specific routing. Experiments on PopQA, GSM8K, and BIG-Bench Hard show that positive steering is knowledge-selective: it increases searches mainly for obscure questions the model cannot answer from parametric memory, while reducing unnecessary calculator or Python calls on solvable tasks. With live search on PopQA, the method raises accuracy from 0.29 without searches to 0.56 at about 1.1 searches per question, tracing a cost-accuracy Pareto frontier. The shared direction also transfers to 6 unseen tools, including translation, weather, SQL, and email, and works across 5 models spanning dense, mixture-of-experts, and multimodal architectures, including Qwen3, Gemma-4-E4B, and OpenAI’s gpt-oss-20b. The strongest intervention occurs around layer 22 in Qwen3-4B. The approach could offer lightweight inference-time control for large-scale agent deployments involving platforms such as NVIDIA and OpenAI, although aggressive positive steering can produce malformed calls and does not improve tool-execution quality.
Original abstract
Deciding whether to call a tool is a core competence of an LLM agent, and a costly one to get wrong: needless calls add latency, accrue cost, and may trigger irreversible side effects, while missing calls leave the model confidently wrong on questions it could only answer through tool-calls. Models manage this balance poorly, both over-using and under-using tools. Existing methods such as post-training and prompt engineering are expensive and difficult to modify at inference time. We show that whether an instruction-tuned model calls a tool can be controlled by a single linear direction in its residual stream, extracted without any training from the model's own tool-use preference signal and turned into an inference-time intervention with no prompt change. Adding the direction with strength $α$ moves the call rate monotonically from near $0\% $ to over $90\%$ while keeping calls well-formed. The steering works in both directions: dialing it down suppresses calls, and dialing it up induces new calls that land precisely on the questions the model cannot answer from its own knowledge. We also show that the direction generalizes to unseen tools with strength comparable to each tool's own direction and without favoring any specific tool choice. With live tool execution, a single sweep of the steering traces a cost/accuracy Pareto frontier and nearly doubles open-domain QA accuracy ($0.29 \! \rightarrow \! 0.56$); the same recipe transfers across a diverse range of models spanning dense, MoE, and multimodal architectures, without any training. Our code is publicly available at https://github.com/YuqiChen4188/Steering-Tool-Use-Propensity.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.