NTH

Tunable Tool-Call Rates in LLM Agents via Representation Steering

AuthorsYuqi Chen, Vincent Siu, Yang Liu, Dawn Song, Chenguang Wang

September 4, 2026 3 min read
Watch on YouTube
The one-line take

A simple inference-time steering vector lets LLM agents smoothly dial tool use up or down, improving the tradeoff between answer quality, cost, and unwanted actions.

Key results

0.29
Unsteered PopQA accuracy

Accuracy with no searches before representation steering.

0.56
Steered PopQA accuracy

Accuracy reached with live search under positive steering.

1.1
Search cost at improved accuracy

Average searches per question at 0.56 PopQA accuracy.

5
Evaluated models

Models tested across dense, MoE, and multimodal architectures.

6
Held-out tools

Unseen tools used to test tool-generalization.

What the paper found

This paper introduces a training-free way to control when an LLM agent uses external tools by steering one linear direction in its residual stream. The direction is extracted with a difference-of-means contrast over the model’s own tool-call opener probability, requiring no gradients, labels, retraining, or prompt changes; adjusting its coefficient α suppresses or induces calls while generally preserving valid tool formatting and model-specific routing. Experiments on PopQA, GSM8K, and BIG-Bench Hard show that positive steering is knowledge-selective: it increases searches mainly for obscure questions the model cannot answer from parametric memory, while reducing unnecessary calculator or Python calls on solvable tasks. With live search on PopQA, the method raises accuracy from 0.29 without searches to 0.56 at about 1.1 searches per question, tracing a cost-accuracy Pareto frontier. The shared direction also transfers to 6 unseen tools, including translation, weather, SQL, and email, and works across 5 models spanning dense, mixture-of-experts, and multimodal architectures, including Qwen3, Gemma-4-E4B, and OpenAI’s gpt-oss-20b. The strongest intervention occurs around layer 22 in Qwen3-4B. The approach could offer lightweight inference-time control for large-scale agent deployments involving platforms such as NVIDIA and OpenAI, although aggressive positive steering can produce malformed calls and does not improve tool-execution quality.

Original abstract

Deciding whether to call a tool is a core competence of an LLM agent, and a costly one to get wrong: needless calls add latency, accrue cost, and may trigger irreversible side effects, while missing calls leave the model confidently wrong on questions it could only answer through tool-calls. Models manage this balance poorly, both over-using and under-using tools. Existing methods such as post-training and prompt engineering are expensive and difficult to modify at inference time. We show that whether an instruction-tuned model calls a tool can be controlled by a single linear direction in its residual stream, extracted without any training from the model's own tool-use preference signal and turned into an inference-time intervention with no prompt change. Adding the direction with strength $α$ moves the call rate monotonically from near $0\% $ to over $90\%$ while keeping calls well-formed. The steering works in both directions: dialing it down suppresses calls, and dialing it up induces new calls that land precisely on the questions the model cannot answer from its own knowledge. We also show that the direction generalizes to unseen tools with strength comparable to each tool's own direction and without favoring any specific tool choice. With live tool execution, a single sweep of the steering traces a cost/accuracy Pareto frontier and nearly doubles open-domain QA accuracy ($0.29 \! \rightarrow \! 0.56$); the same recipe transfers across a diverse range of models spanning dense, MoE, and multimodal architectures, without any training. Our code is publicly available at https://github.com/YuqiChen4188/Steering-Tool-Use-Propensity.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis