NTH

Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

AuthorsYu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley

June 8, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches an LLM to think more efficiently by steering its reasoning step-by-step with a controller that balances accuracy and token budget.

Key results

59.5%
DeepSeek-R1-Distill-Qwen-1.5B token savings on GPQA Diamond

ACTS reduces total tokens by 59.5% on GPQA Diamond with DeepSeek-R1-Distill-Qwen-1.5B, while reaching 30.1% accuracy.

47.7%
DeepSeek-R1-Distill-Qwen-7B token savings on GPQA Diamond

ACTS reduces total tokens by 47.7% on GPQA Diamond with DeepSeek-R1-Distill-Qwen-7B, while reaching 46.8% accuracy.

1%
ACTS inference throughput vs Vanilla on Qwen3-8B

In asynchronous serving, ACTS matches Vanilla's throughput within 1% on Qwen3-8B.

What the paper found

Agentic Chain-of-Thought Steering, or ACTS, from University of California San Diego and Intuit AI Research, reframes LLM reasoning control as a Markov decision process in which a controller agent steers a frozen reasoner step by step under a thinking-token budget. Instead of only shortening or truncating chains of thought, ACTS chooses an explicit high-level strategy—UNDERSTAND, PLAN, EXECUTE, EXPLORE, CHECK, SUMMARIZE, or CONCLUDE—plus a short natural-language steering phrase that opens the next reasoning step, preserving generation continuity while making inference-time control explicit. The controller is initialized from synthetic steering trajectories extracted from OpenR1-Math traces generated by DeepSeek-R1, then improved with multi-budget supervised finetuning and GRPO reinforcement learning using budget-conditioned reward shaping that penalizes both overthinking and premature termination. Across MATH-500, AIME24, AMC, OlympiadBench, and GPQA Diamond, ACTS matches or beats vanilla full-thinking baselines while cutting tokens by 34.6% to 64.1% on DeepSeek-R1-Distill-Qwen-1.5B and 19.8% to 57.0% on Qwen3-8B; on GPQA Diamond it reaches 30.1% accuracy with 59.5% fewer tokens on the 1.5B model and 46.8% accuracy with 47.7% fewer tokens on the 7B model. The method also transfers to a different model family, Qwen3-8B, and to science QA, showing that structured steering can suppress long, confused traces without collapsing reasoning depth.

Original abstract

Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reasoning methods control thinking length by shortening, early-stopping, or compressing traces, leaving how the model thinks implicit. In this paper, we propose Agentic Chain-of-Thought Steering (ACTS), which formulates reasoning steering as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference. At each step, the controller observes the reasoning trace and remaining thinking budget, then issues a steering action consisting of a reasoning strategy and a steering phrase that initiates the next reasoner step. This enables budget-aware strategy control for efficient reasoning while preserving the reasoner's generation continuity. We initialize the controller agent from our constructed synthetic steering trajectories with multi-budget augmentation, and further optimize it via reinforcement learning with budget-conditioned reward shaping. Experiments across multiple benchmarks show that ACTS matches full-thinking performance with substantial token savings, and enables controllable accuracy-efficiency trade-offs across different reasoners and tasks. The code is available at https://github.com/Andree-9/ACTS.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis