Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning
AuthorsYu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley
This paper teaches an LLM to think more efficiently by steering its reasoning step-by-step with a controller that balances accuracy and token budget.
Key results
ACTS reduces total tokens by 59.5% on GPQA Diamond with DeepSeek-R1-Distill-Qwen-1.5B, while reaching 30.1% accuracy.
ACTS reduces total tokens by 47.7% on GPQA Diamond with DeepSeek-R1-Distill-Qwen-7B, while reaching 46.8% accuracy.
In asynchronous serving, ACTS matches Vanilla's throughput within 1% on Qwen3-8B.
What the paper found
Agentic Chain-of-Thought Steering, or ACTS, from University of California San Diego and Intuit AI Research, reframes LLM reasoning control as a Markov decision process in which a controller agent steers a frozen reasoner step by step under a thinking-token budget. Instead of only shortening or truncating chains of thought, ACTS chooses an explicit high-level strategy—UNDERSTAND, PLAN, EXECUTE, EXPLORE, CHECK, SUMMARIZE, or CONCLUDE—plus a short natural-language steering phrase that opens the next reasoning step, preserving generation continuity while making inference-time control explicit. The controller is initialized from synthetic steering trajectories extracted from OpenR1-Math traces generated by DeepSeek-R1, then improved with multi-budget supervised finetuning and GRPO reinforcement learning using budget-conditioned reward shaping that penalizes both overthinking and premature termination. Across MATH-500, AIME24, AMC, OlympiadBench, and GPQA Diamond, ACTS matches or beats vanilla full-thinking baselines while cutting tokens by 34.6% to 64.1% on DeepSeek-R1-Distill-Qwen-1.5B and 19.8% to 57.0% on Qwen3-8B; on GPQA Diamond it reaches 30.1% accuracy with 59.5% fewer tokens on the 1.5B model and 46.8% accuracy with 47.7% fewer tokens on the 7B model. The method also transfers to a different model family, Qwen3-8B, and to science QA, showing that structured steering can suppress long, confused traces without collapsing reasoning depth.
Original abstract
Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reasoning methods control thinking length by shortening, early-stopping, or compressing traces, leaving how the model thinks implicit. In this paper, we propose Agentic Chain-of-Thought Steering (ACTS), which formulates reasoning steering as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference. At each step, the controller observes the reasoning trace and remaining thinking budget, then issues a steering action consisting of a reasoning strategy and a steering phrase that initiates the next reasoner step. This enables budget-aware strategy control for efficient reasoning while preserving the reasoner's generation continuity. We initialize the controller agent from our constructed synthetic steering trajectories with multi-budget augmentation, and further optimize it via reinforcement learning with budget-conditioned reward shaping. Experiments across multiple benchmarks show that ACTS matches full-thinking performance with substantial token savings, and enables controllable accuracy-efficiency trade-offs across different reasoners and tasks. The code is available at https://github.com/Andree-9/ACTS.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.