When is Your LLM Steerable?
AuthorsChenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou
Resources
This paper shows that you can predict whether an LLM steering trick will work by looking at its early hidden states, then use that prediction to search for better steering settings much more cheaply.
Key results
labeled steered generations in the dataset
steering concepts covered by ASTEER
prompts reused across concepts
SteerBoost performance on unseen concepts
first decoded tokens carrying over 75% of importance mass
success rate recovered by SteerBoost-guided search at K=20
What the paper found
This paper from the University of Maryland and MBZUAI asks a precise question: when is an LLM steerable, and can that be predicted before a full generation finishes? The authors build ASTEER, a large testbed of 1.42M labeled steered generations covering 150 concepts, 50 prompts, 3 LLMs, and 2 steering methods, DiffMean and Probe, with outcomes annotated into UnderSteer, SuccSteer, and OverSteer using GPT-5-nano and validated against human judgment with Cohen’s κ of 0.83. They then propose SteerBoost, a Gradient Boosting Decision Trees classifier that reads only the first few decoded tokens and compares steered versus unsteered hidden states across a token-layer grid, using interpretable geometry and decoding-dynamics features such as SteeringAffinity and DeviationAlignment. On held-out unseen concepts, SteerBoost reaches around 0.7 macro-F1, with DiffMean performing strongest at about 0.80 on in-distribution concepts and about 0.72 out-of-distribution, while Probe is harder to predict at roughly 0.68 to 0.74 ID and 0.65 to 0.69 OOD. The model shows that the first 2 decoded tokens carry over 75% of the feature importance mass, and it enables steering-strength search that recovers about 98% of item-level oracle success while using only about 11% of the decoded tokens required by exhaustive grid search, making steerability prediction a practical shortcut for cheaper, near-optimal activation steering.
Original abstract
Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration. Finding the regime and boundaries of successful steering typically requires expensive grid searches and post-hoc evaluation of full autoregressive rollouts. In this work, we investigate whether steerability can be predicted from the model's internal states at the beginning of the generation process, e.g., after generating the first few tokens, and how to leverage such a predictor to improve steering success rate. To this end, we first introduce ASTEER, a testbed including 1.4M steered generations, spanning 150 concepts with each steering success/failure labeled. Leveraging this testbed, we analyze the model's early decoding dynamics by extracting features that compare hidden states before and after steering across layers and initial decoding steps. These features help us understand how steering's effects propagate along layers and token positions, which provide key information for steerability prediction. We then train a Gradient Boosting Decision Trees (GBDT) classifier on these features to predict whether an intervention will under-steer, succeed, or over-steer without requiring full rollout. Our predictor achieves around 0.7 macro-F1 score on unseen concepts, demonstrating that early hidden states encode substantial, structured information about eventual steering efficacy. We further leverage this steerability predictor as guidance for steering strength searching, achieving near-optimal performance with a small fraction of decoding cost.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.