NTH

When is Your LLM Steerable?

AuthorsChenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou

June 28, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that you can predict whether an LLM steering trick will work by looking at its early hidden states, then use that prediction to search for better steering settings much more cheaply.

Key results

1.42M
ASTEER generations

labeled steered generations in the dataset

150
concepts

steering concepts covered by ASTEER

50
prompts

prompts reused across concepts

0.7
macro-F1

SteerBoost performance on unseen concepts

2
first tokens importance

first decoded tokens carrying over 75% of importance mass

98%
oracle recovery

success rate recovered by SteerBoost-guided search at K=20

What the paper found

This paper from the University of Maryland and MBZUAI asks a precise question: when is an LLM steerable, and can that be predicted before a full generation finishes? The authors build ASTEER, a large testbed of 1.42M labeled steered generations covering 150 concepts, 50 prompts, 3 LLMs, and 2 steering methods, DiffMean and Probe, with outcomes annotated into UnderSteer, SuccSteer, and OverSteer using GPT-5-nano and validated against human judgment with Cohen’s κ of 0.83. They then propose SteerBoost, a Gradient Boosting Decision Trees classifier that reads only the first few decoded tokens and compares steered versus unsteered hidden states across a token-layer grid, using interpretable geometry and decoding-dynamics features such as SteeringAffinity and DeviationAlignment. On held-out unseen concepts, SteerBoost reaches around 0.7 macro-F1, with DiffMean performing strongest at about 0.80 on in-distribution concepts and about 0.72 out-of-distribution, while Probe is harder to predict at roughly 0.68 to 0.74 ID and 0.65 to 0.69 OOD. The model shows that the first 2 decoded tokens carry over 75% of the feature importance mass, and it enables steering-strength search that recovers about 98% of item-level oracle success while using only about 11% of the decoded tokens required by exhaustive grid search, making steerability prediction a practical shortcut for cheaper, near-optimal activation steering.

Original abstract

Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration. Finding the regime and boundaries of successful steering typically requires expensive grid searches and post-hoc evaluation of full autoregressive rollouts. In this work, we investigate whether steerability can be predicted from the model's internal states at the beginning of the generation process, e.g., after generating the first few tokens, and how to leverage such a predictor to improve steering success rate. To this end, we first introduce ASTEER, a testbed including 1.4M steered generations, spanning 150 concepts with each steering success/failure labeled. Leveraging this testbed, we analyze the model's early decoding dynamics by extracting features that compare hidden states before and after steering across layers and initial decoding steps. These features help us understand how steering's effects propagate along layers and token positions, which provide key information for steerability prediction. We then train a Gradient Boosting Decision Trees (GBDT) classifier on these features to predict whether an intervention will under-steer, succeed, or over-steer without requiring full rollout. Our predictor achieves around 0.7 macro-F1 score on unseen concepts, demonstrating that early hidden states encode substantial, structured information about eventual steering efficacy. We further leverage this steerability predictor as guidance for steering strength searching, achieving near-optimal performance with a small fraction of decoding cost.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis