On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
AuthorsEtienne Casanova, Rafal Kocielnik, R. Michael Alvarez
Resources
The paper shows that LLMs can be stubborn annotators: once they form a confident answer, extra prompting often fails to change their mind, and what matters most is whether the model’s internal concept matches the task definition.
Key results
Positive association with zero-shot accuracy after controlling for dataset on the N = 54 panel
Fraction of zero-shot errors corrected by prompting
Rescue probability for zero-shot errors with confidence above 0.9
What the paper found
This paper examines how model-internalized priors constrain LLM annotation on toxicity detection, using 9 instruction-tuned models including OpenAI’s GPT-4o-mini, Meta’s Llama-3.1/3.3, Mistral, DeepSeek-V3, and Qwen-2.5-72B across 5 datasets spanning Twitter, OLID, GameTox, Fox News, and Jigsaw. The central result is that Definition-Specific Familiarity (DSF)—semantic alignment between a model’s own concept description and the task definition—predicts zero-shot accuracy after controlling for dataset difficulty, with partial r = +0.41 on an N = 54 model–dataset panel, while three memorization metrics, ROUGE-L, BERTScore, and embedding similarity, show no positive association. Prompting has limited corrective power: the overall rescue rate is only 34.8%, meaning nearly two-thirds of zero-shot errors persist even under aligned definitions, few-shot examples, and DSPy optimization. High-confidence errors are especially sticky, with rescue probability falling to 20.8% for confidence above 0.9, and confidence remains similarly high under misaligned definitions, exposing a calibration failure. Misaligned prompts can shift prediction bias by about −12.1% to +13.1% depending on definition scope, yet models stay overconfident rather than signaling uncertainty. The practical takeaway is that annotation quality depends more on definition alignment than on text memorization or model scale, and the paper recommends DSF-style checks before deployment.
Original abstract
Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions. We investigate three dimensions of this interaction: (1) how an LLM's familiarity with data and task definitions affects performance, (2) the extent to which additional information in prompts can correct zero-shot errors ("decision stickiness"), and (3) model susceptibility to misaligned task definitions. Through experiments on toxicity detection across diverse datasets (spanning social media, gaming, news, and forums) using both dense and mixture-of-experts models, we find that nearly two-thirds of zero-shot errors are resistant to correction, with an overall rescue rate (fraction of initial errors corrected by prompting) of only 34.8%. High-confidence errors prove especially resistant to correction. When given misaligned definitions, LLMs follow them while maintaining confidence levels unchanged from the aligned condition. Crucially, we introduce Definition-Specific Familiarity (DSF), which measures alignment between a model's internal concept and the task definition. After controlling for dataset-level confounds, DSF shows a positive association with model performance (partial r = +0.41), while three distinct memorization metrics (ROUGE-L, BERTScore, and embedding cosine similarity) all fail to show a positive association. These findings show the limitations of prompt-based correction in annotation tasks, highlighting the importance of definition alignment over text-level memorization.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.