NTH

On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance

AuthorsEtienne Casanova, Rafal Kocielnik, R. Michael Alvarez

July 9, 2026 2 min read
Watch on YouTube
The one-line take

The paper shows that LLMs can be stubborn annotators: once they form a confident answer, extra prompting often fails to change their mind, and what matters most is whether the model’s internal concept matches the task definition.

Key results

0.41
DSF partial correlation

Positive association with zero-shot accuracy after controlling for dataset on the N = 54 panel

34.8%
Overall rescue rate

Fraction of zero-shot errors corrected by prompting

20.8%
High-confidence rescue probability

Rescue probability for zero-shot errors with confidence above 0.9

What the paper found

This paper examines how model-internalized priors constrain LLM annotation on toxicity detection, using 9 instruction-tuned models including OpenAI’s GPT-4o-mini, Meta’s Llama-3.1/3.3, Mistral, DeepSeek-V3, and Qwen-2.5-72B across 5 datasets spanning Twitter, OLID, GameTox, Fox News, and Jigsaw. The central result is that Definition-Specific Familiarity (DSF)—semantic alignment between a model’s own concept description and the task definition—predicts zero-shot accuracy after controlling for dataset difficulty, with partial r = +0.41 on an N = 54 model–dataset panel, while three memorization metrics, ROUGE-L, BERTScore, and embedding similarity, show no positive association. Prompting has limited corrective power: the overall rescue rate is only 34.8%, meaning nearly two-thirds of zero-shot errors persist even under aligned definitions, few-shot examples, and DSPy optimization. High-confidence errors are especially sticky, with rescue probability falling to 20.8% for confidence above 0.9, and confidence remains similarly high under misaligned definitions, exposing a calibration failure. Misaligned prompts can shift prediction bias by about −12.1% to +13.1% depending on definition scope, yet models stay overconfident rather than signaling uncertainty. The practical takeaway is that annotation quality depends more on definition alignment than on text memorization or model scale, and the paper recommends DSF-style checks before deployment.

Original abstract

Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions. We investigate three dimensions of this interaction: (1) how an LLM's familiarity with data and task definitions affects performance, (2) the extent to which additional information in prompts can correct zero-shot errors ("decision stickiness"), and (3) model susceptibility to misaligned task definitions. Through experiments on toxicity detection across diverse datasets (spanning social media, gaming, news, and forums) using both dense and mixture-of-experts models, we find that nearly two-thirds of zero-shot errors are resistant to correction, with an overall rescue rate (fraction of initial errors corrected by prompting) of only 34.8%. High-confidence errors prove especially resistant to correction. When given misaligned definitions, LLMs follow them while maintaining confidence levels unchanged from the aligned condition. Crucially, we introduce Definition-Specific Familiarity (DSF), which measures alignment between a model's internal concept and the task definition. After controlling for dataset-level confounds, DSF shows a positive association with model performance (partial r = +0.41), while three distinct memorization metrics (ROUGE-L, BERTScore, and embedding cosine similarity) all fail to show a positive association. These findings show the limitations of prompt-based correction in annotation tasks, highlighting the importance of definition alignment over text-level memorization.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis