NTH

Re-Centering Humans in LLM Personalization

AuthorsLechen Zhang, Jiarui Liu, Tal August

June 20, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that LLM personalization looks much better on synthetic tests than on real human conversations, and introduces human-grounded data and evaluation methods to expose the gap.

Key results

550
real conversations

human conversation dataset size used in the study

5949
attribute validity judgments

human annotations for Stage 1 attribute extraction

11919
attribute-prompt relevance judgments

human annotations for Stage 2 relevance matching

1101
response-preference judgments

human annotations for Stage 3 personalized response generation

53.9%
overgeneralization failures

most common reason uncertain attributes were rejected or flagged

0.726
RoBERTa F1

best attribute verifier performance in Stage 1

What the paper found

Re-Centering Humans in LLM Personalization, from the University of Illinois Urbana-Champaign and Carnegie Mellon University, argues that synthetic personas and LLM judges systematically overstate how well models personalize for real people. The paper decomposes personalization into three stages: extracting stable user attributes from conversation history, matching those attributes to a new prompt, and generating a response that beats a generic baseline. Using 550 real conversations from WildChat and 5,949 attribute-validity judgments, 11,919 attribute–prompt relevance judgments, and 1,101 response-preference judgments, the authors show that real conversations are harder than synthetic data: extracted attributes from human dialogue contain many more uncertain or rejected cases, with overgeneralization accounting for 53.9% of failures. In relevance matching, humans mark only about 20% of attributes as relevant, while LLM judges mark 40–60%, with average human–LLM agreement dropping to κ≈0.30. Lightweight interventions help: a trained RoBERTa verifier reaches 0.726 F1 for attribute verification, and a Qwen3-4B model optimized with GRPO reaches 0.641 F1 for relevance selection. But the third stage remains difficult: humans judge 54.6% of personalized responses as no better than generic, and LLM judges show only modest alignment with human ratings, with Spearman correlation peaking around 0.376. The core message is that personalization quality must be grounded in human judgments, not just synthetic benchmarks or surface-level attribute mentions.

Original abstract

Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data. It remains unclear how well current personalization systems work for real users. In this paper, we study the gap in LLM personalization performance when using synthetic versus human data. We collect human conversations (550 conversations) and judgments across three stages of personalization: extracting user attributes from conversations (5,949 judgments), pairing relevant attributes with new prompts (11,919), and incorporating relevant attributes into a personalized response (1,101). Incorporating human data reveals system limitations at each stage. Models struggle to extract attributes from human conversations, disagree with human judgments on relevant attributes, and generate personalized responses that humans judge no better than generic responses (though that LLM judges widely rate as better). We introduce two lightweight training-based interventions that shift automated personalization evaluation closer to human data in our first two stages. However, in our third stage we find that learned reward models achieve only modest correlation with human ratings, suggesting that human-aligned personalization quality judgments are difficult to model directly. Our collected data provides a foundation for studying how models should extract, select, and incorporate user information in ways that humans find useful.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis