Re-Centering Humans in LLM Personalization
AuthorsLechen Zhang, Jiarui Liu, Tal August
Resources
This paper shows that LLM personalization looks much better on synthetic tests than on real human conversations, and introduces human-grounded data and evaluation methods to expose the gap.
Key results
human conversation dataset size used in the study
human annotations for Stage 1 attribute extraction
human annotations for Stage 2 relevance matching
human annotations for Stage 3 personalized response generation
most common reason uncertain attributes were rejected or flagged
best attribute verifier performance in Stage 1
What the paper found
Re-Centering Humans in LLM Personalization, from the University of Illinois Urbana-Champaign and Carnegie Mellon University, argues that synthetic personas and LLM judges systematically overstate how well models personalize for real people. The paper decomposes personalization into three stages: extracting stable user attributes from conversation history, matching those attributes to a new prompt, and generating a response that beats a generic baseline. Using 550 real conversations from WildChat and 5,949 attribute-validity judgments, 11,919 attribute–prompt relevance judgments, and 1,101 response-preference judgments, the authors show that real conversations are harder than synthetic data: extracted attributes from human dialogue contain many more uncertain or rejected cases, with overgeneralization accounting for 53.9% of failures. In relevance matching, humans mark only about 20% of attributes as relevant, while LLM judges mark 40–60%, with average human–LLM agreement dropping to κ≈0.30. Lightweight interventions help: a trained RoBERTa verifier reaches 0.726 F1 for attribute verification, and a Qwen3-4B model optimized with GRPO reaches 0.641 F1 for relevance selection. But the third stage remains difficult: humans judge 54.6% of personalized responses as no better than generic, and LLM judges show only modest alignment with human ratings, with Spearman correlation peaking around 0.376. The core message is that personalization quality must be grounded in human judgments, not just synthetic benchmarks or surface-level attribute mentions.
Original abstract
Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data. It remains unclear how well current personalization systems work for real users. In this paper, we study the gap in LLM personalization performance when using synthetic versus human data. We collect human conversations (550 conversations) and judgments across three stages of personalization: extracting user attributes from conversations (5,949 judgments), pairing relevant attributes with new prompts (11,919), and incorporating relevant attributes into a personalized response (1,101). Incorporating human data reveals system limitations at each stage. Models struggle to extract attributes from human conversations, disagree with human judgments on relevant attributes, and generate personalized responses that humans judge no better than generic responses (though that LLM judges widely rate as better). We introduce two lightweight training-based interventions that shift automated personalization evaluation closer to human data in our first two stages. However, in our third stage we find that learned reward models achieve only modest correlation with human ratings, suggesting that human-aligned personalization quality judgments are difficult to model directly. Our collected data provides a foundation for studying how models should extract, select, and incorporate user information in ways that humans find useful.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.