ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions
AuthorsChuanyang Jin, Binze Li, Haopeng Xie, Cathy Mengying Fang, Tianjian Li, Shayne Longpre, Hongxiang Gu, Maximillian Chen, Tianmin Shu
Resources
This paper introduces a new dataset of real user chats paired with what people were actually thinking, aiming to help AI systems better infer hidden goals, preferences, and reactions.
Key results
The dataset pairs real-world human-AI conversations with self-reported thoughts from 1,058 participants.
ThoughtTrace contains 2,155 timestamped multi-turn conversations.
Across the collected conversations, the paper reports 17,058 interaction turns.
Each conversation includes private user thoughts, totaling 10,174 annotations.
The conversations were collected across 20 language models, including frontier and open-weight systems.
Adding thought annotations at inference time raises average next-user-message prediction semantic similarity from 21.6 to 30.6, a 41.7% relative gain.
What the paper found
ThoughtTrace introduces the first large-scale dataset that pairs 1,058 real users’ multi-turn human–AI conversations with 10,174 self-reported “thought” annotations, covering 2,155 conversations and 17,058 turns across 20 language models including GPT-5.4, Gemini 3.1 Pro Preview, Claude Opus 4.6, and open-weight systems. The key novelty is that each turn is augmented with two latent-cognition signals: reasons for sending a message and reactions to an assistant reply. Analysis shows these thoughts are not redundant with transcripts: LLM-based semantic coverage scores are only 3.22/5 for reasons and 2.00/5 for reactions, and frontier models infer them poorly, with average semantic similarity of 2.93 for reasons and 2.54 for reactions. Thoughts are also highly structured: reasons concentrate in Task Motivation & Goal (36.9%), Task Continuation (21.4%), and Context Grounding & Constraints (13.1%), while reactions are dominated by Explicit Affirmation (72.2%) but also expose dissatisfaction in Content Relevance, Presentation Style, and Scope Fit. Compared with WildChat and LMSYS-Chat-1M, ThoughtTrace conversations are longer, with an 8-turn median versus 2, and 57.0% of user turns extend prior tasks. Downstream, adding thoughts boosts next-user-message prediction from 21.6 to 30.6 semantic similarity, a 41.7% relative gain, and thought-guided DPO rewrites on Qwen3.5-4B raise Arena-Hard style-controlled win rate by 25.6% over the base model and 4.5% over message-guided rewrites, showing that explicit latent-cognition data can improve both user modeling and alignment.
Original abstract
Conversational AI has now reached billions of users, yet existing datasets capture only what people say, not what they think. We introduce ThoughtTrace, the first large-scale dataset that pairs real-world multi-turn human--AI conversations with users' self-reported thoughts: their reasons for sending prompts and reactions to assistant responses. ThoughtTrace comprises 1,058 users, 2,155 conversations, 17,058 turns, and 10,174 thought annotations collected across 20 language models. Our analysis shows that ThoughtTrace captures long-horizon, topically diverse interactions, and that thoughts are semantically distinct from messages, difficult for frontier LLMs to infer from context, diverse in content, and tied to conversation stages. We further demonstrate the utility of thoughts for downstream modeling. First, thoughts improve user-behavior prediction as inference-time context. Second, thought-guided rewrites provide fine-grained alignment signals for training personalized assistants. Together, ThoughtTrace establishes user thoughts as a new data modality for studying the cognitive dynamics behind human--AI interaction and provides a foundation for building assistants that better understand and adapt to users' latent goals, preferences, and needs.
Read the original paper