NTH
AI research

ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions

AuthorsChuanyang Jin, Binze Li, Haopeng Xie, Cathy Mengying Fang, Tianjian Li, Shayne Longpre, Hongxiang Gu, Maximillian Chen, Tianmin Shu

May 21, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a new dataset of real user chats paired with what people were actually thinking, aiming to help AI systems better infer hidden goals, preferences, and reactions.

Key results

1,058
users

The dataset pairs real-world human-AI conversations with self-reported thoughts from 1,058 participants.

2,155
conversations

ThoughtTrace contains 2,155 timestamped multi-turn conversations.

17,058
interaction turns

Across the collected conversations, the paper reports 17,058 interaction turns.

10,174
thought annotations

Each conversation includes private user thoughts, totaling 10,174 annotations.

20 language models
model coverage

The conversations were collected across 20 language models, including frontier and open-weight systems.

21.6 to 30.6
next-message prediction gain

Adding thought annotations at inference time raises average next-user-message prediction semantic similarity from 21.6 to 30.6, a 41.7% relative gain.

What the paper found

ThoughtTrace introduces the first large-scale dataset that pairs 1,058 real users’ multi-turn human–AI conversations with 10,174 self-reported “thought” annotations, covering 2,155 conversations and 17,058 turns across 20 language models including GPT-5.4, Gemini 3.1 Pro Preview, Claude Opus 4.6, and open-weight systems. The key novelty is that each turn is augmented with two latent-cognition signals: reasons for sending a message and reactions to an assistant reply. Analysis shows these thoughts are not redundant with transcripts: LLM-based semantic coverage scores are only 3.22/5 for reasons and 2.00/5 for reactions, and frontier models infer them poorly, with average semantic similarity of 2.93 for reasons and 2.54 for reactions. Thoughts are also highly structured: reasons concentrate in Task Motivation & Goal (36.9%), Task Continuation (21.4%), and Context Grounding & Constraints (13.1%), while reactions are dominated by Explicit Affirmation (72.2%) but also expose dissatisfaction in Content Relevance, Presentation Style, and Scope Fit. Compared with WildChat and LMSYS-Chat-1M, ThoughtTrace conversations are longer, with an 8-turn median versus 2, and 57.0% of user turns extend prior tasks. Downstream, adding thoughts boosts next-user-message prediction from 21.6 to 30.6 semantic similarity, a 41.7% relative gain, and thought-guided DPO rewrites on Qwen3.5-4B raise Arena-Hard style-controlled win rate by 25.6% over the base model and 4.5% over message-guided rewrites, showing that explicit latent-cognition data can improve both user modeling and alignment.

Original abstract

Conversational AI has now reached billions of users, yet existing datasets capture only what people say, not what they think. We introduce ThoughtTrace, the first large-scale dataset that pairs real-world multi-turn human--AI conversations with users' self-reported thoughts: their reasons for sending prompts and reactions to assistant responses. ThoughtTrace comprises 1,058 users, 2,155 conversations, 17,058 turns, and 10,174 thought annotations collected across 20 language models. Our analysis shows that ThoughtTrace captures long-horizon, topically diverse interactions, and that thoughts are semantically distinct from messages, difficult for frontier LLMs to infer from context, diverse in content, and tied to conversation stages. We further demonstrate the utility of thoughts for downstream modeling. First, thoughts improve user-behavior prediction as inference-time context. Second, thought-guided rewrites provide fine-grained alignment signals for training personalized assistants. Together, ThoughtTrace establishes user thoughts as a new data modality for studying the cognitive dynamics behind human--AI interaction and provides a foundation for building assistants that better understand and adapt to users' latent goals, preferences, and needs.

Read the original paper