NTH

Efficient Test-Time Adaptation through Human-AI Interaction

AuthorsZora Zhiruo Wang, Apurva Gandhi, Rulin Shao, Aspen Chen, Jonas Mueller, Zhiqi Liang, Jett Chen, Michael Ryan, Qianou Ma, Luxi He, Zhoujun Cheng, Andre He, Seungone Kim, Jiayi Geng, Mingqian Zheng, Weiwei Sun, Zheyuan Zhang, Xinran Zhao, Yike Wang, Abe Hou, Liwei Jiang, Pang Wei Koh, Diyi Yang, Graham Neubig, Daniel Fried

September 7, 2026 2 min read
Watch on YouTube
The one-line take

The work shows how agents can learn a user’s evolving standards through repeated interaction, becoming substantially better at personalized writing and visual-creation tasks.

Key results

600
Evaluation tasks

Total tasks completed across writing and data visualization experiments

30
Adapted individuals

Number of users whose preferences and workflows were modeled

20.9%
Maximum solo success improvement

Improvement from weight adaptation on data visualization tasks

22.3%
Rubric failure-capture improvement

Upper end of the gain over language-model-only or human-only rubrics

88.3%
Visualization input-token reduction

Reduction from weight adaptation compared with context adaptation

What the paper found

Efficient Test-Time Adaptation through Human-AI Interaction introduces TAHI, a framework that learns a user’s tacit preferences from ongoing collaboration rather than relying only on pretraining or one-off feedback. Built on the Qwen3.6-35B-A3B agent backbone, with Anthropic’s claude-sonnet-4-6 supporting context induction and rubric evolution, TAHI captures four interaction channels: messages, plan edits, direct artifact edits, and verification-criterion changes. It adapts through either editable memory and procedural skills or Direct Preference Optimization on LoRA adapters, training the agent to produce the user-preferred result in one pass. Across 30 individuals and 600 tasks covering paper abstract writing and HTML data visualization, adaptation improved solo task success by 4.5–20.9% within only 20 task sessions; the largest gain was 20.9% for weight adaptation on visualization. On held-out tasks, the learned preferences generalized across tasks, while weight adaptation reduced input-token requirements by 62.6% for writing and 88.3% for visualization relative to context adaptation. TAHI also evolves evaluation rubrics from interaction traces: these rubrics captured 16.0–22.3% more failures than rubrics produced by language models or humans alone. The results show that agents can absorb both personalized strategies and shared community standards, although implicit visual judgment, especially color and emphasis, remains difficult to transfer.

Original abstract

AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis