NTH
AI research

Personal Visual Context Learning in Large Multimodal Models

AuthorsZihui Xue, Ami Baid, Sangho Kim, Mi Luo, Kristen Grauman

May 18, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a benchmark for testing whether multimodal AI can remember and use a user’s personal visual history, and proposes a memory-bank baseline that helps models answer personalized queries better.

Key results

2,255
Personal-VCL-Bench size

The benchmark contains 2,255 clean context-query instances spanning 7 tasks across persons, objects, behavior, and EgoWearer identification.

55.36%
Gemma-4-31B EgoWearer visual-context accuracy

On the EgoWearer identification task, Gemma-4-31B achieves 55.36% with visual-context prompting, illustrating the modality paradox.

52.99%
Gemma-4-31B EgoWearer language-context accuracy

On the same EgoWearer task, language-context prompting reaches 52.99% on Gemma-4-31B, showing raw pixels do not provide a reliable advantage.

61.60%
Agentic Context Bank EgoWearer accuracy

The structured Agentic Context Bank improves Gemma-4-31B on EgoWearer identification to 61.60%.

59.46% to 66.22%
Behavior error detection improvement

On behavior error detection, the Agentic Context Bank raises performance from 59.46% to 66.22%.

What the paper found

Personal Visual Context Learning (Personal VCL) defines a new capability for large multimodal models: using user-specific visual history at inference time to answer personalized queries. The paper introduces Personal-VCL-Bench, a 2,255-example benchmark built from EgoLife, Ego4D, and CaptainCook4D, spanning seven tasks across persons, objects, behavior, and the capstone EgoWearer identification. On frontier models including Qwen3-VL-8B, Gemma-4-31B, and Gemini-3-Flash, the authors find a modality paradox and a scaling paradox: raw visual context often underperforms textual descriptions, and adding more context frequently fails to improve accuracy. For example, on Gemma-4-31B, visual-context prompting on EgoWearer identification reaches 55.36%, while language context reaches 52.99%, showing no reliable advantage for raw pixels. To address this, they propose the Agentic Context Bank, an inference-time memory system that extracts typed entries such as APPEARANCE, OWNED_OBJECTS, and BEHAVIOR, merges them with ADD/CONFIRM/REVISE/RETRACT operations, then performs query-adaptive evidence selection with a two-stage LMM call. This structured baseline improves Gemma-4-31B from 55.36% to 61.60% on EgoWearer identification and raises overall benchmark performance, including behavior error detection from 59.46% to 66.22% and object detection from 71.74% language-context accuracy to 73.19%.

Original abstract

As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the evolution of these models into true personal assistants hinges on visual personalization: the ability to reason over visual information unique to the wearer. We formalize this capability as Personal Visual Context Learning (Personal VCL), the prompt-time capability of using user-specific visual context to resolve personalized queries. To systematically evaluate this, we present Personal-VCL-Bench, a comprehensive benchmark capturing the personal visual world across persons, objects, and behaviors. Our analysis of frontier LMMs identifies a profound context utilization gap, revealing that the mechanisms for leveraging visual evidence, as well as aggregating multiple visual observations, remain critically understudied. Motivated by these findings, we propose the Agentic Context Bank, a strong inference-time baseline that structures a user's visual context into a self-refining memory bank and employs query-adaptive evidence selection. Our baseline approach consistently improves over standard context prompting regimes across tasks and evaluated backbones, demonstrating a practical path towards future personalized LMMs.

Read the original paper