NTH

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

AuthorsJiangning Zhang, Haojun Chen, Yong Liu

August 28, 2026 3 min read
Watch on YouTube
The one-line take

This survey explains how smart glasses could evolve into trustworthy embodied AI systems that continuously perceive, remember, and act alongside their users.

Key results

3000
Ego4D scale

Ego4D provides 3000 hours of egocentric video for long-form first-person learning and evaluation.

5
Capability levels

The L0-L5 framework spans recording, perception, assistance, persistent state, governed action, and embodied coupling.

What the paper found

This survey reframes smart glasses as hardware-constrained first-person intelligence platforms rather than cameras, displays, or voice assistants. Its central unit of analysis is a closed perception-state-interaction-action loop that must remain temporally valid, correctable, energy-efficient, privacy-aware, and safe under real deployment conditions. The framework consolidates devices into eight verifiable hardware capability axes and seven interdependent capabilities, spanning egocentric perception, multimodal context, persistent spatial state, auditable personal memory, situated agentic action, embodied data interfaces, and deployment constraints. It introduces L0-L5 capability levels, from recording and reactive perception through contextual assistance, persistent state, governed action, and cross-embodiment coupling. Evidence is organized across nine application scenes and nine coupled deployment dimensions, including hardware, runtime, inference, memory, feedback, action, recovery, governance, and reproducibility. Datasets and systems such as Ego4D, with 3000 hours of egocentric video, EPIC-KITCHENS-100, HoloAssist, Project Aria, SuperGlasses, EgoSAT, EGOSTREAM, and EgoPoint-Bench illustrate the field’s shift from offline recognition toward streaming interaction, memory, grounding, and intervention. Commercial routes including Google Glass Enterprise Edition 2, Ray-Ban Meta Gen 2, Meta Ray-Ban Display, Microsoft HoloLens 2, and Apple Vision Pro expose different trade-offs, while multimodal models such as Gemini 2.5 and GPT-4V provide components rather than proof of deployability. The survey’s key conclusion is that capability claims must be conditioned on hardware, runtime, temporal horizon, state persistence, action authority, stakeholders, operating environment, system version, and evidence quality; robot transfer involving systems such as RT-1 requires downstream physical validation rather than human-video performance alone.

Original abstract

Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop.} This survey is \textit{the \textbf{first} to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis