From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
AuthorsJiangning Zhang, Haojun Chen, Yong Liu
Resources
This survey explains how smart glasses could evolve into trustworthy embodied AI systems that continuously perceive, remember, and act alongside their users.
Key results
Ego4D provides 3000 hours of egocentric video for long-form first-person learning and evaluation.
The L0-L5 framework spans recording, perception, assistance, persistent state, governed action, and embodied coupling.
What the paper found
This survey reframes smart glasses as hardware-constrained first-person intelligence platforms rather than cameras, displays, or voice assistants. Its central unit of analysis is a closed perception-state-interaction-action loop that must remain temporally valid, correctable, energy-efficient, privacy-aware, and safe under real deployment conditions. The framework consolidates devices into eight verifiable hardware capability axes and seven interdependent capabilities, spanning egocentric perception, multimodal context, persistent spatial state, auditable personal memory, situated agentic action, embodied data interfaces, and deployment constraints. It introduces L0-L5 capability levels, from recording and reactive perception through contextual assistance, persistent state, governed action, and cross-embodiment coupling. Evidence is organized across nine application scenes and nine coupled deployment dimensions, including hardware, runtime, inference, memory, feedback, action, recovery, governance, and reproducibility. Datasets and systems such as Ego4D, with 3000 hours of egocentric video, EPIC-KITCHENS-100, HoloAssist, Project Aria, SuperGlasses, EgoSAT, EGOSTREAM, and EgoPoint-Bench illustrate the field’s shift from offline recognition toward streaming interaction, memory, grounding, and intervention. Commercial routes including Google Glass Enterprise Edition 2, Ray-Ban Meta Gen 2, Meta Ray-Ban Display, Microsoft HoloLens 2, and Apple Vision Pro expose different trade-offs, while multimodal models such as Gemini 2.5 and GPT-4V provide components rather than proof of deployability. The survey’s key conclusion is that capability claims must be conditioned on hardware, runtime, temporal horizon, state persistence, action authority, stakeholders, operating environment, system version, and evidence quality; robot transfer involving systems such as RT-1 requires downstream physical validation rather than human-video performance alone.
Original abstract
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop.} This survey is \textit{the \textbf{first} to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.