Towards Expert-level Medical AI for Real-time Video Consultations
AuthorsMahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu
Resources
A Gemini-based multimodal medical AI reportedly matches or exceeds physicians on several simulated consultation tasks while remaining weaker in rapport, subtle affect, and fine-grained physical assessment.
Key results
AMIE Video overall case-specific score versus 68% for physicians across 100 clinical scenarios
AMIE Video matched the reference diagnosis in the first differential position, versus 77% for physicians
Patient rating for AMIE Video versus 79% for AMIE Text
Seconds after asynchronous orchestration, reduced from 21.4 seconds
Full multi-agent system score versus 71% for the Talker-only ablation
What the paper found
This paper presents AMIE (Video), a real-time medical consultation system built on Google’s Gemini 3 Flash and Gemini 3.1 Pro, with asynchronous Talker, Planner, and Perception agents that separate rapid dialogue from deeper clinical reasoning and continuous audio-visual analysis. In a randomized OSCE involving 100 clinical scenarios, professional patient actors, and primary-care physicians, independent evaluators gave AMIE an overall case-specific score of 83%, compared with 68% for physicians. Its top-ranked diagnosis was correct in 91% of cases, versus 77% for physicians, while its strongest advantages were physical observation and guided examination. Compared with text-only AMIE, video improved patient-rated communication effectiveness from 79% to 89%, convenience from 71% to 88%, and feeling understood from 81% to 90%. Asynchronous orchestration reduced mean conversational turn latency from 21.4 seconds to 2.6 seconds, while multi-agent ablations showed that the full system improved single-turn evaluation from 42% for the Talker-only baseline to 59% and multi-turn consultation quality from 71% to 87%. The system outperformed or matched clinicians across history-taking, reasoning, treatment planning, and communication, although patients still favored physicians for rapport and partnership. Important limitations remain: AMIE struggles with fine anatomical detail, subtle affect, tremors, nystagmus, high-frequency movement, and natural overlapping conversation. The study also compares against earlier systems including OpenAI’s GPT-Realtime, but its simulated-patient design means AMIE is not ready for deployment with real patients.
Original abstract
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.