NTH

Towards Expert-level Medical AI for Real-time Video Consultations

AuthorsMahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu

August 17, 2026 3 min read
Watch on YouTube
The one-line take

A Gemini-based multimodal medical AI reportedly matches or exceeds physicians on several simulated consultation tasks while remaining weaker in rapport, subtle affect, and fine-grained physical assessment.

Key results

83%
OSCE overall score

AMIE Video overall case-specific score versus 68% for physicians across 100 clinical scenarios

91%
Top-1 diagnostic accuracy

AMIE Video matched the reference diagnosis in the first differential position, versus 77% for physicians

89%
Video communication effectiveness

Patient rating for AMIE Video versus 79% for AMIE Text

2.6
Mean turn latency

Seconds after asynchronous orchestration, reduced from 21.4 seconds

87%
Multi-turn consultation quality

Full multi-agent system score versus 71% for the Talker-only ablation

What the paper found

This paper presents AMIE (Video), a real-time medical consultation system built on Google’s Gemini 3 Flash and Gemini 3.1 Pro, with asynchronous Talker, Planner, and Perception agents that separate rapid dialogue from deeper clinical reasoning and continuous audio-visual analysis. In a randomized OSCE involving 100 clinical scenarios, professional patient actors, and primary-care physicians, independent evaluators gave AMIE an overall case-specific score of 83%, compared with 68% for physicians. Its top-ranked diagnosis was correct in 91% of cases, versus 77% for physicians, while its strongest advantages were physical observation and guided examination. Compared with text-only AMIE, video improved patient-rated communication effectiveness from 79% to 89%, convenience from 71% to 88%, and feeling understood from 81% to 90%. Asynchronous orchestration reduced mean conversational turn latency from 21.4 seconds to 2.6 seconds, while multi-agent ablations showed that the full system improved single-turn evaluation from 42% for the Talker-only baseline to 59% and multi-turn consultation quality from 71% to 87%. The system outperformed or matched clinicians across history-taking, reasoning, treatment planning, and communication, although patients still favored physicians for rapport and partnership. Important limitations remain: AMIE struggles with fine anatomical detail, subtle affect, tremors, nystagmus, high-frequency movement, and natural overlapping conversation. The study also compares against earlier systems including OpenAI’s GPT-Realtime, but its simulated-patient design means AMIE is not ready for deployment with real patients.

Original abstract

Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis