NTH

Towards Autonomous and Auditable Medical Imaging Model Development

AuthorsShengyuan Liu, Jia-Xuan Jiang, Boyun Zheng, Cheng Wang, Zipei Wang, Wentao Pan, Hongtao Wu, Houwen Peng, Yu Gu, Lichao Sun, Yixuan Yuan

July 21, 2026 2 min read
Watch on YouTube
The one-line take

AMID uses collaborating AI agents to automatically design, test, verify, and audit medical-imaging models across many different tasks.

Key results

20
ReX-MLE tasks

Medical-imaging challenge tasks spanning segmentation, detection, classification, quality assessment, and enhancement.

19
Tasks improved over baseline

AMID improved over the strongest listed autonomous MLE baseline on 19 tasks.

0.91
SEG.A Dice

AMID’s accepted Dice score on the SEG.A segmentation challenge.

0.89
SEG.A baseline gain

Absolute improvement in SEG.A Dice over the strongest listed baseline.

0.49
DENTEX AP

AMID’s average precision on DENTEX, matching the human reference.

24
Evaluation budget

Wall-clock hours allocated per challenge in the main experiments.

What the paper found

AMID, from researchers at The Chinese University of Hong Kong, the Chinese Academy of Sciences, Lehigh University, and Microsoft Research, is an autonomous multi-agent system for building auditable medical-imaging models. Its first contribution, Data-Conditioned Method Planning, profiles modality, geometry, labels, patient organization, metrics, and submission requirements, then converts them into executable method lanes grounded in resources such as nnU-Net, MONAI, MedSAM, and pathology encoders UNI and Phikon-v2. Its second contribution, Verification-Guided Two-Stage Optimization, uses broad behavior-gated exploration followed by selective exploitation, while independent reviewers verify data splits, metric calculations, prediction schemas, checkpoints, and traceability before evidence can influence promotion. On the ReX-MLE benchmark’s 20 medical-imaging challenges, using a 24-hour budget, one NVIDIA RTX A6000 GPU, and Codex with GPT-5.5, AMID produced accepted results on every task and improved over the strongest listed autonomous baseline on 19 tasks. It achieved 0.91 Dice on SEG.A, a 0.89 absolute gain over the best baseline, and 0.49 AP on DENTEX, matching the human reference; task-specific gains came from combining YOLOv8 with anatomical calibration and using foundation-model features for pathology. Focused experiments also used Claude Code, which reached 0.64 Dice on PUMA-T1-Seg. The report emphasizes that AMID’s advantage comes from organizing and auditing the search process, not merely from the underlying coding agent, while noting unresolved costs, backend variability, and the absence of clinical validation.

Original abstract

Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts. Here we introduce AMID, an autonomous multi-agent framework for medical imaging model development. AMID first proposes Data-Conditioned Method Planning, which refines coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources. It then develops Verification-Guided Two-Stage Optimization, moving from broad early exploration of diverse method lanes to selective exploitation of promising candidates while enforcing strict verification of validation protocols, metric computation, and prediction artifacts throughout the optimization. Across 20 medical imaging challenge tasks spanning diverse modalities and prediction types, AMID outperformed evaluated general-purpose MLE systems and, on several tasks, approached or matched strong human-designed challenge solutions. These results suggest that AMID can turn task-specific medical imaging model development from bespoke manual engineering into an agentic workflow for producing high-performing and auditable model artifacts across heterogeneous tasks.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis