Towards Autonomous and Auditable Medical Imaging Model Development
AuthorsShengyuan Liu, Jia-Xuan Jiang, Boyun Zheng, Cheng Wang, Zipei Wang, Wentao Pan, Hongtao Wu, Houwen Peng, Yu Gu, Lichao Sun, Yixuan Yuan
Resources
AMID uses collaborating AI agents to automatically design, test, verify, and audit medical-imaging models across many different tasks.
Key results
Medical-imaging challenge tasks spanning segmentation, detection, classification, quality assessment, and enhancement.
AMID improved over the strongest listed autonomous MLE baseline on 19 tasks.
AMID’s accepted Dice score on the SEG.A segmentation challenge.
Absolute improvement in SEG.A Dice over the strongest listed baseline.
AMID’s average precision on DENTEX, matching the human reference.
Wall-clock hours allocated per challenge in the main experiments.
What the paper found
AMID, from researchers at The Chinese University of Hong Kong, the Chinese Academy of Sciences, Lehigh University, and Microsoft Research, is an autonomous multi-agent system for building auditable medical-imaging models. Its first contribution, Data-Conditioned Method Planning, profiles modality, geometry, labels, patient organization, metrics, and submission requirements, then converts them into executable method lanes grounded in resources such as nnU-Net, MONAI, MedSAM, and pathology encoders UNI and Phikon-v2. Its second contribution, Verification-Guided Two-Stage Optimization, uses broad behavior-gated exploration followed by selective exploitation, while independent reviewers verify data splits, metric calculations, prediction schemas, checkpoints, and traceability before evidence can influence promotion. On the ReX-MLE benchmark’s 20 medical-imaging challenges, using a 24-hour budget, one NVIDIA RTX A6000 GPU, and Codex with GPT-5.5, AMID produced accepted results on every task and improved over the strongest listed autonomous baseline on 19 tasks. It achieved 0.91 Dice on SEG.A, a 0.89 absolute gain over the best baseline, and 0.49 AP on DENTEX, matching the human reference; task-specific gains came from combining YOLOv8 with anatomical calibration and using foundation-model features for pathology. Focused experiments also used Claude Code, which reached 0.64 Dice on PUMA-T1-Seg. The report emphasizes that AMID’s advantage comes from organizing and auditing the search process, not merely from the underlying coding agent, while noting unresolved costs, backend variability, and the absence of clinical validation.
Original abstract
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts. Here we introduce AMID, an autonomous multi-agent framework for medical imaging model development. AMID first proposes Data-Conditioned Method Planning, which refines coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources. It then develops Verification-Guided Two-Stage Optimization, moving from broad early exploration of diverse method lanes to selective exploitation of promising candidates while enforcing strict verification of validation protocols, metric computation, and prediction artifacts throughout the optimization. Across 20 medical imaging challenge tasks spanning diverse modalities and prediction types, AMID outperformed evaluated general-purpose MLE systems and, on several tasks, approached or matched strong human-designed challenge solutions. These results suggest that AMID can turn task-specific medical imaging model development from bespoke manual engineering into an agentic workflow for producing high-performing and auditable model artifacts across heterogeneous tasks.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.