Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases
AuthorsCheng Liang, Pengcheng Qiu, Ya Zhang, Yanfeng Wang, Chaoyi Wu, Weidi Xie
Resources
This paper introduces MedSP1000, a new interactive benchmark that tests whether large language models can handle dynamic, real-world-style clinical decision-making instead of just answering isolated questions.
Key results
executible SP cases in MedSP1000
trajectory-level peer-reviewed rubric items
clinical specialties covered by MedSP1000
best model rubric completion rate on MedSP1000
strongest medically specialized model rubric completion rate on MedSP1000
no model exceeded this completion rate in Practice-Based Learning and Improvement
What the paper found
This paper introduces MedSP1000, an SP-derived interactive benchmark built from MedEdPORTAL standardized-patient teaching cases to evaluate large language models as dynamic clinical agents rather than static question-answering systems. The benchmark contains 1,638 executable cases spanning 17 specialties and 24,602 peer-reviewed rubric items mapped to the six ACGME competencies, with a closed-loop setup in which a clinician model interacts with a patient agent and an environment controller before an evaluator scores the full trajectory. The authors report that performance on static medical benchmarks does not transfer reliably: GPT-5.5 is the strongest model on MedSP1000 but reaches only 60.4% rubric completion, while the best medical-specialized model reaches 40.0%, a 20.4-point gap. Practice-Based Learning and Improvement is the weakest dimension across all systems, with no model exceeding 30%, and additional test-time compute does not materially help: GPT-5.5 changes only from 67.1% to 67.8% with Best-of-5 self-consistency and to 68.0% with a five-specialist MedAgents pipeline. The study’s key contribution is showing that process-level, standardized-patient evaluation exposes clinically relevant failures in sequential decision-making that single-turn benchmarks miss.
Original abstract
Large language models (LLMs) are increasingly proposed as clinical agents, yet static, single-turn benchmarks cannot capture how a model dynamically delivers care across an encounter: gathering information, planning treatment, and adapting longitudinal management across successive patient states. Medical education has long addressed an analogous challenge through standardized patients (SPs): trained actors who consistently portray clinical cases, enabling realistic practice and objective, scripted assessment. Here we introduce MedSP1000, an SP-derived interactive benchmark for clinical-agent evaluation, including 1,638 SP cases with 24,602 trajectory-level peer-reviewed rubrics. MedSP1000 converts peer-reviewed SP teaching cases into executable scenarios with defined SP case scripts, clinical environment contexts, and human-validated structured rubric. In each simulation evaluation run, a clinical agent interacts in closed loop with a patient agent and an environment controller, and its behaviour is scored throughout the encounter against expert criteria specified in the original materials. Applying MedSP1000 to a range of general-purpose and medically specialized LLMs, we find that performance on static benchmarks does not reliably translate to such educational scenarios. The best-performing model, GPT-5.5, completes only 60.4% of expert-defined rubric items, whereas the strongest medically specialized model reaches 40.0%; increasing test-time compute produces no measurable gain. These results suggest that current LLMs, including agentic systems tuned for medicine, are not yet reliable enough to be safely integrated into actual clinical practice. More broadly, MedSP1000 shows how process-level, SP-style evaluation can reveal clinically relevant failure modes that single-turn benchmarks miss.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.