Towards a Deterministic Math Solver for Clinical Language Models
AuthorsFelipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange, Rafi Al Attrach, Sahil Kapadia, Zakaria Laouabdia Sellami, Angelo Antonio Talio, Leo Anthony Celi
AffiliationsMIT Critical Data · UNC Chapel-Hil · University of Pavia · Humanitas University
Resources
Instead of trusting language models to do clinical math, the system has them write code for a restricted executor—and shows that this helps larger models but is no universal cure for medical calculation errors.
Key results
MedCalc-Bench Verified evaluation cases
Clinical calculators represented in the benchmark
Full-set accuracy with formula text and gold variables
Full-set accuracy with formula text and gold variables
Percentage-point improvement over 83.47% open-book arithmetic
Accuracy when the 22-calculator library abstains outside its supported cases
What the paper found
This paper evaluates Program-Solve, an interface in which a clinical language model writes case-specific Python and a restricted local executor performs the arithmetic deterministically, rather than asking the model to calculate directly. On MedCalc-Bench Verified, covering 1,100 cases and 55 calculators, the method is compared with open-book arithmetic and a hand-written library using Qwen2.5-7B and Qwen2.5-32B-AWQ. With formulas, gold variables, and the full note supplied, Program-Solve reaches 75.31% accuracy at 7B versus 72.02% for direct arithmetic, but the clustered uncertainty interval crosses zero; at 32B it reaches 90.53% versus 83.47%, a reliable +7.05 percentage-point gain. The 22-calculator library is exact on its 440 supported cases but abstains elsewhere, producing only 40.0% overall accuracy, so much of Program-Solve’s apparent advantage comes from broader coverage rather than superior execution. When formula and variable access are removed, accuracy falls to 28.04% at 7B and 44.71% at 32B, showing that formula selection and clinical-variable extraction remain dominant failure modes. Results are checkpoint-dependent: Mistral-7B trails arithmetic while Microsoft’s Phi-3.5-mini improves over it, so scale alone does not predict benefit. The executor uses allow-listed Python with 256 MB memory and 5-second limits, but it is not a complete security sandbox. The study, supported by Anthropic’s AI for Science program with NVIDIA compute, concludes that verified calculators should be preferred, with program generation as a fallback and explicit abstention when formulas or extracted variables are uncertain.
Original abstract
Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is to hardcode each calculator as a validated function, one at a time. We test an alternative: the model does not calculate. Instead, it writes case-specific Python that a restricted local executor runs as a deterministic solver, and the model's task reduces to deciding how to use it. We evaluate this Program-Solve interface on MedCalc-Bench Verified (1,100 cases, 55 calculators) against direct model arithmetic and a hand-written 22-calculator library, using Qwen2.5-7B and Qwen2.5-32B-AWQ, after auditing the benchmark's formulas against current clinical guidelines and flagging 16 of 55 with version, use or coefficient concerns. With formulas and gold variables supplied and both routes reading the whole note, handing off to the solver is not a reliable advantage at 7B (75.31% against 72.02%, a paired +3.29 points with a 95% calculator-cluster interval of [-3.49, 10.38]) but is one at 32B (90.53% against 83.47%, +7.05 [0.47, 14.60], clear of zero). The hand-written library is exact on its 440 supported cases but abstains elsewhere (40.0% overall). Adding an executor thus helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.