NTH

When Do Biological Reasoning Models Use Their Biological Inputs?

AuthorsAda Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

AffiliationsHarvard University · Massachusetts Institute of Technology

October 7, 2026 3 min read
Watch on YouTube
The one-line take

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Key results

0.847
BioReason shuffled-DNA accuracy

Disease prediction accuracy with intact and shuffled DNA was both 0.847.

97.9%
BioReason text preference

In evidence conflicts, BioReason followed gene and pathway text rather than Evo2 representations in 97.9% of cases.

99.7%
BioReason-Pro text preference

BioReason-Pro followed GO-GPT text rather than ESM3 representations in 99.7% of conflicts.

0.736
Evo2 linear-probe accuracy

A linear probe predicted genome-dependent disease labels from Evo2 representations at 0.736 accuracy.

0.446
Cell2Sentence-Scale scale-up accuracy

The 27B Cell2Sentence-Scale model reached 0.446 cell-type accuracy, versus 0.366 for the 2B model.

0.671
Auxiliary edited-base accuracy

Auxiliary sequence supervision produced 0.671 edited-base accuracy with intact DNA, falling to 0.122 after shuffling.

What the paper found

This paper tests whether biological reasoning models actually use the biological inputs they receive, rather than relying on predictive text or shortcuts. Across six models spanning DNA, protein, and single-cell tasks, the study combines input perturbations, evidence conflicts, linear probes, and reasoning-trace audits. In Qwen3-based BioReason, shuffling DNA leaves disease accuracy unchanged at 0.847 versus 0.847, while its Evo2 representation loses conflicts to gene and pathway text in 97.9% of cases; BioReason-Pro similarly follows GO-GPT text over ESM3 in 99.7% of conflicts. Linear probes nevertheless recover genome-dependent disease labels from Evo2 at 0.736 accuracy versus 0.708 for a text-only predictor, showing that encoded information is not necessarily used by the reasoning model. In contrast, ChatNT, Llama-3.1-based Prot2Text-V2, Mistral-based CellWhisperer, and Google Gemma-2-based Cell2Sentence-Scale show meaningful dependence on biological representations or gene inputs. Scaling Cell2Sentence-Scale from 2B to 27B raises cell-type accuracy from 0.366 to 0.446 and increases the differential effect of removing marker genes from 0.223 to 0.319. Reinforcement learning improves BioReason and BioReason-Pro benchmark scores without consistently increasing foundation-model input contribution, and generated rationales can misstate nucleotide substitutions or mention protein functions absent from final predictions. Auxiliary supervision makes BioReason more sequence-sensitive, raising edited-base accuracy to 0.671 with intact DNA but only 0.122 after shuffling, with most gains concentrated on genomes seen during training. The central conclusion is that benchmark accuracy, post-training gains, and plausible explanations do not establish faithful biological input use.

Original abstract

Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

nnFoundation: 3D Foundation Models for Radiology

Constantin Ulrich Harsy, Tassilo Wald, Karol Gotkowski, Yannick Kirchhoff, Marcel Knopp, Maximilian Rokuss, Elisa Stegmeier, Philipp Schader, Dasha Trofimova, Raphael Stock, Kim-Celine Kahl, Stephen Schaumann, Selen Erkan, David Zimmerer, Stefan Denner, Moritz Langenberg, Sebastian Ziegler, Katharina Eckstein, Maximilian Fischer, Jonathan Suprijadi, Bálint Kovács, Benjamin Hamm, Anand Deshpande, Dimitrios Bounias, Nico Disch, Shuhan Xiao, Jessica Kächele, Jan Sellner, Rajesh Baidya, Jeremias Traub, Lars Krämer, Maximilian Zenk, Tim Rädsch, Stefan Dvoretskii, Robin Peretzke, Jonathan Deissler, Alexandra Ertl, Partha Ghosh, Kris Dreher, Stefan Dinkelacker, Annika Reinke, Evangelia Christodoulou, Numan Saeed, Yoland Savriama, Santiago Estrada, David Kügler, Laura Alexandra Daza Barragan, Cristina Isabel Gonzalez Osorio, Jan Peeken, Michael Baumgartner, Marvin Teichmann, Guillaume Chabin, Matthias Kirchler, Valentin Koch, for the ALFA study, Markus Hohenhaus, Dimitri Koslov, Nina ...

nnFoundation trains complementary 3D convolutional and transformer models on 2.1 million medical scans and shows that matching architecture to the task can improve transfer across diverse radiology applications.

Read analysis