NTH

How Post-Training Shapes Biological Reasoning Models

AuthorsLukas Fesser, Hanlin Zhang, Michelle M. Li, Eric Wang, Bryan Perozzi, Shekoofeh Azizi, Sham M. Kakade, Marinka Zitnik

June 29, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that in biological AI, more training is not always better: the way you mix pretraining, fine-tuning, and reinforcement learning strongly changes whether models generalize or overfit.

Key results

100
Models trained/evaluated

Controlled biological reasoning checkpoints across modalities

200K/5K
CPT data split

FineFineWeb biology subset used for continued pre-training

0.907
DNA ID accuracy

Qwen3-1.7B after 16 SFT epochs on pathway prediction

0.736
DNA OOD accuracy peak

Qwen3-1.7B pathway prediction OOD peak during SFT

0.911
RNA ID accuracy

Qwen3-4B after 16 SFT epochs on target identification

0.956
Protein OOD F1

Qwen3-4B protein function prediction after RL

What the paper found

How Post Training Shapes Biological Reasoning Models, led by Harvard University researchers with Google DeepMind and Google Research collaborators, studies more than 100 controlled multimodal biological reasoning checkpoints across DNA, RNA, and proteins using Qwen3-1.7B, Qwen3-4B, and Gemma 4 E2B backbones coupled to Evo2-1B, TranscriptFormer, or ESM-3 encoders. The paper shows that post-training stages are not additive: continued pre-training on a 200K/5K biology split from FineFineWeb improves downstream adaptation to biological language, supervised fine-tuning sharply raises in-domain accuracy but over-specializes and suppresses out-of-domain robustness, and reinforcement learning from a strong supervised checkpoint recovers transfer, especially OOD. In DNA pathway prediction, Qwen3-1.7B SFT lifts ID accuracy from 0.562 to 0.907 by 16 epochs while OOD peaks at 0.736 around 4 epochs and then falls to 0.687; in RNA target identification, Qwen3-4B reaches 0.911 ID accuracy but only 0.627 OOD before declining. RL reverses this pattern: Qwen3-4B RNA rises from 0.775/0.582 ID/OOD to 0.914/0.745 by 16 epochs, and protein function prediction with propagated GO-F1 reaches 0.956 OOD after RL. Under fixed post-training budgets, the best trade-off comes from brief SFT followed by larger RL, with mixed schedules such as 20% SFT and 80% RL reaching 0.9470 ID F1 and 0.9685 OOD F1 on proteins. The central takeaway is that biological reasoning improves most when compute, data, and adaptation capacity are allocated asymmetrically across stages rather than scaled uniformly.

Original abstract

Scientific reasoning models for biology combine language models with foundation models trained on multimodal biological data, including DNA, RNA, and proteins. These models are built through post-training, yet how each stage shapes reasoning and generalization remains poorly understood. We study when post-training improves performance and when it induces over-specialization. Across genomics, transcriptomics, and proteins, we train and evaluate more than 100 biological reasoning models under controlled variation in backbone, continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL), measuring both in-domain (ID) and out-of-domain (OOD) performance. We find that each post-training stage reshapes generalization in a distinct way rather than contributing uniform gains. CPT improves downstream performance by aligning models with biological language. SFT consistently increases ID performance but causes OOD performance to peak early and decline as models fit the training distribution. RL, when applied to strong SFT checkpoints with aligned rewards, improves OOD performance and partially recovers generalization. These results show that biological reasoning does not improve monotonically with additional supervision or compute. Instead, performance depends on how training stages are composed. Under fixed post-training budgets, the strongest ID-OOD trade-off comes from brief SFT, larger RL allocations, and asymmetric adaptation capacity across stages.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis