NTH

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

AuthorsShentong Mo, Yatao Bian

August 5, 2026 2 min read
Watch on YouTube
The one-line take

APO uses physical consistency rather than expensive labeled structures to teach generative models how to produce plausible 3D atomic configurations.

Key results

63.05%
MP-20 Match Rate with APO

OT plus VE APO Match Rate, compared with 62.47% for supervised FlowDPO

21.14%
MPTS-52 Match Rate with APO

OT plus VE APO result, compared with 20.27% for FlowDPO

3.21
Antibody CDR-H3 Cα RMSD

Angstrom RMSD for OT Path plus APO, versus 3.32 for OT Path plus DPO

2.23%
Spectral reward ablation drop

Match Rate reduction when Spectral Consistency is removed on MP-20

What the paper found

APO, or Atomic Policy Optimization, replaces FlowDPO’s dependence on ground-truth 3D coordinates with fully unsupervised alignment for crystal and antibody structure prediction. It adapts Group Relative Policy Optimization, a strategy used in DeepSeekMath, to flow-matching models by sampling candidate structures in groups and updating the policy from relative rewards. Its dual reward combines a Spectral Consistency Score, computed through eigen-decomposition of candidate-embedding similarities to identify dominant structural modes, with a Crystal Entropy Proxy that penalizes disordered packing, atomic clashes, and thermodynamically implausible geometries. Across Perov-5, MP-20, MPTS-52, and SAbDab, APO improves both structural fidelity and match rates without labels during alignment. On MP-20 with an OT plus VE path, APO reaches a 63.05% Match Rate versus 62.47% for supervised FlowDPO; on the harder MPTS-52 benchmark, it achieves 21.14% versus 20.27%. For antibody CDR-H3 prediction using an OT path, APO lowers Cα RMSD to 3.21 Å from FlowDPO’s 3.32 Å. An ablation shows that removing spectral consistency reduces Match Rate by 2.23%, while removing entropy worsens geometric accuracy, supporting the complementary roles of global manifold consensus and local physical regularity. The authors also report that APO straightens flow trajectories, particularly for Optimal Transport paths, potentially reducing inference difficulty, although its group-based sampling increases training cost and its entropy proxy does not model detailed electronic interactions.

Original abstract

Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) have recently shown promise in this domain, their performance relies heavily on alignment with ground-truth coordinates via supervised preference learning. However, obtaining experimental labels for novel crystal phases or de novo proteins is prohibitively expensive, creating a bottleneck for structural modeling in data-scarce regimes. In this work, we propose (Atomic Policy Optimization), a fully unsupervised alignment framework that eliminates the need for ground-truth reference structures. APO adapts group-relative policy optimization to 3D atomic environments, utilizing a novel dual-reward mechanism: (i) a that reinforces the policy's dominant latent structural modes through eigen-decomposition of sample similarities, and (ii) a that enforces thermodynamic stability. Our framework enables the model to ``self-correct'' by identifying physically plausible configurations within sampled groups. Extensive benchmarks on crystal and antibody structure prediction demonstrate that APO consistently outperforms fully supervised baselines, achieving a new state-of-the-art in match rates and structural fidelity. Furthermore, we show that APO effectively straightens probability paths, significantly improving inference efficiency. Our results suggest that intrinsic physical consistency can serve as a superior guide for alignment compared to noisy, supervised coordinate matching.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis