Formalizing Mathematics at Scale
AuthorsAhmad Rammal, Niket Patel, Fabian Gloeckle, Amaury Hayat, Julia Kempe, Remi Munos, Charles Arnal, Vivien Cabannes
Resources
This work turns thousands of AI agents loose on real math textbooks to build a massive, machine-checked Lean library, pushing automated formalization from a niche tool toward practical scale.
Key results
AutoformBot was applied to 26 open-access mathematical textbooks.
ATLAS contains over 45,000 verified Lean 4 declarations.
Across all 26 books, 2,855 of 4,007 target statements were formalized.
In the 600M-token ablation budget, the full pipeline reached 77% of targets on Richard Stanley’s Algebraic Combinatorics.
What the paper found
FAIR at Meta’s paper “Formalizing Mathematics at Scale” presents AutoformBot, a multi-agent Lean 4 formalization system that coordinates frontier LLMs through git worktrees, pull-request-style review, dependency-aware task DAGs, and a post-merge evaluation harness that checks for `sorry`, axioms, faithfulness, and proof integrity via declaration-dependency graphs. Using Claude Opus 4.6 as the main model, the system autoformalized 26 open-access textbooks across analysis, algebra, topology, combinatorics, probability, geometry, number theory, PDEs, and theoretical computer science, producing ATLAS: more than 45,000 verified Lean 4 declarations and about 500,000 lines of code. The corpus is 71.3% complete overall, with high-coverage books such as Real Analysis at 175/177 targets and Complex Variables at 37/38, while harder areas like Lie Groups and Boolean Functions remained around 40% coverage. Ablations on Richard Stanley’s Algebraic Combinatorics show the full pipeline reaches 77% of targets within a 600M-token budget; removing the orchestrator plateaus at 64%, removing the supervisor drops to 51%, and removing the trace analyzer to 57%, demonstrating that long-horizon planning, target-level verification, and failure memory are all necessary at scale. The authors argue that the main bottlenecks are not just theorem proving but coordination, context degradation, and adversarial verification circumvention.
Original abstract
We present AutoformBot, a multi-agent system for building an Autoformalized Textbook Library At Scale (Atlas) in Lean 4. AutoformBot orchestrates thousands of LLM agents, equipped with formal verification tools, dependency-aware task scheduling, and collaborative version control, to translate informal textbook prose into machine-checked definitions and proofs. We apply our methods to a corpus of 26 open-access textbooks spanning analysis, algebra, topology, combinatorics, and probability, producing Atlas: a verified library of over 45,000 Lean 4 declarations and 500 thousand lines of code. We release two artifacts: (i) AutoformBot, the open-source multi-agent framework; and (ii) Atlas, the resulting formal library. Our results suggest that autoformalizing the core content of graduate-level mathematics at scale is now economically and technically feasible. This opens the door to the automated verification of both human- and machine-generated mathematics at a research level.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.