NTH

Specula: Scaling formal specifications for autonomous model checking of system code

AuthorsQian Cheng, Saad Mohammad Rafid Pial, Ruize Tang, Yiming Su, Emilie Ma, Finn Hackett, Ivan Beschastnikh, Yu Huang, Tianyin Xu

August 4, 2026 2 min read
Watch on YouTube
The one-line take

Specula uses self-improving LLM agents to automatically write formal specifications and uncover deep bugs in real-world system software.

Key results

48
Systems checked

Open-source concurrent and distributed systems evaluated by Specula.

249
Bugs found

Total bugs identified across the 48 evaluated systems.

200
Model-checking bugs

Bugs surfaced through TLC model checking.

9
Median counterexample length

Median number of steps in model-checking counterexamples.

62
Specula bug-finding comparison

Bugs found by Specula in the controlled comparison, versus 2 for Agent-Raw and 3 for Agent-TLA+.

$57
Median checking cost

Median token cost to check one system end to end.

What the paper found

Specula, developed by researchers from Nanjing University, the University of Illinois Urbana-Champaign, Microsoft Research Asia, and the University of British Columbia, is a push-button agentic system for model checking real-world concurrent and distributed software. Using coding agents such as Claude Code, Codex, and Copilot CLI—defaulting to Claude Opus-4.8—it generates TLA+ invariants and implementation-level models, then combines TLC model checking with automated code instrumentation, trace validation, bidirectional repair, and self-evolving loops. This design addresses hallucinated specifications and reward hacking by requiring models to admit code traces while rejecting states that violate invariants. Across 48 open-source systems, including MongoDB, ScyllaDB, GCC libgomp, and SONiC, Specula found 249 bugs; 200 emerged through model checking, with breadth-first search finding 187 and a median counterexample length of 9 steps. On SysMoBench, Specula with Opus-4.8 achieved 100% on syntax, runtime correctness, model-code conformance, invariant satisfaction, and overall quality, while its controlled comparison found 62 bugs versus 2 for raw Claude Code and 3 for Claude Code equipped only with TLA+ tools. The system reproduced 98% of violations judged to be real bugs, required a median of 3.69 hours per system, and had a median token cost of $57. The results show that autonomous formal reasoning can scale beyond handcrafted specifications, although it remains dependent on strong LLMs and does not provide complete guarantees against modeling omissions.

Original abstract

Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications for highly effective model checking and bug finding. Specula employs large language model (LLM) based coding agents to autonomously develop TLA+ specifications, including invariants that describe correctness properties of the target system and formal models that describe the system implementation with the right level of abstractions. Specula is fully autonomous and thus eliminates the barrier of applying formal methods to real-world system code (as in traditional human-centric approaches). Meanwhile, Specula addresses limitations of LLM-driven techniques like reward hacking and hallucinations through self-evolving loops that iteratively improve specification quality by enabling the agents to deepen their understanding of system code and its behaviors. We have used Specula to check 48 open-source system projects; Specula found 249 bugs including many deep bugs that are hard to find by existing approaches. Specula has been used by several companies and is maintained at https://github.com/specula-org/Specula.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis