NTH

EnigmaForge: The Question Is Hidden in the Story

AuthorsDaniel Eisner

AffiliationsSep 2026

October 6, 2026 2 min read
Watch on YouTube
The one-line take

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Key results

600
Evaluation instances

Procedurally generated instances evaluated across matched conditions.

17,400
Scored records

Total model-condition records included in the evaluation.

22
Intuition separation

Fold spread in unguided task success across frontier models.

0.966
GPT-6 Sol fact F1

Highest reported world-reconstruction fact F1.

80.1
GPT-6 Astra intuition

Highest implicit-story task-success score.

585
Claude Fable 5.1 filtered attempts

Attempts blocked by the model's content filter out of 600.

What the paper found

EnigmaForge evaluates whether language models can discover a problem before solving it: each instance provides only a procedurally generated story assembled from letters, receipts, and logbook fragments, with no stated question. The model must infer the hidden task, reconstruct a finite-domain constraint world, use external knowledge bridges, and execute a policy-validated action. Generation begins with a known solution, strengthens and minimizes its evidence, then verifies uniqueness with a DPLL SAT solver, brute-force cross-checking, and per-clue ablation certificates showing that removing any clue creates a second solution. Seed-based generation supports renewable, contamination-resistant evaluation, with multiple prose realizations across five genres and six difficulty levels. Across 600 instances, 120 families, and 17,400 scored records, 25 frontier models and four deterministic baselines were tested in explicit-question, formal-world, and implicit-story conditions. The results separate reconstruction from intuition: unguided task success ranges from 3.6 to 80.1, a 22-fold spread, while fact F1 spans only 1.6-fold. OpenAI's GPT-6 Sol achieves the strongest fact F1 at 0.966 and 75.0 implicit task success; OpenAI's GPT-6 Astra leads intuition at 80.1. Anthropic's Claude Opus 5.5 and DeepSeek-V4 Pro reconstruct worlds well but rank far lower on intuition, while Google's Gemini models generally lose more performance as difficulty rises. xAI's Grok-4.6 shows a 0.000 explicit-to-implicit task gap, whereas GPT-6 Sol performs significantly better without the question, at −0.142. Content filters also distort measurement: Claude Fable 5.1 blocks 585 of 600 attempts, demonstrating that refusal handling must be reported separately from capability.

Original abstract

Most benchmarks hand the model a question. EnigmaForge hands it a stack of old documents and no question at all. Buried in the letters, receipts, and logbook margins is a small logic puzzle whose solution is unique - proved by a SAT solver at generation time, with an ablation certificate showing every clue is load-bearing. Because instances are generated rather than collected, the corpus renews forever. The headline measure is intuition: task success when handed only the story, with world reconstruction as the secondary axis. Twenty-five frontier models ran over 600 instances (17,400 scored records) under three matched conditions. Intuition reshuffles the leaderboard: a 22x spread where fact recovery spans 1.6x, the second-best fact-recoverer ranks fourteenth, one model is indifferent to being told the question, and another is significantly better without it. Several models were blocked by their own content filters before reaching the puzzle - any benchmark scoring refusals as failure is quietly measuring filter behavior.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
03Benchmark

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan

GameHorizon is a large benchmark and dataset designed to test whether AI systems can understand instructions and execute gameplay tasks over short and long time horizons.

Read analysis