EnigmaForge: The Question Is Hidden in the Story
AuthorsDaniel Eisner
AffiliationsSep 2026
Resources
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Key results
Procedurally generated instances evaluated across matched conditions.
Total model-condition records included in the evaluation.
Fold spread in unguided task success across frontier models.
Highest reported world-reconstruction fact F1.
Highest implicit-story task-success score.
Attempts blocked by the model's content filter out of 600.
What the paper found
EnigmaForge evaluates whether language models can discover a problem before solving it: each instance provides only a procedurally generated story assembled from letters, receipts, and logbook fragments, with no stated question. The model must infer the hidden task, reconstruct a finite-domain constraint world, use external knowledge bridges, and execute a policy-validated action. Generation begins with a known solution, strengthens and minimizes its evidence, then verifies uniqueness with a DPLL SAT solver, brute-force cross-checking, and per-clue ablation certificates showing that removing any clue creates a second solution. Seed-based generation supports renewable, contamination-resistant evaluation, with multiple prose realizations across five genres and six difficulty levels. Across 600 instances, 120 families, and 17,400 scored records, 25 frontier models and four deterministic baselines were tested in explicit-question, formal-world, and implicit-story conditions. The results separate reconstruction from intuition: unguided task success ranges from 3.6 to 80.1, a 22-fold spread, while fact F1 spans only 1.6-fold. OpenAI's GPT-6 Sol achieves the strongest fact F1 at 0.966 and 75.0 implicit task success; OpenAI's GPT-6 Astra leads intuition at 80.1. Anthropic's Claude Opus 5.5 and DeepSeek-V4 Pro reconstruct worlds well but rank far lower on intuition, while Google's Gemini models generally lose more performance as difficulty rises. xAI's Grok-4.6 shows a 0.000 explicit-to-implicit task gap, whereas GPT-6 Sol performs significantly better without the question, at −0.142. Content filters also distort measurement: Claude Fable 5.1 blocks 585 of 600 attempts, demonstrating that refusal handling must be reported separately from capability.
Original abstract
Most benchmarks hand the model a question. EnigmaForge hands it a stack of old documents and no question at all. Buried in the letters, receipts, and logbook margins is a small logic puzzle whose solution is unique - proved by a SAT solver at generation time, with an ablation certificate showing every clue is load-bearing. Because instances are generated rather than collected, the corpus renews forever. The headline measure is intuition: task success when handed only the story, with world reconstruction as the secondary axis. Twenty-five frontier models ran over 600 instances (17,400 scored records) under three matched conditions. Intuition reshuffles the leaderboard: a 22x spread where fact recovery spans 1.6x, the second-best fact-recoverer ranks fourteenth, one model is indifferent to being told the question, and another is significantly better without it. Several models were blocked by their own content filters before reaching the puzzle - any benchmark scoring refusals as failure is quietly measuring filter behavior.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan
GameHorizon is a large benchmark and dataset designed to test whether AI systems can understand instructions and execute gameplay tasks over short and long time horizons.