LEGO-Anything: Coding Agents for 3D Scene Reconstruction
AuthorsXirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
AffiliationsUniversity of Maryland, College Park · AWS
Resources
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
Key results
Simulator-grounded benchmark inputs spanning indoor and outdoor reconstruction.
Diverse scenes used to generate the benchmark inputs.
Best reported LEGO-Bench overall score on indoor scenes.
Best reported LEGO-Bench overall score on outdoor scenes.
Largest training-free improvement from the plugin on the Office subset.
Detection readout from frozen reconstructed scenes.
What the paper found
LEGO-Anything reframes single-image 3D reconstruction as Image-to-Code: a coding agent iteratively writes and executes Blender programs, renders the scene, compares it with the reference, and produces an editable, executable, queryable scene program. Its LEGO-Bench benchmark contains 208 RGB inputs from 104 simulator-grounded indoor and outdoor scenes, separately measuring artifact validity, visible-surface geometry, and rendered appearance. Using OpenAI’s Codex harness, GPT-6-astra achieves the strongest overall scores—53.4% indoors and 39.6% outdoors—although high artifact validity masks much weaker geometric and visual fidelity. Trajectory analysis identifies weak initialization, regressive edits, and unreliable self-evaluation as the main failure modes. LEGO-Plugin addresses them with VGGT-based Enhanced Initialization, SAM 3 and Depth Anything V2 evidence for Grounded Refinement, and transaction-based Version Control; without training, it delivers relative gains of up to 62.7%. Finally, LEGO-World derives detection, segmentation, and relative-depth readouts from frozen reconstructed scenes: LEGO-Anything reaches 30.14 box AP on COCO, 14.75 mask AP on LVIS, and 0.1554 AbsRel on ETH3D, substantially behind specialist models such as DINO, SAM 3, and Depth Anything 3. The result is a promising executable representation, but current coding agents remain insufficiently precise for faithful natural-image reconstruction.
Original abstract
A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.
Read the original paperMore in AI Agents
Browse all 56 papers →MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.
Atria Dawn: The Dawn of Agentic Superintelligence
Honglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu, Zixian Huang, Jiajie Jin, Alexander Lam, Yining Li, Jiahang Lin, Yanjiang Liu, Xinyu Lu, Haijun Lv, Junlin Shang, Qisheng Su, Guoqiang Wang, Rui Wang, Zhecan Wang, Hao Xiang, Xinchen Xie, Shuhao Xing, Xiaoyu Xing, Wanghan Xu, Xinyu Yang, Yajie Yang, Chengfeng Zhao, Haoran Zhao, Ruojun Zhou, Yunhua Zhou, Yicheng Zou, Kun Cai, Qiye Cai, Xinmeng Che, Haodong Chen, Jiabei Chen, Jiahao Chen, Jiayi Chen, Yujia Chen, Lizhi Cui, Youheng Dai, Xin Deng, Yi Dong, Shihan Dou, Chenya Gu, Xu Guo, Ding Han, Feiyang Hao, Haotan He, Jie Hou, Binze Hu, Zijian Hu, Junhao Huang, Huicheng Jiang, Jiazhen Jiang, Shufan Jiang, Jiahao Kuang, Bowen Lai, Bo Li, Jiaqiang Li, Peng Li, Qilong Li, Zhuoqun Li, Jiaxiang Liu, Shuainan Liu, Tong Liu, Yi Liu, Zhonghang Lu, Jianwen Luo, Yanyi Luo, Huijie Lv, Ningsheng Ma, Zerun Ma, Houcheng Min, Chengjun Pan, Qiyuan Peng, Xiaoxuan Peng, Jianmin Qian, Jiantao Qiu, Wanying Ren, Huayu Sha, Jifei Shan...
Atria Dawn explores how research-oriented AI agents can move beyond completing tasks to partnering with humans on scientific discovery while preserving human oversight.