NTH

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

AuthorsXu Xu, Jinxiu Liu, Zhangbo Qiao, Jiaxing Lu, Xiangyu Zhang, Yubin Gu, Fangwei Ning, Yan Shi

AffiliationsBeihang University · Overview of OmniHarness and its diverse visual generation capabilities. 1 · The Chinese University of Hong Kong · National University of Singapore

September 20, 2026 2 min read
Watch on YouTube
The one-line take

OmniHarness turns successful visual-agent behavior into reusable symbolic skills that can be practiced, adapted, and plugged into new generation tasks.

Key results

50
Inquiry iterations

Self-directed inquiry iterations used to acquire reusable policies before downstream evaluation

43
Learned symbolic policies

Policies stored after 50 inquiry iterations

95.0%
ComfyBench Creative Resolve

Resolve rate achieved by Codex GPT-4o with OmniHarness

27.5
Creative baseline improvement

Percentage-point improvement over the strongest ComfyBench Creative baseline

0.997
GenEval overall

Overall compositional text-to-image score

0.86
WISE overall

Overall world-knowledge-informed text-to-image score

What the paper found

OmniHarness is a visual-generation harness that turns verified ComfyUI executions into reusable symbolic policies: instance-specific prompts and images are removed, while procedures, preconditions, dependencies, and failure remedies remain available for retrieval, adaptation, and composition. Before receiving downstream objectives, it performs self-directed inquiry, selecting practice tasks with capability novelty and a competence-frontier heuristic; during execution, ordered planning, intermediate verification, and localized recovery prevent errors from propagating, while policy updates occur with model weights frozen. After 50 inquiry iterations, the system learned 43 symbolic policies and was evaluated across ComfyBench, GenEval, GenEval2, Reason-Edit, WISE, and KRIS-Bench using multiple agent frameworks and models including OpenAI’s GPT-4o, Google’s Gemini-2.5-Flash, DeepSeek-V3, and GPT-Image-1. On ComfyBench Creative tasks, Codex GPT-4o with OmniHarness achieved a 95.0% Resolve rate, 27.5 percentage points above the strongest baseline; overall Resolve reached 92.5%. It scored 0.997 on GenEval and 0.86 on WISE, while image-only policies transferred to unseen video tasks with an 85.9% overall Resolve rate. Removing self-directed inquiry reduced Creative Resolve from 95.0% to 72.5%, and disabling online policy updates reduced overall Resolve to 88.5%, supporting both proactive skill acquisition and continual feedback-driven refinement. A frozen policy snapshot also improved ComfyAgent, ComfyMind, and SymbOmni without fine-tuning, indicating portability across visual-agent architectures and modalities.

Original abstract

Unified multimodal large language models (MLLMs) and multi-agent systems have advanced visual generation. However, three limitations remain. (1) Existing methods often distill task-specific experience with limited generalizability. (2) Reflection is often deferred until task completion. (3) Knowledge is often acquired only in response to downstream task demands. To address these limitations, we introduce OmniHarness, a framework for generalizable visual generation via symbolic policy learning. OmniHarness abstracts verified executions into symbolic policies for visual generation task families, capturing shared procedures and applicability conditions while removing instance-specific inputs. The harness instantiates, adapts, and composes these policies for new tasks. Intermediate verification guides refinement and failure recovery during execution. Through self-directed inquiry, OmniHarness autonomously generates and executes practice tasks near its capability limits before downstream objectives are specified. Execution feedback continually refines the policies while model parameters remain fixed. Experiments across six benchmarks, three MLLM backbones, and three visual agent frameworks demonstrate strong performance and continual capability expansion. On ComfyBench's Creative tasks, OmniHarness achieves a 95.0% resolve rate, exceeding the strongest baseline by 27.5 percentage points. Frozen policy snapshots improve existing visual agent systems through plug-and-play reuse.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis