NTH

EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers

AuthorsZongyuan Yang, Mingjing Yi, Wanli Ma, Chenzhuo Fan, Bocheng Li, Baolin Liu, Yuke Lou, Yingde Song, Yongping Xiong, Zhengdong Guo, Shimu Wang

June 30, 2026 2 min read
Watch on YouTube
The one-line take

EVA01 is a multimodal model that treats 3D meshes as a first-class input/output modality, enabling text-to-3D generation and interactive 3D editing in one unified system.

Key results

2B
Backbone size

EVA01 is built on Qwen3-VL-2B-Instruct.

1.2M
Static asset corpus

Raw 3D assets aggregated for training.

400K
Premium asset subset

High-quality filtered subset used for later training stages.

3M
Procedural editing trajectories

Synthesized multi-turn procedural edit sequences.

300K
Semantic editing trajectories

Synthesized multi-turn semantic edit sequences.

35.72
Text-to-3D CLIP

Toys4K text-to-3D result for EVA01.

What the paper found

EVA01, from the SeeleAI Team, is a native 3D multimodal model built on a Qwen3-VL-2B-Instruct Mixture-of-Transformers backbone that treats textured 3D meshes as first-class tokens rather than external outputs. The key idea is to split the system into an Understanding Expert and a Generation Expert with hard modality routing and shared global self-attention, so semantic reasoning in text, image, and mesh space can directly guide sparse 3D flow-matching generation. The paper’s main technical novelty is a structured sparse voxel representation at 512^3 resolution, paired with 3D Interleaved MRoPE and a five-stage curriculum that aligns mesh captioning, image-conditioned initialization, semantic modality alignment, context-aware instruction tuning, and high-quality finetuning. Trained on 1.2M raw 3D assets, a 400K premium subset, 3M procedural editing trajectories, and 300K semantic editing trajectories, EVA01 produces state-of-the-art native text-to-3D fidelity on Toys4K, reaching 35.72 CLIP, 122.48 FD, 1.18 KD, and 70.4% user preference, while also enabling mask-free multi-turn editing with 93.75% preference and strong identity preservation across turns. On PointLLM-200, the alignment stage achieves the best reference-style captioning scores, and the final model shifts toward broader semantic coverage, with GPT-based judge scores of 59.095 and 65.910. Overall, the work argues that decoupling semantic understanding from geometric synthesis is essential for 3D-native multimodal systems, especially when long-context editing and geometric consistency matter more than one-shot reconstruction.

Original abstract

This paper addresses the challenge of integrating 3D meshes as a native modality within Multimodal Large Language Models (MLLMs). Diffusion-based large reconstruction models decouple semantic understanding from geometric reasoning, operating as stateless reconstructors conditioned on dense 2D pixel priors. Recent MLLM-based methods treat the 3D modality as an external output rather than a native component of the multimodal sequence, making incremental adaptations without a systematic analysis of how geometric manifolds align with MLLM feature spaces. We introduce EVA01, a unified framework that extends the modality boundary of MLLMs to natively incorporate 3D mesh understanding, generation, and context-aware editing. Built upon a Mixture-of-Transformers (MoT) architecture, EVA01 decouples the model into a pre-trained Understanding Expert ($E_{\mathrm{und}}$) and a structurally mirrored Generation Expert ($E_{\mathrm{gen}}$), coupled through shared global self-attention with hard modality routing. This design aligns the semantic latent space of the MLLM backbone with the geometric manifold, enabling direct transfer of multimodal priors without intermediate 2D representations. Results show that EVA01 achieves state-of-the-art native text-to-3D generation fidelity and unlocks robust long-context multi-turn geometric editing with identity preservation, a capability fundamentally inaccessible to stateless reconstruction pipelines. Our findings further offer architectural insights for integrating 2D foundation models with 3D tasks, informing the design of 3D-native multimodal systems. Project Page: https://www.seeles.ai/research/pages/EVA01

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis