NTH

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

AuthorsSen Liang, Cong Wang, Zhentao Yu, Fengbin Guan, Zhengguang Zhou, Teng Hu, Youliang Zhang, Yuan Zhou, Xin Li, Qinglin Lu, Zhibo Chen

July 5, 2026 2 min read
Watch on YouTube
The one-line take

Goku introduces a million-scale dataset and benchmark for instruction-based video editing, plus a model that better handles both appearance changes and structural motion edits.

Key results

2M
dataset scale

Goku video editing pairs

1M
source clips

Koala-36M clips retained for synthesis

88%
synthesized samples filtered

Removed by the progressive post-synthesis validation stage

1000
benchmark cases

Human-verified Goku-Bench test cases

7
metrics

Specialized evaluation metrics in Goku-Bench

8%
instruction following gain

Best improvement over other open-source models on Goku-Bench

What the paper found

Goku, developed by the University of Science and Technology of China and Tencent Hunyuan, pushes instruction-based video editing beyond narrow appearance transforms by releasing a 2 million-pair dataset that explicitly covers basic edits, camera movement, subject movement, reference-guided editing, and 2-to-5-task multi-edit compositions. The paper’s synthesis pipeline starts from 1 million Koala-36M clips, then uses Gemini2.5-Pro for instruction generation and progressive filtering, Grounded-SAM2 for temporal masks, VACE, Flux, Minimax-Remover, Wan2.2, and RecamMaster for task-specific synthesis, and discards about 88% of synthesized candidates to preserve alignment, temporal stability, and photorealism. On top of this data, the authors propose Goku-Edit, a dual-branch model built on Wan2.2-5B with a frozen Qwen3VL-8B MLLM text encoder, a dedicated mask-prediction branch, RoPE-aligned spatial cross-attention, and SpatialCFG for stronger structural control. They also introduce Goku-Bench, a human-verified benchmark with 1,000 test cases and 7 metrics spanning physical rule fidelity, spatial accuracy, instruction following, and task-specific subject motion, camera motion, and style transfer. On Goku-Bench, Goku-Edit delivers up to an 8% gain in instruction following over other open-source models, and ablations show that Goku scaling, filtering, and the spatial modules all contribute measurably to the final performance.

Original abstract

Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge this gap, we present Goku, a large-scale dataset featuring 2 million high-quality, instruction-aligned video editing pairs, which is the first to extend task boundaries from basic appearance editing to multi-task and structural manipulations(e.g., precise control of subject movement). To tackle the data synthesis challenges inherent in these complex tasks, we design an efficient data synthesis pipeline that decomposes complex edits into controllable sub-problems and introduce a progressive filtering system for data reliability throughout the whole process. Furthermore, we explore the optimal network structures on Goku, and propose Goku-Edit. To deeply comprehend complex editing instructions, Goku-Edit leverages an MLLM as its text encoder and adopts a decoupled dual-branch design: a dedicated mask branch handles structural control, freeing the main branch for appearance rendering. A comprehensive video editing benchmark, Goku-Bench, is also proposed with 1,000 human-verified test cases and 7 novel editing-specific metrics. Evaluated on Goku-Bench, Goku-Edit obtains up to +8% improvement on other open-source models in terms of instruction following.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis