NTH

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

AuthorsBojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun, Lihan Zhang, Bin Liang, Kam-Fai Wong, Haibin Huang, Chi Zhang, Xuelong Li

September 3, 2026 2 min read
Watch on YouTube
The one-line take

RefVideo-6M is a massive, quality-controlled dataset that teaches video-editing models to follow both instructions and visual references more reliably.

Key results

5M
Video editing pairs

ReferenceVideo-6M contains 5M video editing pairs.

1M
Image editing pairs

The dataset adds 1M reference-based image editing pairs.

6M
Visual references

Approximately 6M visual references support location- and appearance-guided editing.

26
Editing tasks

The video subset spans 26 editing tasks.

50%
RefMoT training computation reduction

RefMoT reduces reference-stage training computation by 50%.

4.16
Gemini-3-Pro reference-editing score

RefMoT’s overall score on reference-based editing exceeds UniVideo’s 3.98.

What the paper found

RefVideo-6M addresses a central weakness in instructional video editing: training targets generated by editing models often contain artifacts, while text-only supervision cannot specify appearance or editing location precisely. The dataset contains 5M video pairs and 1M image pairs, covering 26 video editing tasks at 720p and 81 to 129 frames, plus 6M visual references such as bounding boxes, circles, masks, styles, textures, objects, lighting directions, backgrounds, and clothing. Its reliability comes from using artifact-free real videos as targets and, for most tasks, reversing edited and original videos so synthetic edits become inputs rather than ground truth; GPT-5.2 and DINOv3 filter artifacts, mismatched instructions, motion failures, and unchanged edits. The accompanying RefMoT model adapts a HunyuanVideo1.5 instruction editor through a Mixture-of-Tokens reference branch, freezing the main branch and training reference-specific projections, which reduces reference-stage training computation by 50% while keeping inference cost comparable. On reference-based editing, RefMoT reaches a Gemini-3-Pro overall score of 4.16, exceeding UniVideo’s 3.98, while on instruction-based editing it reaches 4.56 versus the previous-best 4.28 and receives 54.7% user preference. The pipeline also uses OpenAI’s GPT-5.2 and GPT-5.5, Google’s Gemini-3-Pro, Black Forest Labs’ FLUX.2-Klein-9B, and HunyuanVideo1.5, demonstrating that carefully constructed multimodal supervision—not only larger architectures—can substantially improve controllability, visual quality, temporal consistency, and reference fidelity.

Original abstract

Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, target videos are commonly produced by automatic editing models, which may introduce visible artifacts and unreliable supervision signals. Second, most public datasets rely primarily on textual instructions, while lacking visual references that are crucial for precise, identity-preserving, and controllable editing. To address these limitations, we introduce RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples. To ensure reliable supervision, our dataset uses a construction pipeline that treats artifact-free real videos as editing targets and generates quality-filtered input conditions with multiple editing experts. In addition, it provides approximately 6 million visual references, covering diverse reference types and editing scenarios, thereby enabling models to learn fine-grained visual correspondence beyond text-only instructions. Based on RefVideo-6M, we further train a reference-guided video editing model, Ref-MoT, to evaluate the effectiveness and scalability of the proposed dataset. Extensive experiments demonstrate that RefVideo-6M provides substantially more reliable supervision than existing datasets and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency. The open-source dataset is available at https://huggingface.co/datasets/RefVideo6M/RefVideo6M.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis