RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
AuthorsBojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun, Lihan Zhang, Bin Liang, Kam-Fai Wong, Haibin Huang, Chi Zhang, Xuelong Li
Resources
RefVideo-6M is a massive, quality-controlled dataset that teaches video-editing models to follow both instructions and visual references more reliably.
Key results
ReferenceVideo-6M contains 5M video editing pairs.
The dataset adds 1M reference-based image editing pairs.
Approximately 6M visual references support location- and appearance-guided editing.
The video subset spans 26 editing tasks.
RefMoT reduces reference-stage training computation by 50%.
RefMoT’s overall score on reference-based editing exceeds UniVideo’s 3.98.
What the paper found
RefVideo-6M addresses a central weakness in instructional video editing: training targets generated by editing models often contain artifacts, while text-only supervision cannot specify appearance or editing location precisely. The dataset contains 5M video pairs and 1M image pairs, covering 26 video editing tasks at 720p and 81 to 129 frames, plus 6M visual references such as bounding boxes, circles, masks, styles, textures, objects, lighting directions, backgrounds, and clothing. Its reliability comes from using artifact-free real videos as targets and, for most tasks, reversing edited and original videos so synthetic edits become inputs rather than ground truth; GPT-5.2 and DINOv3 filter artifacts, mismatched instructions, motion failures, and unchanged edits. The accompanying RefMoT model adapts a HunyuanVideo1.5 instruction editor through a Mixture-of-Tokens reference branch, freezing the main branch and training reference-specific projections, which reduces reference-stage training computation by 50% while keeping inference cost comparable. On reference-based editing, RefMoT reaches a Gemini-3-Pro overall score of 4.16, exceeding UniVideo’s 3.98, while on instruction-based editing it reaches 4.56 versus the previous-best 4.28 and receives 54.7% user preference. The pipeline also uses OpenAI’s GPT-5.2 and GPT-5.5, Google’s Gemini-3-Pro, Black Forest Labs’ FLUX.2-Klein-9B, and HunyuanVideo1.5, demonstrating that carefully constructed multimodal supervision—not only larger architectures—can substantially improve controllability, visual quality, temporal consistency, and reference fidelity.
Original abstract
Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, target videos are commonly produced by automatic editing models, which may introduce visible artifacts and unreliable supervision signals. Second, most public datasets rely primarily on textual instructions, while lacking visual references that are crucial for precise, identity-preserving, and controllable editing. To address these limitations, we introduce RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples. To ensure reliable supervision, our dataset uses a construction pipeline that treats artifact-free real videos as editing targets and generates quality-filtered input conditions with multiple editing experts. In addition, it provides approximately 6 million visual references, covering diverse reference types and editing scenarios, thereby enabling models to learn fine-grained visual correspondence beyond text-only instructions. Based on RefVideo-6M, we further train a reference-guided video editing model, Ref-MoT, to evaluate the effectiveness and scalability of the proposed dataset. Extensive experiments demonstrate that RefVideo-6M provides substantially more reliable supervision than existing datasets and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency. The open-source dataset is available at https://huggingface.co/datasets/RefVideo6M/RefVideo6M.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.