NTH
Research collection

Generative Models research

Research on learning to generate new data, from images and video to structured outputs. Explore modeling choices, evaluation methods, and practical tradeoffs.

63 papers · Latest edition October 1, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Generative Models papers

Newest editions first.

01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis
04Generative Model

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

Jiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang, Keyang Xu, Jieru Mei, Hongliang Fei, Ruogu Fang, Wei Shao, Cihang Xie, Yuyin Zhou

MovieGrid turns long videos into jointly modeled spatial grids, helping generative models produce more coherent stories with many connected shots.

Read analysis
05Generative Model

SAM3D-Part: Interactive Part Selection and Generation from 3D Objects

Jiahao Chang, Dong Du, Wanhu Sun, Yujian Zheng, Chuanyu Pan, Bowen Zhao, Chongjie Ye, Yuanming Hu, Xiaoguang Han

SAM3D-Part lets users ask for specific components of a 3D object and generates complete, correctly placed meshes rather than decomposing the entire object.

Read analysis
06Generative Model

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li, Dechuang Chen, Ming Lin, Jingjiang Zhou, Haopeng Jin, Qi Jia, Xiaohang Wang, Yaole Wang, Zhanqiang Zhang, Ran Li, Zhengkun Huang, Shuyue Xiong, Yuji Wang, Zikun Dai, Hui He, Yang Luo, Mang Ning, Weiqi Feng, Chengyang Ye, Xinyue Lin, Min Zhao, Hongzhou Zhu, Hengkai Tan, Zeyuan Wang, Chendong Xiang, Kaiwen Zheng, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu

Vidu S2 brings real-time, interactive, editable, and potentially spatial AI-generated video to live applications.

Read analysis
07Generative Model

Kirin: Animal Motion Generation from In-the-Wild Video

Brian Nlong Zhao, Zhuoyang Pan, James M. Rehg, Jiajun Wu, Shangzhe Wu

Kirin turns large collections of animal videos into a multimodal system that can generate and animate realistic movements for diverse species.

Read analysis
08Generative Model

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen, Muntasir Wahed, Nabeel Bashir, Xiaona Zhou, Tianjiao Yu, Vedant Shah, Ismini Lourentzou

EditVid is a training-free video editor that combines several clever mechanisms to preserve motion, identity, and locality across many kinds of edits.

Read analysis
09Generative Model

TGR: Advancing Industrial Recommendation from Generative-Paradigm Ranking toward Unified Generation and Reasoning

TGR Team, Lei Cheng, Haonan Hu, Beibei Kong, Yudong Li, Zang Li, Yunsheng Pang, Hongyang Su, Jianchao Tu, Yunlong Wang, Bing Wen, Junzhang Zhu, Shaojie Zhu, Chengxiang Zhuo

Tencent’s TGR framework replaces fragmented recommendation pipelines with generative ranking, slate generation, and offline reasoning, delivering measurable gains at massive production scale.

Read analysis
10Generative Model

EditaLive! Unified Character Video Editing for Live Streaming

Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun

EditaLive turns human video editing into a fast, live-streaming experience by adapting and distilling a pretrained animation model while preserving facial expressions and character appearance.

Read analysis
13Generative Model

InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

Yunze Tong, Mushui Liu, Canyu Zhao, Shiyi Zhang, Didi Zhu, Peng Zhang, Wanggui He, Jinlong Liu, Ying Chen, Hao Jiang, Pipei Huang, Bo Zheng

InfinityEdit lets video generators apply ongoing edits to an endless stream while maintaining temporal continuity and stability over time.

Read analysis
16Generative Model

Beyond Pixels: From Video Priors to 4D Worlds

Zihao Liu, Xiaolong Shen, Zhenglin Zhou, Ruijie Quan, Yi Yang

Latent-to-4D turns the hidden representations of video generators directly into stable dynamic 3D worlds without passing through RGB.

Read analysis
18Generative Model

V-RAE: Rethinking Video Latent Spaces for Generation

Minghui Guo, Shengqiong Wu, Hao Fei

V-RAE makes video generation more efficient and semantically meaningful by compressing videos into latents derived from pretrained vision models rather than pixel-focused autoencoders.

Read analysis
19Generative Model

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Smriti Joshi, Apostolia Tsirikoglou, Daniel M. Lang, Richard Osuala, Noah Márquez Varaa, Alejandro Guzman, Grzegorz Skorupko, Sebastian Ibarra Arregui, Lidia Garrucho, Akane Ohashi, Dimitra Ntoula, Eugen Divjak, Oğuz Lafcı, Jan C. Peeken, Julia A. Schnabel, Fredrik Strand, Oliver Diaz, Karim Lekadir

A single-pass generative model synthesizes realistic, time-resolved breast MRI contrast without gadolinium and shows promising robustness and clinical utility.

Read analysis
21Generative Model

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

Xiangbo Gao, Siyuan Yang, Ping He, Mingyang Wu, Yuheng Wu, Yushen Zuo, Jiongze Yu, Ryan Cui, Hongyuan Hua, Devin Ma, Xiao Jin, Yubo Yuan, Qing Yin, Jie Yang, Zhengzhong Tu

Visko Orbis aims to make long, high-quality video generation interactive in real time, letting users change the story or visual direction while the video is still being produced.

Read analysis
24Generative Model

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng

VideoCoCo uses executable Blender code as chain-of-thought to plan physical dynamics before a generative video model turns them into realistic videos.

Read analysis
25Generative Model

Generative AI floods and dilutes the market for books

Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg, Paramveer Dhillon

AI-generated books may be lower quality, but their sheer volume is reshaping Amazon’s market and squeezing revenue from human-authored books.

Read analysis
26Generative Model

Qwen-Music Technical Report

Jin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zihan Liu, Zijian Lin, Dake Guo, Hangrui Hu, Lei Xie, Linhan Ma, Wei Xue, Wenxiang Guo, Xinfa Zhu, Xipin Wei, Yangze Li, Yuanjun Lv, Yuxuan Wang, Yunfei Chu, Zhiyong Wu

Qwen-Music combines melody planning, large-scale multilingual training, and waveform rendering to generate and reinterpret high-quality songs with vocals.

Read analysis
29Generative Model

CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training

Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert

This paper trains a 3D latent generative model to synthesize realistic chest CT scans with controllable clinical findings, then uses reinforcement learning to make the requested attributes appear more reliably.

Read analysis
30Generative Model

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

Paul Engstler, Iro Laina, Christian Rupprecht, Andrea Vedaldi

SynCity 3000 turns image-to-3D generation into a scalable scene-level diffusion system that can build large, coherent 3D worlds from prompts and layouts.

Read analysis
31Generative Model

GEAR: Guided End-to-End AutoRegression for Image Synthesis

Bin Lin, Zheyuan Liu, Chenguo Lin, Sixiang Chen, Yunyang Ge, Yunlong Lin, Jianwei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Li Yuan

GEAR lets an image generator and its tokenizer learn together, making autoregressive image synthesis train faster and produce better visual representations.

Read analysis
33Generative Model

PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation

Chunshi Wang, Haohan Weng, Junliang Ye, Biwen Lei, Yang Li, Zibo Zhao, Zeqiang Lai, Kaiyi Zhang, Yunhan Yang, Zhuo Chen, Chunchao Guo, Yawei Luo

PolyFlow turns mesh generation into a parallel continuous process by embedding topology into a learnable state space, making artist-style 3D mesh synthesis faster and more controllable.

Read analysis
34Generative Model

DanceOPD: On-Policy Generative Field Distillation

Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, Lixue Gong, Yongyuan Liang, Meng Chu, Leigang Qu, Lingdong Kong, Wei Liu, Tat-Seng Chua

DanceOPD teaches a single image generator to juggle text-to-image creation, editing, and guidance signals by learning from its own rollout states, helping one model do many generation tasks without the usual capability conflicts.

Read analysis
35Generative Model

Perceptual Flow Matching for Few-Step Generative Modeling

Chuyang Zhao, Yifei Song, Hongfa Wang, Jianlong Yuan, Yuan Zhang, Siming Fu, Zhineng Chen, Huilin Deng, Haoyang Huang, Nan Duan

This paper makes generative models faster by training them in perceptual feature space, letting them produce good samples in just a few steps instead of dozens.

Read analysis
36Generative Model

Representation Distribution Matching for One-Step Visual Generation

Lan Feng, Wuyang Li, Eloi Zablocki, Matthieu Cord, Alexandre Alahi

This paper shows how to train strong one-step image generators by matching feature distributions from multiple frozen encoders, achieving state-of-the-art quality and even compressing a four-step model into a single step.

Read analysis
37Generative Model

Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

Xin Jin, Huanqia Cai, Zhen Li, Zechao Zhan, Dengyang Jiang, Aiming Hao, Yuming Jiang, Chunle Guo, Peng Gao, Ming-Ming Cheng, Steven C. H. Hoi

This paper upgrades reward models for image generation by teaching them to reason about score distributions instead of predicting a single scalar, making them both more accurate and easier to deploy.

Read analysis
41Generative Model

Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models

Yanming Zhang, Yihan Bian, Jingyuan Qi, Yuguang Yao, Lifu Huang, Tianyi Zhou

This paper teaches mask diffusion models to “think again” by revisiting and locally revising their own outputs across multiple turns, enabling stronger reasoning and refinement without starting over.

Read analysis
44Generative Model

Sumi: Open Uniform Diffusion Language Model from Scratch

Mengyu Ye, Keito Kudo, Wataru Ikeda, Ryosuke Matsuda, Keisuke Sakaguchi, Jun Suzuki

Sumi is a fully open 7B diffusion language model trained from scratch on 1.5T tokens, offering the first large-scale reference point for studying uniform diffusion generation in practice.

Read analysis
46Generative Model

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, Yan Zhou, Zijie Meng, Zhimin Zhang, Yawen Luo, Guoxin Zhang, Yu-Shen Liu, Pengfei Wan

OmniDirector makes video generation follow complex multi-shot camera moves by turning camera paths into visual motion grids and training a large controllable diffusion system on them.

Read analysis
47Generative Model

SwiftVR: Real-Time One-Step Generative Video Restoration

Jiaqi Yan, Xiangyu Chen, Xinlin Zhong, Haibin Huang, Chi Zhang, Jie Liu, Jiantao Zhou, Xuelong Li

SwiftVR makes high-quality video restoration run in real time on consumer GPUs by redesigning one-step generative modeling for efficient streaming inference.

Read analysis
48Generative Model

Channel-wise Vector Quantization

Wei Song, Tianhang Wang, Yitong Chen, Tong Zhang, Zuxuan Wu, Ming Li, Jiaqi Wang, Kaicheng Yu

This paper proposes a new way to turn images into discrete tokens by quantizing channels instead of patches, enabling a fresh autoregressive image generator that paints details progressively from coarse structure to fine texture.

Read analysis
50Generative Model

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Huaisong Zhang, Hao Yu, Yuxuan Zhang, Jiahe Wang, Xinrui Chen, Haoxiang Cao, Feng Lu, Wendong Zhang, Changqian Yu, Chun Yuan

This paper turns image-quality debugging for text-to-image models into a structured task that pinpoints what is wrong, where it is wrong, why, and how important it is, then uses those defects to improve the generator.

Read analysis
52Generative Model

Streaming Video Generation with Streaming Force Control

Hanhui Wang, Yiming Xie, Haiwen Feng, Zhaoyang Lv, Shenlong Wang, Huaizu Jiang

StreamForce is a real-time video generation system that can instantly react to continuous force inputs, aiming to make generated motion both physically responsive and visually realistic.

Read analysis
54Generative Model

Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation

Yuxuan Bian, Zeyue Xue, Songchun Zhang, Shiyi Zhang, Weiyang Jin, Yaowei Li, Junhao Zhuang, Haoran Li, Jie Huang, Haoyang Huang, Nan Duan, Qiang Xu

This paper introduces a new memory system that lets video generators keep producing coherent frames for extremely long durations in real time, pushing toward truly infinite video generation.

Read analysis
55Generative Model

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

Dong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, Jiawei Zhang, Jinjing Zhao, Sirui Zhang, Yang Yue, Zhiyang Liang, Baining Guo, Chong Luo, Jianmin Bao, Ji Li, Lei Shi, Qinhong Yang, Xiuyu Wu, Xuelu Feng, Yan Lu, Yanchen Dong, Yitong Wang, Yunuo Chen

Lens is a more compute-efficient text-to-image model that uses richer data, smarter batching, and post-training tricks to match or beat larger systems while generating images faster.

Read analysis
56Generative Model

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

Jianzong Wu, Hao Lian, Jiongfan Yang, Dachao Hao, Ye Tian, Yunhai Tong, Jingyuan Zhu, Biaolong Chen, Qiaosong Qi, Aixi Zhang, Wanggui He, Mushui Liu, Jinlong Liu, Hao Jiang

LoomVideo is a faster, smaller video generation-and-editing model that uses a clever conditioning trick to avoid expensive token concatenation while still handling multimodal inputs.

Read analysis
57Generative Model

VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral, Adil Kaan Akan, Kaan Oktay, Hoda Eldardiry, Pinar Yanardag

VideoMLA makes minute-scale autoregressive video diffusion much cheaper in memory by compressing the KV cache with a low-rank latent representation, enabling longer video rollouts with better throughput and strong quality.

Read analysis
61Generative Model

Generative Modeling by Value-Driven Transport

Pablo Moreno-Muñoz, Adrian Müller, Gergely Neu

This paper turns generative modeling into a control problem and uses value functions to learn fast, straight transport paths for sampling data efficiently.

Read analysis
62Generative Model

Let EEG Models Learn EEG

Yifan Wang, Yijia Ma, Wen Li, Chenyu You

This paper teaches a transformer to generate EEG as a continuous signal, using flow matching and signal-aware constraints to better preserve the brain-wave patterns that traditional denoising methods often miss.

Read analysis