SwiftVR: Real-Time One-Step Generative Video Restoration
AuthorsJiaqi Yan, Xiangyu Chen, Xinlin Zhong, Haibin Huang, Chi Zhang, Jie Liu, Jiantao Zhou, Xuelong Li
Resources
SwiftVR makes high-quality video restoration run in real time on consumer GPUs by redesigning one-step generative modeling for efficient streaming inference.
Key results
SwiftVR throughput at 1920×1080 on one H100-80G under causal streaming
SwiftVR throughput at 2560×1440 on one H100-80G under causal streaming
SwiftVR throughput at 3840×2160 on one H100-80G under causal streaming
SwiftVR throughput at 1920×1080 on a consumer RTX 5090
MFSWA speedup over the full-attention teacher
Baseline VAE size compared with ReAE
What the paper found
SwiftVR, from TeleAI at China Telecom and the University of Macau, is a streaming one-step generative video restoration framework built on the Wan2.2-TI2V-5B backbone that targets live, high-resolution restoration under strict latency and memory limits. Its core novelty is mask-free shifted-window self-attention, which encodes window locality with deterministic indexing so attention stays on standard dense SDPA kernels rather than mask-based or sparse fallbacks, and a lightweight Restoration-aware Autoencoder that cuts decoding overhead. Trained in three stages—latent flow matching, window-attention distillation, and joint pixel-space fine-tuning with adversarial supervision—SwiftVR reaches a 1.62× throughput gain over its full-attention teacher. On a single H100-80G under causal streaming, it runs at 54.42 FPS for 1920×1080, 31.32 FPS for 2560×1440, and 13.84 FPS for 3840×2160, while all compared diffusion-based baselines run out of memory at 4K. On a consumer RTX 5090, it achieves 26 FPS at 1080p, which the paper presents as the first real-time 1080p generative video restoration on consumer hardware. Quantitatively, SwiftVR is the most efficient one-step method at 2560×1440, and the ReAE reduces the autoencoder bottleneck from 704.69M parameters in Wan2.2-VAE to 40.95M while preserving substantially better reconstruction than a tiny autoencoder. Across SPMCS, UDM10, YouHQ40, and VideoLQ, SwiftVR leads no-reference perceptual metrics such as MUSIQ, CLIP-IQA, and MANIQA in several settings, emphasizing perceptual realism over pure PSNR.
Original abstract
Real-time video restoration (VR) for live streams requires high-resolution outputs under strict per-frame latency constraints. Existing one-step diffusion-based VR models remain difficult to deploy on consumer-grade GPUs due to two main bottlenecks: quadratic spatial attention at high resolutions and the latency-memory overhead of large video autoencoders. We present SwiftVR, a streaming one-step generative VR framework that reduces both bottlenecks under a causal chunk-wise protocol. For attention, mask-free shifted-window self-attention gathers each spatial window into a dense tensor via deterministic indexing, keeping all attention calls on the dense scaled dot-product attention path without masks, cyclic shifts, padding, or hardware-specific sparse kernels. Because SwiftVR uses only standard dense SDPA calls, the trained model transfers to consumer GPUs without retraining or custom kernels. For autoencoding, a lightweight Restoration-aware Autoencoder enables fast chunk-wise decoding while preserving reconstruction quality. On a single H100, SwiftVR sustains 31~FPS at 2560x1440 and 14~FPS at 3840x2160, whereas all compared diffusion-based VR baselines exceed the memory limit at 4K. On a consumer RTX~5090, SwiftVR reaches 26~FPS at 1920x1080. To our knowledge, SwiftVR is the first generative VR model to achieve real-time 1080p streaming on a consumer-grade GPU, while attaining strong no-reference perceptual quality with lower inference cost. Project is available at https://h-oliday.github.io/SwiftVR.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.