OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning
AuthorsYunyang Ge, Xianyi He, Zezhong Zhang, Bin Lin, Bin Zhu, Xinhua Cheng, Li Yuan
Resources
OSP-Next makes text-to-video generation faster and more hardware-friendly by mixing sparse attention, smarter parallelism, quantization, and post-training RL while keeping quality high.
Key results
OSP-Next on VBench, outperforming the Wan2.1 baseline
SSP reduces communication volume versus Ulysses sequence parallelism
OSP-Next speedup on NVIDIA H200 under the 5-second 768P setting
OSP-Next speedup on NVIDIA H200 under the 5-second 768P setting
Quality drop for OSP-Next-HiF8 relative to the baseline
OSP-Next-HiF8 speedup on Ascend 950PR under the 5-second 768P setting
What the paper found
OSP-Next, from the Open-Sora Plan Team at Peking University with collaborators at Nanyang Technological University and Rabbitpre AI, is a text-to-video Diffusion Transformer that targets the core bottleneck of quadratic attention by combining Skiparse-2D Attention, Sparse Sequence Parallelism (SSP), HiF8 quantization, and Mix-GRPO reinforcement learning. The paper’s key novelty is to apply fixed-pattern sparse attention separately along height and width, making the sparsity layout natively compatible with FlashAttention while better matching video locality than Skiparse-1D. On the systems side, SSP exploits local equivalence to switch sparse patterns with a single All-to-All, cutting communication volume by 75% versus Ulysses sequence parallelism. On the optimization side, HiF8 is used with coarse per-tensor scaling, and the sparse model is post-trained with Mix-GRPO using VideoAlign rewards for visual quality, motion quality, and text alignment. Initialized from Wan2.1-T2V-14B and trained on an internal high-quality video corpus at 81 frames and 720×1280, OSP-Next reaches a VBench total score of 83.73%, exceeding the Wan2.1 baseline. In speed tests on NVIDIA H200, it delivers up to 1.64× single-GPU and 1.52× eight-GPU speedup, while OSP-Next-HiF8 preserves performance with only a 0.4% VBench drop and reaches 2.27× speedup on Ascend 950PR.
Original abstract
Diffusion Transformers achieve strong video generation quality, but the quadratic cost of full attention limits efficiency. We introduce OSP-Next, an efficient text-to-video generation model that integrates sparse attention, parallelism, quantization, and reinforcement learning. OSP-Next uses a hybrid full-sparse attention architecture, where the sparse component is implemented with Skiparse-2D Attention. This fixed-pattern mechanism applies token-wise and group-wise sparse attention along spatial dimensions, leveraging locality while maintaining native compatibility with FlashAttention kernels. Based on the local equivalence of rearrangement in Skiparse-2D Attention, we further propose Sparse Sequence Parallelism (SSP), which partitions subsequences across ranks and switches sparse patterns through a single All-to-All communication. Compared with Ulysses Sequence Parallelism (SP), SSP provides a native parallel strategy for sparse attention and reduces communication volume by 75%. OSP-Next also incorporates HiF8 quantization to enable stable joint training with 8-bit quantization and sparse fine-tuning, and applies Mix-GRPO post-training to improve the performance of the sparse model. Experiments show that OSP-Next achieves a VBench total score of 83.73%, surpassing the Wan2.1 baseline. Under the 5-second 720P and 5-second 768P settings, OSP-Next achieves up to 1.64$\times$ single-GPU speedup and over 1.52$\times$ eight-GPU speedup on NVIDIA H200 GPUs. In addition, with only a 0.4% drop in VBench total score, OSP-Next-HiF8 achieves 1.69$\times$ and 2.27$\times$ speedups under the two settings on a single Ascend 950PR, demonstrating the efficiency and performance of OSP-Next across hardware platforms.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.