NTH

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

AuthorsXiangbo Gao, Siyuan Yang, Ping He, Mingyang Wu, Yuheng Wu, Yushen Zuo, Jiongze Yu, Ryan Cui, Hongyuan Hua, Devin Ma, Xiao Jin, Yubo Yuan, Qing Yin, Jie Yang, Zhengzhong Tu

August 7, 2026 2 min read
Watch on YouTube
The one-line take

Visko Orbis aims to make long, high-quality video generation interactive in real time, letting users change the story or visual direction while the video is still being produced.

Key results

4K
Output resolution

Progressive streaming super-resolution delivers 4K video.

24
Frame rate

The system generates and serves video at 24 FPS in real time.

1
Prompt response latency

Prompt updates become visible in under 1 second on average.

74
Evaluation cases

The automated benchmark evaluates 74 interactive long-video cases.

1838
Overall Arena Elo

Orbis achieves the highest overall human-preference Elo among nine systems.

1940
Temporal stability Elo

Orbis achieves the highest temporal-stability rating in the human Arena study.

What the paper found

Visko Orbis 1.0, from Team Visko, proposes a Live Model for video generation: unlike offline systems such as OpenAI’s Sora and Google DeepMind’s Veo, it remains active, accepts prompt changes during generation, and applies them to future uncommitted chunks without restarting. A chunk-wise latent-flow generator supports text-to-video, image-to-video, video continuation, multilingual prompts, and hour-scale rollouts through bounded multi-scale memory that preserves recent detail while compressing older history. Event-aligned temporal captions, rollout-corrupted history training, guidance distillation, self-forcing distribution matching, GRPO reward optimization, and a latent world model target continuity, prompt alignment, and physically plausible motion. The serving stack combines KV-cache reuse, compiled and fused transformer execution, Ulysses-style sequence parallelism, progressive decoding, and a reference-aware video upscaler to deliver 4K video at 24 FPS, with prompt updates visible in under 1 second on average. On a 74-case benchmark containing one- to three-minute outputs, Orbis records a DOVER aesthetic score of 0.8101 and a technical score of 0.5572, while its long-form human Arena evaluation gives it the highest overall Elo of 1838 and temporal-stability Elo of 1940 among nine systems. The results suggest its main advantage is not only image quality, but maintaining a steerable, coherent stream as events and instructions change.

Original abstract

We present Visko Orbis 1.0, a Live Model for real-time, interactive long-video generation. Users can change the prompt at any moment during generation, and the update becomes visible in real time. Visko Orbis 1.0 supports long-form text-to-video, image-to-video, and video continuation, with multilingual prompts and prompt switching while generation is in progress. A bounded multi-scale memory preserves subjects, scenes, and style across chunks, sustaining hour-scale rollouts without evident quality or color drift. Built on a distilled chunk-wise streaming generator and a streaming video upscaler, Visko Orbis 1.0 delivers real-time 4K video generation at 24 FPS using an optimized GPU serving engine. In long-form Arena comparisons, Visko Orbis 1.0 obtains the highest overall-preference and temporal-stability ratings among state-of-the-art real-time interactive video-generation systems.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis