NTH

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

AuthorsDingyu Yao, Junhao Zhou, Chenxu Yang, Chuanyu Qin, Haowen Hou, Zheming Liang, Congcong Wang, Yuhang Cao, Shenglong Ye, Shuai Xie, Shuhuan Gu, Haoyang Huang, Qingyi Si, Nan Duan, Jiaqi Wang

June 26, 2026 3 min read
Watch on YouTube
The one-line take

This paper introduces an always-on vision-language assistant that watches live video, decides when to speak or stay silent, and can delegate hard tasks to a background model.

Key results

8B
model scale

Open vision-driven interaction model size

4M
interaction clips

Time-aligned streaming clips used for training

16
AdaCodec tokens per predictable frame

Approximate visual token cost for predictable frames

77.6%
win rate vs Doubao

Overall human preference in 58-case evaluation

87.9%
win rate vs Gemini

Overall human preference in 58-case evaluation

2
continuous video runtime

System supports roughly two hours of continuous video

What the paper found

JoyAI-VL-Interaction is an open-source 8B vision-language interaction model from JD.com designed to watch a continuous video stream and decide every second whether to stay silent, speak, or delegate a hard subtask to a background model. Unlike turn-based systems such as OpenAI GPT-Realtime-2, Google Gemini 3.1 Flash Live, and ByteDance Doubao’s video-call assistant, its decision to act is learned inside the model rather than triggered by polling or user turns. The model is built on JoyAI-VL 1.0 with Qwen3-8B and Qwen3-VL ViT, and it uses AdaCodec to compress predictable frames to about 16 visual tokens, enabling hours of continuous streaming at sub-second latency. Training uses more than 4M time-aligned streaming clips across six task families, with a weighted supervised objective and GRPO reinforcement learning for timing, silence, and delegation. In human head-to-head evaluation across 58 cases, JoyAI-VL-Interaction wins 77.6% against Doubao and 87.9% against Gemini, with especially strong results on monitoring and alerting, real-time counting, and real-time translation, where it reaches 100% against Gemini. The system release also includes pluggable ASR/TTS, long-horizon memory, a background bridge, and a deployable vLLM runtime for sustained real-time interaction.

Original abstract

Many moments in the real world do not wait for a user to ask. A fire starts on a security monitor, an expression flickers across a video call, or a product a viewer wants flashes by in a livestream. Yet today's large models remain mostly turn-based by design: they answer only when addressed, and even video-call apps that appear interactive still operate as question-answer systems, reacting only when polled or prompted. We argue for a different paradigm: a model that is present in the world like a person. It continuously watches what is happening now, decides on its own whether to speak or stay silent, interacts in real time, and delegates to a background model when the problem is hard. To advance interaction models and their adoption across domains, we make two fully open-sourced contributions. First, we release JoyAI-VL-Interaction, an 8B-scale, vision-first VL-interaction model. The model makes the response decision internally, choosing each second to stay silent, respond, or delegate to a background model, and it excels at vision-triggered responsiveness and time awareness. We pair it with a transferable training recipe, from which capabilities we never trained for emerge, such as guiding a shopper through changing app screens or improvising a lecture from a slide deck. Second, we release a complete, deployable system built around that model. The system streams any ongoing video into the model, making it genuinely present in the world. All other components are pluggable, including ASR/TTS modules, memory, visualization UI, and a background brain that can connect to any API or agent. Across six real-world scenarios, human raters prefer JoyAI-VL-Interaction over the in-app video-call assistants of Doubao and Gemini by a wide margin. To our knowledge, this is the first open, vision-driven interaction model released together with its training recipe, data, and complete deployable system.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis