JoyAI-VL-Interaction
The first open, vision-driven real-time interaction model β it watches a live video stream and decides on its own when to speak, stay silent, or delegate. While enabling online, real-time interaction, this release also delivers powerful offline video understanding, making it the most comprehensive open-source model for video-related capabilities in the 8B parameter class.
π Paper Β· π Project Page & Demos Β· π» GitHub Β· π€ Paper Page
Overview
Most large models today are turn-based: they answer only when you ask. But many moments in the real world don't wait for a question β a fire starts on a security feed, someone falls, a product flashes by in a livestream. Once missed, the moment is gone.
JoyAI-VL-Interaction is built for exactly these moments. It is an 8B-scale, vision-first interaction model that continuously watches a live video stream and, every second, decides on its own to take one of three actions:
- Speak β respond when something is worth saying
- Stay silent β keep watching when nothing warrants a response (a first-class, trained action)
- Delegate β hand a hard subtask to a background model/agent, keep watching, and weave the result back in when it returns
The decision of when to act is learned inside the model (from second-by-second time-aligned data + RL), not bolted on by an external turn-detector or polling loop. Vision is the first-class driver; speech (ASR/TTS) is treated as pluggable I/O.
To our knowledge, this is the first open, vision-driven interaction model released together with its training recipe, data, and a complete deployable system.
Model Performance
Offline Video Evaluation
Across 26 standard video understanding benchmarks, JoyAI-VL-Interaction achieves an average score of 57.53, outperforming Qwen3-VL-8B-Instruct at 54.16 by 3.37 points.
Demo
https://github.com/user-attachments/assets/2853fc95-ad21-4972-8206-5f3d19798b14
Citation
If you find our work helpful, feel free to give us a cite.
@misc{yao2026joyaivlinteractionrealtimevisionlanguageinteraction,
title={JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence},
author={Dingyu Yao and Junhao Zhou and Chenxu Yang and Chuanyu Qin and Haowen Hou and Zheming Liang and Congcong Wang and Yuhang Cao and Shenglong Ye and Shuai Xie and Shuhuan Gu and Haoyang Huang and Qingyi Si and Nan Duan and Jiaqi Wang},
year={2026},
eprint={2606.14777},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.14777},
}
- Downloads last month
- -