NTH
Research collection

Multimodal AI research

Explore models that connect language, images, audio, and other signals. Read findings on cross-modal understanding, generation, and evaluation.

61 papers · Latest edition October 2, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Multimodal AI papers

Newest editions first.

02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis
04Multimodal

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi

ModaLens tests whether medical VLMs truly look at the X-ray when they already have the report, revealing that reports can substantially suppress measurable image sensitivity.

Read analysis
09Multimodal

Training-Free Speech-Centric Omni Understanding with Frozen VLMs

Ankan Deria, Hanoona Rasheed, Xilin He, Fahad Shahbaz Khan, Salman Khan

This work shows that frozen vision-language models can gain strong multilingual speech and audio-visual abilities simply by connecting them to Whisper, without retraining the models themselves.

Read analysis
10Multimodal

A Multimodal Foundation Model for Longitudinal Patient Representation and Scalable Insight Generation in Oncology

Eugene Vorontsov, Yi Kan Wang, Alican Bozkurt, Adam Casson, Ludmila Tydlitatova, Michal Zelechowski, Ezra E. W. Cohen, Jyoti D. Patel, Max Banaszak, Caitlin McWilliams, Shane Colley, Kate Sasser, Ryan Fukushima, Eric Lefkofsky, Razik Yousfi, Siqi Liu

This work builds a multimodal oncology foundation model that learns evolving patient representations from clinical records, genomics, and pathology to improve outcome prediction and treatment selection.

Read analysis
15Multimodal

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge

LAION-BVD opens access to 10 million hours of multimodal video data, aiming to make large-scale video, audio, and image pre-training more accessible.

Read analysis
17Multimodal

Towards Expert-level Medical AI for Real-time Video Consultations

Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu

A Gemini-based multimodal medical AI reportedly matches or exceeds physicians on several simulated consultation tasks while remaining weaker in rapport, subtle affect, and fine-grained physical assessment.

Read analysis
19Multimodal

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

Wanshun Su, Yang Shi, Feihu Liu, Ziwen Yu, Yan Min, Zhuoran Zhang, Qixun Wang, Haotian Wang, Shixuan Liu, Yuanxing Zhang, Peng Wu, Chengfu Huo, Liang Ding

OmniPack makes omni-modal LLMs far more efficient by compressing redundant audio-visual tokens while preserving most of their task-relevant information.

Read analysis
22Multimodal

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang

ClinFusion is a vision-focused medical multimodal LLM that combines 2D and 3D image understanding with radiologist-aligned evaluation for more reliable clinical reports and reasoning.

Read analysis
23Multimodal

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Senqiao Yang, Kaichen Zhang, Zhaoyang Jia, Jinghao Guo, Yifei Shen, Xinjie Zhang, Xiaoyi Zhang, Haoqing Wang, Xiao Li, Peng Zhang, Xiang An, Yin Xie, Zhening Liu, Xun Guo, Jiahao Li, Shicheng Zheng, Jinglu Wang, Zongyu Guo, Wenxuan Xie, Zihan Zheng, Yuxuan Luo, Bin Li, Yan Lu

Mage-VL turns video codecs and motion-aware tokenization into a faster multimodal model designed for real-time visual understanding.

Read analysis
24Multimodal

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang

See2Think tests whether multimodal models truly reason with their sketches and visual feedback, finding that they often choose the right visual actions but struggle to render and use the results faithfully.

Read analysis
25Multimodal

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping

AV-Flamingo is an open multimodal language model that learns to reason across audio, visuals, and long videos while grounding its explanations in time.

Read analysis
26Multimodal

Let RGB Be the Language of Vision

Timing Yang, Jinrui Yang, Xinlong Li, Yuhan Wang, Haoran Li, Yanqing Liu, Guoyizhe Wei, Jixuan Ying, Chen Wei, Rama Chellappa, Yuyin Zhou, Cihang Xie, Alan Yuille, Feng Wang

RINO treats every visual signal as an image, allowing one shared model to understand and generate across tasks like segmentation, depth estimation, and pose-guided synthesis.

Read analysis
27Multimodal

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

Yuhan Zhu, Changlian Ma, Xiangyu Zeng, Xinhao Li, Zhiqiu Zhang, Songze Li, Jun Zhang, Tianxiang Jiang, Yuandong Yang, Ziang Yan, Zikang Wang, Xinyu Chen, Haoran Chen, Shaowei Zhang, Limin Wang

TimeLens2 helps video-language models identify all the moments supporting an answer, using better multi-span supervision and a reward designed for flexible temporal evidence sets.

Read analysis
28Multimodal

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen

MedPMC is a large-scale pipeline that turns biomedical papers into cleaner, more useful image-text training data, boosting medical multimodal models and clinical retrieval.

Read analysis
29Multimodal

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

Yuliang Liu, Zhang Li, Ziyang Zhang, Shuo Zhang, Qiang Liu, Jiajun Song, Zidun Guo, Xinhan Wang, Handong Zheng, Yang Liu, Dongliang Luo, Zhiyin Ma, Jiarui Zhang, Xiang Bai

MonkeyOCRv2 turns document images into a specialized visual foundation model that reads text, preserves layout, and powers compact high-performing document AI systems.

Read analysis
31Multimodal

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang

VideoChat3 is a fully open, efficient 4B-parameter video-language model designed to understand everything from short clips to long and streaming videos.

Read analysis
32Multimodal

Evidence-Backed Video Question Answering

Shijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu, Caiming Xiong, Silvio Savarese, Chen Sun, Juan Carlos Niebles

E-VQA teaches video-language models to answer questions while showing exactly when and where in the video the evidence appears.

Read analysis
34Multimodal

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teiletche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, Sewon Min

PixelRAG shows that for web-based retrieval-augmented generation, using screenshots instead of extracted text can improve both answer quality and efficiency across several QA and agent tasks.

Read analysis
35Multimodal

Vision as Unified Multimodal Generation

Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen, Xuanke Shi, Sihan Wang, Boxuan Li, Linyan Wang, Siyi Xie, Xin You, Jinsheng Quan, Zhongang Cai, Haiwen Diao, Ziwei Liu, Lei Yang, Dahua Lin, Quan Wang

This paper shows how a single multimodal model can handle detection, segmentation, depth, OCR, and more by turning all vision tasks into text-and-image generation.

Read analysis
36Multimodal

Gemma 4 Technical Report

Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, Mayank Chaturvedi, Victor Cotruta, Alice Coucke, Phil Culliton, Robert Dadashi, Lucas Dixon, Mohamed Elhawaty, Utku Evci, Clément Farabet, Johan Ferret, Filippo Galgani, Sertan Girgin, Jean-Bastien Grill, Maarten Grootendorst, Jiaxian Guo, Cassidy Hardin, Yanzhang He, Steven M. Hernandez, Omri Homburger, Léonard Hussenot, Juyeong Ji, Armand Joulin, Aishwarya Kamath, Parnian Kassraie, Olivier Lacombe, Preethi Lahoti, Gaël Liu, Gus Martins, Luciano Martins, Tatiana Matejovicova, Ramona Merhej, Nikola Momchev, Sneha Mondal, Ryan Mullins, Sindhu Raghuram Panyam, Shreya Pathak, Sarah Perrin, André Susano Pinto, Etienne Pot, Angéline Pouget, Alexandre Ramé, Sabela Ramos, Douglas Reid, David Rim, Morgane Rivière, Karsten Roth, Louis Rouillard, Omar Sanseviero, Pier Giuseppe Sessa, Shane Settle, Danila Sinopalnikov, Sara Smoot, Pi...

Gemma 4 is Google’s new open-weight multimodal model family, pushing stronger reasoning, faster inference, and better long-context performance across text, image, and audio tasks.

Read analysis
37Multimodal

Unified Audio Intelligence Without Regressing on Text Intelligence

Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim, Boxin Wang, Zihan Liu, Sungwon Kim, Yang Chen, Arushi Goel, Rajarshi Roy, Wenliang Dai, Zhuolin Yang, Yangyi Chen, Dongfu Jiang, Sreyan Ghosh, Tuomas Rintamaki, Andrew Tao, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping

This paper builds a single model that can understand and generate both audio and text, while trying hard not to lose the reasoning and alignment skills of its text-only LLM core.

Read analysis
38Multimodal

DataComp-VLM: Improved Open Datasets for Vision-Language Models

Matteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian Böther, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan Hammoud, Thomas De Min, Simone Caldarella, Jehanzeb Mirza, Sedrick Keh, Mehdi Cherti, Hilde Kuehne, Bernt Schiele, Serena Yeung-Levy, Muhammad Ferjad Naeem, Federico Tombari, Ana Klimovic, Elisa Ricci, Matthias Bethge, Sewoong Oh, Ameya Prabhu, Alessio Tonioni, Jenia Jitsev, Massimiliano Mancini, Ludwig Schmidt, Nikhil Parthasarathy

This paper creates a large open benchmark for testing how to build better vision-language training data, and shows that mixing the right kinds of data can matter more than filtering.

Read analysis
39Multimodal

EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers

Zongyuan Yang, Mingjing Yi, Wanli Ma, Chenzhuo Fan, Bocheng Li, Baolin Liu, Yuke Lou, Yingde Song, Yongping Xiong, Zhengdong Guo, Shimu Wang

EVA01 is a multimodal model that treats 3D meshes as a first-class input/output modality, enabling text-to-3D generation and interactive 3D editing in one unified system.

Read analysis
40Multimodal

Multimodal Music Recommendation System using LLMs

Srikar Prabhas Kandagatla, Sreehitha R. Narayana, Chandana Magapu, Swetha Mohan, Shamanth Kuthpadi, Hongjie Chen, Ryan A. Rossi, Franck Dernoncourt, Nesreen Ahmed

This paper builds a music recommender that uses song audio, lyrics, and AI-generated descriptions together with LLMs to better predict what people will want to hear next.

Read analysis
42Multimodal

From Pixels to Words -- Towards Native One-Vision Models at Scale

Haiwen Diao, Jiahao Wang, Penghao Wu, Yuhao Dong, Yuwei Niu, Yue Zhu, Zhongang Cai, Weichen Fan, Linjun Dai, Silei Wu, Xuanyu Zheng, Mingxuan Li, Yuanhan Zhang, Bo Li, Hanming Deng, Huchuan Lu, Quan Wang, Lei Yang, Lewei Lu, Dahua Lin, Ziwei Liu

This paper introduces a native vision-language model that learns pixel-to-word connections end-to-end without separate encoders or adapters, aiming to make multimodal systems more unified and competitive at scale.

Read analysis
44Multimodal

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

Dingyu Yao, Junhao Zhou, Chenxu Yang, Chuanyu Qin, Haowen Hou, Zheming Liang, Congcong Wang, Yuhang Cao, Shenglong Ye, Shuai Xie, Shuhuan Gu, Haoyang Huang, Qingyi Si, Nan Duan, Jiaqi Wang

This paper introduces an always-on vision-language assistant that watches live video, decides when to speak or stay silent, and can delegate hard tasks to a background model.

Read analysis
45Multimodal

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

Junke Wang, Xiao Wang, Jiacheng Pan, Xuefeng Hu, Feng Li, Jingxiang Sun, Chaorui Deng, Zilong Chen, Yunpeng Chen, Kaibin Tian, Matthew Gwilliam, Hao Chen, Danhui Guan, Kun Xu, Weilin Huang, Zuxuan Wu, Haoqi Fan, Yu-Gang Jiang, Zhenheng Yang

ARM is a multimodal AI model that turns images into discrete tokens so one autoregressive system can understand, generate, and edit visuals, with reinforcement learning improving both quality and instruction following.

Read analysis
46Multimodal

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang

HYDRA-X is a new unified multimodal model that uses one visual tokenizer for both images and videos, improving understanding, generation, and editing by doing more of the work inside the tokenizer itself.

Read analysis
48Multimodal

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

Hao Li, Ganlong Zhao, Yufei Liu, Haotian Hou, Guoquan Ye, Tongyan Fang, Chunxiao Liu, Siyuan Huang, Jianbo Liu, Xiaogang Wang, Hongsheng Li

ACE-Ego-0 makes robot learning more scalable by turning egocentric human videos into usable action supervision and training vision-language-action models on both human and robot data together.

Read analysis
50Multimodal

Avatar V: Scaling Video-Reference Avatar Video Generation

Benjamin Liang, Ce Chen, Desmond Lin, Ivan Somov, Jiajun Zhao, Jiewei Yuan, Jingfeng Zhang, Junhao Huang, Nik Nolte, Pedram Haqiqi, Penghan Wang, Rong Yan, Rui Zhang, Sam Prokopchuk, Sivan Wang, Viktor Goriachko, Yi Ren, Yuanming Li, Yutao Chen, Zhenhui Ye, Zhibin Hong, Zilong Nie, Zujin Guo

Avatar V is a large-scale system that uses reference videos, not just photos, to generate more realistic talking avatars with better identity, gestures, and lip-sync.

Read analysis
51Multimodal

InstructSAM: Segment Any Instance with Any Instructions

Yuqian Yuan, Wentong Li, Zhaocheng Li, Yutong Lin, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang, Wenqiao Zhang

InstructSAM lets users give natural-language instructions to segment multiple instances in one shot by bridging a vision-language model with SAM3 through learnable instance queries.

Read analysis
53Multimodal

Your Embedding Model is SMARTer Than You Think

Jianrui Zhang, Hyun Jung Lee, Sukanta Ganguly, Tae-Eui Kam, Donghyun Kim, Yong Jae Lee

SMART turns ordinary single-vector embedding models into stronger multimodal retrievers by reusing their hidden states for late-interaction retrieval without needing full retraining.

Read analysis
58Multimodal

A Signal-Language Foundation Model for Broad-Spectrum Cardiovascular Assessment from Routine Electrocardiography

Ziqing Yu, Yuhui Tao, Jiayu Huo, Lei Pan, Zilong Xiao, Juecheng Chen, Xiao Li, Jianxuan Li, You Zhou, Zhixing Li, Cong Wang, Beijian Zhang, Chen Chen, Hongyang Lu, Konstantinos Patlatzoglou, Daniel B. Kramer, Jonathan W. Waks, Yangang Su, Fu Siong Ng, Shuo Wang, Yixiu Liang, Junbo Ge

A CLIP-like foundation model for ECGs learns from millions of ECG-report pairs and can help detect not only common rhythm problems but also rare and subtle cardiovascular diseases.

Read analysis
59Multimodal

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Xiang An, Yin Xie, Feilong Tang, Yunyao Yan, Huajie Tan, Didi Zhu, Changrui Chen, Xiuwei Zhao, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Kaichen Zhang, Wenkang Zhang, Zheng Cheng, Nansen Zhang, Chunsheng Wu, Chunjiang Ge, Zimin Ran, Dehua Song, Chunyuan Li, Shikun Feng, Ming Hu, Zhangquan Chen, Junbo Niu, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng

LLaVA-OneVision-2 is a next-generation multimodal model that compresses video more intelligently to better understand, ground, and reason over long visual content.

Read analysis
61Multimodal

Bernini: Latent Semantic Planning for Video Diffusion

Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan Yuan

Bernini uses a multimodal language model to plan video semantics and a diffusion model to turn that plan into realistic video, aiming to improve generation and editing with strong generalization.

Read analysis