Gemma 4 Technical Report
AuthorsGemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, Mayank Chaturvedi, Victor Cotruta, Alice Coucke, Phil Culliton, Robert Dadashi, Lucas Dixon, Mohamed Elhawaty, Utku Evci, Clément Farabet, Johan Ferret, Filippo Galgani, Sertan Girgin, Jean-Bastien Grill, Maarten Grootendorst, Jiaxian Guo, Cassidy Hardin, Yanzhang He, Steven M. Hernandez, Omri Homburger, Léonard Hussenot, Juyeong Ji, Armand Joulin, Aishwarya Kamath, Parnian Kassraie, Olivier Lacombe, Preethi Lahoti, Gaël Liu, Gus Martins, Luciano Martins, Tatiana Matejovicova, Ramona Merhej, Nikola Momchev, Sneha Mondal, Ryan Mullins, Sindhu Raghuram Panyam, Shreya Pathak, Sarah Perrin, André Susano Pinto, Etienne Pot, Angéline Pouget, Alexandre Ramé, Sabela Ramos, Douglas Reid, David Rim, Morgane Rivière, Karsten Roth, Louis Rouillard, Omar Sanseviero, Pier Giuseppe Sessa, Shane Settle, Danila Sinopalnikov, Sara Smoot, Pi...
Resources
Gemma 4 is Google’s new open-weight multimodal model family, pushing stronger reasoning, faster inference, and better long-context performance across text, image, and audio tasks.
Key results
Gemma 4 dense model family size upper bound; the suite spans 2.3B to 31B parameters
Gemma 4 26B-A4B MoE uses 3.8B activated parameters
Global KV cache footprint reduced through cache reuse and attention design
Gemma 4 31B score on MMLU Pro
Gemma 4 31B score on AIME 2026 without tools
Gemma 4 31B Elo on Chatbot Arena Text
What the paper found
Google DeepMind’s Gemma Team presents Gemma 4, an open-weight, natively multimodal model family spanning dense and Mixture-of-Experts variants from 2.3B to 31B parameters, plus a 26B MoE with 3.8B active parameters. The report’s main novelty is a unified multimodal design that combines stronger vision and audio encoders with a 12B encoder-free model that ingests raw 40ms audio chunks and image patches directly into the language backbone. Gemma 4 also adds a thinking mode for explicit reasoning traces, speculative decoding via an autoregressive multi-token prediction drafter, quantization-aware training for low-memory deployment, and long-context optimizations including a 5:1 local-to-global attention pattern and KV-cache reuse that cuts global cache footprint by up to 37.5%. On evaluation, Gemma 4 31B reaches 85.2 on MMLU Pro, 89.2 on AIME 2026 without tools, 80.0 on LiveCodeBench v6, and 2150 Elo on Codeforces, while also leading the dense open category on Chatbot Arena with 1451 Elo. The 12B model shows competitive multimodal gains without separate encoders, and the audio stack achieves a 78% reduction in encoder footprint from 390 MB to 87 MB, alongside relative improvements of 12% and 10% on CoVoST translation and 17% and 12% on FLEURS ASR for E2B and E4B. Across long-context tasks, Gemma 4 scales to 128k and even ~256k token evaluations, with RULER accuracy of 96.4 at 128k for 31B and 89.8 for 26B-A4B, demonstrating that Google DeepMind has pushed an efficient open multimodal system into frontier-adjacent territory.
Original abstract
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.