NTH

Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations

AuthorsKevin Y. Li, Asher Trockman, Ananda Theertha Suresh, Ziteng Sun

May 31, 2026 3 min read
Watch on YouTube
The one-line take

Oryx is a hybrid language model that dynamically alternates between attention and recurrent mixing within a sequence, aiming to get both strong long-context performance and efficient generation.

Key results

90%
Shared parameters

Oryx ties key and value projections across mixers, with more than 90% of parameters shared across mixer modes.

100B
Pretraining tokens

All reported model scales were pretrained on 100B FineWeb-Edu tokens.

1.4B
Model scale

The largest Oryx models evaluated in the paper are at the 1.4B scale.

0.7
LM improvement

At the 1.4B scale, Oryx outperforms corresponding single-mixer baselines by at least 0.7 percentage points on averaged language modeling tasks.

8.6
Real-world retrieval gain

With mixed inference, Oryx-TG surpasses its linear baseline by at least 8.6 percentage points on real-world retrieval tasks.

38.6
Needle-in-a-haystack gain

With mixed inference, Oryx-TM surpasses its linear baseline by at least 38.6 percentage points on needle-in-a-haystack tests.

What the paper found

Google Research and Carnegie Mellon University introduce Oryx, a sequence-axis hybrid language model that can switch at inference time between quadratic softmax attention and linear recurrent mixers such as Mamba-2 and Gated DeltaNet, while sharing more than 90% of parameters through tied key and value projections. The central idea is that attention and recurrent layers can update compatible internal states from a shared query-key-value representation, so the model can process some chunks with attention for retrieval-heavy computation and others with linear recurrence for cheaper generation. Trained on 100B FineWeb-Edu tokens at 130M, 380M, 810M, and 1.4B scales with chunked mixed-mode training, Oryx matches or beats single-mixer baselines under fixed token budgets; at 1.4B, its averaged downstream language modeling scores improve by at least 0.7 percentage points over the corresponding baselines. On retrieval, Oryx-TM and Oryx-TG reach Transformer-level performance even when less than 10% of tokens are processed in attention mode, and mixed inference boosts linear baselines by at least 8.6 points on real-world retrieval and 38.6 points on needle-in-a-haystack tasks. The paper’s key empirical result is that chunk-level mixed-mode training makes mode switching robust across non-chunk boundaries and multiple switches, supporting the claim that attention and linear recurrent mechanisms can share representations rather than require separate architectures.

Original abstract

Softmax attention is the cornerstone of modern large language models, but its memory scales linearly and compute quadratically with sequence length. Linear recurrent models, such as linear attention and state space models, have become widely studied as alternatives to attention due to their linear compute and constant memory. While these sub-quadratic token mixing methods, or mixers, achieve promising efficiency gains and competitive results on a wide range of benchmarks, current linear recurrent models still lag behind on tasks that require long-context retrieval or in-context learning. A growing body of work studies hybrid architectures that attempt to mitigate these trade-offs by statically interleaving or merging attention and recurrent blocks. In this work, we explore a new axis of developing hybrid models: across the token sequence. We propose Oryx, a hybrid model that can, throughout a sequence, flexibly switch between different mixers, for example quadratic attention for rich context utilization and linear recurrences for efficient generation. Oryx ties at least 90% of its parameters across mixers, enabling attention and recurrent modes to operate over shared internal representations. We validate our design with Mamba-2 and Gated DeltaNet variants, up to 1.4B models. Under fixed token budgets and a mixed-training strategy, Oryx achieves comparable or better performance than its single-mixer baselines. At the 1.4B scale, all instances of Oryx outperform their respective baselines by at least 0.7 percentage points on averaged language modeling tasks. On retrieval tasks, Oryx achieves performance comparable to the Transformer baseline even when processing only a tiny fraction (<10%) of the tokens in attention mode. These results suggest that attention and linear recurrent models can share internal representations, and motivate sequence-axis hybridization as a promising direction.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis