NTH

TiRex-2: Generalizing TiRex to Multivariate Data and Streaming

AuthorsPatrick Podest, Marco Pichler, Elias Bürger, Levente Zólyomi, Bernhard Voggenberger, Wilhelm Berghammer, Daniel Klotz, Sebastian Böck, Günter Klambauer, Sepp Hochreiter

July 31, 2026 2 min read
Watch on YouTube
The one-line take

TiRex-2 is a recurrent time-series foundation model designed to forecast many interacting variables continuously without repeatedly recomputing the entire history.

Key results

38.4M
Univariate active parameters

Active parameter count when TiRex-2 processes univariate inputs.

44.1M
Additional multivariate parameters

Parameters activated for multivariate forecasting.

32M
Streaming evaluation length

Cumulative streamed steps over which forecast quality remained stable.

700,000
Pretraining steps

Optimizer steps used for the main pretraining phase.

1.542
Full-model fev-bench MASE

Mean MASE of the full TiRex-2 model on fev-bench.

0.220
Grouped-attention ablation delta

MASE increase after removing grouped attention.

What the paper found

TiRex-2, developed by Sepp Hochreiter’s team at JKU Linz and NXAI GmbH, extends the recurrent TiRex foundation model to multivariate forecasting with past and future-known covariates. Its xLSTM backbone alternates a bidirectional time mixer with asymmetric grouped attention across variates: future covariates can be processed in both temporal directions, while the mask prevents future target information from leaking backward, preserving strict target causality. This recurrent design has linear cost for full-context processing and constant cost per patch during streaming, unlike Transformer models such as Chronos-2, whose cached attention cost grows with context length. To train cross-variate behavior from abundant univariate data, the authors introduce on-the-fly synthetic coupling using functional, linear-mixing, cointegration, and causal mechanisms, plus missingness and discretization effects. TiRex-2 achieves state-of-the-art zero-shot results on GIFT-Eval and fev-bench, remains stable while streaming to 32M steps, and uses 38.4M active parameters in univariate mode with an additional 44.1M for multivariate forecasting. The model was pretrained for 700,000 optimizer steps on 2 NVIDIA H100 GPUs; on fev-bench, the full model reaches 1.542 MASE, while removing grouped attention worsens error by 0.220 MASE.

Original abstract

We introduce TiRex-2, a recurrent xLSTM-based time series foundation model that generalizes the univariate TiRex to multivariate forecasting with both past and future covariates. Real-world forecasting is inherently sequential: observations arrive continuously, variables evolve jointly, and a subset of covariates is known ahead of time. Existing Transformer-based time series foundation models capture cross-variate dependencies but incur quadratic complexity in context length and require full-history recomputation as new observations arrive. TiRex-2 addresses these limitations through a memory-centric recurrent design that operates at constant per-patch cost under streaming. The model combines a bidirectional time mixer with an asymmetric grouped-attention variate mixer, enabling the integration of future-known covariates while preserving strict causality over target variables. To our knowledge, this is the first time series foundation model that achieves this combination of properties. To support scalable multivariate pretraining, we propose a synthetic coupling pipeline that composes diverse multivariate samples on the fly from large univariate corpora. Empirically, TiRex-2 achieves state-of-the-art zero-shot performance on GIFT-Eval and fev-bench, remains stable when streamed to arbitrary context lengths, and maintains constant inference cost per patch. The model uses 38.4M active parameters in univariate mode, with an additional 44.1M parameters activated for multivariate forecasting.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis