TiRex-2: Generalizing TiRex to Multivariate Data and Streaming
AuthorsPatrick Podest, Marco Pichler, Elias Bürger, Levente Zólyomi, Bernhard Voggenberger, Wilhelm Berghammer, Daniel Klotz, Sebastian Böck, Günter Klambauer, Sepp Hochreiter
Resources
TiRex-2 is a recurrent time-series foundation model designed to forecast many interacting variables continuously without repeatedly recomputing the entire history.
Key results
Active parameter count when TiRex-2 processes univariate inputs.
Parameters activated for multivariate forecasting.
Cumulative streamed steps over which forecast quality remained stable.
Optimizer steps used for the main pretraining phase.
Mean MASE of the full TiRex-2 model on fev-bench.
MASE increase after removing grouped attention.
What the paper found
TiRex-2, developed by Sepp Hochreiter’s team at JKU Linz and NXAI GmbH, extends the recurrent TiRex foundation model to multivariate forecasting with past and future-known covariates. Its xLSTM backbone alternates a bidirectional time mixer with asymmetric grouped attention across variates: future covariates can be processed in both temporal directions, while the mask prevents future target information from leaking backward, preserving strict target causality. This recurrent design has linear cost for full-context processing and constant cost per patch during streaming, unlike Transformer models such as Chronos-2, whose cached attention cost grows with context length. To train cross-variate behavior from abundant univariate data, the authors introduce on-the-fly synthetic coupling using functional, linear-mixing, cointegration, and causal mechanisms, plus missingness and discretization effects. TiRex-2 achieves state-of-the-art zero-shot results on GIFT-Eval and fev-bench, remains stable while streaming to 32M steps, and uses 38.4M active parameters in univariate mode with an additional 44.1M for multivariate forecasting. The model was pretrained for 700,000 optimizer steps on 2 NVIDIA H100 GPUs; on fev-bench, the full model reaches 1.542 MASE, while removing grouped attention worsens error by 0.220 MASE.
Original abstract
We introduce TiRex-2, a recurrent xLSTM-based time series foundation model that generalizes the univariate TiRex to multivariate forecasting with both past and future covariates. Real-world forecasting is inherently sequential: observations arrive continuously, variables evolve jointly, and a subset of covariates is known ahead of time. Existing Transformer-based time series foundation models capture cross-variate dependencies but incur quadratic complexity in context length and require full-history recomputation as new observations arrive. TiRex-2 addresses these limitations through a memory-centric recurrent design that operates at constant per-patch cost under streaming. The model combines a bidirectional time mixer with an asymmetric grouped-attention variate mixer, enabling the integration of future-known covariates while preserving strict causality over target variables. To our knowledge, this is the first time series foundation model that achieves this combination of properties. To support scalable multivariate pretraining, we propose a synthetic coupling pipeline that composes diverse multivariate samples on the fly from large univariate corpora. Empirically, TiRex-2 achieves state-of-the-art zero-shot performance on GIFT-Eval and fev-bench, remains stable when streamed to arbitrary context lengths, and maintains constant inference cost per patch. The model uses 38.4M active parameters in univariate mode, with an additional 44.1M parameters activated for multivariate forecasting.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.