$t_0$: A Time-Series Foundation Model for Forecasting with Context
AuthorsLucas Meyer, Claudio Sole, Huikan Xiang, Nicolas Li, Lucas Franceschino, Arnau Quera-Bofarull, Maarten P. Scholl, Joachim Fainberg, Geoffrey Négiar
Resources
t0 is an open-weight time-series foundation model that uses historical and future context to deliver strong zero-shot probabilistic forecasts across diverse datasets.
Key results
Size of the first released model.
Size of the successor model.
t0-beta aggregate probabilistic forecasting score.
t0-beta skill score across the benchmark.
Percentage-point skill improvement for t0-alpha across 30 tasks.
Reduction versus the lagged-price baseline for both t0 models.
What the paper found
The paper introduces t0, an open-weights time-series foundation model family designed to forecast targets using both historical data and multivariate context, including past and known-future covariates. Its architecture is a decoder-style patch Transformer that alternates causal attention across time with permutation-equivariant attention across variates, while a monotone quantile head produces probabilistic forecasts at five native levels. Pretraining combines the 89-dataset GIFT-Eval corpus with synthetic kernel, causal-graph, covariate-effect, and mixing generators, exposing the model to explicit covariate-to-target dependencies. The released t0-alpha has 102M parameters, while t0-beta has 256M. On GIFT-Eval, t0-beta reaches 0.4738 CRPS and 0.6865 MASE, ranking third among evaluated zero-shot models and competing with TimesFM-3.0, Chronos-2, and Toto-2.0. On fev-bench, t0-beta achieves 46.7 skill, again ranking third. For t0-alpha, supplying known-future covariates raises skill by 6.3 percentage points across 30 tasks, while past covariates add 2.7 points. The model supports single-pass forecasts up to 1024 steps and uses quantile-path rollout for longer horizons. In an independent evaluation of hourly ERCOT electricity prices over 29 months, both t0 models reduce mean absolute error versus a lagged-price baseline by 38%, although their nominal 80% intervals cover only 74–75% of realized prices.
Original abstract
We present $t_0$, a family of open-weights foundation models for forecasting with multivariate context. We release its first two members: $\texttt{t0-alpha}$ and $\texttt{t0-beta}$, respectively 102M and 256M parameters. Both condition their forecasts on target history, past covariates, and known-future covariates, without task-specific retraining. Their transformer layers alternate attention along time and across variates. They produce probabilistic forecasts through quantile predictions. Pretraining combines curated public data with synthetic generator families constructed to contain covariate-to-target dependencies. On GIFT-Eval, $\texttt{t0-alpha}$ reaches an aggregate CRPS of 0.4941, and $\texttt{t0-beta}$ a CRPS of 0.4738 and a MASE of 0.6865, third on both and within 4.0% of the best zero-shot TSFM. On fev-bench they score 42.2 and 46.7 in skill, the latter third again and 2.0 points behind the leader. We analyze $\texttt{t0-alpha}$ in depth. Known-future covariates raise its skill by 6.3 percentage points across 30 tasks. The report also examines its calibration, its rollout strategy on long horizons, and its robustness to missing data. On the Victoria electricity-demand benchmark, $\texttt{t0-beta}$ is among the most accurate models with a context of nearly a year. In an independent Macrocosm evaluation of hourly ERCOT prices over 29 months, both cut the MAE of the lagged-price baseline by 38%.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.