NTH

Forecast Collapse in Time-Series Foundation Models

AuthorsShu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen, Huan Liu

August 22, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that time-series models can make deceptively flat forecasts and introduces a method that improves stock ranking without sacrificing calibration.

Key results

1000
Finance1K equity count

Hourly US equity panel used for matched return and volume forecasting.

0.027
MSE return amplitude

Raw forecast amplitude relative to realized return amplitude under MSE training.

0.046
MSE return IC

Cross-sectional information coefficient for MSE-trained return forecasts.

25.260
IC-only return amplitude

Raw amplitude produced when optimizing cross-sectional IC without calibration control.

0.126
CalibRank return IC

Cross-sectional IC achieved by CalibRank on the controlled return-forecasting experiment.

1.836
CalibRank return amplitude

Raw amplitude relative to the target under CalibRank.

What the paper found

“Forecast Collapse in Time-Series Foundation Models” examines why models such as TimesFM and Chronos can produce nearly flat forecasts for difficult financial targets. On Finance1K, an hourly panel covering 1,000 US equities and 28,510 time steps, mean-squared-error training predicts stock returns at only 0.027 of target amplitude, with cross-sectional information coefficient, or IC, of 0.046. The paper identifies two mechanisms: low predictability shrinks the amplitude of calibrated point forecasts, while per-series objectives fail to identify the cross-series coupling required for ranking assets. Optimizing IC directly raises return-ranking performance to 0.124 but produces forecasts at 25.260 times the target amplitude, revealing a severe calibration-ranking tradeoff. The proposed CalibRank objective combines MSE with differentiable per-timestamp Pearson correlation; on the controlled backbone it reaches IC 0.126 and raw amplitude 1.836, nearly tripling IC while remaining close to the target scale. Across 12 forecasting architectures, CalibRank improves IC for every model, showing that the failure is not specific to a Transformer design. A matched volume experiment provides a negative control: because volume is more predictable, MSE forecasts retain substantial scale and ranking quality. Across 97 GIFT-Eval configurations, raw amplitude also tracks achieved predictability, with correlations of 0.88 for Chronos and 0.87 for TimesFM. The central implication is that time-series benchmarks must evaluate per-series accuracy together with amplitude and the cross-sectional structure used by downstream decisions.

Original abstract

When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, while per-series objectives leave cross-series structure unidentified. These findings reveal a calibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizing cross-sectional correlation improves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduce CalibRank, a simple objective that balances calibration and ranking. On Finance1K, CalibRank nearly triples cross-sectional correlation while keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis