Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen, Huan Liu
5 min
When applying time-series foundation models (TSFMs) to hourly equity returns, researchers observe an unexpected phenomenon termed "forecast collapse": predictions become nearly flat and exhibit poor cross-sectional ranking, as measured by the information coefficient (IC). Surprisingly, this collapse largely disappears when forecasting trading volume under identical experimental protocols. This paper investigates forecast collapse across TSFMs, twelve deep-learning models, and 97 public benchmark configurations, linking the issue directly to target predictability. The study exposes a critical blind spot in conventional time-series evaluation: per-series metrics can completely hide failures in cross-series structure that are vital for downstream decision-making.
The authors identify two distinct theoretical reasons behind forecast collapse. First, low target predictability inherently limits the amplitude of calibrated point forecasts. When a target has a very low signal-to-noise ratio, calibrated models shrink their predictions toward zero relative to the realized target. Second, standard per-series objectives leave cross-series structure completely unidentified. Because pointwise risks only evaluate individual series against their own targets, forecasts with identical per-series risk can yield drastically different cross-sectional correlations.
These mechanisms establish a strict calibration-ranking tradeoff. Optimizing squared error (MSE) leads to flat predictions with controlled scale but poor ranking. Conversely, directly optimizing cross-sectional correlation improves ranking but can inflate forecast amplitude by more than an order of magnitude. Rescaling a forecast can alter its magnitude, but it cannot recover missing ranking information.
To overcome the calibration-ranking tradeoff, the authors introduce CalibRank, a simple composite objective that jointly optimizes MSE and cross-sectional correlation using a balancing parameter. Evaluated on Finance1K—a new public panel of 1,000 US equities with aligned return and volume targets—CalibRank nearly triples cross-sectional correlation while keeping forecast amplitude close to the true target scale. Furthermore, it consistently improves ranking performance across all tested baseline models.
When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, while per-series objectives leave cross-series structure unidentified. These findings reveal a calibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizing cross-sectional correlation improves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduce CalibRank, a simple objective that balances calibration and ranking. On Finance1K, CalibRank nearly triples cross-sectional correlation while keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.
Sam: So a model could score well on individual errors while completely failing at the thing investors actually care about — which stock will do better than another.
Alex: Precisely. That gap between individual accuracy and cross-sectional ranking is the hidden flaw. You can have a model that looks decent on standard benchmarks and is still useless for building a portfolio.
Sam: So how do the authors propose fixing this?
Alex: They introduce a composite training objective. Think of it like a report card with two grades instead of one. The first grade is the traditional one — how close were your individual predictions to the actual values? The second grade is new — did you correctly rank the stocks relative to each other at each point in time? The model has to do well on both to pass.
Sam: And does adding that ranking grade actually help?
Alex: Across twelve different model architectures, adding the ranking component meaningfully improved the models' ability to correctly order stocks. The individual prediction errors did increase slightly — there's a genuine tradeoff — but the improvement in ranking was substantial, and the error stayed far lower than if you'd optimized for ranking alone without any accuracy anchor.
Sam: So you're finding a middle ground between two competing goals.
Alex: That's a good way to put it. There's a frontier of possible operating points, and practitioners can choose where on that frontier they want to sit depending on what matters more for their specific use case.
Sam: Are there limitations to keep in mind here?
Alex: Several important ones. The empirical results come primarily from one specific financial dataset, so it's not yet clear how broadly these findings generalize to other markets or other types of time-series data. The ranking measure they use — Pearson correlation — captures linear relationships between stocks, but it doesn't capture more complex dependencies, like what happens during a market crash when everything moves together in unusual ways. And the authors are candid that future work will need to address those richer joint behaviors.
Sam: So this is a meaningful step forward, but with real boundaries on what it claims.
Alex: That's a fair summary. The paper identifies a structural reason why sophisticated models fail on low-signal data, and it offers a principled fix that works across a range of architectures. But the authors are careful not to overstate it. The field will likely need to extend these ideas to handle tail risks and more complex market dynamics before this becomes a complete solution.
Sam: It's a good reminder that even the most advanced models are only as useful as the objectives they're trained on.
Alex: Well put. The model learns exactly what you ask it to learn — and if you ask the wrong question, you get the wrong answer, no matter how powerful the system is. Thanks for listening to ResearchPod.