Foundation Model Forecasts: Form and Function
Alvaro Perez-Diaz, James C. Loach, Danielle E. Toutoungi, Lee Middleton
TL;DR
The paper tackles the challenge that forecast accuracy alone does not guarantee practical utility for TSFMs. It introduces a formal four-type taxonomy—point, quantile, parametric per-step marginals, and trajectory ensembles—and proves that trajectory ensembles are strictly the most expressive, as they can be marginalized to yield other forms but not vice versa without extra structure. It then maps six canonical operational tasks to minimally sufficient forecast types, develops a convertibility framework with formal impossibility results for path-dependent questions from marginals, and proposes a task-aligned evaluation regimen that couples marginal and joint metrics. Through a survey of over 50 TSFMs, the work highlights a practical mismatch: most models yield point or parametric outputs, while many applications require joint horizon distributions, threshold-crossing analyses, and scenario generation. The findings guide both TSFM developers and practitioners toward forecast type selection and evaluation that align with real-world decision tasks, moving beyond accuracy as the sole performance criterion.
Abstract
Time-series foundation models (TSFMs) achieve strong forecast accuracy, yet accuracy alone does not determine practical value. The form of a forecast -- point, quantile, parametric, or trajectory ensemble -- fundamentally constrains which operational tasks it can support. We survey recent TSFMs and find that two-thirds produce only point or parametric forecasts, while many operational tasks require trajectory ensembles that preserve temporal dependence. We establish when forecast types can be converted and when they cannot: trajectory ensembles convert to simpler forms via marginalization without additional assumptions, but the reverse requires imposing temporal dependence through copulas or conformal methods. We prove that marginals cannot determine path-dependent event probabilities -- infinitely many joint distributions share identical marginals but yield different answers to operational questions. We map six fundamental forecasting tasks to minimal sufficient forecast types and provide a task-aligned evaluation framework. Our analysis clarifies when forecast type, not accuracy, differentiates practical utility.
