There is No "apple" in Timeseries: Rethinking TSFM through the Lens of Invariance
Arian Prabowo, Flora D. Salim
TL;DR
The paper argues that current timeseries foundation models lag behind NLP/CV baselines because timeseries data are not semantically complete on web-scale corpora. It critiques the 'scrape everything' paradigm as effective for text/images but insufficient for time series, since many concepts and dynamics are not captured in TS data. It introduces a Timeseries Invariance Ontology, formalizing invariances as transformations with $T(g \cdot x) = T(x)$ and enumerating classes such as spectral, amplitude, shape, elastic, distributional, and parametric invariances to guide data collection and model design. The authors advocate deliberate dataset curation and synthetic generation to achieve a world-complete representation of temporal dynamics, enabling robust generalisation, reasoning, and emergent behaviour in TSFMs.
Abstract
Timeseries foundation models (TSFMs) have multiplied, yet lightweight supervised baselines and even classical models often match them. We argue this gap stems from the naive importation of NLP or CV pipelines. In language and vision, large web-scale corpora densely capture human concepts i.e. there are countless images and text of apples. In contrast, timeseries data is built to complement the image and text modalities. There are no timeseries dataset that contains the concept apple. As a result, the scrape-everything-online paradigm fails for TS. We posit that progress demands a shift from opportunistic aggregation to principled design: constructing datasets that systematically span the space of invariance that preserve temporal semantics. To this end, we suggest that the ontology of timeseries invariances should be built based on first principles. Only by ensuring representational completeness through invariance coverage can TSFMs achieve the aligned structure necessary for generalisation, reasoning, and truly emergent behaviour.
