A Multi-Layer Machine Learning and Econometric Pipeline for Forecasting Market Risk: Evidence from Cryptoasset Liquidity Spillovers
Yimeng Qiu, Feihuang Fang
TL;DR
This work addresses whether core cryptoasset liquidity and volatility signals generate spillovers that forecast market-wide risk. It combines a three-layer evidence chain (Layer A: LV/returns; Layer B: LV/returns PCs; Layer C: volatility PCs predicting cross-sectional crowding) with a leakage-safe ML pipeline (XGBoost with SHAP) and interpretable VAR/HAR-X analyses on a 74-asset, 1462-observation daily panel from 2021–2025. Key contributions include a cross-sectional volatility crowding target, a leakage-aware forecasting protocol, and reproducible artifacts, with empirical findings showing significant Layer-A and Layer-C links and moderate out-of-sample predictive performance (e.g., $R^2=0.53$, ROC-AUC $=0.74$, PR-AUC $\approx 0.47$). The approach offers an interpretable early-warning tool for market risk in cryptoassets, balancing predictive performance with transparency through SHAP and VAR-based diagnostics.
Abstract
We study whether liquidity and volatility proxies of a core set of cryptoassets generate spillovers that forecast market-wide risk. Our empirical framework integrates three statistical layers: (A) interactions between core liquidity and returns, (B) principal-component relations linking liquidity and returns, and (C) volatility-factor projections that capture cross-sectional volatility crowding. The analysis is complemented by vector autoregression impulse responses and forecast error variance decompositions (see Granger 1969; Sims 1980), heterogeneous autoregressive models with exogenous regressors (HAR-X, Corsi 2009), and a leakage-safe machine learning protocol using temporal splits, early stopping, validation-only thresholding, and SHAP-based interpretation. Using daily data from 2021 to 2025 (1462 observations across 74 assets), we document statistically significant Granger-causal relationships across layers and moderate out-of-sample predictive accuracy. We report the most informative figures, including the pipeline overview, Layer A heatmap, Layer C robustness analysis, vector autoregression variance decompositions, and the test-set precision-recall curve. Full data and figure outputs are provided in the artifact repository.
