Table of Contents
Fetching ...

A Multi-Layer Machine Learning and Econometric Pipeline for Forecasting Market Risk: Evidence from Cryptoasset Liquidity Spillovers

Yimeng Qiu, Feihuang Fang

TL;DR

This work addresses whether core cryptoasset liquidity and volatility signals generate spillovers that forecast market-wide risk. It combines a three-layer evidence chain (Layer A: LV/returns; Layer B: LV/returns PCs; Layer C: volatility PCs predicting cross-sectional crowding) with a leakage-safe ML pipeline (XGBoost with SHAP) and interpretable VAR/HAR-X analyses on a 74-asset, 1462-observation daily panel from 2021–2025. Key contributions include a cross-sectional volatility crowding target, a leakage-aware forecasting protocol, and reproducible artifacts, with empirical findings showing significant Layer-A and Layer-C links and moderate out-of-sample predictive performance (e.g., $R^2=0.53$, ROC-AUC $=0.74$, PR-AUC $\approx 0.47$). The approach offers an interpretable early-warning tool for market risk in cryptoassets, balancing predictive performance with transparency through SHAP and VAR-based diagnostics.

Abstract

We study whether liquidity and volatility proxies of a core set of cryptoassets generate spillovers that forecast market-wide risk. Our empirical framework integrates three statistical layers: (A) interactions between core liquidity and returns, (B) principal-component relations linking liquidity and returns, and (C) volatility-factor projections that capture cross-sectional volatility crowding. The analysis is complemented by vector autoregression impulse responses and forecast error variance decompositions (see Granger 1969; Sims 1980), heterogeneous autoregressive models with exogenous regressors (HAR-X, Corsi 2009), and a leakage-safe machine learning protocol using temporal splits, early stopping, validation-only thresholding, and SHAP-based interpretation. Using daily data from 2021 to 2025 (1462 observations across 74 assets), we document statistically significant Granger-causal relationships across layers and moderate out-of-sample predictive accuracy. We report the most informative figures, including the pipeline overview, Layer A heatmap, Layer C robustness analysis, vector autoregression variance decompositions, and the test-set precision-recall curve. Full data and figure outputs are provided in the artifact repository.

A Multi-Layer Machine Learning and Econometric Pipeline for Forecasting Market Risk: Evidence from Cryptoasset Liquidity Spillovers

TL;DR

This work addresses whether core cryptoasset liquidity and volatility signals generate spillovers that forecast market-wide risk. It combines a three-layer evidence chain (Layer A: LV/returns; Layer B: LV/returns PCs; Layer C: volatility PCs predicting cross-sectional crowding) with a leakage-safe ML pipeline (XGBoost with SHAP) and interpretable VAR/HAR-X analyses on a 74-asset, 1462-observation daily panel from 2021–2025. Key contributions include a cross-sectional volatility crowding target, a leakage-aware forecasting protocol, and reproducible artifacts, with empirical findings showing significant Layer-A and Layer-C links and moderate out-of-sample predictive performance (e.g., , ROC-AUC , PR-AUC ). The approach offers an interpretable early-warning tool for market risk in cryptoassets, balancing predictive performance with transparency through SHAP and VAR-based diagnostics.

Abstract

We study whether liquidity and volatility proxies of a core set of cryptoassets generate spillovers that forecast market-wide risk. Our empirical framework integrates three statistical layers: (A) interactions between core liquidity and returns, (B) principal-component relations linking liquidity and returns, and (C) volatility-factor projections that capture cross-sectional volatility crowding. The analysis is complemented by vector autoregression impulse responses and forecast error variance decompositions (see Granger 1969; Sims 1980), heterogeneous autoregressive models with exogenous regressors (HAR-X, Corsi 2009), and a leakage-safe machine learning protocol using temporal splits, early stopping, validation-only thresholding, and SHAP-based interpretation. Using daily data from 2021 to 2025 (1462 observations across 74 assets), we document statistically significant Granger-causal relationships across layers and moderate out-of-sample predictive accuracy. We report the most informative figures, including the pipeline overview, Layer A heatmap, Layer C robustness analysis, vector autoregression variance decompositions, and the test-set precision-recall curve. Full data and figure outputs are provided in the artifact repository.
Paper Structure (11 sections, 5 equations, 6 figures, 2 tables)

This paper contains 11 sections, 5 equations, 6 figures, 2 tables.

Figures (6)

  • Figure 1: End-to-end pipeline: factor construction $\rightarrow$ layered causality (A/B/C) and interpretable VAR/HAR-X $\rightarrow$ leakage-safe ML evaluation.
  • Figure 2: Layer A heatmap ($-\log_{10} p$). File: output/run_20250808_221301/heatmap_layerA.png .
  • Figure 3: Layer C robustness across RS windows and fixed lags ($-\log_{10}p$). File: output/robust_output/compare_20251001_013842/robust_summary_layerC_heatmap.png .
  • Figure 4: VAR (main) FEVD. File: output/run_20250808_221301/var_main_fevd.png .
  • Figure 5: Test ROC curve (fixed threshold chosen on validation). File: output/run_20250808_221301/cls_test_roc_curve.png .
  • ...and 1 more figures