FootFormer: Estimating Stability from Visual Input
Keaton Kraiger, Jingjing Li, Skanda Bharadwaj, Jesse Scott, Robert T. Collins, Yanxi Liu
TL;DR
FootFormer introduces a cross-modality architecture that directly predicts foot pressure distributions, foot contact maps, and 3D center of mass from video-based pose sequences. By employing a GCN-based pose encoder, a spatiotemporal transformer, and cross-attentive decoders, it jointly optimizes three outputs to enable stability analysis. Across PSU-TMM100, MMVP, UnderPressure, and Ordinary Movements, FootFormer achieves state-of-the-art or comparable performance on pressure, contact, and stability metrics, with strong generalization to unseen everyday movements. The work highlights the feasibility of vision-driven stability estimation and provides a unified model that surpasses previous multi-model baselines while offering a public codebase.
Abstract
We propose FootFormer, a cross-modality approach for jointly predicting human motion dynamics directly from visual input. On multiple datasets, FootFormer achieves statistically significantly better or equivalent estimates of foot pressure distributions, foot contact maps, and center of mass (CoM), as compared with existing methods that generate one or two of those measures. Furthermore, FootFormer achieves SOTA performance in estimating stability-predictive components (CoP, CoM, BoS) used in classic kinesiology metrics. Code and data are available at https://github.com/keatonkraiger/Vision-to-Stability.git.
