Doctor Rashomon and the UNIVERSE of Madness: Variable Importance with Unobserved Confounding and the Rashomon Effect
Jon Donnelly, Srikar Katta, Emanuele Borgonovo, Cynthia Rudin
TL;DR
UNIVERSE introduces a theoretically grounded framework to bound VI in the presence of unobserved confounding and the Rashomon effect by leveraging Rashomon sets extended to unobserved features and finite-sample corrections. The approach yields high-probability bounds on the VI of the true conditional mean function $g^*$ and, via a VI-drift parameter, accounts for distributional shifts induced by unobserved variables. It provides finite-sample guarantees and demonstrates tight, useful bounds on VI through semi-synthetic experiments across multiple datasets and a FICO credit-risk case study. The results show that accounting for model uncertainty, VI estimation uncertainty, and drift is essential to achieve reliable bounds, with practical implications for model interpretation and data collection decisions. The framework is general across model classes and VI metrics, with future work aimed at scaling to complex models and broader problem domains.
Abstract
Variable importance (VI) methods are often used for hypothesis generation, feature selection, and scientific validation. In the standard VI pipeline, an analyst estimates VI for a single predictive model with only the observed features. However, the importance of a feature depends heavily on which other variables are included in the model, and essential variables are often omitted from observational datasets. Moreover, the VI estimated for one model is often not the same as the VI estimated for another equally-good model - a phenomenon known as the Rashomon Effect. We address these gaps by introducing UNobservables and Inference for Variable importancE using Rashomon SEts (UNIVERSE). Our approach adapts Rashomon sets - the sets of near-optimal models in a dataset - to produce bounds on the true VI even with missing features. We theoretically guarantee the robustness of our approach, show strong performance on semi-synthetic simulations, and demonstrate its utility in a credit risk task.
