Robust Estimation of Polyserial Correlation
Max Welz
TL;DR
This work addresses estimating the polyserial correlation $\rho$ when the partially-latent normality assumption is violated for a fraction $\varepsilon$ of observations (partial misspecification). It introduces a robust estimator based on minimum density power divergence with tuning parameter $\alpha$, downweighting observations poorly described by the model and yielding weights that diagnose misspecification. The estimator is Fisher-consistent, consistent, and asymptotically normal, with near-ML efficiency under correct specification and practical computation (default $\alpha=0.5$ achieves substantial robustness with only modest efficiency loss). Implemented in the free R package $\texttt{robcat}$ and demonstrated on a personality-psychology dataset, the method identifies outliers and provides more reliable inference for mixed data in SEM-like analyses.
Abstract
The association between a continuous and an ordinal variable is commonly modeled through the polyserial correlation model. However, this model, which is based on a partially-latent normality assumption, may be misspecified in practice, due to, for example (but not limited to), outliers or careless responses. We demonstrate that the typically used maximum likelihood (ML) estimator is highly susceptible to such misspecification: One single observation not generated by partially-latent normality can suffice to produce arbitrarily poor estimates. As a remedy, we propose a novel estimator of the polyserial correlation model designed to be robust against the adverse effects of observations discrepant to that model. The estimator achieves robustness by implicitly downweighting such observations; the ensuing weights constitute a useful tool for pinpointing potential sources of model misspecification. We show that the proposed estimator generalizes ML and is consistent as well as asymptotically Gaussian. As price for robustness, some efficiency must be sacrificed, but substantial robustness can be gained while maintaining more than 98% of ML efficiency. We demonstrate our estimator's robustness and practical usefulness in simulation experiments and an empirical application in personality psychology where our estimator helps identify outliers. Finally, the proposed methodology is implemented in free open-source software.
