CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
Binbin Huang, Haobin Duan, Yiqun Zhao, Zibo Zhao, Yi Ma, Shenghua Gao
TL;DR
Cupid addresses the challenge of reconstructing 3D objects from a single image with unknown viewpoint by jointly modeling the object and camera pose in a probabilistic framework. It introduces a two-stage cascaded flow model: first, occupancy and UV-pose generation provides a robust pose estimate via PnP; second, a pose-aligned refinement injects pixel-level cues to recover accurate geometry and appearance. The approach achieves state-of-the-art monocular geometry fidelity (outperforming existing methods by over 3 dB PSNR and reducing Chamfer Distance) and extends naturally to multi-view and scene-level reconstruction without post-hoc optimization. By unifying 3D generation priors with explicit pose reasoning, Cupid enables faithful reconstructions and scalable scene composition for single-view and beyond.
Abstract
We introduce Cupid, a generative 3D reconstruction framework that jointly models the full distribution over both canonical objects and camera poses. Our two-stage flow-based model first generates a coarse 3D structure and 2D-3D correspondences to estimate the camera pose robustly. Conditioned on this pose, a refinement stage injects pixel-aligned image features directly into the generative process, marrying the rich prior of a generative model with the geometric fidelity of reconstruction. This strategy achieves exceptional faithfulness, outperforming state-of-the-art reconstruction methods by over 3 dB PSNR and 10% in Chamfer Distance. As a unified generative model that decouples the object and camera pose, Cupid naturally extends to multi-view and scene-level reconstruction tasks without requiring post-hoc optimization or fine-tuning.
