Joint Multi-Condition Representation Modelling via Matrix Factorisation for Visual Place Recognition
Timur Ismagilov, Shakaiba Majeed, Michael Milford, Tan Viet Tuyen Nguyen, Sarvapali D. Ramchurn, Shoaib Ehsan
TL;DR
This paper tackles multi-reference visual place recognition under appearance and viewpoint changes by proposing a training-free, descriptor-level mapping that jointly models multiple reference observations. Each place is represented as a subspace formed from its multi-condition descriptors via QR decomposition, enabling projection-based residual matching without retraining. Empirical results show recall gains up to about 18% on multi-appearance data, up to about 10% on multi-viewpoint data, and around 5% on unstructured references, while maintaining low computational overhead and descriptor-backbone agnosticism. The SotonMV benchmark further enables systematic evaluation of multi-view representations, highlighting the method's practical relevance for robust, edge-friendly localisation in robotics and autonomous systems.
Abstract
We address multi-reference visual place recognition (VPR), where reference sets captured under varying conditions are used to improve localisation performance. While deep learning with large-scale training improves robustness, increasing data diversity and model complexity incur extensive computational cost during training and deployment. Descriptor-level fusion via voting or aggregation avoids training, but often targets multi-sensor setups or relies on heuristics with limited gains under appearance and viewpoint change. We propose a training-free, descriptor-agnostic approach that jointly models places using multiple reference descriptors via matrix decomposition into basis representations, enabling projection-based residual matching. We also introduce SotonMV, a structured benchmark for multi-viewpoint VPR. On multi-appearance data, our method improves Recall@1 by up to ~18% over single-reference and outperforms multi-reference baselines across appearance and viewpoint changes, with gains of ~5% on unstructured data, demonstrating strong generalisation while remaining lightweight.
