SARL: Spatially-Aware Self-Supervised Representation Learning for Visuo-Tactile Perception

Gurmeher Khurana; Lan Wei; Dandan Zhang

SARL: Spatially-Aware Self-Supervised Representation Learning for Visuo-Tactile Perception

Gurmeher Khurana, Lan Wei, Dandan Zhang

TL;DR

SARL addresses the need for geometry-aware representations in contact-rich manipulation by leveraging fused visuo-tactile data and augmenting BYOL with three map-level losses that preserve spatial structure. SAL, PPDA, and RAM enforce attentional, semantic, and geometric consistency on intermediate feature maps, yielding richer representations than global invariance alone. Across six downstream tasks and nine SSL baselines, SARL demonstrates substantial gains, particularly on geometry-sensitive tasks, and shows strong transfer to unseen visuo-tactile datasets, underscoring the value of spatial equivariance for manipulation-ready perception.

Abstract

Contact-rich robotic manipulation requires representations that encode local geometry. Vision provides global context but lacks direct measurements of properties such as texture and hardness, whereas touch supplies these cues. Modern visuo-tactile sensors capture both modalities in a single fused image, yielding intrinsically aligned inputs that are well suited to manipulation tasks requiring visual and tactile information. Most self-supervised learning (SSL) frameworks, however, compress feature maps into a global vector, discarding spatial structure and misaligning with the needs of manipulation. To address this, we propose SARL, a spatially-aware SSL framework that augments the Bootstrap Your Own Latent (BYOL) architecture with three map-level objectives, including Saliency Alignment (SAL), Patch-Prototype Distribution Alignment (PPDA), and Region Affinity Matching (RAM), to keep attentional focus, part composition, and geometric relations consistent across views. These losses act on intermediate feature maps, complementing the global objective. SARL consistently outperforms nine SSL baselines across six downstream tasks with fused visual-tactile data. On the geometry-sensitive edge-pose regression task, SARL achieves a Mean Absolute Error (MAE) of 0.3955, a 30% relative improvement over the next-best SSL method (0.5682 MAE) and approaching the supervised upper bound. These findings indicate that, for fused visual-tactile data, the most effective signal is structured spatial equivariance, in which features vary predictably with object geometry, which enables more capable robotic perception.

SARL: Spatially-Aware Self-Supervised Representation Learning for Visuo-Tactile Perception

TL;DR

Abstract

SARL: Spatially-Aware Self-Supervised Representation Learning for Visuo-Tactile Perception

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (4)