A Geometric Approach to Steerable Convolutions
Soumyabrata Kundu, Risi Kondor
TL;DR
This work addresses extending CNNs to full rotation and translation equivariance by proposing a geometric derivation of steerable convolutions in arbitrary $d$-dimensional space, grounded in the action of the special Euclidean group $SE(d)$ and Fourier-space representations. It develops a unified framework for equivariant linear maps, derives first- and higher-layer steerable convolution equations, and introduces a practical interpolation-based implementation that reduces parameters and enhances robustness to noise. The approach naturally explains the emergence of spherical harmonics and Clebsch–Gordan decompositions as intrinsic to the symmetry structure, and provides theoretical bounds on equivariance loss due to interpolation and discretization. Empirically, interpolation-based steerable filters outperform Cartesian-based methods across multiple 2D and 3D datasets and exhibit stronger resilience to perturbations, suggesting substantial practical impact for rotation-aware vision tasks and medical imaging applications.
Abstract
In contrast to the somewhat abstract, group theoretical approach adopted by many papers, our work provides a new and more intuitive derivation of steerable convolutional neural networks in $d$ dimensions. This derivation is based on geometric arguments and fundamental principles of pattern matching. We offer an intuitive explanation for the appearance of the Clebsch--Gordan decomposition and spherical harmonic basis functions. Furthermore, we suggest a novel way to construct steerable convolution layers using interpolation kernels that improve upon existing implementation, and offer greater robustness to noisy data.
