Quantitative Bounds for Sorting-Based Permutation-Invariant Embeddings
Nadav Dym, Matthias Wellershoff, Efstratios Tsoukanis, Daniel Levy, Radu Balan
TL;DR
The paper studies sorting-based permutation-invariant embeddings $β_{\mathbf{A}}$ for $\mathbb{R}^{n\times d}$, seeking injectivity (orbit separation) and bi-Lipschitz control with minimal output dimension. It derives sharp embedding-dimension bounds, showing injectivity for $β_{\mathbf{A}}$ when $D \ge n(d-1)+1$ for full-spark $\mathbf{A}$ and a nearly optimal lower bound $D \gtrsim d\log n$; it also proves injectivity for projection-based variants with embedding dimension $(2n-1)d$. On distortion, the authors construct $\mathbf{A}$ with projective-uniformity achieving bi-Lipschitz distortions scaling as $O(n^2)$ (independent of $d$) at $D \asymp n^2 d$, and establish a universal lower bound $\Omega(\sqrt{n})$ on distortion. They further show that dimension-reduction mappings $β_{\mathbf{A},L}$ and $\delta_{\mathbf{A},\mathbf{B}}$ retain injectivity with near-optimal embedding dimensions and comparable distortions, and connect the results to Wasserstein-distance interpretations via sliced-Wasserstein. Numerical experiments illustrate gaps between theory and practice for small parameters and highlight practical potential for permutation-invariant learning in graph-structured data.
Abstract
We study the sorting-based embedding $β_{\mathbf A} : \mathbb R^{n \times d} \to \mathbb R^{n \times D}$, $\mathbf X \mapsto {\downarrow}(\mathbf X \mathbf A)$, where $\downarrow$ denotes column wise sorting of matrices. Such embeddings arise in graph deep learning where outputs should be invariant to permutations of graph nodes. Previous work showed that for large enough $D$ and appropriate $\mathbf A$, the mapping $β_{\mathbf A}$ is injective, and moreover satisfies a bi-Lipschitz condition. However, two gaps remain: firstly, the optimal size $D$ required for injectivity is not yet known, and secondly, no estimates of the bi-Lipschitz constants of the mapping are known. In this paper, we make substantial progress in addressing both of these gaps. Regarding the first gap, we improve upon the best known upper bounds for the embedding dimension $D$ necessary for injectivity, and also provide a lower bound on the minimal injectivity dimension. Regarding the second gap, we construct matrices $\mathbf A$, so that the bi-Lipschitz distortion of $β_{\mathbf A} $ depends quadratically on $n$, and is completely independent of $d$. We also show that the distortion of $β_{\mathbf A}$ is necessarily at least in $Ω(\sqrt{n})$. Finally, we provide similar results for variants of $β_{\mathbf A}$ obtained by applying linear projections to reduce the output dimension of $β_{\mathbf A}$.
