FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss
Parthasaarathy Sudarsanam, Sebastian Braun, Hannes Gamper
TL;DR
The paper addresses the challenge of efficiently encoding spatial audio represented in First-Order Ambisonics (FOA) at ultra-low bitrate while preserving directional cues. It introduces FOA-VQGAN, the first discrete neural spatial audio codec for FOA, by extending the WavTokenizer framework to four-channel FOA signals and adding a spatial consistency loss based on intensity-vector cosine similarity. Key contributions include a 75-token-per-second, 0.9 kbps FOA representation with a 4096-code VQ codebook, a four-channel FOA encoder/decoder with iSTFT, and a GAN-based training regime that enforces spatial fidelity; the approach shows strong acoustic and spatial reconstruction across simulated and real-room datasets and supports downstream spatial tasks like SELD on STARSS23. The work enables efficient spatial audio transmission and provides a foundation for generative, token-based spatial audio models, with future work aiming to exploit interchannel FOA structure for further gains.
Abstract
Neural audio codecs have been widely studied for mono and stereo signals, but spatial audio remains largely unexplored. We present the first discrete neural spatial audio codec for first-order ambisonics (FOA). Building on the WavTokenizer architecture, we extend it to support four-channel FOA signals and introduce a novel spatial consistency loss to preserve directional cues in the reconstructed signals under a highly compressed representation. Our codec compresses 4-channel FOA audio at 24 kHz into 75 discrete tokens per second, corresponding to a bit rate of 0.9 kbps. Evaluations on simulated reverberant mixtures, non-reverberant clean speech, and FOA mixtures with real room impulse responses show accurate reconstruction, with mean angular errors of 13.76°, 3.96°, and 25.83°, respectively, across the three conditions. In addition, discrete latent representations derived from our codec provide useful features for downstream spatial audio tasks, as demonstrated on sound event localization and detection with STARSS23 real recordings.
