Table of Contents
Fetching ...

Resounding Acoustic Fields with Reciprocity

Zitong Lan, Yiduo Hao, Mingmin Zhao

TL;DR

This work tackles resounding, the estimation of room impulse responses $h(t)$ at arbitrary emitter positions from sparse measurements, to enable dynamic, realistic spatial audio in AR/VR. It introduces Versa, a reciprocity-inspired framework that leverages emitter/listener pose exchanges (Versa-ELE) and a self-supervised scheme (Versa-SSL) to enforce reciprocity under realistic gain-pattern asymmetries. Across simulated and real datasets, Versa significantly improves impulse-response accuracy and perceptual realism, with marked gains when emitter patterns differ, and a perceptual study confirming enhanced spatial audio comfort. The approach highlights a physics-grounded learning paradigm that integrates fundamental wave reciprocity into training, offering robust generalization under sparse supervision and potential applicability beyond acoustics to other wave phenomena.

Abstract

Achieving immersive auditory experiences in virtual environments requires flexible sound modeling that supports dynamic source positions. In this paper, we introduce a task called resounding, which aims to estimate room impulse responses at arbitrary emitter location from a sparse set of measured emitter positions, analogous to the relighting problem in vision. We leverage the reciprocity property and introduce Versa, a physics-inspired approach to facilitating acoustic field learning. Our method creates physically valid samples with dense virtual emitter positions by exchanging emitter and listener poses. We also identify challenges in deploying reciprocity due to emitter/listener gain patterns and propose a self-supervised learning approach to address them. Results show that Versa substantially improve the performance of acoustic field learning on both simulated and real-world datasets across different metrics. Perceptual user studies show that Versa can greatly improve the immersive spatial sound experience. Code, dataset and demo videos are available on the project website: https://waves.seas.upenn.edu/projects/versa.

Resounding Acoustic Fields with Reciprocity

TL;DR

This work tackles resounding, the estimation of room impulse responses at arbitrary emitter positions from sparse measurements, to enable dynamic, realistic spatial audio in AR/VR. It introduces Versa, a reciprocity-inspired framework that leverages emitter/listener pose exchanges (Versa-ELE) and a self-supervised scheme (Versa-SSL) to enforce reciprocity under realistic gain-pattern asymmetries. Across simulated and real datasets, Versa significantly improves impulse-response accuracy and perceptual realism, with marked gains when emitter patterns differ, and a perceptual study confirming enhanced spatial audio comfort. The approach highlights a physics-grounded learning paradigm that integrates fundamental wave reciprocity into training, offering robust generalization under sparse supervision and potential applicability beyond acoustics to other wave phenomena.

Abstract

Achieving immersive auditory experiences in virtual environments requires flexible sound modeling that supports dynamic source positions. In this paper, we introduce a task called resounding, which aims to estimate room impulse responses at arbitrary emitter location from a sparse set of measured emitter positions, analogous to the relighting problem in vision. We leverage the reciprocity property and introduce Versa, a physics-inspired approach to facilitating acoustic field learning. Our method creates physically valid samples with dense virtual emitter positions by exchanging emitter and listener poses. We also identify challenges in deploying reciprocity due to emitter/listener gain patterns and propose a self-supervised learning approach to address them. Results show that Versa substantially improve the performance of acoustic field learning on both simulated and real-world datasets across different metrics. Perceptual user studies show that Versa can greatly improve the immersive spatial sound experience. Code, dataset and demo videos are available on the project website: https://waves.seas.upenn.edu/projects/versa.
Paper Structure (22 sections, 8 equations, 10 figures, 11 tables)

This paper contains 22 sections, 8 equations, 10 figures, 11 tables.

Figures (10)

  • Figure 1: Estimated acoustic field at novel emitter position. For simulated scenes (five training emitters each), we show loudness distribution (red: loud, blue: quiet) and phase patterns at 1 kHz (periodic blue-red). Purple arrows mark emitter poses. Compared to the baseline (AVR lan2024acoustic), Versa-ELE improves phase map, and Versa-SSL achieves accurate energy map with proper directivity.
  • Figure 2: Left: Reciprocity in acoustic propagation. The path impact function $\Gamma(\cdot)$ is invariant to swapping the emitter and listener poses, because the local acoustic transfer function $f(\cdot)$ at each point is unchanged when the incident and outgoing signal directions are reversed. Right: Impact of gain patterns. Although the propagation path and path impact function remain the same, differences in emitter and listener gain patterns $G$ produce distinct impulse responses $h(t; \cdot)$.
  • Figure 3: Leveraging reciprocity for modeling acoustic fields. a) The vanilla method uses direct supervision with measured impulse responses. b) Versa-ELE enforces response invariance under pose exchange to create physically valid virtual samples. c) Versa-SSL aligns emitter and listener gain pattern to maintain consistency under pose exchange and to enable reciprocity-based self-supervision.
  • Figure 4: Comparison of acoustic field predictions across baseline methods with and without Versa-ELE. We visualize loudness distribution and phase patterns for three simulated scenes (each with five training emitter positions and identical emitter/listener gain patterns). Purple stars mark emitter positions. Versa-ELE enhances all baseline methods' predictions, particularly in modeling energy distribution around emitters. Ground truth shown in the bottom row.
  • Figure 5: Paired (left) vs. unpaired (right) impulse responses.
  • ...and 5 more figures