Skip to content

Mask-Based Speech Enhancement for Spatial Audio: A Comparison of Ambisonics, Beamforming, and Microphone Channels

Sheli Hendel, Boaz Rafaely, Dorothea Kolossa

eess.ASarXiv:2609.18532

Abstract

Mask-based speech enhancement is widely used for suppressing noise and interference, but its performance in spatial audio algorithms with multichannel output has not been studied extensively. In such settings, speech enhancement must improve speech quality while preserving spatial cues that are essential for localization, spatial awareness, and spatial release from masking. In this work, we systematically compare time frequency masking applied to three signal representations: microphone signals, beamformer outputs, and Ambisonics signals. Performance is evaluated in terms of speech quality, intelligibility, binaural cue preservation, and reverberation preservation. Results reveal a clear trade-off between enhancement and spatial fidelity: beamformer-domain masking achieves the highest speech enhancement scores, while Ambisonics-domain masking better preserves the spatial attributes of the residual interference. All methods preserve the target's localization cues.

Create a lesson