U-PAST: A Phase-Aware Audio Spectrogram Transformer-U-Net for Single-Channel Speech Enhancement
Cao Duong Ly, Jörn Anemüller
Abstract
Convolutional neural networks (CNNs), used widely and successfully in audio enhancement, capture long-range time-frequency dependencies only indirectly, through successive convolution and pooling. Here, we present U-PAST, a hybrid transformer-U-Net architecture that addresses this limitation through self-attention dependency-modeling in the complex spectrogram domain. U-PAST tokenizes a complex STFT representation, similarly to the magnitude spectrogram tokenization of the Audio Spectrogram Transformer (AST), applies a multi-layer transformer encoder, and reconstructs the enhanced complex spectrogram with a U-Net-style decoder. We evaluate four architectural variants with between 1.17M and 2.40M parameters on the DNS Challenge, VoiceBank-DEMAND, and LibriMix corpora under matched, acoustic mismatch, and two-dataset mismatch conditions. U-PAST attains the best SI-SDR of any evaluated model under acoustic mismatch and closely trails substantially larger convolutional and time-domain baselines by 0.26 dB to 0.63 dB SI-SDR under the remaining three conditions while achieving the strongest perceptual (DNSMOS) quality under dataset mismatch. The largest evaluated configuration, U-PAST-H (2.40M parameters), is consistently the strongest variant of the family, offering an attractive performance-to-cost trade-off at a small parameter footprint.
Create a lesson
Related papers
VibeVoice-ASR-Streaming Technical Report
Yujie Tu, Zhiliang Peng, Jianwei Yu et al.
VAANI Noise Event Dataset: A curated spontaneous speech dataset annotated with timestamps for noise events
Pavan Kumar J, Agneedh Basu, Pranav Bhat et al.
Sensing Bone-Conducted Speech with Earbuds
Christoph Weyer, Peter Jax
TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models
Yuhang Dai, Xin Shu, Zengxi Li et al.
Ontology-based Target Sound Extraction
Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai et al.
Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR
Jiasheng Kuang, Linru Zheng, Hongjin Song et al.