LAST: Looped Audio Spectrogram Transformer
Haider Al-Tahan, Sean O'Brien, Anastasia Razdaibiedina, N. Apurva Ratan Murty
Abstract
Increasing depth of transformer models improves recognition, but it comes at a substantial cost. Each additional layer requires more parameters, which makes the process computationally inefficient. We ask whether additional processing can focus on integrating features already computed. Looped Audio Spectrogram Transformer (LAST) first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, thereby making later passes inexpensive. On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. Further evaluations show improved robustness to temporal masking and various other auditory augmentations, with better generalization on classification tasks with music, environmental, and event sounds.
Create a lesson
Related papers
From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment
Liwei Lin, Gus Xia
Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?
Serli Kopar, Alkis Koudounas, Roshan P. Rane et al.
Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization
Jeeyoung Yun, Seohwan Yun, Sungwoong Kim
Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling
Mohammed Hafsati, Ahmed Loughzali
AudioJev: Direct Audio Decisions with Order-Calibrated Probabilities
Sihan Lv, Zhen Li, Zhiqi Cao et al.
Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR
Hwayeon Kim, Youngwon Choi, Hyeonyu Kim