GrainSpeech: Less Context, More Detail for Compact Speech Synthesis
Zitao Liang, Chang Gao
Abstract
Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction. Guided by this finding, we introduce a fixed-receptive-field convolutional encoder that reduces the respective prediction errors by 36.0%, 17.3%, and 3.4%. We further show that directly transferring image-domain gradient-variance supervision restores fine-scale variation but degrades predicted quality, motivating a Mel-specific formulation with axis-specific gradients, overlapping local statistics, and log-domain variance matching. GrainSpeech contains only 264.8K parameters and achieves 17.9x real-time Mel generation on a microcontroller (MCU), while attaining UTMOS scores comparable to substantially larger models with less than 1.5% of their parameters. Source code and demos are available at https://github.com/lab-emi/GrainSpeech.
Create a lesson
Related papers
Absolute Quality Ratings of Speech Enhancement Systems by Listeners of Different Ages and Degrees of Hearing Loss
Matteo Torcoli, Chih-Wei Wu, Andrea Esposito et al.
Mask-Based Speech Enhancement for Spatial Audio: A Comparison of Ambisonics, Beamforming, and Microphone Channels
Sheli Hendel, Boaz Rafaely, Dorothea Kolossa
Reviving Etter method for autoregressive inpainting: Generalization, evaluation, implementation
Ondřej Mokrý, Matěj Hrdlička, Pavel Rajmic
Correlation-Guided Encoder Selection for Multi-Encoder Large Audio-Language Models
Pei-Jun Liao, Hung-Shin Lee, Wenze Ren et al.
Task-oriented neural FOA encoding for SELD from irregular microphone arrays
Jiachen Liu, Yin Cao, Ming Wu et al.
G-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement
Guo-Ruei Tseng, Hung-Shin Lee, Hsin-Min Wang et al.