Sensing Bone-Conducted Speech with Earbuds
Christoph Weyer, Peter Jax
Abstract
Clear capture of the wearer's own voice (OV) is essential when using earbuds for mobile communication. However, OV capture remains challenging in noisy environments. Bone-conducted (BC) speech, which can be sensed as vibrations of the earbud housing, can be used to improve OV capture. However, neither bandwidth nor spatial characteristics of OV-induced earbud vibrations have been analyzed in detail, despite both characteristics being relevant, e.g., for sensor choice and placement. This study investigates both characteristics, based on measurements with two earbud models. Spectrally, results indicate that OV-induced earbud vibrations exhibit a low-pass characteristic, with a steep roll-off of -93 dB per decade above 400 Hz. Thus, sensors with comparatively low noise floors are required to sense the vibrations above 1. Spatially, results indicate that the earbuds mainly vibrate in and out of the ear canal entrance, with high consistency between subjects and fits. Simulations confirm that this enables capture of the high-power vibrations below 400 Hz by a single-axis sensor with less than 1.5 dB mean attenuation.
Create a lesson
Related papers
VibeVoice-ASR-Streaming Technical Report
Yujie Tu, Zhiliang Peng, Jianwei Yu et al.
VAANI Noise Event Dataset: A curated spontaneous speech dataset annotated with timestamps for noise events
Pavan Kumar J, Agneedh Basu, Pranav Bhat et al.
TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models
Yuhang Dai, Xin Shu, Zengxi Li et al.
Ontology-based Target Sound Extraction
Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai et al.
U-PAST: A Phase-Aware Audio Spectrogram Transformer-U-Net for Single-Channel Speech Enhancement
Cao Duong Ly, Jörn Anemüller
Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR
Jiasheng Kuang, Linru Zheng, Hongjin Song et al.