Deep Learning for Real-Time Sound Order Recognition in Human-Robot Interaction
Rezaul Tutul, Usaid Khan, Andre Jakob, Ilona Buchem
Abstract
Recognizing the temporal order of overlapping sounds is an underexplored challenge in human-robot interaction (HRI), with direct relevance to applications such as first responder detection systems. This paper presents a deep learning framework for real-time sound order recognition using recordable buzzers that emit distinct non-verbal sounds (cat meows, dog barks, helicopter noises). A multi-branch convolutional neural network (CNN) processes Mel spectrograms, Mel-frequency cepstral coefficients (MFCCs), and short-time Fourier transform (STFT) features, with an attention-based fusion mechanism to emphasize critical temporal cues. Experiments were conducted under same-amplitude, varied-amplitude, and unseen sound conditions. The proposed system achieved 99% accuracy in balanced overlaps, 91% under amplitude variation, and 74% on unseen test data with normalization. These results demonstrate that deep learning can reliably recognize sound order in overlapping conditions, supporting practical HRI scenarios. While experiments were conducted on carefully controlled synthetic overlaps, we additionally report latency benchmarks demonstrating real-time feasibility and provide an extended discussion on generalization, ecological validity, and deployment challenges in real-room environments.
Create a lesson
Related papers
FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
Chengxian Hu, Zhiming Ma, Mingjun Pan et al.
TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
Huiyuan Liu, Zhiming Ma, Yanxing Liu et al.
Multi-Teacher Distillation for Cross-Domain Streaming Electrolaryngeal Speech Encoding
Benedikt Mayrhofer, Enrique Orozco Olivares, Franz Pernkopf et al.
Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification
Seungmin Seo, Oleg Aulov, P. Jonathon Phillips et al.
TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking
Giorgia Adorni, Michela Papandrea, Battista Rimoldi et al.
VoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval
Aaron Yee, Fengjie Lu, Jiarui Hai et al.