Language Recognition using Time Delay Deep Neural Network

Abstract

This work explores the use of a monolingual Deep Neural Network (DNN) model as an universal background model (UBM) to address the problem of Language Recognition (LR) in I-vector framework. A Time Delay Deep Neural Network (TDDNN) architecture is used in this work, which is trained as an acoustic model in an English Automatic Speech Recognition (ASR) task. A logistic regression model is trained to classify the I-vectors. The proposed system is tested with fourteen languages with various confusion pairs and it can be easily extended to include a new language by just retraining the last simple logistic regression model. The architectural flexibility is the major advantage of the proposed system compared to the single DNN classifier based approach.

0

Turn this paper into a lesson

ArcXiv compiles a structured reading guide from this paper's metadata: plain-English importance, contributions, prerequisite concepts, which sections to read first, flashcards, and a quiz. Grounded in the abstract, never invented.

Discussion (0)

Sign in to join the discussion.

Loading comments…