A High Quality Text-To-Speech System Composed of Multiple Neural Networks
Orhan Karaali, Gerald Corrigan, Noel Massey, Corey Miller, Otto Schnurr, Andrew Mackie
Abstract
While neural networks have been employed to handle several different text-to-speech tasks, ours is the first system to use neural networks throughout, for both linguistic and acoustic processing. We divide the text-to-speech task into three subtasks, a linguistic module mapping from text to a linguistic representation, an acoustic module mapping from the linguistic representation to speech, and a video module mapping from the linguistic representation to animated images. The linguistic module employs a letter-to-sound neural network and a postlexical neural network. The acoustic module employs a duration neural network and a phonetic neural network. The visual neural network is employed in parallel to the acoustic module to drive a talking head. The use of neural networks that can be retrained on the characteristics of different voices and languages affords our system a degree of adaptability and naturalness heretofore unavailable.
Create a lesson
Related papers
ANTShapes Benchmarking Datasets for Event-Based Neuromorphic Object Classification
M. Middleton, H. Kayan, B. Sen Bhattacharya et al.
Bug Localization from Bug Reports: A Multi-Objective Approach
Waleed Ahmad, Mehtab Kiran Suddle, Maryam Bashir
Synthesis of Hopfield Neural Network: Novel Results
Garimella Rama Murthy
Homo-RAG: Homology-Guided Retrieval-Augmented Generation for Cross-Species Gene Function Prediction
Azrin Sultana
Learning Whom to Trust : Decision-Generated Credibility in Social Learning
Gabriel Bontemps, Abhishek Banerjee
On Scaling Coordinate-Based Neuroevolution: The Quadtree Bottleneck in ES-HyperNEAT
Romain Claret, Michael O'Neill, Paul Cotofrei et al.