HQSTS · High-Quality voice model for STatistical parametric speech Synthesis
„Хоризонт 2020“ — Действия „Мария Склодовска-Кюри“
- Период
- 2015-10-01 → 2017-12-31
- Финансиране от ЕС
- 183 455 €
- Участници
- 1
- Схема
- MSCA-IF-EF-ST
Линиите свързват координатора с партньорите.
Накратко на български
Методи за анализ и синтез на гласа изследват как човешката реч се превръща в параметри, за да се възпроизведе отново, например при превръщането на текст в реч. Това подобрява качеството на синтезираните гласове и премахва изкуственото бръмчене, особено в тиха среда.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
High-Quality voice model for STatistical parametric speech Synthesis
A speech analysis/synthesis method aims at representing a speech waveform, produced by a person speaking, as a time sequence of parameters. Based on this time sequence, the speech waveform can be resynthesized. The analysis/synthesis methods are cornerstones for many speech technologies (e.g. text-to-speech, telecommunications, voice restoration). For the majority of applications, these methods need to have two key properties: (i) a high perceived quality of the speech sound, and, (ii) a statistical characterization of the parameters' sequence necessary for statistical approaches. The current analysis/synthesis methods exhibit however a lack of perceived quality. This issue does not pose a problem for noisy environments, but prohibits the use of statistical approaches in quiet environments, where the listener is fully aware of all the details of the sound. Recent phase processing tools allowed the description of the phase spectrum and noise properties in a way that shows the drawbacks and limits of current analysis/synthesis methods. Additionally, these same tools are also promising means for modeling the phase and noise information, which is paramount for good quality. The primary goal of the HQSTS project is to create a high-quality analysis/synthesis method that will broaden the applications of statistical approaches of speech technologies in quiet environments, where a high-quality is an absolute necessity. A new high-quality analysis/synthesis method has been developed called Pulse Model in Log domain (PML). From a practical point of view, it prevents the buzziness often present in synthetic voices and thus clearly improved the overall quality. From a theoretical point of view, the new and simple approach offers a better control of the sound characteristics and ease the developments of further quality improvements. A full training system for speech synthesis has also been realized during this research work that implements state-of-the-art Artificial Neural Nets techniques (ANN). This system has been made open-source in order to constitute a solid anchor for researchers and developers that need working implementation at hand in the current fast pace developments of ANN. The analysis/synthesis method as well as the training system for speech synthesis are available on GitHub.com at: https://github.com/gillesdegottex
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
A speech analysis/synthesis method aims at representing a speech waveform, produced by a person speaking, as a time sequence of parameters. Based on this time sequence, the speech waveform can be resynthesized. The analysis/synthesis methods are cornerstones for many speech technologies (e.g. text-to-speech, telecommunications, voice restoration). For the majority of applications, these methods need to have two key properties: (i) a high perceived quality of the speech sound, and, (ii) a statistical characterization of the parameters' sequence necessary for statistical approaches, which have attracted great interest during the last decades in speech technologies. The current analysis/synthesis methods that provide a statistical characterization exhibit however a lack of perceived quality. This issue does not pose a problem in applications designed for noisy environments (e.g. navigation devices, smart-phone applications, announcements in train stations). On the contrary, it prohibits the use of statistical approaches in quiet environments, e.g. in the music, cinema and video game industries, where the listener is fully aware of all the details of the sound. This problem is mainly due to the lack of an accurate representation of the phase information and its correlation with the amplitude information. Indeed, recent phase processing tools allowed the description of the phase spectrum properties in a way that shows the drawbacks and limits of current analysis/synthesis methods. Additionally, these same tools are also promising means for modeling the phase information, which is paramount for good quality. The primary goal of the HQSTS project is to create a high-quality analysis/synthesis method that will broaden the applications of statistical approaches of speech technologies in quiet environments, where a high-quality is an absolute necessity.
Оригинален текст от CORDIS (на английски).
Участници
- THE CHANCELLOR MASTERS AND SCHOLARS OF THE UNIVERSITY OF CAMBRIDGE · CAMBRIDGEКоординаторОбединеното кралство
Връзки
- Виж в CORDIS
- DOI: 10.3030/655764
- http://gillesdegottex.eu/Demos/HQSTS/
- https://arquivo.pt/wayback/20201221140429/http://gillesdegottex.eu/Demos/HQSTS/
Данни: CORDIS, © Европейски съюз
