H2020Individual fellowship2019–2021

ETE SPEAKER · Robust End-To-End SPEAKER recognition based on deep learning and attention models

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2019-06-01 → 2021-01-31
EU contribution
€120,817
Participants
1
Scheme
MSCA-IF-EF-ST

Lines connect the coordinator with its partners.

Results in brief

Robust End-To-End SPEAKER recognition based on deep learning and attention models

Automatic speaker recognition is the task performed by a machine of identifying the person speaking in a given recording. There are several closely related tasks such as language recognition, where the system determines which language is being spoken; voice activity detection (VAD), where segments containing actual speech are separated from other unwanted information in the signal (silence, music); speaker diarization (SD), where the system determines speaker turns in a recording; and automatic speech recognition (ASR), where the system processes the speech segment in order to transcribe the message contained on it. The complexity of these tasks lies in the wide variety of nuisance variability contained in the speech signal (recording device, acoustic conditions, etc.), which the system needs to disentangle from the information that is relevant for the target task. These challenges are faced by automatic systems and also by humans. For instance, while humans are relatively good at discriminating speakers known to them, it is a real challenge when it involves unknown voices. Thus, automatic systems are able to outperform humans for a large number of unknown speakers and take advantage of the information available in large datasets with thousands of hours of speech. Speaker recognition as well as other related tasks have several applications in real-world scenarios, especially nowadays when more and more devices are operated by humans just with their voice. For instance, voice-driven bank applications should grant access only to the authorized person, for which robust text-dependent speaker recognitions systems are essential. Moreover, obtaining a robust speaker representation improves notably speaker diarization and all its relevant applications such as indexing audiovisual resources (internet, companies and institution meetings, court sessions, parliament sessions) or support for hearing-impaired people with speaker-colored subtitles on TV or speaker specific models for more accurate automatic transcriptions. It is also a very relevant task for production of linguistic resources useful for research and development. The ETE SPEAKER project aims to improve speaker recognition systems to make them robust to different scenarios and specific tasks, with special focus on deep learning-based approaches. These systems are able to learn the information needed to represent and discriminate between speakers directly from data, similar to what humans do during their learning process. In this line, we have explored deep learning methods that extract information from the recordings encoding both speaker identity and message content in the context of text-dependent speaker recognition, improving existing techniques and analyzing the behavior of different modules on the system (bottleneck feature extractors, neural embedding x-vector and i-vector extractors, etc.). Furthermore, we have developed speaker diarization systems based on attention models and trained them in an end-to-end way. This way, the system performs the whole diarization task, which implies learning the separation of speaker turns, VAD and even overlapped speech where more than one speaker is speaking (which is a limitation of traditional approaches to this task) entirely from data.

Data: CORDIS, © European Union

Project objective

This project focuses on automatic speaker recognition (SID), the task of determining the identity of the speaker in a speech recording. Disentangling the speaker specific information from the rest of nuisance variability requires complex models. Deep neural networks (DNNs) have recently showed their potential for this, as the popular x-vector learnt by a DNN.Here, we aim for end-to-end SID where the system is optimized as a whole for the target task. Despite several attempts in this line of research, many aspects still remain unexplored or not explored thoroughly.We also propose to explore recurrent approaches, suitable for dealing with temporal signals, as well as different pooling methods to obtain a fixed-length representation from a variable length input sequence of speech features.Next, we want to explore different flavors of attention mechanisms, which make the DNN to focus on relevant parts of the input, providing a way to quantify how much evidence has been collected about the speaker identity and the uncertainty of the obtained representation, which is a critical issue when making (Bayesian) decisions in SID.Finally, some other approaches such as using the raw signal (instead of features) or other advances that might arise will be also explored for SID and related tasks.To achieve our goals, we will start from theory, implement the proposed approaches and test on public SID benchmarks such as NIST SREs. The outcomes are intended to benefit both scientific community and speech processing industry.The applicant Dr. Alicia Lozano-Diez is an excellent female researcher, who has done her Ph.D. at Audias (Universidad Autonoma de Madrid, Spain), a respected research lab. The host group Speech@FIT from Brno University of Technology (Czechia) has a top-class track on speech processing research. Thus, we expect the combination of both the researcher and the host to boost the researcher career and benefit the host group (and its industrial European partners).

Original text from CORDIS.

Participants

  • VYSOKE UCENI TECHNICKE V BRNE · BRNO STREDCoordinatorCzechia

Links

Data: CORDIS, © European Union