SPEAKER DICE · Robust SPEAKER DIariazation systems using Bayesian inferenCE and deep learning methods
Horizon 2020 — Marie Skłodowska-Curie Actions
- Duration
- 2017-03-01 → 2019-02-28
- EU contribution
- €142,721
- Participants
- 1
- Scheme
- MSCA-IF-EF-ST
Lines connect the coordinator with its partners.
Results in brief
Robust SPEAKER DIariazation systems using Bayesian inferenCE and deep learning methods
The SPEAKER DICE project dealt with the Speaker Diarization (SD) task. The SD is a task, which consists in automatically finding speaker turns in an audio utterance, or as it is commonly stated, finding “who spoke when?” Although being apparently easy for humans, diarization is a highly challenging task for machines, as it deals with the complex task of Speaker Recognition, it needs to find the (unknown) number of speakers in the utterance, it has to do segmentation of speech into speaker turns (finding boundaries between speakers) and needs to deal with overlapped speech (cross-talk). One of the main applications of Speaker Diarization is the indexing of audiovisual resources with speakers. This indexing allows a structured search and access to resources depending on the speaker of interest. This feature can be very useful in a wide range of scenarios. First of all it would be very valuable for public institutions, allowing the indexing of sessions of parliaments, courts, etc. The indexing can also be helpful for companies allowing, for example, access to specific parts of meetings or seminars. Besides, TV, internet and radio broadcasters would benefit from such system, as they could provide a more versatile access to their contents. The indexing of TV broadcasters is of a special interest, as it would allow the automatic colouring of subtitles according to the speaker, which would make the media more accessible for hearing-impaired people. In addition to the direct applications speaker diarization, the diarization systems are also helpful and relevant for other related tasks. To list a few, it can be used for speaker adaptation in Automatic Speech Recognition (ASR). Also, it is a very important part of the system pipeline for Speaker Recognition (SR) in wild scenarios in which several speakers are present but only one is of interest. Moreover, it is relevant for the production of linguistic resources, as it allows collecting language utterances avoiding speaker repetitions. This project focuses on improving, extending current and developing new approaches to enhance the performance of Speaker Diarization systems. For that purpose, we set three main objectives: first, optimize the current Bayesian models which have strong mathematical foundation to achieve better performance. Second, driven by the success of the artificial Neural Network (NN) based techniques for the related speaker recognition task, we aim to integrate NN modules into the diarization pipeline. Finally, the third objective is to make the system applicable to the general case, so that it generalizes to any kind of speech and environment.
Data: CORDIS, © European Union
Project objective
The proposed project deals with Speaker Diarization (SD) which is commonly defined as the task of answering the question “who spoke when?” in a speech recording. The first objective of the proposal is to optimize the Bayesian approach to SD, which has shown to be promising for the tasks. For Variational Bayes (VB) inference, that is very sensitive to initialization, we will develop new fast ways of obtaining a good starting point. We will also explore alternative inference methods, such as collapsed VB or collapsed Gibbs Sampling, and investigate into alternative priors similar to those introduced for Bayesian speaker recognition models.The second part of the proposal is motivated by the huge performance gains that, in recent years, have been brought to other recognition tasks by Deep Neural Networks (DNNs). In the context of SD, DNNs have been used in the computation of i-vectors, but their potential was never explored for other stages of SD. We will study ways of integrating DNNs in the different stages of SD systems.The objectives of the proposal will be achieved by theoretical work, implementation, and careful testing on real speech data. The outcomes of the project are intended not only for scientific publications, but eagerly awaited by European speech data mining industry (for example Czech Phonexia or Spanish Agnitio).The project is proposed by an excellent female researcher, Dr. Mireia Diez, having finished her thesis in the GTTS group of University of the Basque Country, one of the most important European labs dealing with speaker recognition and diarization. The proposed host is the Speech@FIT group of Brno University of Technology, with a 20-year track of top speech data mining research. The proposed research training and combination of skills of Dr. Diez and the host institution have chances to advance the state-of-the-art in speaker diarization, provide the applicant with improved career opportunities and benefit European industry.
Original text from CORDIS.
Participants
- VYSOKE UCENI TECHNICKE V BRNE · BRNO STREDCoordinatorCzechia
Links
Data: CORDIS, © European Union
