TalkingHeads · Audiovisual Speech Recognition in-the-wild
Horizon 2020 — Marie Skłodowska-Curie Actions
- Duration
- 2016-06-01 → 2018-05-31
- EU contribution
- €183,455
- Participants
- 1
- Scheme
- MSCA-IF-EF-ST
Lines connect the coordinator with its partners.
Results in brief
TalkingHeads: Audiovisual Speech Recognition in-the-wild
Audio-visual (AV) Automatic Speech Recognition (ASR) in unconstrained (in-the-wild) videos collected from real-world multimedia databases (outdoor conversation/interviews, TV shows with multiple speakers) using novel deep learning methodologies and architectures. IMPORTANCE FOR SOCIETY There are numerous applications associated to visual speech recognition, ranging from medical application (Aphonia, Dysphonia, hearing aids, a.o.) to AV-ASR for noisy environments, and from audiovisual biometrics to surveillance. OBJECTIVES 1. To improve the state-of-the-art in AV ASR by using novel deep learning methods. 2. To introduce new applications related to AV ASR. 3. To transfer knowledge between Speech Recognition and Computer Vision. OUR CONCLUSIONS 1. Visual speech recognition is now a mature technology. It is capable of increasing the accuracy of audio-only speech recognition in both clean and noisy environments, and even when the speaker' mouth area in not captured in high resolution. 2. Deep learning is currently the dominant machine learning paradigm for addressing tasks related to AV ASR, and novel, large and in-the-wild databases are the key towards deploying deep learning methods. 4. We believe that the output of this TalkingHeads constitutes a significant contribution to the development of AV ASR applications.
Data: CORDIS, © European Union
Project objective
Audio-visual speech recognition refers to the problem of recognizing speech using both audio and video information. Speech is not a purely auditory process but the way that the listener perceives it is also through the recognition of the visual patterns associated with the mouth movement. This correlation of the audio-visual information has been occasionally explored in literature in order to develop more robust automatic speech recognition systems for cases in which the auditory environment is noisy (e.g. background noise, multiple speakers). However, the problem of audio-visual speech recognition has been mainly studied in controlled, laboratory conditions. TalkingHeads proposes, for the first time, the problem of audio-visual speech recognition in unconstrained (in-the-wild) videos collected from real-world multimedia databases and a set of methodologies that will work well under the assumed in-the-wild setting. TalkingHeads brings together a talented but experienced researcher (ER) with expertise in speech analysis (diarization and recognition) and the Supervisor with large research experience in Computer Vision for face analysis in-the-wild (recognition, detection, alignment and tracking, and facial expression analysis). TalkingHeads will establish the ER as an independent and internationally recognized researcher in the area of audio-visual fusion and speech recognition. Through TalkingHeads’ achievable work plan, the ER will attain a high level of research maturity by (a) complementing his expertise on speech analysis through extensive training in Computer Vision, (b) conducting research on a challenging research problem (audio-visual speech recognition in-the-wild) with significant career opportunities in both the academia and the industry, (c) publishing at high impact factor conferences and journals, (d) establishing a network of research collaborators, and (e) enhancing personal skills (e.g. supervisory experience, leadership and management skills).
Original text from CORDIS.
Participants
- THE UNIVERSITY OF NOTTINGHAM · NottinghamCoordinatorUnited Kingdom
Links
- View on CORDIS
- DOI: 10.3030/706668
- https://web.archive.org/web/20180207183028/http://www.talking-heads.eu/
Data: CORDIS, © European Union
