PercQualAVS · A Model for Predicting Perceived Quality of Audio-visual Speech based on Automatic Assessment of Intermodal Asynchrony
7РП — „Хора“ (Действия „Мария Кюри“)
- Период
- 2011-05-01 → 2013-04-30
- Финансиране от ЕС
- 155 542 €
- Участници
- 1
- Схема
- MC-IEF
Линиите свързват координатора с партньорите.
Накратко на български
Разминаването между движението на устните и звука на гласа се анализира, за да се предвиди как хората възприемат качеството на видеото. Това помага за автоматичното измерване на синхронизацията в аудио-визуалните сигнали.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
A Model for Predicting Perceived Quality of Audio-visual Speech based on Automatic Assessment of Intermodal Asynchrony
The aim of the project was to develop a model for predicting the perceived quality of asynchronous audio-visual speech. The project was conceived as an inherently multidisciplinary endeavour that would encompass aspects of computer vision, speech processing, cognitive science, machine learning and statistical computing. The project can be conceived of as having the following main components: 1. A computer vision and speech processing component that would be used to extract useful audio-visual features from an input signal. These features would form the basis to which an automatic asynchrony detection measure would be applied. 2. A data gathering component, which would oversee the acquisition of new primary data (mainly in the form of a new audio-visual corpus for use in the project) as well as the gathering of subjective perceptual response data, through the design and execution of several perceptual experiments. 3. An analytics component that involved the exploration and analysis of the perceptual responses gathered through experimentation. The results of these analyses would form the basis for an initial model of audio-visual asynchrony perception. 4. A machine learning component, which would seek to use the model that resulted from component 3 above as the basis for an automatic model for measuring asynchrony in some audio-visual input and making a prediction regarding how a human being might perceive such asynchronous input in terms of quality. The technical implementation and execution of the project can be said to be its most successful component. The fellow succeeded in developing computer vision based feature extractors, which automatically tracked the lips in real time and extracted relevant and useful features from them. In conjunction with acoustic feature extractors he developed using standard speech processing toolkits, the fellow was able to generate useful digital representations of the audio-visual speech inputs which formed the basis of the project's analysis. The fellow also developed software that processed these extracted features automatically and applied several techniques for measuring the degree of synchrony or asynchrony between the audio and visual components. The resulting asynchrony score was mapped onto the same perceptual response scales used during the perceptual experiments and provided a means through which the automatic asynchrony detection could be directly assessed in terms of accuracy, as it could be directly compared with the experimental data. This direct means of comparison between the perceptual responses gathered experimentally and the automatically rated asynchrony also allowed the fellow to develop a learned model of audio-visual asynchrony perception using techniques from the area of machine learning. The project was considerably less successful in terms of its dissemination and research outputs, particularly publications. The fellow wrote and submitted regularly to conferences, but the results (as highlighted by the reviewers) were generally felt to be intermediate and therefore the papers were rejected. Unfortunately, one paper written towards the closing of the project was accepted but had to be withdrawn on the basis that notice of acceptance had come after the project funding had closed. Although significant progress was made in terms of developing software, tools and processes for each of the outlined components above, it is probably true that in the initial stages of the project the focus of the fellow lay too heavily on these areas. Looking back, it would have been more productive to have focused more on more visible results, such as concentrating on more focused experimental studies (rather than the rather narrow focus of acquiring perceptual data). This broader outlook would probably have been rewarded with a more favourable acceptance-to-submission ratio, in terms of publications. However, the technical and professional gains made by the fellow in terms of acquiring extensive skills in new and useful areas such as data analysis, machine learning and computer vision, should not be overlooked, and can perhaps offset in some small way, the disappointing performance of the project in other areas.
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
In recent years, there has been a marked increase in communication technologies and computer interfaces that operate within the audio-visual speech domain, (e.g. video-telephony, synthesised avatars, etc). Faithful synchrony between the visual and acoustic speech elements of such technologies is of great importance in ensuring that they are perceived by end-users as operating at high and optimal quality levels. The effect of intermodal asynchrony on user-perceived quality is typically assessed using subjective evaluation techniques. A system for automatically assessing asynchrony levels, and predicting quality degradation on that basis, would therefore be both desirable and useful, and will have direct application to techniques for automatic synchrony adjustment.The proposed project will examine audio-visual speech as both spoken naturally by humans and as artificially synthesised by machines, and will employ subjective assessment techniques and machine learning in a combined iterative semi-automatic strategy for producing a Quality Prediction Model. Different levels of intermodal asynchrony will first be assessed by human subjects, who will be required to score the effect of the asynchrony levels on perceived speech quality using standardisedtechniques that will be modified for use with multimodal speech. Asynchrony patterns and their corresponding subjective assessment scores will be automatically learned by machines, resulting in an initial Quality Prediction Model. The initial model will be tested using data that will be simultaneously assessed by humans, using the subjective assessment techniques, above. Theoutput from the prediction model will be directly compared with the subjective scores, providing an initial evaluation of the model's performance. The model will be adjusted on this basis, and re-trained using new data. The process of re-train, re-test, re-score, will be repeated iteratively, leading to a more robust quality prediction model.
Оригинален текст от CORDIS (на английски).
Участници
- TECHNISCHE UNIVERSITAT BERLIN · BerlinКоординаторГермания
Връзки
Данни: CORDIS, © Европейски съюз
