FP6Individual fellowship2004–2006

SPEECHSEG · Production and perception of prosodic cues to speech segmentation: multisensorial aspects

FP6 — Marie Curie Actions (Human Resources and Mobility)

Duration
2004-12-29 → 2006-12-28
EU contribution
€150,804
Participants
1
Scheme
IIF

Lines connect the coordinator with its partners.

Results in brief

Final Activity Report Summary - SPEECHSEG (Production and perception of prosodic cues to speech segmentation: multisensorial aspects)

Among the questions addressed was how people engaged in a conversation know where words begin and end. Speech is continuous; it does not have the convenient spaces that separate words in written language. Nevertheless, we generally have no problem in finding the words in our native language. Listeners use many cues in the non-trivial task of speech segmentation. These include a variety of prosodic cues and other language-specific patterns. For example, if an English-speaking listener hears 'mn', she knows that there must be a word break, since English words cannot start with 'mn'. French listeners use the intonation or distinctive pitch patterns of their language as cues of segmentation. The presence of a rise in fundamental frequency (F0) helps them find content words, such as nouns, verbs, etc. F0, the rate of vocal fold vibration, is the primary pitch correlate. The temporal relationship of high and low points, i.e. tones, in the F0 pattern with units on the segmental level like consonants, vowels, and syllables, i.e. tonal alignment, is also important. Dr Welby and her colleagues investigated French tonal alignment, finding consistent patterns, which were not observed in other languages. They also studied tonal alignment from an articulatory point of view, examining the relationship between lip and tongue gestures and F0. They also examined acoustic, i.e. intonational, formant and duration, cues to speech segmentation, using pairs like 'l' affiche and the poster' or 'la fiche and the sheet', which had the same sequences of consonants and vowels. For example, there was sometimes an F0 rise at the start of 'l' affiche and the poster'. Listeners distinguished between pairs like these even before the end of the sequence, suggesting that acoustic differences helped listeners in speech segmentation, even at the earliest stages of processing. Furthermore, they examined whether word boundaries cues such as F0 rises were enhanced in noisy environments, in which listeners might have difficulties in finding word boundaries. This was important because many conversations take place in some sort of noise, such as children playing, cars passing and wind whistling. However, speakers adapt, generally speaking louder, slower, and in a higher pitch, and some of these changes make speech easier to understand. Dr Welby and her colleagues recorded speech in quiet conditions, in white noise and in 'cocktail party' noise and examined intonational and articulatory characteristics of the speech. The results showed that some speakers produced more intonational rises in noisy conditions, suggesting that they might have altered their speech to provide their listeners with cues to content word, e.g. noun, verb, etc, beginnings. In addition, the differences in F0 between quiet and noisy conditions observed for French differed from those observed in another study on Dutch. This difference pointed to the importance of taking into account language-specific differences when studying speech in noise (Lombard speech). Dr Welby and her colleagues used a special lip-tracking system to examine the articulation of speech in noise. They found evidence of hyper-articulation, at least in some conditions. This could be useful to the listener and viewer in segmenting speech. We know that listeners use visual cues as well as auditory cues in other areas of speech perception and that this allows speech to be a robust medium, even in difficult speaking conditions. For example, even in a noisy train station, listeners can distinguish the word mom from the word Tom by the closing of the lips at the beginning of the word. The idea that listeners might use similar articulatory or visual cues to help them segment speech seemed plausible, given that researchers had preliminary evidence that some types of intonation patterns had articulatory or visual correlates. Beyond their contribution to theoretical issues, the results had several potential practical applications, for example in the development of 'smart' speech technologies that detected the presence of noise and adapted to it.

Data: CORDIS, © European Union

Project objective

When we speak, how do we know where one word ends and the next word begins? This task, speech segmentation, is a critical topic in linguistic and psycholinguistic research, since speech segmentation precedes all other linguistic processes, such as syntactic parsing.We know that listeners use a range of cues in their native language to help them in speech segmentation. The proposed project focuses on prosodic cues, e.g. intonational and durational cues, to speech segmentation in French.The project examines a number of outstanding questions including:- Do speakers produce more cues in situations in which listeners may experience difficulty in segmenting speech (e.g. noisy environments)?- What is the time course of the use of prosodic cues to speech segmentation - do listeners use these cues in online speech processing in natural settings?- Are there articulatory / visual correlates to prosodic events that act as cues to speech segmentation?Techniques will include a 'Map Task', which elicits natural, but controlled speech; eye-tracking, which examines a natural response (eye gaze) in speech processing; and tracking of speech articulators (e.g., lips and tongue). The results will make theoretical and applied contributions to several fields, including linguistics, psycholinguistics, phonetics, and engineering.On a theoretical level, the project will add to our knowledge about the nature of cues used in speech segmentation. The results will also have applications to the improvement of speech technology, such as text-to-speech synthesis and automatic speech recognition.The project addresses a number of Community goals and project objectives, including enhancing Europe's international competitiveness (innovative theoretical research and its application to technology), increasing mobility of researchers and ties with third country institutions, and providing equality of opportunity for the disabled (for whom high quality speech technology can improve quality of life).

Original text from CORDIS.

Participants

  • INSTITUT NATIONAL POLYTECHNIQUE DE GRENOBLE · GRENOBLECoordinatorFrance

Links

Data: CORDIS, © European Union