1,721,231 research outputs found

    Études linguistiques et phonétiques du code-switching français-arabe : analyses de grands corpus et traitement automatique de la parole

    No full text
    Cette thèse présente des recherches linguistiques et phonétiques sur le code-switching Français-Arabe Algérien. Un corpus de 7h30 de parole (5h de parole spontané et 2h30 de parole lue) a été constitué en enregistrant 20 hommes et femmes parlant le français et l'arabe algérien. Cette thèse présente également les méthodes de traitement des données orales du code-switching telles que la segmentation de la parole, la segmentation des énoncés de code-switching ainsi que la transcription du français et du dialecte arabe algérien. Cette thèse présente également des méthodes d'alignement automatique de ces données bilingues ainsi qu'un alignement combiné de deux alignements monolingues. Nous avons mené des expériences basées sur l'alignement automatique avec des variations qui traitent de la question de l'influence d'un système phonologique d'une langue A sur des productions phonétiques en code-switching du français et de l'arabe algérien. Nous avons d'abord abordé la variation en réalisant une étude sur la variation des voyelles, dans des productions en langue française et en arabe algérien. Nous avons aussi abordé les consonnes emphatiques et l'emphatisation des deux langues. Enfin, nous avons également travaillé sur les géminées et la gémination dans les productions langagières en code-switching. Les résultats ont montré que le code-switching FR-AA se caractérise par des changements de langues très courts qui sont un réel défi pour l’identification des langues dans le code-switching. Le code-switching a un impact sur la variation phonétique des voyelles et des consonnes. La parole du code-switching permet au locuteur de produire moins de variation de voyelles et de consonnes que la parole monolingue.This thesis proposes linguistic and phonetic investigations of French-Algerian Arabic code-switching. A corpus of 7h30 of speech (5h of spontaneous speech and 2h30 of read speech) has been designed with 20 males and females French-Algerian Arabic speakers.This thesis also proposes code-switching speech data processing methods such as language segmentation, code-switching utterance segmentation and transcription of French and Algerian Arabic dialect. Automatic speech alignment methods of the code-switching data are proposed with combined alignment of two monolingual alignments. We conducted experiments based on language automatic identification and automatic alignment with variations that deals with the question of the influence of a phonological system of a language A on code-switching speech in phonetic productions of French and Algerian Arabic. We dealt first with identifying the language change boundaries. We performed also a variation study on vowel variation, in both French and Arabic productions. Finally, we dealt with three types of consonant variation in the code-switching speech: gemination, emphatization and voicing consonant as variants in production. The results shown that the code-switching French-Algerian Arabic is characterized by very short language switches witch constitute a big challenge to the code-switching languages identification . The code-switching has an impact of the phonetic variation in both vowel and consonants. The code-switching allows the speakers to produce less vowel and consonant variation than the monolingual speech

    Re-ranking of candidates answers of a question-answering system.

    No full text
    L’objectif de cette thèse a été de proposer une approche robuste pour traiter le problème de la recherche dela réponse précise à une question.Notre première contribution a été la conception et la mise en œuvre d’un modèle de représentation robuste de l’informationet son implémentation. Son objectif est d’apporter aux phrases des documents et aux questions de l’informationstructurelle, composée de groupes de mots typés (segments typés) et de relations entre ces groupes. Ce modèle a été évalué sur différents corpus (écrits, oraux, web) et a donné de bons résultats, prouvant sa robustesse.Notre seconde contribution a consisté en la conception d’une méthode de réordonnancement des candidats réponsesretournés par un système de questions-réponses. Cette méthode a aussi été conçue pour des besoins de robustesse, ets’appuie sur notre première contribution. L’idée est de comparer une question et le passage d’où a été extraite une réponse candidate, et de calculer un score de similarité, en s’appuyant notamment sur une distance d’édition.Le réordonnanceur a été évalué sur les données de différentes campagnes d’évaluation. Les résultats obtenus sontparticulièrement positifs sur des questions longues et complexes. Ces résultats prouvent l’intérêt de notre méthode, notreapproche étant particulièrement adaptée pour traiter les questions longues, et ce quel que soit le type de données. Leréordonnanceur a ainsi été évalué sur l’édition 2010 de la campagne d’évaluation Quaero, où les résultats sont positifs.The objective of this work is to introduce a new robust approach to treat the problem of finding the correctanswer to a question.Our first contribution is the design and implementation of a robust representation model for information. The aim is torepresent the structural information of sentences of documents and questions structural information. This representation iscomposed of typed groups of words (typed segments) and relations between these groups. This model has been evaluatedon several corpus (written, oral, web) and achieved good resultats, which proves his robustness.Our second contribution consisted is the design of a re-ranking method of a set of the candidate answers output by thequestion-answering system. This re-ranking method is based on the structural information representation. The general ideais to compare a question and a passage from where a candidate answer was extracted, and to compute a similarity score by using a modified edit distance we proposed.Our re-ranking method has been evaluated on the data of several evaluation campaigns. The results are quite goodon long and complex questions. These results show the interest of our method : our approach is quite adapted to treatlong question, whatever the type of the data. The re-ranker has been officially evaluated on the 2010 edition of the Quaeroevaluation campaign, with positives results

    Des corpus arborés à l’induction de structures syntaxiques partielles

    No full text
    Nos travaux portent sur les treebanks, ces corpus de textes dotés d’annotations de structures syntaxiques. Ils sont très utiles dans de nombreux domaines, de la linguistique au traitement automatique de la langue. Après une introduction portant sur leur rôle dans des domaines variés, nous plongeons dans l’histoire de leur création, depuis les pratiques d’annotation manuelle de textes vers les treebanks modernes avec l’avènement des technologiques. Le chapitre 3 montre les méthodes de création de ces treebanks. Le chapitre 4 discute des problématiques liées à la constitution des guides d’annotation, et mets en évidences certaines de ces problématiques au travers de deux études, la première portant sur traitement des expressions multi-mots, la seconde sur la constitution d’un treebank dans une langue peu pourvue en ressources, le Naija langue parlée au Nigéria étudiée dans le cadre du projet ANR NaijaSynCor. Le chapitre 5 présente l’outil Arborator-Grew, conçu pour faciliter l’annotation collaborative des treebanks. Le chapitre 6 étudie comment des lois linguistiques fondamentales comme la loi de Menzerath-Altmann et le Heavy Constituent Shift interagissent. Il propose également plusieurs procédures pour générer des arbres artificiels, permettant de contraster leurs propriété avec celles des arbres syntaxiques. Enfin, le chapitre 7 vise à utiliser des techniques statistiques pour découvrir la structure sous-jacente des phrases dans un texte. En résumé, ce travail montre l’importance des treebanks dans notre compréhension des langues, et leur rôle dans le développement des technologies linguistiques en soulignant l’innovation continue dans ce domaine.This document focuses on treebanks, textual corpora with syntactic annotations. These treebanks are invaluable in numerous fields, ranging from linguistic studies to natural language processing. First, we explore how these treebanks aid researchers in multiple domains. Next, we dive into the history of treebank development. Before the computer age, researchers began to manually create collections of annotated texts, which evolved with the advent of computers into modern treebanks. Chapter 3 focuses on the challenges and methods of creating these treebanks. Chapter 4 addresses challenges relating to developing annotation guidelines. These discussions bring us to two case studies, the first relating to how to best handle complex multi-word expressions, the second retracing the development of a treebank for a low-resource language, the Naija pidgin-creole spoken spoken in Nigeria and its analysis as part of the ANR NaijaSynCor project. Chapter 5 introduces a new tool, Arborator- Grew, designed to facilitate collaborative annotation of treebanks. Chapter 6 studies how linguistic laws such as the Menzerath-Altmann law and the Heavy Constituent Shift interact. It also introduces several tree generation algorithms, which we use to contrast the properties of syntactic and artificial trees. Finally, Chapter 7 aims to use statistical techniques to latent structure of sentences in a text. In summary, this work highlights the importance of treebanks in our understanding of languages and their significant role in the development of language technologies. It also emphasizes continuous innovation in this field, opening new avenues for the study and analysis of languages

    Segmental reduction in spoken French through different speech styles : contributions of large speech corpora and automatic speech processing on schwa, /ʁ/ and reduction of multiple segments

    No full text
    Ce travail sur la réduction segmentale (i.e. délétion ou réduction temporelle) en français spontané nous a permis non seulement de proposer deux méthodes de recherche pour les études en linguistique, mais également de nous interroger sur l'influence de différents facteurs de variation sur divers phénomènes de réduction et d'apporter des connaissances sur la propension à la réduction des segments. Nous avons appliqué la méthode descendante qui utilise l'alignement forcé avec variantes lorsqu’il s’agissait de phénomènes de réduction spécifiques. Lorsque ce n'était pas le cas, nous avons utilisé la méthode ascendante qui examine des segments absents et courts. Trois phénomènes de réduction ont été choisis : l'élision du schwa, la chute du /ʁ/ et la propension à la réduction des segments. La méthode descendante a été utilisée pour les deux premiers. Les facteurs en commun étudiés sont le contexte post-lexical, le style, le sexe et la profession. L’élision du schwa en syllabe initiale de mots polysyllabiques et la chute du /ʁ/ post-consonantique en finale de mots ne sont pas toujours influencées par les mêmes facteurs. De même, l’élision du schwa lexical et celle du schwa épenthétique ne sont pas conditionnées par les mêmes facteurs. L’étude sur la propension à la réduction des segments nous a permis d'appliquer la méthode ascendante et d’étudier la réduction des segments de manière générale. Les résultats suggèrent que les liquides et les glides résistent moins à la réduction que les autres consonnes et que les voyelles nasales résistent mieux à la réduction que les voyelles orales. Parmi les voyelles orales, les voyelles hautes arrondies ont tendance à être plus souvent réduites que les autres voyelles orales.This study on segmental reduction (i.e. deletion or temporal reduction) in spontaneous French allows us to propose two research methods for linguistic studies on large corpora, to investigate different factors of variation and to bring new insights on the propensity of segmental reduction. We applied the descendant method using forced alignment with variants when it concerns a specific reduction phenomena. Otherwise, we used the ascendant method using absent and short segments as indicators. Three reduction phenomena are studied: schwa elision, /ʁ/ deletion and the propensity of segmental reduction. The descendant method was used for analyzing schwa elision and /ʁ/ deletion. Common factors used for the two studies are post-lexical context, speech style, sex and profession. Schwas elision at initial syllable position in polysyllabic words and post-consonantal /ʁ/ deletion at word final position are not always conditioned by the same variation factors. Similarly, lexical schwa and epenthetic schwa are not under the influence of the same variation factors. The study on the propensity of segmental reduction allows us to apply the ascendant method and to investigate segmental reduction in general. Results suggest that liquids and glides resist less the reduction procedure than other consonants and nasal vowels resist better reduction procedure than oral vowels. Among oral vowels, high rounded vowels tend to be reduced more often than other oral vowels

    La production et la perception de l'allemand chez les apprenants francophones : analyse de corpus de parole, électroéxncephalographie et enseignement

    No full text
    Ce projet de recherche vise à étudier la production et la perception de la parole chez les apprenants francophones de l’allemand. Un corpus de parole de 7 heures correspondant à trois tâches (imitation, lecture, description) a été enregistré. Il comprend des germanophones natifs et des apprenants francophones. Nous avons analysée les productions des segments intéressants d'après le cadre du SLM. Une étude de perception en EEG utilisant [h-ʔ], [ʃ-ç] et les voyelles courtes et longues a été réalisée sur des germanophones natifs et des apprenants francophones. Enfin, l'impact de l'enseignement sur l'amélioration des production et perception a été examiné à travers une étude longitudinale. L'étude de production montre que, suivant les tâches, les apprenants produisent le [h] en début de mot sans problème majeur. De même, ils peuvent produire des voyelles de durée contrastive. Cependant, pour les trois tâches, les apprenants ont plus de difficultés pour la production de la qualité vocalique, de [ç] et [ŋ]. Fait notable, la perception ne reflète pas toujours la production. Les apprenants tendent à ne pas percevoir le [h] en début de mot alors que la production de ce segment en répétition est bonne. À l'inverse, les apprenants perçoivent le contraste [ʃ-ç] mais sa production reste difficile. Seulement dans les voyelles courtes et longues, la perception reflète la production.L'étude d'enseignement montre que la conscience linguistique affecte différemment perception et production : une conscience linguistique accrue permet d'affiner la perception de phonèmes à contenu acoustique complexe et la production des phonèmes faciles à produire du point de vue articulatoire.This research project proposes to investigate the production and perception of German speech in French learners of German. A 7h speech corpus containing three production tasks (imitation, reading, description) produced by German natives and French learners was recorded. Segmental productions of challenging vowels and consonants were analysed according to the SLM. A perception experiment involving [h-ʔ], [ʃ-ç] and short and long vowels using EEG was carried out on German natives and French learners. Finally, the impact of pronunciation teaching on improved speech production and perception was investigated. Undergraduates following a stand-alone pronunciation class were recorded and performed perception tests before and at the end of the course. The production study showed that French learners may produce word-initial [h] faithfully. With regard to short and long vowels, contrasting vowel duration is produced. However, French learners encounter more difficulties with respect to vowel quality. This holds for the production of [ç] and [ŋ]. Interestingly, perception does not always mirror production. The EEG results showed that the perception of word-initial [h] is poor in French learners whereas production accuracy is good. On the contrary, French learners perceive the [ʃ-ç] contrast but its production remains difficult. Only in short and long vowels, perception mirrored production. The teaching study showed that the increased linguistic awareness may affect non-native speech perception and production in different ways: phones that are easy to produce from an articulatory point of view can benefit from teaching. Increased awareness helps to better perceive phones with rich acoustic information

    Evaluation d'unites de decision pour la reconnaissance de la parole continue

    No full text
    SIGLECNRS T Bordereau / INIST-CNRS - Institut de l'Information Scientifique et TechniqueFRFranc

    Phonetic corpora and big data

    No full text
    International audienceDuring the last years, 'big data' has emerged as a trendy, highly promising portmanteau term in economics and high-tech domains, such as information technology and speech processing. Big data are often described using a 3V scheme: volume, variety, velocity: a huge volume of data, a large variety of possibly unstructured, heterogeneous data sources, a high frequency or velocity of data generation over time. In this Glasgow ICPhS 2015 discussant session, we will question the 'big data' term with respect to phonetics and speech sciences at large. In this context, big data typically refer to huge, generally unstructured collections of speech or audio-visual data, pre-existing any phoneticians' investigation hypotheses. Can such data become beneficial to phonetic sciences

    IMPACT OF DURATION AND VOWEL INVENTORY SIZE ON<br />FORMANT VALUES OF ORAL VOWELS: AN AUTOMATED<br />FORMANT ANALYSIS FROM EIGHT LANGUAGES.

    No full text
    International audienceEight languages (Arabic, English, French, German,Italian, Mandarin Chinese, Portuguese, Spanish)with 6 differently sized vowel inventories wereanalysed in terms of vowel formants. A tendencyto phonetic reduction for vowels of short acousticdurations clearly emerges for all languages. Thedata did not provide evidence for an effect ofinventory size on the global acoustic space andonly the acoustic stability of quantal vowel /i/ isgreater than that of other vowels in many cases.Keywords: vowel formants, adaptive dispersion,reduction, quantal theory

    Analyses formantiques automatiques en français : périphéralité des voyelles orales en fonction de la position prosodique.

    No full text
    National audienceThe aim of the present study is to highlight peripheralityof French vowels in two prosodic positions: (i) word-finalsyllables (as compared to word-initial syllables), and (ii)in the vicinity of (before and after) pauses. The LIMSIspeech alignment system is used [1] and formant valuesof oral vowels are automatically measured in a total of25000 segments from two hours of journalistic broadcastspeech in French. A tendency to reduction for all vowels(in terms of the shrinking of the vocalic triangle formedby F1 and F2 values) of short duration was clearlyobserved in a former study (Gendrot & Adda-Decker [2]).We show that at some extent, a similar relationship holdsfor vowels in both word-final and word-initial syllables
    corecore