Publikationsserver des Instituts für Deutsche Sprache
Not a member yet
    11061 research outputs found

    Semiautomatic data generation for academic Named Entity Recognition in German text corpora

    Get PDF
    An NER model is trained to recognize three types of entities in academic contexts: person, organization, and research area. Training data is generated semiautomatically from newspaper articles with the help of word lists for the individual entity types, an off-the-shelf NE recognizer, and an LLM. Experiments fine-tuning a BERT model with different strategies of post-processing the automatically generated data result in several NER models achieving overall F1 scores of up to 92.45%

    Preface

    Get PDF

    Data modelling

    Get PDF
    In this chapter, we will first discuss different data formats in which structured content can be formally represented, explaining their respective advantages and disadvantages, and how suitable query languages can be used to retrieve information from these data structures. The third section covers the core issues of data modelling – how to describe the structure of specific lexicographical content, e.g. which “boxes” the lexicographic content should be put in – both in abstract terms and with reference to the data structures introduced before. put in, including the associated advantages and disadvantages. There are many lexicographic projects that face largely similar challenges. For this reason, initiatives have been launched oriented towards developing standardised solutions for modelling lexicographic data, similar to a set of guidelines for sorting Lego bricks. We report on these in the fourth section

    Understanding technology use in face-to-face interaction

    Get PDF
    The contributions of this special issue will look at video data documenting how participants use and adapt to technologies in and for social interaction. This special issue therefore focuses on multimodal, i.e., both verbal and nonverbal, communication practices of participants in a selection of social settings, both day-to-day and institutional, where various technologies, both well-known and more innovative, are present

    Minority languages through the lens of the linguistic landscape

    No full text
    When we first started the project of looking at minority languages through a linguistic landscape lens, we felt that the visibility of minority languages in public space had been insufficiently dealt with in traditional minority language research. A linguistic landscape approach, as it had developed over the last years, would constitute a valuable path to explore, by looking at the ‘same old issues’ of language contact and language conflict from a specific angle. We were convinced that fresh linguistic landscape data would be able to provide innovative and useful insights into ‘patterns of language […] use, official language policies, prevalent language attitudes, [and] power relations between different linguistic groups’ (Backhaus 2007, p. 11). The linguistic landscape approach, as presented by the different authors in this volume, has clearly proven to be a heuristic appropriate and relevant for a wide range of minority language situations. More specifically, the ideas and analyses in the different chapters do contribute to a further understanding of minority languages and their speakers. They deepen our comprehension of language policies, power relations and ideologies in minority language settings

    Literary psycholinguistics and the poem

    Get PDF
    This dissertation examines the processes and mechanisms of language comprehension that characterise the reception of literary texts. It reports a series of experiments that use behavioural and electrophysiological measures to study the effects of readers' genre conceptions on the processing and evaluation of verbal stimuli. Focusing on different aspects and stages of literary reception, all of these experiments contrast poetry-specific processing and evaluation routines with routines appropriate for processing everyday language or prose texts, while strictly controlling for linguistic variables. Chapter 1 reports two experiments that employed sentence judgments (Experiment 1a) and event-related brain potentials (Experiment 1b) to investigate the influence of genre categorization (poetry vs. no categorization) on the evaluation and the online comprehension of single sentences. Specifically, we examined whether external genre categorization modulates the use of prosodic information in sentence reading and whether it affects the processing and the perceived meaningfulness of semantically (in)congruent utterances. The use of single sentences rather than entire texts allowed us to isolate a priori effects of genre categorization, i.e., processing adjustments triggered by the category and not by the (con)text. Chapter 2 presents additional analyses of the EEG data collected in Experiment 1b. Aiming to pin down the effects of genre categorization on anticipatory attentional states, these analyses (Experiment 1c) focused on pre-stimulus alpha power as an index of selective attention prior to the first linguistic input as well as on further early (<350ms) ERP effects of genre categorization. Chapter 3 reports a study that combined eye tracking and speech recordings (Experiment 2) to contrast the processing strategies for literary prose vs. poetry during oral text reading. Moving beyond aprioristic effects of genre categorization, this study specifically aimed to reveal genre-specific dynamics by studying the interaction of text category and (con)text during literary processing. Chapter 4 presents two experiments that used sentence judgments to examine the influence of morpho-syntactic and prosodic variables on the grammatical and literary-aesthetic evaluation of poetic verse (Experiment 3a) and of regular sentences (Experiment 3b). These studies aimed to demonstrate that experienced readers systematically associate the genre of poetry with selected grammatical structures and features. Chapter 5 summarizes the results of the presented experiments, relates them to distinct stages of the reception process, discusses some limitations of the present work and identifies possible directions for future research in literary psycholinguistics.Die vorliegende Dissertation untersucht Prozesse und Mechanismen der Sprachwahrnehmung, die charakteristisch für die Rezeption literarischer Texte sind. Sie berichtet eine Reihe von Studien, die mittels behavioraler und elektrophysiologischer Messmethoden den Einfluss von Gattungskonzeptionen auf die Verarbeitung und Evaluation sprachlicher Reize untersuchen. Was diese Studien verbindet, ist die systematische Gegenüberstellung gedichtspezifischer und alltagssprachlicher bzw. prosaspezifischer Verarbeitungs- und Evaluationsroutinen unter Beibehaltung des sprachlichen Materials. Was diese Studien unterscheidet, sind die untersuchten Aspekte des Rezeptionsprozesses und die dazu verwendeten, der Psycholinguistik entlehnten Methoden. Kapitel 1 berichtet zwei Studien, die mithilfe systematisch gesammelter Leserintuitionen (Experiment 1a) und ereigniskorrelierter Hirnpotentiale (EKPs; Experiment 1b) untersuchen, welchen Einfluss Gedicht- und Verskonzeptionen auf das Echtzeitverstehen und die Beurteilung einzelner Sätze haben, die gattungstypische formale und semantische Merkmale aufweisen. Kapitel 2 berichtet weitere Analysen der in Experiment 1b gesammelten EEG-Daten. Mittels Zeit-Frequenz- und EKP-Analysen untersucht Experiment 1c den Einfluss von Gattungszuschreibung auf antizipatorische Aufmerksamkeit vor dem Lesen und auf die frühe Echtzeitverarbeitung geschriebener Sprache. Kapitel 3 berichtet eine Studie, die mithilfe kombinierter Blickbewegungsmessungen und Sprachaufnahmen (Experiment 2) untersucht, wie Gattungszuschreibung (Gedicht vs. literarische Prosa) das (Vor)lesen unbekannter Texte beeinflusst, und so behaviorale und akustische Marker dieser literarischen Lesemodi identifiziert. Im Gegensatz zu früheren Untersuchungen zielt diese Studie explizit darauf ab, distinktive Dynamiken gattungspezifischen Lesens zu erfassen. Kapitel 4 berichtet zwei Studien, die sich systematisch gesammelter Leserintuitionen bedienen, um den Einfluss syntaktischer und prosodischer Variablen auf die grammatische und literarisch-ästhetische Evaluation einzelner Verse (Experiment 3a) und Sätze (Experiment 3b) zu untersuchen. Diese Studien zeigen, dass erfahrene Leser die Textsorte Gedicht mit spezifischen grammatischen Strukturen und Merkmalen verbinden. Ich argumentiere, dass das Auftreten dieser Textmerkmale charakteristische Verarbeitungsschritte erfordert bzw. erleichtert, die von Lesern als spezifisch poetische Qualitäten des Textes und der Leseeerfahrung wahrgenommen werden. Kapitel 5 fasst die Resultate der berichteten Studien zusammen, bezieht sie auf einzelne Stufen des Rezeptionsprozesses und diskutiert die Grenzen ihrer Generalisierbarkeit. Das Kapitel schließt mit einem Ausblick auf mögliche zukünftige Forschungsfelder literarischer Psycholinguistik. Auf der Grundlage der gesammelten Beobachtungen ergibt sich folgende Liste gedichtspezifischer Anpassungen der Sprachverarbeitung und -evaluation. Die Konzeption der Textsorte Gedicht... 1. beeinflusst die Erwartungen und die Aufmerksamkeit der Leser bereits vor dem Lesen 2. bewirkt jedoch nicht, dass Leser prosodischen Rekurrenzen gesteigerte Aufmerksamkeit schenken 3. führt zu gattungsspezifischen Anpassungen des Leseverhaltens und der Blickbewegungsroutinen 4. bestimmt Strategien zur gattungsadäquaten phonetischen Realisierung sprachlicher Reize 5. beeinflusst zwar nicht die frühe Verarbeitung (scheinbarer) semantischer Inkongruenz, veranlasst Leser jedoch, sowohl während des Lesens als auch danach größeren Interpretationsaufwand zu betreiben, um zu einer kohärenten Bedeutungsrepräsentation zu gelangen 6. beeinflusst nicht die strategische Nutzung verarbeiteter prosodischer Regelmäßigkeiten (Metrum) beim Lesen 7. macht Leser resilienter gegenüber gattungstypischen historischen Wortformen, die ohne Kenntnis der Gattung kurzzeitig die verfügbaren kognitiven Ressourcen während des Lesens reduzieren 8. beinhaltet gattungsspezifische Evaluationskriterien für grammatische und semantische Merkmale sprachlicher Reize Gemeinsam betrachtet zeigen diese Resultate deutlich, dass literarische Gattungen selbst in der Konzeption literarischer Laien mit ausgewählten sprachlichen Konstruktionen und Merkmalen verknüpft sind, sowie mit adäquaten Anpassungen der Sprachverarbeitung und -evaluation. Diese Anpassungen betreffen sämtliche Stufen des literarischen Rezeptionsprozesses: die Aufmerksamkeit des Lesers vor dem Lesen, die Verarbeitungsroutinen während des Lesens, und die Evaluationskriterien für sprachliche Reize nach dem Lesen

    Komposita mit den relationalen Zweitgliedern "Gatte" und "Gattin" – eine korpusbasierte Studie aus genderlinguistischer Perspektive

    Get PDF
    In diesem Beitrag werden Komposita mit den relationalen Zweitgliedern Gatte und Gattin aus genderlinguistischer Perspektive untersucht, basierend auf manuell annotiertem zeitungssprachlichen Korpusmaterial. Frauen werden im analysierten Korpus ca. 12-mal häufiger in ihrer ehelichen Rolle versprachlicht als Männer. Statistische Analysen zeigen, dass sie dabei systematisch in ein possessives Verhältnis zum Ehemann gesetzt werden (Arztgattin = Gattin eines Arztes), während Ehemänner in den untersuchten Komposita tendenziell doppelt individualisiert werden (Arztgatte = Gatte, der Arzt ist). Neben den Zweitgliedern geben auch die Genera der beiden Konstituenten Aufschluss über die kodierte Bedeutungsrelation: Genusgleichheit (Kanzlergatte) führt zu einer qualifizierenden, Genusdivergenz (Kanzleringatte) zu einer possessiven Lesart. Die Analyse belegt außerdem die Existenz movierter Kompositumserstglieder – diese sind sogar die häufigste Form zur Benennung weiblicher Personen im Erstglied. Trotzdem herrscht bei der Bezugnahme auf Frauen eine größere Formenvarianz als bei Männern, welche fast ausschließlich mit maskulinen Erstgliedern versprachlicht werden. Damit zeigt die Studie, wie genderlinguistische Perspektiven auch im Bereich der Wortbildung einen neuen Analysezugang bilden.This study examines compounds with the relational second elements Gatte (‘husband’) and Gattin (‘wife’) from a gender-linguistic perspective, based on manual annotations of material from a presscorpus. In the analysed corpus, women are referred to in their marital roles 12 times more often than men. Statistical analyses show that they are systematically put into a possessive relation to their husband (Arztgattin = ‘a doctor’s wife’), while husbands tend to be individualised twice within the analysed compounds (Arztgatte = ‘a husband who is also a doctor’). Besides the second elements, the grammatical gender of the two constituents provides information about the meaning relations within the compounds: if both have the same grammatical gender (Kanzlergatte), a qualifying meaning is encoded; if they have differing grammatical genders (Kanzleringatte), a possessive reading is triggered. The analysis also provides evidence for the existence of feminised first elements –they are even the most common form to refer to women with the first element. However, reference to women is subject to a high degree of formal variance compared to men, who are almost exclusively referred to with masculine forms. Thus, the study shows how gender-linguistic perspectives can contribute new analytical approaches to word formation processes

    Languages and parliaments. The impact of decentralisation on minority languages

    No full text
    The central question of Marten's volume is how languages and parliaments interact, and what role a parliamentary institution can play within language policy. This question is addressed in particular in the context of minority languages and language revitalisation processes. Based on in-depth research of parliamentary documents and interviews with policy makers, scholars, and language activists from Scotland and Norway, the study investigates how the establishment of the decentralised Scottish Parliament and the parliamentary assembly for the Sámi population in Norway, the Sameting, have generated increased efforts of language maintenance of the Gaelic and Sámi languages respectively. For this purpose, Marten on the one hand contrasts the situations before and after the establishment of these two parliaments in 1999 and 1989 respectively, and on the other hand compares the developments in the two countries in the light of the different political structures in Scotland and Norway. The study illustrates how negotiations take place between supportive and reluctant policy makers in the two parliamentary contexts and shows how they have eventually resulted in a higher level of empowerment of the two speech communities. As a result, the volume therefore shows that a decentralisation of parliaments can indeed lead to increased language maintenance efforts, albeit within certain limits. Parliamentary decentralisation is thus identified to be one piece within the large puzzle of minority language policy. As such, it is related to the theoretical literature on minority languages by suggesting an additional component in the evaluation of minority language situations

    9,397

    full texts

    11,061

    metadata records
    Updated in last 30 days.
    Publikationsserver des Instituts für Deutsche Sprache is based in Germany
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇