Publikationsserver des Instituts für Deutsche Sprache
Not a member yet
11061 research outputs found
Sort by
Sequence, gaze, and modal semantics: modal verb selection in German permission inquiries
Two modal verbs of German are regularly used to express deontic possibility: können (‘can’) and dürfen (‘may’). We examine how speakers select between them, focusing on modal inquiries for permission to carry out some action (darf/kann ich das machen, ‘may/can I do this’). Our data are video-recordings of everyday face-to-face interaction, which we analyze sequentially, drawing on interactional-linguistic methods. We find that local sequential context and aspects of visible turn-design systematically enter into the accomplishment of deontic meaning. (1) Position of the modal inquiry within a course of action informs verb selection: speakers select kann to nominate an action as coming “out of the blue” and initiating a new course of action; darf to nominate an action as sequentially occasioned: a solution to an already known-in-common problem. (2) Bodily behavior (gaze, body posture) guides the interpretation of modal flavor, moving a kann-inquiry towards a deontic or a circumstantial interpretation, and moving a darf-inquiry towards a deontic or a bouletic interpretation. Overall, the study demonstrates the systematic contributions of sequential position and body behavior in the accomplishment of modal meanings
Unlocking the corpus: enriching metadata with state-of-the-art NLP methodology and linked data
In research data management, descriptive metadata are indispensable to describing data and are a key element in preparing data according to the FAIR principles (Wilkinson et al., 2016). Extracting semantic metadata from textual research data is currently not part of most metadata workflows, even more so if a research data set can be subdivided into smaller parts, such as a newspaper corpus containing multiple newspaper articles. Our approach is to add semantic metadata at the text level to facilitate the search over data. We show how to enrich metadata with three NLP methods: named entity recognition, keyword extraction, and topic modeling. The goal is to make it possible to search for texts that are about certain topics or described by certain keywords, or to identify people, places, and organisations mentioned in texts without actually having to read them and at the same time facilitate the creation of task-tailored subcorpora. To enhance this usability of the data we explore options based on the German Reference Corpus DeReKo, the largest linguistically motivated collection of German language material (Kupietz & Keibel, 2009; Kupietz et al., 2010, 2018), which contains multiple newspapers, books, transcriptions, etc., and enrich its metadata on the level of subportions, i.e. newspaper articles. We received access to a number of data files in DeReKo’s native XML format, I5. To develop the methodology, we focus on a single XML file containing all issues of one newspaper of a whole year. The following sections only give an overview of our approach, we intend, however, to provide a detailed description of the experiments and the selection of data in a subsequent longer contribution
The technological context for internet lexicography
Computer technology is becoming ever smaller and cheaper, both to acquire and operate, and its processing and storage performance is increasing exponentially. This is one of the technological requirements for making dictionaries available online, but so too is the infrastructure of the Internet, which makes it possible to exchange information and data simply and reliably between billions of interconnected computers. This chapter is devoted to the fundamental technological preconditions for present-day Internet lexicography. First, we outline what actually happens “behind” the user interfaces that are visible on the screen when a user accesses a dictionary online and how these processes can be recorded in log data for the purposes of documenting them. Second, we discuss how the identity and long-term availability of content can be maintained in view of the possibility of online material being constantly updated
Internet lexicography. An introduction
The Internet has become the central publication platform for dictionaries. This profound change in the dictionary landscape gives rise to a whole range of new questions for lexicographic practice and dictionary research.
This volume provides for the first time an introduction to the central fields of work in Internet lexicography and presents the current state of scientific research and lexicographic practice. The chapters cover key aspects of dictionary creation, such as the technical framework, data modeling, and lexicographic process, linking dictionary content, access and navigation structures, automatic extraction of lexicographic information, user participation, and research on dictionary use.
The aim of this volume is to provide students and teachers (at universities) with an introductory and easy-to-read overview on Internet lexicography, thus anchoring this important and innovative field of research and practice in university teaching. All chapters convey the basic concepts and methods in a comprehensible way and are enriched by references to further and more in-depth reading
Der Krisendiskurs im Kontext des russisch-ukrainischen Krieges
In der gegenwärtigen Zeit wird die Gesellschaft durch eine Vielzahl von Krisen geprägt, was sich deutlich in der Sprache niederschlägt. Der russisch-ukrainische Krieg stellt einen komplexen Krisendiskurs dar, der sowohl die ukrainische Gesellschaft als auch die internationale Gemeinschaft erheblich beeinflusst (Tripps, Vogel 2022). Seit 2022 zeigen sich im Krisendiskurs in der Ukraine und in Deutschland sowohl Gemeinsamkeiten als auch signifikante Unterschiede, die auf historische und politische Kontexte, mediale Darstellungen, kulturelle Spezifika, unterschiedliche Wahrnehmungen sowie sprachliche Besonderheiten zurückzuführen sind
Programmieren für Germanist*innen (digGer - Digitale Germanistik). Elektronische Ressource
Probleme in kleinere Teilprobleme zu zerlegen, zu systematisieren und die Auswertung zu algorithmisieren sind wichtige kulturelle Fertigkeiten. Durch die stattfindende Weiterentwicklung der Geisteswissenschaften hin zu ‚Digital Humanities‘ eröffnen sich neuen Ansätze, Methoden, Gegenstände und Arbeitsmittel – Programmierung, eine wichtige Kulturtechnik, wird gegenwärtig jedoch nicht in der Breite des Faches gelehrt. Denn obwohl der Bedarf dafür vorhanden ist, fehlt es an adäquaten Lehr- und Lernmaterialien.
Die geplante Lerneinheitengruppe soll sich diesem Desiderat annehmen und leicht verständliche, auf die Gegenstände der Germanistik fokussierte Ressourcen entwickeln. In der Germanistik hat sich in den letzten Jahren insbesondere die Programmiersprache Python als verbreitete Einstiegssprache etabliert. Python bietet viele Programm-Bibliotheken für häufige Aufgabenbereiche und gleichzeitig existiert eine gute Interoperabilität zu anderen Programmiersprachen wie z. B. R. Die Lerneinheitengruppe adressiert neben dem sprachtypischen Wissen (Syntax, verfügbare Bibliotheken etc.) auch allgemeines Wissen zu Programmiersprachen (Refactoring, Wie findet man Hilfe?, Was ist ein guter Programmierstil?)
Editorial: Understanding technology use in face-to-face interaction: A conversation analytic perspective
Bürgerbriefe im Nationalsozialismus - Reaktionsweisen des Regimes
Der erste Teil des Beitrags umreißt zu diesem Zweck in groben Zügen die Dimensionen des Eingabewesens während des Nationalsozialismus. Neben dem Versuch einer quantitativen Einordnung wird gezeigt, dass das NS-Regime grundsätzlich darum bemüht war, sich als empfänglich für „berechtigte“ Anliegen und Sorgen der „Volksgenossinnen“ und „-genossen“ zu präsentieren. Der Fokus des längeren zweiten Abschnitts liegt anschließend auf der Bandbreite an Reaktionsweisen der angeschriebenen Instanzen. Anhand einzelner Fallbeispiele wird eine Typologie entworfen, die systematisch darlegt, welche unterschiedlichen Reaktionsweisen auf Bittgesuche und Beschwerdeschreiben zu verzeichnen sind und welche Verläufe Kommunikationsprozesse im Anschluss an Eingaben nehmen konnten
Schwankungen zwischen schwacher und starker Substantivflexion
The present corpus-study investigates fluctuations between the so called “weak” and “strong” nominal inflection classes in German. Corroborating previous studies (Köpcke 1995, Schäfer 2019), we show that the tendency for traditionally weak masculine nouns to join the strong pattern is strongest for nouns displaying phonotactic and semantic properties that are atypical of weak nouns. Conversely, we show that among strong masculine nouns attested in a weak form at least once, nouns with phonotactic and semantic properties typical of weak nouns are overrepresented compared to masculine nouns not attested in weak forms. Using logistic regression, we show, among other things, that non-canonical forms are more likely to occur in informal texts (represented by internet forum discussions) than in (more formal) newspaper texts. While this is fully expected for the traditionally weak nouns, where the use of the non-canonical (strong) forms involves a loss of case suffixes (viz. the loss of -(e) n in the accusative and dative singular), it is perhaps more noteworthy with respect to the traditionally strong nouns, where, conversely, the use of the non-canonical (weak) forms leads to additional case morphology. Moreover, we find that the shift from strong to weak differs from the shift from weak to strong with respect to grammatical case (accusative / dative vs. genitive). We propose that is connected to the status of the genitive as a marker of formal / written style