Publikationsserver des Instituts für Deutsche Sprache
Not a member yet
11061 research outputs found
Sort by
Using LLMs for experimental stimulus pretests in linguistics. Evidence from semantic associations between words and social gender
Whether large language models (LLMs) can validly complement or substitute human participants in experimental research remains an open question. Focusing on language cognition, we assess the suitability of GPT-4o and LLaMA 3.1 models (70B Instruct and 8B Instruct) for performing a semantic-association task in German. LLMs labeled noun phrases by social-gender association and rated association strength, mirroring a human participant task. Overall, LLM ratings aligned with human data, but item-level analyses revealed systematic deviations in response patterns
Federated content search for Lexical Resources (LexFCS): Specification
The landscape of digital lexical resources is often characterized by dedicated local portals and proprietary interfaces as primary access points for scholars and the interested public. In addition, legal and technical restrictions are potential issues that can make it difficult to efficiently query and use these valuable resources. As part of the research data consortium Text+, solutions for the storage and provision of digital language resources are being developed and provided in the context of the unified cross-domain German research data infrastructure NFDI. The specific topic of accessing lexical resources in a diverse and heterogenous landscape with a variety of participating institutions and established technical solutions is met with the development of the federated search and query framework LexFCS. The LexFCS extends the established CLARIN Federated Content Search that already allows accessing spatially distributed text corpora using a common specification of technical interfaces, data formats, and query languages. This paper describes the current state of development of the LexFCS, gives an insight into its technical details, and provides an outlook on its future development
“Oh, Now I have to Speak” older adults’ first encounters with voice-based applications in smartphone courses
This chapter deals with the question of what we can learn from interaction in institutional settings about the usability and learnability of everyday technologies such as voice-based Intelligent Personal Assistants (IPAs), especially for older adults or, more generally, less-expert technology users. Based on an analysis of video recordings made during smartphone courses in adult education centers in Germany, this contribution provides a qualitative and micro-analytical perspective on non-expert adult users’ processes of discovering and exploring voice-based technologies. Using the framework of multimodal conversation analysis, both linguistic formats and embodied actions are examined, revealing the participants’ situated and dynamic understandings of how one type of IPA (as a smartphone app or widget) works and operates. The analysis of these either guided or accidental discoveries of a new technology can provide new insights regarding the specific challenges associated with handling IPAs and instructing new users how to do so. Based on these observations, this chapter also provides some general thoughts on teaching digital skills to less-expert user
Human languages trade off complexity against efficiency
From a cross-linguistic perspective, language models are interesting because they can be used as idealised language learners that learn to produce and process language by being trained on a corpus of linguistic input. In this paper, we train different language models, from simple statistical models to advanced neural networks, on a database of 41 multilingual text collections comprising a wide variety of text types, which together include nearly 3 billion words across more than 6,500 documents in over 2,000 languages. We use the trained models to estimate entropy rates, a complexity measure derived from information theory. To compare entropy rates across both models and languages, we develop a quantitative approach that combines machine learning with semiparametric spatial filtering methods to account for both language- and document-specific characteristics, as well as phylogenetic and geographical language relationships. We first establish that entropy rate distributions are highly consistent across different language models, suggesting that the choice of model may have minimal impact on cross-linguistic investigations. On the basis of a much broader range of language models than in previous studies, we confirm results showing systematic differences in entropy rates, i.e. text complexity, across languages. These results challenge the long-held notion that all languages are equally complex. We then show that higher entropy rate tends to co-occur with shorter text length, and argue that this inverse relationship between complexity and length implies a compensatory mechanism whereby increased complexity is offset by increased efficiency. Finally, we introduce a multi-model multilevel inference approach to show that this complexity-efficiency trade-off is partly influenced by the social environment in which languages are used: languages spoken by larger communities tend to have higher entropy rates while using fewer symbols to encode messages
Muster domänenbezogener Indexikalität von Kommunikationsverben im gesprochenen Deutsch. Ein korpuslinguistischer Beschreibungsansatz
Gebrauchsbasierte Sprachmodelle gehen davon aus, dass Sprecher/innen auf der Basis ihres sprachlichen Inputs ein Musterwissen ausbilden und zwar auch in Bezug auf die Assoziationen sprachlicher Mittel zu Merkmalen des situativen Kontextes. Korpuslinguistisch sind statistisch belegbare Assoziationen von Ausdrucksmitteln zu im Korpus erfassten Kontextmerkmalen (Indizierungspotenziale) erschließbar und können in den Mustern ihrer Verteilung betrachtet werden. Es wird eine Untersuchung vorgestellt, die diesen Ansatz anhand der Assoziationen von Kommunikationsverben zu Interaktionsdomänen exploriert. Dabei wird das FOLK-Korpus als Modell des gesprochenen Deutsch behandelt, für das Typen domänenbezogener Indizierungspotenziale ermittelt und Gesprächskonstellationen nach Ähnlichkeit ihrer Indizierungspotenzialprofile gruppiert werden. Der Beitrag zeigt exemplarisch, wie sich Konstellationen des Lebensbereichs Bildung aus dieser Perspektive beschreiben lassen
Gesprochenes Deutsch in den Regionen. Eine Standortbestimmung für die Bundesrepublik Deutschland
Der Beitrag skizziert regionale Unterschiede im Spannungsgefüge von Standard und Dialekt in den Regionen Deutschlands. Es wird deutlich, welche Konfigurationen der Standard-Dialekt-Achse sich im intergenerationellen Vergleich nachweisen lassen, welche individuellen Bedingungen hinter diesen Konfigurationen stehen und welche kommunikativen Potentiale sich daraus ergeben. Die Datenbasis liefert das Akademie-Forschungsprojekt REDE mit einer flächendeckenden Dokumentation der regionalen Varietäten in Deutschland. Der Beitrag greift die Erhebungen von Personen dreier Generationen in unterschiedlichen Explorationskontexten auf und analysiert den sich daraus ergebenden regionalsprachlichen Möglichkeitsraum des gesprochenen Deutsch in der Bundesrepublik Deutschland
Disruptive und diskursive Ereignisse. Ein Vorschlag zur Ausdifferenzierung mit Beispielen aus dem feministischen Abtreibungsdiskurs
Vor dem Hintergrund von Disruption werden in diesem Artikel geltende Konzepte von diskursivem Ereignis hinterfragt und miteinander in Beziehung gesetzt. Die davon abgeleiteten disruptiven Ereignisse werden als eine Subkategorie von diskursivem Ereignis verstanden. Am Beispiel des feministischen Abtreibungsdiskurses wird diesem Ansatz gefolgt und mittels einer Analyse von Praxen des Widersprechens ermittelt, inwieweit durch feministische Akteur*innen die Urteile des Bundesverfassungsgerichts von 1975 und 1993 als disruptive Ereignisse konstruiert werden.Against the background of disruption, this article scrutinizes current concepts of discursive event and relates them to each other. The disruptive events derived from this are understood as a subcategory of discursive event. Using the example of the feminist abortion discourse, this approach is followed and an analysis of practices of contradiction is used to determine the extent to which feminist actors construe the judgments of the Federal Constitutional Court of 1975 and 1993 as disruptive events
Disruptivität des Terrordiskurses. Agonale Aushandlungsprozesse in der Wikipedia
Terroristische Anschläge sind nicht immer noch, sondern immer mehr und immer wieder omnipräsent und eine zentrale Bedrohung gesellschaftlicher Ordnung. Zuletzt verdeutlichte der Terrorangriff der Hamas auf Israel am 7. Oktober 2023, dass Frieden ein fragiles Gut ist und die Grenzen zum Krieg zunehmend verblassen. Terror wirkt sich insofern disruptiv auf bestehende Gesellschaftsordnungen aus, ist aber auch als Diskursgegenstand selbst von Disruption gezeichnet, insofern als sich das Reden über Terror als diskontinuierlich und streitbar erweist. Die in diesem Beitrag vorgenommene Analyse agonaler Aushandlungen im Terrordiskurs verdeutlicht das anhand eines Diskursausschnitts in der Wikipedia exemplarisch. Sie beleuchtet konkrete plattform- und diskursspezifische Merkmale, die die disruptive Kondition des Terrordiskurses hervorrufen oder sich aus dieser ergeben, und exponiert damit eine besondere Wechselwirkung von Diskursthema und Diskursplattform. Damit kann das Definitionsproblem des Terrors zwar nicht gelöst, aber über die Reflexion seiner sprachlichen und diskursiven Konstitution weiter dekodiert werden.Terrorist attacks are not still omnipresent, but are becoming more and more omnipresent and a central threat to social order. Most recently, Hamas' terrorist attack on Israel on October 7, 2023 made it clear that peace is a fragile commodity and the boundaries to war are increasingly fading. In this respect, terror has a disruptive effect on existing social orders, but is also characterized by disruption as a subject of discourse itself, insofar as talking about terror proves to be discontinuous and contentious. The analysis of agonal negotiations in the terror discourse presented in this article illustrates this using a discourse excerpt from Wikipedia as an example. It sheds light on concrete platform- and discourse-specific characteristics that cause or result from the disruptive condition of the terror discourse, and thus exposes a special interaction between discourse topic and discourse platform. Although this does not solve the problem of defining terror, it can be further decoded by reflecting on its linguistic and discursive constitution
Quantitative Analysis of Gendered Assumptions in a Nineteenth-Century Women’s Encyclopedia
This paper quantifies textual patterns relating to gendered assumptions in a fairly unique text, an entire “women’s encyclopedia” from 1830’s Germany, which at 10 volumes and 1,461,000 word tokens was of comparable size to contemporary general encyclopedias, but written and marketed for a female audience. We perform experiments on classifying gender of biographical entries and querying a specific textual feature, calendar dates, with context from comparison 19th-20th century encyclopedias from the EncycNet corpus
Deutsch in der Ukraine
Deutsch tritt in der Ukraine in drei verschiedenen Ausprägungen auf: Deutsch als Minderheitensprache, Deutsch als historische Bezugssprache (z.B. in Zeiten der Habsburgermonarchie) und Deutsch als Fremdsprache. Die deutschsprachige Minderheit in der Ukraine zählt über 30.000 Menschen