Publikationsserver des Instituts für Deutsche Sprache
Not a member yet
11061 research outputs found
Sort by
Metadata for time aligned corpora
For a detailed description of time aligned corpora, for example spoken language corpora and multimodal corpora, specific metadata categories are necessary, extending the scope of traditional metadata categories. We argue that it is necessary to allow metadata on all levels of annotation, i.e. on a general level for catalogues, on the session level for each recording, on the annotation level for multi tier score annotation, even on the level of individual annotation segments. We use existing standards where they allow this distinction and introduce metadata categories for the layer level
Korpusanalyseplattform der nächsten Generation
Für den Zugriff auf die IDS-Korpora wurde Anfang der 1990er Jahre am IDS das Korpusrecherche- und -analysesystem COSMAS (Corpus Search, Management and Analysis System) (al-Wadi 1994) entwickelt, welches sich bereits in seiner ersten seit 1991 bis 2003 im Betrieb befindlichen Version – COSMAS I – in der Praxis bewährt hatte. Unter den zahlreichen Funktionalitäten waren u.a. ‘virtuelle’ Korpuskomposition, statistische Kookkurrenzanalyse und morphologischer Suchassistent besonders innovativ. 2003 wurde COSMAS I durch die neuere Version COSMAS II (Bodmer 2005) ersetzt, welche vor allem für den Umgang mit Mehrfachannotationen entworfen wurde. Die Datenbasis von COSMAS II speist sich heute aus verschiedenen Quellen: Neben DeReKo sind ebenso historische und einige Projektkorpora mittels COSMAS II für die öffentliche Recherche und Analyse zugänglich gemacht worden. Derzeit hat COSMAS II weltweit ca. 19.000 registrierte Nutzer, die auf die angebotenen Ressourcen zugreifen können. Da COSMAS II jedoch bereits Anfang der 1990er Jahre konzipiert wurde und der Arbeitsaufwand, derartige Software zu erweitern, mit steigender Lebensdauer und Komplexität überproportional steigt, wird es zunehmend schwieriger, die Software an die sich rasch wandelnden Bedarfe anzupassen. Indes haben sich sowohl die technischen als auch die wissenschaftlichen Rahmenbedingungen derart stark verändert, dass es sinnvoll erschien, ein neuartiges Analyse-Tool zu entwickeln, welches neuen Anforderungen und Herausforderungen gerecht wird
The NottdeuYTSch Corpus: a corpus of German-language YouTube comments
In diesem Beitrag wird das Nottinghamer Korpus deutscher YouTube-Sprache (das NottDeuYTSch-Korpus) vorgestellt. Das Korpus hat eine Größe von über 33 Millionen Wörtern, die aus etwa 3 Millionen YouTube-Kommentaren gesammelt wurden. Die Kommentare wurden zwischen 2008 und 2018 veröffentlicht und wurden von einer Gruppe von überwiegend jungen Deutschsprachigen geschrieben. Das NottDeuYTSch-Korpus bietet einen authentischen und repräsentativen sprachlichen Schnappschuss junger Deutschsprachiger und ermöglicht umfangreiche Forschungsmöglichkeiten in verschiedenen linguistischen Bereichen wie Lexik, Morphologie, Syntax, Orthografie, Multilingualismus, sowie Gesprächs- und Diskursanalyse.This paper introduces the Nottinghamer Korpus deutscher YouTube-Sprache (‘The Nottingham German YouTube Language Corpus’ - or NottDeuYTSch corpus). The corpus comprises over 33 million words, taken from roughly 3 million YouTube comments published between 2008 and 2018, written by a young, German-speaking demographic. The NottDeuYTSch corpus provides an authentic and representative linguistic snapshot of young German speakers and offers significant opportunities for in-depth research in several linguistic fields, such as lexis, morphology, syntax, orthography, multilingualism, and conversational and discursive analysis
The communicative repertoire in times of globalization
Our current era of globalization is characterized above all by increased mobility, namely by the increasing mobility of people and the development of new communication technologies, including the mobility of linguistic signs and resources. This process raises new theoretical and methodological questions in linguistics, which results in the development of a new sociolinguistics of globalization (Blommaert 2010) in recent years. One of the most obvious ways to trace this new and dynamic development is to analyze individual language repertoires, especially those of migrants. In this essay, I examine aspects of the communicative repertoire of a refugee who fled to Germany in 2015 to escape the civil war in Syria. I draw on two interviews I conducted with him (in the following I refer to him by the pseudonym „Baran“). The first interview with Baran was recorded in 2016, a few months after his arrival in Germany. The second interview is from 2023, seven years later. In both recordings, German was the dominant language of interaction. I will analyze and show the characteristics of his German at the beginning of his immigration, how he resorts to practices of language mixing between German, Turkish and English (which has recently also been referred to as translanguaging) and how his German has developed over the course of the past seven years
What lexical factors drive look-ups in the English Wiktionary?
This study aims to establish what lexical factors make it more likely for dictionary users to consult specific articles in a dictionary using the English Wiktionary log files, which include records of user visits over the course of 6 years. Recent findings suggest that lexical frequency is a significant factor predicting look-up behavior, with the more frequent words being more likely to be consulted. Three further lexical factors are brought into focus: (1) age of acquisition; (2) lexical prevalence; and (3) degree of polysemy operationalized as the number of dictionary senses. Age of acquisition and lexical prevalence data were obtained from recent published studies and linked to the list of visited Wiktionary lemmas, whereas polysemy status was derived from Wiktionary entries themselves. Regression modeling confirms the significance of corpus frequency in explaining user interest in looking up words in the dictionary. However, the remaining three factors also make a contribution whose nature is discussed and interpreted. Knowing what makes dictionary users look up words is both theoretically interesting and practically useful to lexicographers, telling them which lexical items should be prioritized in lexicographic work
Das Schweizer und das deutsche Wort des Jahres 2022. Anmerkungen aus ukrainischer Sicht
Seit 1977 wird in Deutschland jedes Jahr ein Wort bzw. eine Wortsequenz zum „Wort des Jahres“ gekürt. Vorgenommen wird die Wahl von einer Jury, die sich aus Mitgliedern der Gesellschaft für deutsche Sprache (GfdS) zusammensetzt. In der deutschsprachigen Schweiz gibt es eine solche Aktion ebenfalls (seit 2003); inzwischen wird das Wort des Jahres aber nicht mehr nur auf Deutsch, sondern auch auf Französisch, Italienisch und Rätoromanisch gewählt. Wenn im Folgenden vom „Schweizer Wort des Jahres“ die Rede ist, ist damit aber immer nur das Deutschschweizer Jahreswort gemeint. Durchgeführt wird die Aktion von einem Forschungsteam, das an der Zürcher Hochschule für Angewandte Linguistik (ZHAW) tätig ist
Die digitale Hashtag-Kampagne rund um #CoronaEltern und #CoronaElternRechnenAb: Twitter-Positionierungspraktiken in der Pandemie
As kindergartens and schools closed down during the first wave of the COVID-19 pandemic in Germany, two hashtags emerged on Twitter: #CoronaEltern (#CoronaParents) and #CoronaElternRechnenAb (#CoronaParentsDocumentTheCosts). In this paper, we examine the positioning practices around both hashtags as expressions of “digital activism” (Joyce 2010: VIII). One characteristic of the hashtag campaign is that political demands are hardly ever made directly. Rather, the participants resort to five main linguistic patterns: (1) they address different target groups; (2) they refer to different protagonists; (3) in the subcorpus #CoronaEltern specifically, they constitute themselves as a collective through (4) the recurring use of first-person narratives; (5) and generalization and typification. Our findings show that #CoronaParents are not just parents in times of a pandemic: #CoronaParents are only those who see themselves as such, participating in an evolving, at times misunderstood community
Brisante Gegenstände. Zur valenztheoretischen Integrierbarkeit von Konstruktionen
Die nachfolgenden Überlegungen stehen dabei in einem größeren theoretischen Zusammenhang, sie sind Teil des Konzepts der Grammatischen Textanalyse, mit dem eine deszendente syntaktische Analyse vom Text zur Wortgruppe ermöglicht werden soll (Ágel i.Vorb.). In dieser Monografie werden sowohl für klassische Problembereiche der VT, wie z.B. die freien Dative oder die Präpositionalobjekte, als auch für die ‘Lieblingsthemen’ der KxG, wie z.B. Resultativkonstruktionen (inkl. Caused Motion) oder Geräusch(emissions)verben als Fortbewegungsverben, valenztheoretische Lösungen vorgeschlagen. In dem vorliegenden Beitrag müssen jedoch diese größeren Themenbereiche ausgeklammert bleiben
Finite vs. infinite Attributsätze: zu-dass-Alternation bei Substantiven
In German, certain nouns (such as e.g. Versprechen ‘promise’ or Eigenart ‘characteristic’) can take a subordinate clause that can be realised either in finite form (as a dass-clause ‘that-clause’) or in non-finite form (as a zu-infinitive ‘to-infinitive’). We investigate the distribution of the two variants based on samples drawn from the German Reference Corpus (Kupietz et al. 2018) and the German web corpus DECOW16B (Schäfer & Bildhauer 2012), thus covering both conceptually written registers (as typically found in newspaper texts) and less formal registers (as typically found in internet forums). We first identify the conditions under which the two variants are interchangeable in the first place and subsequently investigate the factors that probabilistically govern speakers’ choices in variable contexts. Among other things, we test the hypotheses that the likelihood for the clause to be realised in non-finite form increases i) along with the quality of the control configuration, ii) in clauses not containing a modal verb (vs. those containing a modal verb), iv) in simple clauses (vs. complex clauses) and v) in newspaper texts (vs. internet forums)