Publikationsserver des Instituts für Deutsche Sprache
Not a member yet
    11061 research outputs found

    The Georgian Dialect Corpus: problems and prospects

    Get PDF
    The Georgian Dialect Corpus (GDC) covers a significant segment of the spoken language of Georgia. It is conceived as a sub-corpus of the Georgian National Corpus1 and is designed for wide interdisciplinary research. Since 2006, the project has been fund- ed by the Shota Rustaveli National Science Foundation. With its structure, the GDC represents a wide spectrum of regional, temporal and stylistic variations of the Georgian linguistic reality. It contains texts from all Georgian dialects (including the dialects spread in Iran, Turkey, and Azerbaijan); intensive work on a corpus of Laz texts is underway. Currently, we are working on the elaboration of a morphological annotation concept. In this process, the first step is lemmatization. While automatic lemmatization is an easily solvable and trivial problem in corpora of standard languages with exhaustive morphological descriptions, it is a rather difficult task in a dialect corpus containing a comprehensive collection of texts from up to twenty dialects. Therefore it is under- taken manually in most dialect corpora. In our concept, we effectively apply a lexico- graphical datapool and a standard language parser within a semi-automatic annotation process. The lemmatization process is then based on the standard form, dialect lemmata and standard lemmata being “deemed equal”. The implementation of this presupposes the manual lemmatization of a certain amount of dialect texts

    Grammatisches Wissen, grammatische Bildung und Grammatikunterricht – das digitale Informationssystem grammis zwischen Forschungsvermittlung und forschendem Lernen

    Get PDF
    Die Auseinandersetzung mit den grammatischen Mustern eines Sprachsystems ist sowohl in der Erstsprache (L1) als auch beim Erwerb von Zweit- bzw. Fremdsprachen (L2) die vielleicht größte Hürde. Gründe hierfür können unterschiedlich sein. Grammatik-Lehrwerke überfordern Lernende ohne spezielle Vorkenntnisse zum einen oft mit ihrer Komplexität, zum anderen werden sie gelegentlich als zu abstrakt in dem Sinne wahrgenommen, dass die konkrete Bezugnahme auf Sprachpraxis nicht überzeugend erschlossen werden kann. Das hat naheliegenderweise Folgen: Entweder misslingt der Erwerb von metasprachlichem Wissen ganz oder das erworbene Wissen wird rasch wieder vergessen. Unter anderem aus diesen Gründen ist der Bedarf nach Ressourcen, die eine nachhaltig wirksame, nicht überlastende und weiterfördernde Beschäftigung mit strukturellem Sprachwissen ermöglichen, stärker denn je. Die Erstellung einer solchen Ressource für den Schul- bzw. den Fremd-/Zweitsprachenunterricht verfolgt das Projekt LernGrammis. Um die Bedarfe der mutmaßlich heterogenen Nutzungsgruppen und Bildungskontexte abzuschätzen, führt das Projekt eine breit adressierte Online-Umfrage unter Lehrenden und Lernenden durch und evaluiert die Praxiseinbindung mit Hilfe experimenteller Studien und qualitativer Interviews. Die Umsetzung der konzeptionellen Grundlagen erfolgt im Rahmen von grammis, einem wissenschaftlich basierten Informationssystem zur deutschen Grammatik und Orthografie, das seit vielen Jahren eines der meistgenutzten Online-Angebote des Leibniz-Instituts für Deutsche Sprache darstellt. Es präsentiert aktuelle Forschungsergebnisse, Korpusanalysen, Spezialwörterbücher und empirische Sprachdatensammlungen, weiterhin Erklärungen und Hintergrundwissen für Sprachinteressierte in der ganzen Welt. Als nachhaltig verfügbare digitale Bildungsinfrastruktur entstehen zudem Lernbausteine, Lernpfade und interaktive Übungen für die schulische bzw. universitäre Ausbildung sowie für Deutschlernende.The exploration of grammatical language patterns is perhaps the greatest challenge both in first language acquisition (L1) and second or foreign language acquisition (L2). The reasons for this challenge vary. Grammar textbooks often overwhelm learners without specific prior knowledge, due to their complexity and occasional perception of being too abstract, in the sense that concrete reference to language practice cannot convincingly be established. This has obvious consequences: either the acquisition of metalinguistic knowledge fails completely, or the acquired knowledge is quickly forgotten. For these reasons, the demand for resources that facilitate a sustainably effective, non-overwhelming, and continuously supportive engagement with structural language knowledge is stronger than ever. The LernGrammis project aims to create such a resource for school or second/foreign language instruction. In order to assess the needs of presumptively heterogeneous user groups and educational contexts, the project conducts a broadly targeted online survey among teachers and learners and evaluates practical implementation through experimental studies and qualitative interviews. The conceptual foundations are implemented within the framework of grammis, a scientifically based information system on German grammar and orthography, which has been one of the most widely used online resources of the Leibniz Institute for the German Language for many years. It presents current research findings, corpus analyses, specialized dictionaries, and empirical language data collections, along with explanations and background knowledge for a language-interested audience worldwide. Additionally, as a sustainably available digital educational infrastructure, it comprises learning modules, learning paths, and interactive exercises for school or university education, as well as for German learners

    Text mining in the Humanities - A plea for research infrastructures

    Get PDF
    Research infrastructures for the Humanities can help to share digital resources and content services. In particular, they can help researchers in the Digital Humanities to save time and efforts when developing software to deal with specific research issues. Web services and web applications can be used to build a research infrastructure for sharing data and algorithms. However, the development of such infrastructures and their key software components is a software engineering task that increasingly also poses interesting and challenging research problems for Computer Science

    Evaluating Workflows for Creating Orthographic Transcripts for Oral Corpora by Transcribing from Scratch or Correcting ASR-Output

    Get PDF
    Research projects incorporating spoken data require either a selection of existing speech corpora, or they plan to record new data. In both cases, recordings need to be transcribed to make them accessible to analysis. Underestimating the effort of transcribing can be risky. Automatic Speech Recognition (ASR) holds the promise to considerably reduce transcription effort. However, few studies have so far attempted to evaluate this potential. The present paper compares efforts for manual transcription vs. correction of ASR-output. We took recordings from corpora of varying settings (interview, colloquial talk, dialectal, historic) and (i) compared two methods for creating orthographic transcripts: transcribing from scratch vs. correcting automatically created transcripts. And (ii) we evaluated the influence of the corpus characteristics on the correcting efficiency. Results suggest that for the selected data and transcription conventions, transcribing and correcting still take equally long with 7 times real-time on average. The more complex the primary data, the more time has to be spent on corrections. Despite the impressive latest developments in speech technology, to be a real help for conversation analysts or dialectologists, ASR systems seem to require even more improvement, or we need sufficient and appropriate data for training such systems

    »Ich hab nen Jörg ist einfach leichter zu sagen als ich hab nen Tumor.« Sprachlich realisierte Copingstrategien von Krebspatientinnen und -patienten in digital illness narratives

    No full text
    Durch die gewachsene Bedeutung der Psychoonkologie ist das Themenfeld der Krankheitsverarbeitung (Coping) vermehrt in das Blickfeld der Forschung gerückt. Gleichzeitig entstehen im Web 2.0 neue digitale Formen der intermedialen narrativen Repräsentation von Krankheit, Leid und Krankheitsbewältigung (Cybercoping), wodurch sich für Betroffene neue Möglichkeiten eröffnen, eine Erkrankung durch medienvermittelte Kommunikation und Vergemeinschaftung zu bewältigen und sich eine soziale Identität als chronisch Kranke zu verleihen (vgl. Deppermann 2018). Der Beitrag präsentiert auf theoretischer Basis der Copingforschung sowie der Gesprächsforschung zu narrativer Identitätsbildung eruierte Copingstrategien in Krankheitsnarrativen von Krebspatientinnen und -patienten. Coping wird als kommunikativer Prozess verstanden, der sich in Sprachhandlungen widerspiegelt. Das Untersuchungsmaterial bilden autobiografische Erzählungen in Internetvideos, öffentlich geteilt von zwanzig Betroffenen auf der Social-Media-Plattform YouTube. Copingmechanismen werden in den untersuchten Narrativen in Form von emotionsgeladenen Sprachäußerungen und humoristisch bzw. ironisch gefärbten Sprachhandlungen zur Emotionsregulierung und Entlastung sowie in Gestalt von metaphorischen Deutungsmustern und Personifizierungen der (Tumor-)Erkrankung angezeigt. In den Sprachhandlungen der Erzählenden wird aktives problemorientiertes Coping durch sich selbst und die Community aktivierende Sprache, eine häufig agentivische Selbstdarstellung und -positionierung der Betroffenen und eine durch Positivierung und Neubewertung sinnstiftende Kohärenz sichtbar.Due to the growing importance of psycho-oncology, coping as a mechanism has increasingly gained relevance in research. At the same time, new digital forms of intermedial narrative representation of illness, suffering, and coping emerged in Web 2.0 (cybercoping). This opens up new possibilities for those affected as they can cope through digitally mediated communication in communities and thereby give themselves a social identity as chronically ill persons (cf. Deppermann 2018). On the theoretical basis of coping research and conversation analytical research on narrative identity formation the paper illustrates coping strategies and mechanisms identified in illness narratives of cancer patients. This is based on the assumption that coping is a communicative process that can be reflected in speech acts. The research material consists of autobiographical narratives from twenty affected people in internet videos, publicly shared on the social media platform YouTube. In the narratives, coping mechanisms are displayed in form of emotionally charged speech and humorous or ironic speech acts in order to regulate or relief emotion. Moreover, metaphorical interpretation patterns and personifications of the (tumor) disease can be observed. In the narrators’ speech acts, active problem-oriented coping becomes visible through self- and community-activating speech, a frequent self-representation and self-positioning of the affected persons as active agents and through positivity and re-evaluation of meaning-giving coherence

    Schauraum/Spielraum: Eine standbildbasierte Fallstudie zur Rolle des gebauten Raums in der Interaktion

    No full text
    In dem folgenden Beitrag möchte ich dem Zusammenspiel von gebautem Raum und Interaktion anhand eines Falles nachgehen, in dem die Interaktion in ganz besonderem Maße an den gebauten Raum gebunden ist, in dem sie stattfindet, und zwar derart, dass sie außerhalb dieses speziell für sie hergerichteten, semiotisch besonders reichhaltigen Raums nicht stattfinden kann. Die Rede ist vom gemeinsamen Museumsbesuch

    Das Korpus MIKO. „Mitschreiben in Vorlesungen: Ein multimodales Lehr-Lernkorpus“

    No full text
    Das Projekt Sprache und Studienerfolg bei Bildungsausländer/-innen fokussiert nicht nur die Entwicklung sprachlicher Kompetenzen des Deutschen und ihren direkten Einfluss auf den Studienerfolg internationaler L2-Studierender, sondern untersucht auch zwei ausgewählte, stark sprachgeprägte Handlungen, und zwar das Schreiben von Klausuren und das Mitschreiben in Vorlesungen. Bedarfsanalysen hatten ergeben, dass diese Sprachhandlungen internationale Studierende vor erhebliche Herausforderungen stellen. Zur Untersuchung diesbezüglicher Fragen kam eine Reihe verschiedener Methoden zum Einsatz. Eine wichtige Säule der Untersuchung des Mitschreibens im Studium bildete die Erstellung des nachnutzbaren und öffentlich zugänglichen Korpus MIKO, das in diesem methodologisch ausgerichteten Kapitel gesondert beschrieben wird, während Kapitel 7 stärker inhaltlich auf das Mitschreiben eingeht. MIKO (kurz für: Mitschreiben in Vorlesungen: Ein multimodales Lehr-Lernkorpus) ist ein multimodales, wissenschaftssprachliches Vorlesungskorpus, das sprachlich-fachliche Anforderungen des Mitschreibens in Vorlesungen der Studieneingangsphase fokussiert

    Video-vermittelte Interaktion

    No full text
    Mit dem Aufkommen verschiedener Formen video-vermittelter Interaktion (VMI) ab den 1990er Jahren begann die ethnomethodologisch und konversationsanalytisch ausgerichtete Forschung, damit verbundene neue Kommunikationspraktiken zu untersuchen. Die frühere VMI-Forschung setzte sich hauptsächlich mit hochspezialisierten, institutionellen Settings auseinander und betonte dabei die sensorisch reduzierte und somit potenziell problematische Natur dieser Kommunikationsform. Dies ist im Wesentlichen auch der Schwerpunkt neuerer Forschung geblieben, und bisher berücksichtigen nur wenige Studien auch alltägliche und mobile Formen der VMI oder ihr positives, kommunikativ-soziales Potenzial. Der wachsende Bedarf an Distanz- Kommunikation insbesondere der letzten Jahre hat zu einem verstärkten Interesse an VMI sowohl im beruflichen als auch im privaten Umfeld geführt, so dass dieses Kapitel auch den Beginn einer neuen Ära der interaktionalen VMI-Forschung skizziert. 1 Einleitung 2 Die Entdeckung einer neuen Kommunikationsform 3 Video-vermittelte Interaktion als Teil alltäglicher sozialer Praktiken 4 Eine neue Ära der Forschung zu video-vermittelter Interaktion 5 Fazit: Hybride Zukunft 6 Literatu

    Request for confirmation sequences in German

    Get PDF
    In request for confirmation (RfC) sequences, interlocutors negotiate their social positions regarding access and rights to knowledge. The article presents an overview of a quantitative analysis of 200 RfCs and their responses in German conversations to highlight the relevant linguistic resources speakers of the language deployed to position themselves vis-à-vis a confirmable proposition. In German RfCs, modal particles and tags play an important role in expressing the requester’s epistemic stance; explicit inference marking is used less frequently. Responses usually include response tokens (among others doch as a token specialized for disconfirming negatively formatted RfCs) and an expansion. The article shows that such expansions do important work to tailor the response to the situated informational needs of the requester in a cooperative way beyond the constraints of type-conformity

    9,397

    full texts

    11,061

    metadata records
    Updated in last 30 days.
    Publikationsserver des Instituts für Deutsche Sprache is based in Germany
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇