Berlin-Brandenburg Academy of Sciences and Humanities
Berlin Brandenburgischen Akademie der Wissenschaften DokumentenserverNot a member yet
4539 research outputs found
Sort by
Thesaurus Linguae Aegyptiae: Text Trees of Corpus v20 (2025)
Overview of the hierarchy trees of text objects and texts in the Thesaurus Linguae Aegyptiae, as of Corpus Edition 20, 2025Übersicht über die Hierarchiebäume von Textobjekten und Texten im Thesaurus Linguae Aegyptiae, Stand Corpus-Ausgabe 20, 202
Bericht der Kommission für den Thesaurus linguae Latinae über die Zeit vom 1. April 1915 bis zum 31. März 1916
Studien zur vergleichenden Grammatik der Türksprachen : 2. Stück. Über das Verbum al- "nehmen"als Hilfszeitwort
Choosing Suitable Text Corpora for Identifying Collocations – A Case Study of a Large Reference Dictionary of Contemporary German
This paper investigates the impact of corpora for extracting collocation candidates from large text corpora. We compare a variety of corpora, including best-selling fiction and non-fiction literature, a contemporary German reference corpus, the German Wikipedia, a curated web page monitor corpus, and a large newspaper corpus. All corpora undergo processing with an NLP pipeline for morphological and syntactic annotations. Collocation candidates are then extracted using dependency parse tree patterns.
For our evaluation, we utilized three gold standard collocation datasets for contemporary German with a total amount of appr. 200,000 collocations. Our findings confirm that well-curated, medium-sized corpora, diverse in text types, offer superior coverage compared to opportunistically collected corpora of equivalent size. We could also confirm that well curated newspaper texts perform better than pure web corpora of the same size. Additionally, we
observed significant variation in coverage among the examined corpora, depending on specific syntactic relations.
Future work will focus on training models to classify collocation candidates, leveraging the advancements in LLMs. Further research into larger, more diverse datasets for both training and evaluation could improve collocation ranking and candidate discovery