1,721,036 research outputs found

    Overview of the EvaLatin 2020 Evaluation Campaign

    No full text
    This paper describes the first edition of EvaLatin, a campaign totally devoted to the evaluation of NLP tools for Latin. The two shared tasks proposed in EvaLatin 2020, i. e. Lemmatization and Part-of-Speech tagging, are aimed at fostering research in the field of language technologies for Classical languages. The shared dataset consists of texts taken from the Perseus Digital Library, processed with UDPipe models and then manually corrected by Latin experts. The training set includes only prose texts by Classical authors. The test set, alongside with prose texts by the same authors represented in the training set, also includes data relative to poetry and to the Medieval period. This also allows us to propose the Cross-genre and Cross-time subtasks for each task, in order to evaluate the portability of NLP tools for Latin across different genres and time periods. The results obtained by the participants for each task and subtask are presented and discussed

    Let's Do It Orderly: A Proposal for a Better Taxonomy of Adverbs in Universal Dependencies, and Beyond

    No full text
    With this paper, we hope to highlight the critical issues and “bad practices” that are currently observed in the representation of adverbs in the annotation framework of Universal Dependencies (where they are ideally identified with the universal part of speech ADV), which themselves more generally mirror the conspicuous lack of systematic definitions of this word class in traditional grammars. This fact, on the one side, hampers a useful and meaningful linguistic description of adverbs in Universal Dependencies’ treebanks and beyond, and, on the other side, has unfortunate consequences for linguistic research, such as too confused, impractical results when, for example, querying a treebank for all ADV-tagged words. Therefore, with the aid of minimal data analysis and numerous examples, this contribution tries to raise awareness about this issue, and proposes a revised, more typologically grounded and general framework for the classification of adverbs in Universal Dependencies, and more broadly advocates a more flexible representation of the interplay between word classes and syntactic functions, through the comprehensive concept of “transposition” or “transfer”. While grounded in the Universal Dependencies formalism, the scope of the discussion of this paper is by no means limited to it, and might be of interest to any practitioner of (computationally-oriented) linguistic annotation

    A comparison of graph-based word sense induction clustering algorithms in a pseudoword evaluation framework

    No full text
    This article presents a comparison of different Word Sense Induction (wsi) clustering algorithms on two novel pseudoword data sets of semantic-similarity and co-occurrence-based word graphs, with a special focus on the detection of homonymic polysemy. We follow the original definition of a pseudoword as the combination of two monosemous terms and their contexts to simulate a polysemous word. The evaluation is performed comparing the algorithm’s output on a pseudoword’s ego word graph (i.e., a graph that represents the pseudoword’s context in the corpus) with the known subdivision given by the components corresponding to the monosemous source words forming the pseudoword. The main contribution of this article is to present a self-sufficient pseudoword-based evaluation framework for wsi graph-based clustering algorithms, thereby defining a new evaluation measure (top2) and a secondary clustering process (hyperclustering). To our knowledge, we are the first to conduct and discuss a large-scale systematic pseudoword evaluation targeting the induction of coarse-grained homonymous word senses across a large number of graph clustering algorithms

    UDante: First Steps Towards the Universal Dependencies Treebank of Dante's Latin Works

    Get PDF
    This paper presents the early stages of the development of a new treebank containing all of Dante Alighieri’s Latin works. In particular, it describes the conversion of the original TEI-XML files to CoNLL-U, the creation of a gold standard, the process of training four annotators and the evaluation of the syntactic annotation in terms of inter-annotator agreement and LA, UAS and LAS. The aim is to release a new resource, in view of the celebrations for the 700th anniversary of Dante’s death, which can support the development of the Vocabolario Dantesco

    Overview of the EvaLatin 2022 Evaluation Campaign

    No full text
    This paper describes the organization and the results of the second edition of EvaLatin, the campaign for the evaluation of Natural Language Processing tools for Latin. The three shared tasks proposed in EvaLatin 2022, i. e. Lemmatization, Part-of-Speech Tagging and Features Identification, are aimed to foster research in the field of language technologies for Classical languages. The shared dataset consists of texts mainly taken from the LASLA corpus. More specifically, the training set includes only prose texts of the Classical period, whereas the test set is organized in three sub-tasks: a Classical sub-task on a prose text of an author not included in the training data, a Cross-genre sub-task on poetic and scientific texts, and a Cross-time sub-task on a text of the 15th century. The results obtained by the participants for each task and sub-task are presented and discussed

    Hell Awaits: Building a Universal Dependencies Treebank for Dante Alighieri’s Comedy

    No full text
    In this paper, we describe the creation of a treebank for Dante’s Comedy in Universal Dependencies, the first syntactically annotated text for Old Italian following a dependency-based paradigm. We detail the phase of treebanking the first part of the Comedy, the Inferno, and we discuss some annotation issues, specifically ellipses and comparative structures. Then, we perform an evaluation of automated dependency parsing with models trained on the currently available annotated portion of the text

    UDante. L’annotazione sintattica dei testi latini di Dante

    No full text
    L’articolo descrive il lavoro di realizzazione di UDante, il corpus dei testi latini di Dante Alighieri annotato a livello sintattico in base ai criteri stabiliti dall’iniziativa Universal Dependencies. Dopo avere introdotto e motivato lo stile di annotazione adottato, l’articolo presenta nel dettaglio le fasi di costruzione di UDante, soffermandosi particolarmente sul processo di conversione del formato dei dati e sulla loro annotazione manuale. Viene, quindi, descritta l’integrazione dei testi di UDante nella knowledge base di LiLa, grazie a cui il corpus sintattico dei testi latini di Dante è reso interoperabile con altre risorse linguistiche per il latino. Infine, alcuni esempi d’interrogazione di UDante sono riportati con l’obiettivo di dimostrarne l’utilità in termini di supporto alla compilazione del Vocabolario Dantesco Latino

    Highway to Hell. Towards a Universal Dependencies Treebank for Dante Alighieri’s Comedy

    No full text
    In this paper, we describe the creation in Universal Dependencies of a treebank for Dante’s Comedy, the first syntactically annotated text for Old Italian following a dependency-based schema. We detail the phase of treebanking the first part of the Comedy, the Inferno, and we describe some annotation issues. Then, we perform an evaluation of automated dependency parsing with models trained on the currently available annotated portion of the text

    Verbs in -sc- between Inflection and Derivation. Lexicographic Representation and Theoretical Issues

    No full text
    Latin verbs in -sc- have long intrigued linguists, who have explored their semantics and derivational history. However, their classification as either inflectional or derivational has received much less attention. This paper delves into the lexicographic representation of these verbs and its implications for annotation practices, and critically examines the inflection-derivation distinction, focusing on the remarkable fact that perfectum forms are shared by both sc-verbs and their counterparts without -sc-. This suggests a potential shift toward considering forms in -sc- as part of inflection rather than derivation. Additional evidence in favour of such a view is provided by the fact that, semantically, the ‐sc- suffix appears to relate more to (grammatical) aspect than Aktionsart. The paper discusses these complexities and offers alternative representation options that align with recent theoretical proposals and have practical advantages

    Verba Bestiae: How Latin Conquered Heavy Metal

    No full text
    The presence of Latin in heavy metal music ranges from full texts, intros, song and album titles to band names, pseudonyms, and literary quotations. This chapter sheds light on heavy metal’s fascination with the history and ‘arcane’ sound of Latin, and investigates its patterns of use in lyrics with the help of Natural Language Processing tools and digitally-available linguistic resources. First, the authors collected a corpus of lyrics containing differing amounts of Latin and enhanced it with descriptive metadata. Next, the authors calculated the richness of the vocabulary and the distribution of content words. The authors processed the corpus with a morphological analyser and performed both a manual and a computational search for intertextuality, including allusions, paraphrase and verbatim quotations of literary sources. The authors show that, despite it being a dead language, Latin is very frequently used in metal. Its historical status appears to fascinate bands and lends itself well to those religious, epic and mysterious themes so characteristic of the heavy metal world. The widespread use of Latin in metal lyrics, however, sees many bands simply reusing Latin texts – mostly from the Bible – or even misspelling literary quotations
    corecore