Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Renish Dictionary + Supplements to the Renish Dictionary - RhWB (ELEXIS)

    No full text
    Rheinisches Wörterbuch + Nachträge zum Rheinischen Wörterbuch. With its nine volumes, the Rhenish Dictionary (1928-1971) is the most comprehensive dialect dictionary of West Central German and documents the dialectal vocabulary of the area of the former Prussian Rhine Province. Please note that the Rhenish Dictionary consists of two parts: the main dictionary and the Supplements to the Renish Dictionary

    Slovene Web genre identification corpus GINCO 1.0

    No full text
    The Slovene Web genre identification corpus GINCO 1.0 contains web texts, manually annotated with genre, from two Slovene web corpora, the slWaC 2.0 corpus, crawled in 2014, and a web corpus, crawled in 2021 in the scope of the MaCoCu project. The corpus allows for automated genre identification and genre analyses as well as other web corpora research, and comprises two parts: - subcorpus of suitable texts, containing 1002 texts (478,969 words), manually annotated with 24 genre categories (News/Reporting, Announcement, Research Article, Instruction, Recipe, Call (such as a Call for Papers), Legal/Regulation, Information/Explanation, Opinionated News, Review, Opinion/Argumentation, Promotion of a Product, Promotion of Services, Invitation, Promotion, Interview, Forum, Correspondence, Script/Drama, Prose, Lyrical, FAQ (Frequently Asked Questions), List of Summaries/Excerpts, and Other) - subcorpus of unsuitable texts, containing 123 texts (173,778 words), discarded as not suitable for genre annotation due to reasons, encoded by the labels (Machine Translation, Generated Text, Not Slovene, Encoding Issues, HTML Source Code, Boilerplate, Too Short/Incoherent, Too Long (longer than 5,000 words), Non-Textual (no full sentences, e.g. tables, lists), and Multiple texts). The texts in the suitable subset are annotated with up to three genre categories, where the primary label is the most prevalent, and secondary and tertiary labels denote presence of additional genre(s). They are encoded in three levels of detail, allowing experiments with the full set (24 labels), set of 21 labels (labels with less than 5 instances are merged with label Other) and set of 12 labels (similar labels are merged). Additionally, the corpus contains some metadata about the text (e.g. url, domain, year) and its paragraphs (e.g. near-duplicates and their usefulness for the genre identification)

    Corpus of term-annotated texts RSDO5 1.0

    No full text
    The RSDO5 corpus was compiled in order to serve as a training set for automatic term identification. It consists of 12 texts with 250,000 words and almost 38,000 manually annotated terms. The corpus texts were published between 2000 and 2019, are either PhD theses (3), a scientific book based on a PhD thesis (1), graduate level text books (4), or journal articles (4) and belong to the fields of biomechanics (3), linguistics (3), chemistry (3), or veterinary science (3). Apart from the manually annotated terms, the corpus was automatically annotated with Universal Dependencies annotations, i.e. tokenisation, sentence segmentation, lemmatisation, morpological features and dependency syntax

    The CLASSLA-StanfordNLP model for lemmatisation of standard Slovenian 1.3

    No full text
    The model for lemmatisation of standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and using the Sloleks inflectional lexicon (http://hdl.handle.net/11356/1230). The estimated F1 of the lemma annotations is ~99.7. The difference to the previous version is that the internal lexicon is built on the lexicon training data only, and not on the (automatically XPOS-annoteted) corpus data

    Glossary (EN-LT-NO) of Terms Denoting Phobia Types (ELEXIS)

    No full text
    Trilingual (EN-LT-NO) glossary of terms denoting phobia types extracted from the articles of English "The Guardian", Lithuanian "DELFI", and Norwegian "Dagbladet" news media sites

    The German Dictionary by Jacob and Wilhelm Grimm (first edition) - DWB (ELEXIS)

    No full text
    Deutsches Wörterbuch von Jacob Grimm und Wilhelm Grimm (Erstbearbeitung). Deutsches Wörterbuch by Jacob and Wilhelm Grimm is the largest and most comprehensive dictionary of the German language (1450-1960). The first volume was published in 1854, the last more than 120 years later in 1961

    Kalkar's Dictionary (ELEXIS)

    No full text
    Kalkars Ordbog. Ordbog til det ældre danske sprog. Otto Kalkar's "Dictionary of older Danish" – or just "Kalkar's Dictionary" is a dictionary of the early stages of the Danish language, covering the period 1300-1700. It was first published in four volumes 1881-1907 with a supplementary volume in 1918. It was retrodigitized in 2017. The main information types of a full entry are: Headword (with variants) Part of speech Danish definitions or usage descriptions Quotations with sources Many entries hold only a small subset of this information. Minimal entries (especially from the supplementary volume), e.g. entries just listing a reference or an extra evidence are left out in the ELEXIS upload. The printed dictionary has a nested structure where derivations and compounds are listed under a main headword. In the digital version, main lemmas and sub-lemmas are all marked up as headwords at the same level

    GLOBAL French-Greek Dictionary - MLDS (ELEXIS)

    No full text
    A general language French to Greek dictionary

    GLOBAL French-Japanese Dictionary - MLDS (ELEXIS)

    No full text
    A general language French to Japanese dictionary

    Training corpus ssj500k 2.3

    No full text
    The ssj500k training corpus contains about 500,000 tokens manually annotated on the levels of tokenisation, sentence segmentation, morphosyntactic tagging, and lemmatisation. About half of the corpus is also manually annotated with syntactic dependencies, named entities, and verbal multiword expressions. About a quarter of the corpus is also annotated with semantic role labels. The morphosyntactic tags and syntactic dependencies are included both in the JOS/MULTEXT-East framework, as well as in the framework of Universal Dependencies. The annotations of the ssj500k corpus follow (1) the MULTEXT-East V6 morphosyntactic specifications for Slovene, http://nl.ijs.si/ME/V6/msd/, (2) the JOS dependency schema, http://nl.ijs.si/jos/bib/jos-skladnja-navodila.pdf, the Universal Dependencies morphosyntactic specifications and syntactic dependencies for Slovene-SSJ, https://universaldependencies.org/, (4) the Janes annotation guidelines for Slovenian named entities, http://nl.ijs.si/janes/wp-content/uploads/2017/09/SlovenianNER-eng-v1.1.pdf, and (5) the Guidelines of the PARSEME shared task on verbal multiword expressions, http://parsemefr.lif.univ-mrs.fr/parseme-st-guidelines/1.1/ The vocabulary of (1) and (2) is provided in the back element and (3), (4), and (5) in the teiHeader of the TEI encoded corpus. The semantic role labels are also documented in the teiHeader. In contrast to the previous version 2.2, this version includes the corrected Universal Dependencies relations from UD version 2.8, updates the TEI encoding and adds UD annotations to the vertical file

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇