Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Parallel Corpus (EN-FR-LT) of EU Financial Documents (ELEXIS)

    No full text
    Parallel corpus is comprised of 154 EU legislative documents (English documents and their translations into French and Lithuanian) related to various financial issues and enacted in the period . The documents were extracted from the Official Journal of the European Union available in the open EUR-Lex database. All documents were converted into plain text format and sentence-aligned. The size of the corpora is 1,006,485 words in English, 1,181,647 words in French, 803,845 words in Lithuanian. See also: http://hdl.handle.net/20.500.11821/3

    eSSKJ: Dictionary of the Slovenian standard language (semantic database): school edition

    No full text
    The eSSKJ: Dictionary of the Slovenian Standard Language (semantic database): Educational Edition represents the semantic part of the entries for the purpose of educational language portal Franček. The underlying data stems from the eSSKJ: Dictionary of the Slovenian Standard Language. The content is adjusted for primary- and secondary-school users. The dataset is linked to the "Franček Portal Headword List" (http://hdl.handle.net/11356/1445)

    Corpus of term-annotated texts RSDO5 1.1

    No full text
    The RSDO5 corpus was compiled in order to serve as a training set for automatic term identification. It consists of 12 texts with 250,000 words and almost 38,000 manually annotated terms, each marked to be either in- or out-domain. The corpus texts were published between 2000 and 2019, are either PhD theses (3), a scientific book based on a PhD thesis (1), graduate level text books (4), or journal articles (4) and belong to the fields of biomechanics (3), linguistics (3), chemistry (3), or veterinary science (3). Apart from the manually annotated terms, the corpus was automatically annotated with Universal Dependencies annotations, i.e. tokenisation, sentence segmentation, lemmatisation, morpological features and dependency syntax. As opposed to the previous version, this one adds in- and out-domain marking on terms in the TEI and vertical files

    Basic vocabulary of The Danish Dictionary - DDO (ELEXIS)

    No full text
    Den Danske Ordbog (DDO), Basic vocabulary. DDO describes the vocabulary of modern Danish from 1950 till present on the basis of a large text corpus. Contents: 5.122 entries 4.474 entries holding senses linked to WordNet base concepts in connection with the DanNet project. 648 of the most comprehensive DDO entries not included in the group mentioned above This resource contains the basic content of the basic vocabulary of the online version of DDO (ordnet.dk/ddo). The data has been processed using Elexifier. Information types included (if present). Elements are listed under their source names = the element names in m:e=""...” in the uploaded, Elexified file: headword headword2: Alternative official spelling(s) of the headword POS sense (1 or more) definition (1 or more (#2ff = subsenses)) reference (if the DDO entry has no definition of its own) example: 1st example of the DDO sense english: English translation(s) extracted via DanNet (if a link exists) synonym antony

    PASSWORD English Multilingual Dictionary - KEMD (ELEXIS)

    No full text
    An English multilingual dictionary including a translation equivalent for each sense of the English entry in 42 languages

    Word list of the General Corpus of Lithuanian Language (ELEXIS)

    No full text
    The dataset was compiled in 2016 and is based on the Corpus of the Contemporary Lithuanian Language, version tekstynas.vdu.lt (139 m tokens). It contains wordforms (types) with tokens from the monolingual, general language corpus of Lithuanian. See also: http://hdl.handle.net/20.500.11821/

    Palatinate Dictionary - PfWB (ELEXIS)

    No full text
    Pfälzisches Wörterbuch. The Palatinate Dictionary lists the entire dialectal vocabulary of the Palatinate in use today

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Bulgarian 1.1

    No full text
    This model for morphosyntactic annotation of standard Bulgarian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the BulTreeBank training corpus (http://hdl.handle.net/11495/D93F-C6E9-65D9-2) and using the CoNLL2017 word embeddings (http://hdl.handle.net/11234/1-1989). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~96.8. The difference to the previous version of the model is that the pre-trained embeddings are limited to 250 thousand entries and adapted to the new code base

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Macedonian 1.1

    No full text
    This model for morphosyntactic annotation of standard Macedonian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the 1984 training corpus (to be published) and using the Macedonian CLARIN.SI word embeddings (http://hdl.handle.net/11356/1359). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~97.6. The difference to the previous version of the model is that the pre-trained embeddings are limited to 250 thousand entries and adapted to the new code base

    Corpus of Croatian news portals ENGRI (2014-2018)

    No full text
    The corpus consists of texts collected from the most popular (based on the Reuters Institute Digital News Report for 2018, retrieved from http://www.digitalnewsreport.org in April, 2019) news portals in Croatia in the period from 2014 to 2018: Direktno, Dnevno, Net Hr, Hrt, Index_Hr, Jutarnji, Novilist, Rtl, SlobodnaDalmacija, Večernji, Tportal, Dnevnik. Web browsing and web crawling were used to select and store the texts with their useful HTML information (publication date of the article, its URL, and title). The linguistic processing of the corpus was performed with the CLASSLA package (https://pypi.org/project/classla/) on the levels of tokenization, sentence splitting, morphosyntactic tagging, lemmatization, dependency parsing and named entity recognition. This corpus is a linguistically-processed version of the original corpus published at https://repository.pfri.uniri.hr/islandora/object/pfri%3A2156 and is distributed in the CoNLL-U format (https://universaldependencies.org/format.html)

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇