Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Ladin-German Dictionary (Letter F) (ELEXIS)

    No full text
    Dizionar Ladin-Deutsch (Letter F). This dataset contains the letter F of the Ladin-German dictionary by Giovanni Mischí

    Lemma list of the Dictionary of Contemporary Portuguese - DLPC (ELEXIS)

    No full text
    Dicionário da Língua Portuguesa Contemporânea (DLPC) is a monolingual Portuguese dictionary published by Academia das Ciências de Lisboa (2001). This dictionary also represents the first complete edition of a Portuguese Academy dictionary, from A to Z. The DLPC contains around 70,000 entries. This dictionary is being updated. The dictionary was accessible on-line at www.acad-ciencias.pt/academia/illlp-novo-dicionario. However, this URL no longer works

    Dictionary of Bavarian Dialects of Austria - WBOE (ELEXIS)

    No full text
    Wörterbuch der Bairischen Mundarten in Österreich. Automatically created XML/TEI versions of first 5 volumes of Dictionary of Bavarian Dialects in Austria, published between 1963 and 2015, covering A, B/P, C, D/T and E. The dictionary of Bavarian dialects in Austria (WBÖ) is a large-scale dialect dictionary in which the Bavarian vocabulary in Austria and the neighboring South Tyrol is documented. The first 5 volumes also take into account the historical language of the office and other Bavarian-speaking areas of the former monarchy

    Bilingual Corpus of Underground Mining (ELEXIS)

    No full text
    PodzemniRadovi-sr-en, dvojezični poravnati korpus radova iz oblasti rudarstva. Undeground-mining-sr-en: bilingual texts from the Underground Mining Engineering journal (55 papers from 8 issues), aligned at the sentence level (4831 translation units), represented in the TMX (Translation Memory eXchange) format

    GeologyTerm (ELEXIS)

    No full text
    Podskup GeolISSTerm delimično poravnat sa GeoSciML. Subsets of GeolISSTerm - Dictionary of geological terms and terms used in Geological information system of Serbia (GeolISS) partially aligned with GeoSciML

    Treq Translation Equivalents (ELEXIS)

    No full text
    Data for Treq interface 2.0 derived from the InterCorp parallel corpus release 12

    Basque Dictionary (ELEXIS)

    No full text
    Hiztegi Batua. Standard Basque Dictionary 2010 TEI-XML release

    List of formulaic sequences in standard written Slovenian

    No full text
    This document contains 1,891 formulaic sequences in standard written Slovenian, i.e. frequently recurring strings of two to five words, manually annotated for syntactic structure, pragmatic function, and dictionary relevance. The list of sequences with a minimum frequency threshold of 20/million is based on the Frequency lists of word-level n-grams from lowercase word forms in Gigafida 2.0 (http://hdl.handle.net/11356/1274) and contains the union of top-1,000 formulaic sequences ranked by frequency and five association measures (Dice, t-test, MI, MI3, simple-LL). Note that there exists a related entry "List of formulaic sequences in spoken Slovenian", http://hdl.handle.net/11356/1279

    The CLASSLA-StanfordNLP model for lemmatisation of standard Croatian 1.1

    No full text
    The model for lemmatisation of standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1183) and using the hrLex inflectional lexicon (http://hdl.handle.net/11356/1232). The estimated F1 of the lemma annotations is ~97.6. The difference to the previous version of the model is that it is trained with the lemmatiser padding bug removed, cf. https://github.com/stanfordnlp/stanfordnlp/issues/143

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Bulgarian 1.0

    No full text
    This model for morphosyntactic annotation of standard Bulgarian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the BulTreeBank training corpus (http://hdl.handle.net/11495/D93F-C6E9-65D9-2) and using the CoNLL2017 word embeddings (http://hdl.handle.net/11234/1-1989). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~96.8

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇