Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Norwegian Dictionary Norsk Ordbank (New Norwegian) - NO (ELEXIS)

    No full text
    Norsk Ordbank (New Norwegian) is a database of lemma forms together with paradigms, full forms and morpho-syntactic features

    Systematic Dictionary of the Lithuanian Language (ELEXIS)

    No full text
    Sisteminis lietuvių kalbos žodynas. The dictionary is based on the words (their meanings) presented in the Dictionary of the Modern Lithuanian Language, as well as the academic Dictionary of the Lithuanian Language and etc. In addition to general language words, a number of dialectal words are added to it. All the presented words (their meanings) are divided into 10 chapters with their own names

    Dictionary of Lesser Used Slovenian Words (ELEXIS)

    No full text
    Besedišče slovenskega jezika z oblikoslovnimi podatki (po gradivu za slovar sodobnega knjižnega jezika zbrane besede, ki niso bile sprejete v Slovar slovenskega knjižnega jezika). Dictionary of Lesser Used Slovenian Words contains 178457 headwords not included in the Dictionary of the Slovenian Standard Language. Information on inflection, part of speech and source is included in the entries. See also: http://hdl.handle.net/11356/113

    Sanskrit-English Dictionary - Monier-Williams (1899) (ELEXIS)

    No full text
    Monier-Williams Sanskrit-English Dictionary, C-SALT Lex-0 Edition. Etymologically and philologically arranged Sanskrit-English Dictionary with special reference to cognate Indo-European languages by Monier-Williams, Monier, Ernst Leumann and Carl Cappeller (1899). TEI Lex-0 version created by the Cologne Center for eHumanities University of Cologne as part of the Cologne Digital Sanskrit Dictionaries C-SALT Lex-0 2020

    The CLASSLA-StanfordNLP model for UD dependency parsing of standard Bulgarian 1.0

    No full text
    The model for UD dependency parsing of standard Bulgarian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the UD-parsed portion of the BulTreeBank training corpus (http://hdl.handle.net/11495/D93F-C6E9-65D9-2) and using the CoNLL2017 word embeddings (http://hdl.handle.net/11234/1-1989). The estimated LAS of the parser is ~91.5

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0

    No full text
    This model for morphosyntactic annotation of non-standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1210), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1205). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~95.11

    The CLASSLA-StanfordNLP model for lemmatisation of non-standard Serbian 1.1

    No full text
    The model for lemmatisation of non-standard Serbian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200), the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the hr500k training corpus (http://hdl.handle.net/11356/1183) and the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/), using the srLex inflectional lexicon (http://hdl.handle.net/11356/1233). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~97.62. The difference to the previous version of the lemmatizer is that now it relies solely on XPOS annotations, and not on a combination of UPOS, FEATS (lexicon lookup) and XPOS (lemma prediction) annotations

    The CLASSLA-StanfordNLP model for lemmatisation of non-standard Croatian 1.1

    No full text
    The model for lemmatisation of non-standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1183), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), using the hrLex inflectional lexicon (http://hdl.handle.net/11356/1232). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~97.54. The difference to the previous version of the lemmatizer is that now it relies solely on XPOS annotations, and not on a combination of UPOS, FEATS (lexicon lookup) and XPOS (lemma prediction) annotations

    Annotated corpus of Slovenian language-related news articles MetaLangNEWS-Sl

    No full text
    A comprehensive corpus of news articles on the topic of language, published in major Slovenian daily newspapers and news portals in the five-year period of January 1, 2015 - January 1, 2020. The corpus is designed to facilitate research on metalanguage (‘language about language’), linguistic ideologies, language policy and planning, as well as the specific contemporary debates on language defining, naming, and standardisation, ongoing in post-Yugoslav societies. The corpus has been tagged using the CLASSLA-StanfordNLP models for morphosyntactic annotation and lemmatisation of standard Slovenian. The corpus is available in plain text version, XML with full metadata, and tagged CONLL-U format. MetaLangNEWS-Sl is complemented with a separate corpus of citizen metalanguage comments, i.e. online comments to the news articles, available as MetaLangNEWS-COMMENTS-Sl (http://hdl.handle.net/11356/1362). Parallel versions from Croatia (http://hdl.handle.net/11356/1369) and Serbia (http://hdl.handle.net/11356/1371) are also available

    Latin Valency Lexicon - VALLEX (ELEXIS)

    No full text
    Latin VALLEX is a valency lexicon for Latin. It was built in close connection with the semantic/pragmatic annotation of the Index Thomisticus Treebank and the Latin Dependency Treebank. Data are stored in a single XML file, whose structure is the same of that for the valency lexicon for Czech PDT-VALLEX

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇