Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Dialogue act annotated spoken corpus GORDAN 1.0 (audio/video)

    No full text
    The GORDAN 1.0 corpus contains authentic data of spoken communication, annotated for dialogue acts. This entry contains the complete audio files of the corpus (seven wav files, 1 hour of recording), and video files (four mp4 video files). Video files are provided only for the recordings where video was available in the original dataset (datasets GOS, Gos Videolectures and BERTA). This resource contains only the recordings of the corpus, while the annotated transcriptions are available at http://hdl.handle.net/11356/1291

    The CLASSLA-StanfordNLP model for named entity recognition of non-standard Serbian 1.0

    No full text
    This model for named entity recognition of non-standard Serbian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200), the hr500k training corpus (http://hdl.handle.net/11356/1183), the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240) and the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), using the CLARIN.SI-embed.sr word embeddings (http://hdl.handle.net/11356/1206). The training corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed

    Annotated corpus of Slovenian language-related news comments MetaLangNEWS-COMMENTS-Sl

    No full text
    A comprehensive corpus of user comments on online news articles on the topic of language from major Slovenian daily newspapers and news portals, published in the five-year period of January 1, 2015 - January 1, 2020. The corpus is designed to facilitate research on metalanguage (‘language about language’), linguistic ideologies, language policy and planning, as well as the specific contemporary debates on language defining, naming, and standardisation, from the bottom-up perspective. The corpus has been tagged using the CLASSLA-StanfordNLP models for morphosyntactic annotation and lemmatisation of non-standard Slovenian. The corpus is available in plain text version, XML with full metadata, and tagged CONLL-U format. This collection is complementary to the corpus of news articles MetaLangNEWS-Sl (http://hdl.handle.net/11356/1360). Parallel versions from Croatia (http://hdl.handle.net/11356/1370) and Serbia (http://hdl.handle.net/11356/1372) are also available

    List of word relations from the Sloleks 2.0 lexicon 1.0

    No full text
    This entry consists of a TSV file containing a list of 66,347 Slovene word pairs from the Sloleks Morphological Lexicon of Slovene (v2.0; http://hdl.handle.net/11356/1230) that have been automatically identified as morphologically related according to a number of manually designed morphological relation rules (e.g. "dež" -> "deževen", "pisati" -> "pisatelj", "prijatelj" -> "prijateljica"). Each line in the list contains the following columns: - original lemma (e.g. "pisati"), - related lemma (e.g. "pisatelj"), - original lemma, automatically deconstructed into individual word parts (e.g. "pis_ati"), - related lemma, automatically deconstructed into individual word parts (e.g. "pis_at_elj"), - MTE-6 lexical features of the original lemma (e.g. "G"),* - MTE-6 lexical features of the related lemma (e.g. "Som"),* - ID of the original lemma from Sloleks 2.0, - ID of the related lemma from Sloleks 2.0, - the overlapping or central part (common to both the original and the related lemmas; e.g. "pis") - the ID of the morphological relation rule used to identify the relation (e.g. "G.Som.5.2.1"), - the morphological relation rule (e.g. "[G]_ati -> [G]_at_elj"). * MTE-6 refers to MULTEXT-East Version 6 morphosyntactic specifications for Slovenian, available at http://nl.ijs.si/ME/V6/ Each rule constitutes a pattern to form a morphological relation. For instance, "[G]_ati -> [G]_at_elj" indicates that a verb (G) ending with the word part "ati" is related to the lemma formed by replacing "_ati" with "_at_elj". Note that the list contains no proper nouns and no relations for 38 morphological rules that have been included in the hierarchy of rules (listed in the accompanying file nssss_sloleks_word_relation_rules.tsv), but need to take into account additional rules that have not yet been implemented in the current version of the extraction process (such as irregular conversions in overlapping word parts: "gri_sti" - "griz_enj_e", "sneg" - "snež_ak")

    Lemma list of the Dictionary of the Danish Language - ODS (ELEXIS)

    No full text
    Ordbog over det Danske Sprog (ODS), lemma list. Contents and format: This list contains the headwords of the online version of ODS (and ODS-S) (ordnet.dk/ods). ODS describes the Danish language from 1700 till 1950. The elements are: headword (attributes: entryid, homno (if present)), POS. If an entry has headword variants, every form is listed in it’s own row. These forms are placed alphabetically, but share the ID (as they origin from the same ODS entry). The forms given in this list follow the spelling principles used in ODS. For example å is rendered through aa and nouns are capitalised. Remarks: The forms in the list cannot be expected to follow the official guidelines for Danish orthography, as the dictionary was published between 1918 and 2005. Furthermore some of the variants (historical, regional) described in the entries are included in the list. Special characters sometimes occur as the are printed in the book, sometimes in a normalised form, e.g.: ç → ç/c, ô → ô/o, ï → ï/i. The list reflects the ODS/ODS-S headwords (and their POS, ID, homograph number etc.) at the time of the latest list update. This information is subject to change in later versions (as corrections are continuously being made in the online version). The POS inventory is as in ODS/ODS-S. However, for nouns en and et are changed to sb

    IT-VaLex (ELEXIS)

    No full text
    IT-VaLex is a collection of verbal lexical entries enhanced with valency and subcategorization frames at a syntactic level. IT-VaLex is closely related to the Index Thomisticus Treebank, since it is a corpus-driven valency lexicon automatically induced from the syntactic layer of annotation of the Index Thomisticus Treebank. In the Index Thomisticus Treebank annotation style, verbal arguments are those to which the following tags are assigned: Sb (Subject), Obj (Object), OComp (Object Complement) and Pnom (Predicate Nominal). The difference between IT-VaLex and the valency lexicon Latin VALLEX is that the former is derived from the syntactically annotated data of the Index Thomisticus Treebank, while the latter is connected to the semantic/pragmatic annotation of both Latin treebanks (plus a number of hard-coded entries)

    Basque Monolingual Dictionary - Sarasola (1996) (ELEXIS)

    No full text
    Ibon Sarasola: Euskal Hiztegia, the 1996 edition. By the time of this edition, the most complete monolingual Basque dictionary

    CroSloEngual BERT

    No full text
    Trilingual BERT (Bidirectional Encoder Representations from Transformers) model, trained on Croatian, Slovenian, and English data. State of the art tool representing words/tokens as contextually dependent word embeddings, used for various NLP classification tasks by finetuning the model end-to-end. CroSloEngual BERT are neural network weights and configuration files in pytorch format (ie. to be used with pytorch library)

    The CLASSLA-StanfordNLP model for named entity recognition of standard Serbian 1.0

    No full text
    This model for named entity recognition of standard Serbian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200) and using the CLARIN.SI-embed.sr word embeddings (http://hdl.handle.net/11356/1206)

    The CLASSLA-StanfordNLP model for named entity recognition of standard Bulgarian 1.0

    No full text
    This model for named entity recognition of standard Bulgarian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the BulTreeBank training corpus (http://hdl.handle.net/11495/D93F-C6E9-65D9-2) and using the CoNLL2017 word embeddings (http://hdl.handle.net/11234/1-1989)

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇