Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    The CLASSLA-Stanza model for semantic role labeling of standard Slovenian

    No full text
    The model for semantic role labeling of standard Slovenian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1434). The estimated F1 of the semantic role annotations is ~77.2

    Corpus of Romanian Academic Genres ROGER

    No full text
    The corpus contains academic papers from eight disciplines, written by the Romanian students in native Romanian and English L2. The corpus was collected over a three-year period (2018-2021) with the help of 27 collaborators from nine Romanian universities. The corpus is available for online querying through a dedicated platform developed at the CODHUS research centre from the West University of Timisoara

    Dictionary of Croatian Idioms (ELEXIS)

    No full text
    Frazeološki rječnik hrvatskoga jezika is an open-access dictionary of Croatian idioms based on data from a large electronic corpus. The resulting dictionary will serve as a gateway for a large number of users and researchers to current idiomatic usage of Croatian

    Dictionary of Southern Serbian Dialects (ELEXIS)

    No full text
    Речник говора јужне Србије (Dictionary of Southern Serbian Dialects) is a dialectologica dictionary by Momčilo Zlatanović, which was first published in 1998

    Serbian-English Terminology in the Power Engineering Domain - SrpEngPE (ELEXIS)

    No full text
    SrpEngPE - Serbian-English dictionary with terminology in the power engineering domain: automatically extracted from domain parallel corpus http://jerteh.rs/biblisha/ListaDokumenata.aspx?JCID=13&lng=en , extracted monolingual term candidates and translation pairs were manually evaluated and post-edited

    Bilingual List of German-Serbian Translated Pairs of Lexical Units - SrpNemLex (ELEXIS)

    No full text
    SrpNemLex - Bilingual list of German-Serbian translated pairs of lexical units: automatically extracted from parallel corpus that contains 14 novels http://jerteh.rs/biblisha/ListaDokumenata.aspx?JCID=11&lng=en , extracted monolingual candidates and translation pairs were manually evaluated and post-edited

    Slovene corpus for general relation extraction SloREL 1.0

    No full text
    The SloREL corpus contains annotations for training relation extraction models on Slovene documents. It contains documents from Slovene Wikipedia with annotated entities and relations. We constructed the annotations using a semi-supervised process based on linking the documents to the WikiData knowledge graph. The corpus contains 244,437 sentences from Slovene Wikipedia pages. We also provide 896 additional sentences collected from the 24ur.com news website with annotated and linked entities, which do not contain annotated relations and are meant for additional testing of the models. The entities in our corpus are linked to the entities in the WikiData knowledge graph which is useful for models that take advantage of additional knowledge from a knowledge graph. All together the corpus comprises 245,333 sentences with 813,952 relations and 1,616,193 entities. The corpus comprises of multiple documents: - schema-definition.xsd: defines the structure of the xml documents containing relation annotations. - wikipedia-train.xml: training portion of the wikipedia corpus - wikipedia-test.xml: testing portion of the wikipedia corpus - wikipedia-validation.xml: validation portion of the wikipedia corpus - 24ur.xml: additional sentences from the 24ur.com news article

    Terminological dictionary of artificial intelligence

    No full text
    The terminological dictionary was compiled within the framework of the project Development of Slovene in the Digital Environment. It is an example collection of 413 terms from the field of artificial intelligence, especially from the subfields of machine learning, computer vision, natural language processing, and fuzzy logic. Definitions, English equivalents, and possible synonyms are added to the terms. The dictionary is based on a conceptual approach, according to which terms are perceived as designations for concepts that are related to each other in the conceptual system of the subject field. Consequently, the terms are interrelated in the naming system of the subject field. The dictionary is distributed in XML using the TBX (TermBase eXchange) standard for representing and exchanging information from termbases

    CMC training corpus Janes-Tag 3.0

    No full text
    Janes-Tag is a manually annotated corpus of Slovene Computer-Mediated Communication (CMC) consisting of about 15,000 short texts (190,000 words), mostly tweets but also blogs, forums and news comments. The corpus is meant as a gold-standard training and testing dataset for tokenisation, sentence segmentation, word normalisation, morphosyntactic tagging, lemmatisation and named entity annotation of non-standard Slovene. As the corpus has been carefully manually annotated, it is also suitable for detailed linguistic explorations which require highly accurate and reliable annotations. The corpus is composed of two parts, the older (texts to 2016) and smaller (65,000 words) Janes Tag 2.1, and the tweet-only newer (2022, 125,000 words) Janes RSDO. Only the Janes Tag 2.1 part is annotated with named entities and with classification of the texts according to their estimated technical (T1-T3) and linguistic (L1-L3) standardness. The data is available in the source TEI encoding and in derived CoNLL-U format. Both contain JOS/MULTEXT-East morphosyntactic descriptions as well as Universal Dependencies morphological features. Compared to the previous version, this one corrects some errors, updates the encoding, and adds Janes-RSDO. The first version of this corpus is described in: FIŠER, Darja, LJUBEŠIĆ, Nikola, ERJAVEC, Tomaž. 2020. The Janes project: language resources and tools for Slovene user generated content. Language Resources and Evaluation. https://doi.org/10.1007/s10579-018-9425-z Note that a related corpus, Janes-Norm 3.0 (http://hdl.handle.net/11356/1733), is also available. It contains Janes-Tag 3.0 and an additional subcorpus with manually checked sentences, tokens and normalised words but only automatically assigned lemmas and MULTEXT-East MSDs

    Slovenian parliamentary corpus (1990-2022) siParl 3.0

    No full text
    The siParl corpus contains minutes of the Assembly of the Republic of Slovenia for 11th legislative period 1990-1992, minutes of the National Assembly of the Republic of Slovenia from the 1st to the 8th legislative period 1992-2022, minutes of the working bodies of the National Assembly of the Republic of Slovenia from the 2nd to the 7th legislative period 1996-2018, and minutes of the Council of the President of the National Assembly from the 2nd to the 7th legislative period 1996-2018. The corpus comprises of over 11 thousand sessions, one million speeches and 200 million words. The corpus is encoded according to the Parla-CLARIN schema (https://github.com/clarin-eric/parla-clarin). Each mandate is in one directory, and each session in one file. As opposed to the previous version 2.0, this version adds new data (minutes of the National Assembly of the Republic of Slovenia of the 8th legislative period) and corrects many errors

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇