Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
840 research outputs found
Sort by
The CLASSLA-Stanza model for semantic role labeling of standard Slovenian
The model for semantic role labeling of standard Slovenian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1434). The estimated F1 of the semantic role annotations is ~77.2
Corpus of Romanian Academic Genres ROGER
The corpus contains academic papers from eight disciplines, written by the Romanian students in native Romanian and English L2.
The corpus was collected over a three-year period (2018-2021) with the help of 27 collaborators from nine Romanian universities.
The corpus is available for online querying through a dedicated platform developed at the CODHUS research centre from the West University of Timisoara
Dictionary of Croatian Idioms (ELEXIS)
Frazeološki rječnik hrvatskoga jezika is an open-access dictionary of Croatian idioms based on data from a large electronic corpus. The resulting dictionary will serve as a gateway for a large number of users and researchers to current idiomatic usage of Croatian
Dictionary of Southern Serbian Dialects (ELEXIS)
Речник говора јужне Србије (Dictionary of Southern Serbian Dialects) is a dialectologica dictionary by Momčilo Zlatanović, which was first published in 1998
Serbian-English Terminology in the Power Engineering Domain - SrpEngPE (ELEXIS)
SrpEngPE - Serbian-English dictionary with terminology in the power engineering domain: automatically extracted from domain parallel corpus http://jerteh.rs/biblisha/ListaDokumenata.aspx?JCID=13&lng=en , extracted monolingual term candidates and translation pairs were manually evaluated and post-edited
Bilingual List of German-Serbian Translated Pairs of Lexical Units - SrpNemLex (ELEXIS)
SrpNemLex - Bilingual list of German-Serbian translated pairs of lexical units: automatically extracted from parallel corpus that contains 14 novels http://jerteh.rs/biblisha/ListaDokumenata.aspx?JCID=11&lng=en , extracted monolingual candidates and translation pairs were manually evaluated and post-edited
Slovene corpus for general relation extraction SloREL 1.0
The SloREL corpus contains annotations for training relation extraction models on Slovene documents. It contains documents from Slovene Wikipedia with annotated entities and relations. We constructed the annotations using a semi-supervised process based on linking the documents to the WikiData knowledge graph. The corpus contains 244,437 sentences from Slovene Wikipedia pages. We also provide 896 additional sentences collected from the 24ur.com news website with annotated and linked entities, which do not contain annotated relations and are meant for additional testing of the models. The entities in our corpus are linked to the entities in the WikiData knowledge graph which is useful for models that take advantage of additional knowledge from a knowledge graph. All together the corpus comprises 245,333 sentences with 813,952 relations and 1,616,193 entities.
The corpus comprises of multiple documents:
- schema-definition.xsd: defines the structure of the xml documents containing relation annotations.
- wikipedia-train.xml: training portion of the wikipedia corpus
- wikipedia-test.xml: testing portion of the wikipedia corpus
- wikipedia-validation.xml: validation portion of the wikipedia corpus
- 24ur.xml: additional sentences from the 24ur.com news article
Terminological dictionary of artificial intelligence
The terminological dictionary was compiled within the framework of the project Development of Slovene in the Digital Environment. It is an example collection of 413 terms from the field of artificial intelligence, especially from the subfields of machine learning, computer vision, natural language processing, and fuzzy logic. Definitions, English equivalents, and possible synonyms are added to the terms.
The dictionary is based on a conceptual approach, according to which terms are perceived as designations for concepts that are related to each other in the conceptual system of the subject field. Consequently, the terms are interrelated in the naming system of the subject field.
The dictionary is distributed in XML using the TBX (TermBase eXchange) standard for representing and exchanging information from termbases
CMC training corpus Janes-Tag 3.0
Janes-Tag is a manually annotated corpus of Slovene Computer-Mediated Communication (CMC) consisting of about 15,000 short texts (190,000 words), mostly tweets but also blogs, forums and news comments.
The corpus is meant as a gold-standard training and testing dataset for tokenisation, sentence segmentation, word normalisation, morphosyntactic tagging, lemmatisation and named entity annotation of non-standard Slovene. As the corpus has been carefully manually annotated, it is also suitable for detailed linguistic explorations which require highly accurate and reliable annotations.
The corpus is composed of two parts, the older (texts to 2016) and smaller (65,000 words) Janes Tag 2.1, and the tweet-only newer (2022, 125,000 words) Janes RSDO. Only the Janes Tag 2.1 part is annotated with named entities and with classification of the texts according to their estimated technical (T1-T3) and linguistic (L1-L3) standardness.
The data is available in the source TEI encoding and in derived CoNLL-U format. Both contain JOS/MULTEXT-East morphosyntactic descriptions as well as Universal Dependencies morphological features.
Compared to the previous version, this one corrects some errors, updates the encoding, and adds Janes-RSDO.
The first version of this corpus is described in:
FIŠER, Darja, LJUBEŠIĆ, Nikola, ERJAVEC, Tomaž. 2020. The Janes project: language resources and tools for Slovene user generated content. Language Resources and Evaluation. https://doi.org/10.1007/s10579-018-9425-z
Note that a related corpus, Janes-Norm 3.0 (http://hdl.handle.net/11356/1733), is also available. It contains Janes-Tag 3.0 and an additional subcorpus with manually checked sentences, tokens and normalised words but only automatically assigned lemmas and MULTEXT-East MSDs
Slovenian parliamentary corpus (1990-2022) siParl 3.0
The siParl corpus contains minutes of the Assembly of the Republic of Slovenia for 11th legislative period 1990-1992, minutes of the National Assembly of the Republic of Slovenia from the 1st to the 8th legislative period 1992-2022, minutes of the working bodies of the National Assembly of the Republic of Slovenia from the 2nd to the 7th legislative period 1996-2018, and minutes of the Council of the President of the National Assembly from the 2nd to the 7th legislative period 1996-2018. The corpus comprises of over 11 thousand sessions, one million speeches and 200 million words. The corpus is encoded according to the Parla-CLARIN schema (https://github.com/clarin-eric/parla-clarin). Each mandate is in one directory, and each session in one file.
As opposed to the previous version 2.0, this version adds new data (minutes of the National Assembly of the Republic of Slovenia of the 8th legislative period) and corrects many errors