Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Glossary (EN-LT-DA) of General Data Protection Regulation Terms (ELEXIS)

    No full text
    Trilingual glossary (EN-LT-DA) of the English terms referring to personal data and their equivalents in the Lithuanian and Danish languages

    Dictionary of the German-Lorraine Dialects - LothWB (ELEXIS)

    No full text
    Wörterbuch der deutsch-lothringischen Mundarten. The dialectal vocabulary of the German-speaking parts of Lorraine is recorded in the one-volume Dictionary of German-Lorraine Dialects. It only lists words and phrases that deviate from the written language, so that one can speak here of an idiotikon rather than a dialect dictionary

    Corpus of Slovenian school texts SBSJ 1.0

    No full text
    Corpus of Slovenian school texts is a lemmatized and POS-tagged specialized corpus, which includes 428 short school texts written primarily by primary-school students from 1st to 5th grades from 2017 to 2020. The corpus consists of approximately 95,000 tokens and was designed as one of the resources for the compilation of The School Dictionary of the Slovenian Language, which is being created as part of the project Franček Web Portal, Language Counselling for Slovene Teachers and School Dictionary of the Slovene Language. The corpus was lemmatized and POS-tagged with the Obeliks tagger (http://oznacevalnik.slovenscina.eu/Vsebine/Sl/ProgramskaOprema/Navodila.aspx) using JOS morphosyntactic descriptions. The corpus is written in XML and complies with TEI specifications as given in the CLARIN.SI customisation (https://github.com/clarinsi/TEI-schema). Note that the corpus is intergrated with the CLARIN.SI concordancers, but the corpus available on the concordancers is much larger than the TEI sample available for download

    Slovenian RoBERTa contextual embeddings model: SloBERTa 2.0

    No full text
    The monolingual Slovene RoBERTa (A Robustly Optimized Bidirectional Encoder Representations from Transformers) model is a state-of-the-art model representing words/tokens as contextually dependent word embeddings, used for various NLP tasks. Word embeddings can be extracted for every word occurrence and then used in training a model for an end task, but typically the whole RoBERTa model is fine-tuned end-to-end. SloBERTa model is closely related to French Camembert model https://camembert-model.fr/. The corpora used for training the model have 3.47 billion tokens in total. The subword vocabulary contains 32,000 tokens. The scripts and programs used for data preparation and training the model are available on https://github.com/clarinsi/Slovene-BERT-Tool Compared with the previous version (1.0), this version was trained for further 61 epochs (v1.0 37 epochs, v2.0 98 epochs), for a total of 200,000 iterations/updates. The released model here is a pytorch neural network model, intended for usage with the transformers library https://github.com/huggingface/transformers (sloberta.2.0.transformers.tar.gz) or fairseq library https://github.com/pytorch/fairseq (sloberta.2.0.fairseq.tar.gz

    Slovenian Twitter hate speech dataset IMSyPP-sl

    No full text
    A hand-labeled training (50,000 tweets labeled twice) and evaluation set (10,000 tweets labeled twice) for hate speech on Slovenian Twitter. The data files contain tweet IDs, hate speech type, hate speech target, and annotator ID. For obtaining the full text of the dataset, please contact the first author. Hate speech type: 1. Appropriate - has no target 2. Inappropriate (contains terms that are obscene, vulgar; but the text is not directed at any person specifically) - has no target 3. Offensive (including offensive generalization, contempt, dehumanization, indirect offensive remarks) 4. Violent (author threatens, indulges, desires, or calls for physical violence against a target; it also includes calling for, denying, or glorifying war crimes and crimes against humanity) Hate speech target: 1. Racism (intolerance based on nationality, ethnicity, language, towards foreigners; and based on race, skin color) 2. Migrants (intolerance of refugees or migrants, offensive generalization, call for their exclusion, restriction of rights, non-acceptance, denial of assistance…) 3. Islamophobia (intolerance towards Muslims) 4. Antisemitism (intolerance of Jews; also includes conspiracy theories, Holocaust denial or glorification, offensive stereotypes…) 5. Religion (other than above) 6. Homophobia (intolerance based on sexual orientation and / or identity, calls for restrictions on the rights of LGBTQ persons 7. Sexism (offensive gender-based generalization, misogynistic insults, unjustified gender discrimination) 8. Ideology (intolerance based on political affiliation, political belief, ideology… e.g. “communists”, “leftists”, “home defenders”, “socialists”, “activists for…”) 9. Media (journalists and media, also includes allegations of unprofessional reporting, false news, bias) 10. Politics (intolerance towards individual politicians, authorities, system, political parties) 11. Individual (intolerance toward any other individual due to individual characteristics; like commentator, neighbor, acquaintance ) 12. Other (intolerance towards members of other groups due to belonging to this group; write in the blank column on the right which group it is) Training dataset The training set is sampled from data collected between December 2017 and February 2020. The sampling was intentionally biased to contain as much hate speech as possible. A simple model was used to flag potential hate speech content and additionally, filtering by users and by tweet length (number of characters) was applied. 50,000 tweets were selected for annotation. Evaluation dataset The evaluation set is sampled from data collected between February 2020 and August 2020. Contrary to the training set, the evaluation set is an unbiased random sample. Since the evaluation set is from a later period compared to the training set, the possibility of data linkage is minimized. Furthermore, the estimates of model performance made on the evaluation set are realistic, or even pessimistic, since the evaluation set is characterized by a new topic: Covid-19. 10,000 tweets were selected for the evaluation set. Annotation results Each tweet was annotated twice: In 90% of the cases by two different annotators and in 10% of the cases by the same annotator. Special attention was devoted to evening out the overlap between annotators to get agreement estimates on equally sized sets. Ten annotators were engaged for our annotation campaign. They were given annotation guidelines, a training session, and a test on a small set to evaluate their understanding of the task and their commitment before starting the annotation procedure. Annotator agreement in terms of Krippendorff Alpha is around 0.6. Annotation agreement scores are detailed in the accompanying report files for each dataset separately. The annotation process lasted four months, and it required about 1,200 person-hours for the ten annotators to complete the task

    Corpus of 1968 Slovenian literature Maj68 1.0

    No full text
    Maj68 corpus contains 874 texts published between 1964 and 1972 in the periodicals "Tribuna", "Problemi" and "Problemi. Literatura." The texts contain complete bibliographical data, are classified according to text and language type, degree of presence of non-standard Slovenian, foreign languages, modernism, and visual elements. The data about the authors of the texts are provided with their gender and year of birth. The presence of visual elements is marked in the corpus; note that 39 texts have only visual elements, i.e. do not contain text. The corpus is available as facsimiles (PDFs), in the TEI source encoding, as plain text files accompanied by metadata files, and as the linguistically annotated TEI corpus, and the derived vertical files. The TEI encoding follows the CLARIN.SI TEI customisation (https://github.com/clarinsi/TEI-schema). The automatic linguistic annotation includes lemmas, MULTEXT-East morphosyntactic descriptions and Universal Dependencies morphological features and syntactic annotation

    Q-CAT Corpus Annotation Tool 1.2

    No full text
    The Q-CAT (Querying-Supported Corpus Annotation Tool) is a computational tool for manual annotation of language corpora, which also enables advanced queries on top of these annotations. The tool has been used in various annotation campaigns related to the ssj500k reference training corpus of Slovenian (http://hdl.handle.net/11356/1210), such as named entities, dependency syntax, semantic roles and multi-word expressions, but it can also be used for adding new annotation layers of various types to this or other language corpora. Q-CAT is a .NET application, which runs on Windows operating system. Version 1.1 enables the automatic attribution of token IDs and personalized font adjustments. Version 1.2 supports the CONLL-U format and working with UD POS tags

    ILSP Conceptual Dictionary of Modern Greek (ELEXIS)

    No full text
    ConceptNet-el (Εννοιολογικό Λεξικό της Νέας Ελληνικής ΙΕΛ). ConceptNet-el is a conceptual dictionary of Modern Greek that assumes the form of a linguistic ontology. It comprises c. 35K instances, that is, unique combinations of form and sense. Both single and multi-word entries are included. Information about POS, gender, or any other information at the level of morphology is provided. Alternative forms, morphologically related forms, and lexical semantic information (synonyms, antonyms) is also provided. Instances are mapped onto senses or concepts organized under semantic fields or super-concepts that form an ontology. Concepts are interlinked via a wealth of semantic relations

    GLOBAL French-Russian Dictionary - MLDS (ELEXIS)

    No full text
    A general language French to Russian dictionary

    GLOBAL French-Turkish Dictionary - MLDS (ELEXIS)

    No full text
    A general language French to Turkish dictionary

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇