Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    ASR training dataset for Croatian ParlaSpeech-HR v1.0

    No full text
    The ParlaSpeech-HR dataset is built from parliamentary proceedings available in the Croatian part of the ParlaMint corpus and the parliamentary recordings available from the Croatian Parliament's YouTube channel. The corpus consists of segments 8-20 seconds in length. There are two transcripts available: the original one, and the one normalised via a simple rule-based normaliser. Each of the transcripts contains word-level alignments to the recordings. Each segment has a reference to the ParlaMint 2.1 corpus (http://hdl.handle.net/11356/1432) via utterance IDs. If a segment is based on a single utterance, speaker information for that segment is available as well. There is speaker information available for 381,849 segments, i.e., 95% of all segments. Speaker information consists of all the speaker information available from the ParlaMint 2.1 corpus (name, party, gender, age, status, role). There are all together 309 speakers in the dataset. The dataset is divided into a training, a development, and a testing subset. Development data consist of 500 segments coming from the 5 most frequent speakers, with the goal of not losing speaker variety on dev data. Test data consist of 513 segments that come from 3 male (258 segments) and 3 female speakers (255 segments). There are no segments coming from the 6 test speakers in the two remaining subsets. The 22,076 instances not having speaker information are not assigned to any of the three subsets. The remaining 380,836 instances form the training set

    Slovene-English parallel corpus MaCoCu-sl-en 1.0

    No full text
    The Slovene-English parallel corpus MaCoCu-sl-en 1.0 was built by crawling the ".si" internet top-level domain in 2021, extending the crawl dynamically to other domains as well. All the crawling process was carried out by the MaCoCu crawler (https://github.com/macocu/MaCoCu-crawler). Websites containing documents in both target languages were identified and processed using the tool Bitextor (https://github.com/bitextor/bitextor). Considerable efforts were devoted into cleaning the extracted text to provide a high-quality parallel corpus. This was achieved by removing boilerplate and near-duplicated paragraphs and documents that are not in one of the targeted languages. Document and segment alignment as implemented in Bitextor were carried out, and BicleanerAI (https://github.com/bitextor/bicleaner-ai) and Bifixer (https://github.com/bitextor/bifixer) were used for fixing, cleaning, and deduplicating the final version of the corpus. While the TXT format consists solely of pairs of source and target segments (one or several sentences), each segment pair in the TMX format is accompanied by the following metadata: - source and target document URL; - quality score as provided by the tool BicleanerAI; - translation direction identification: the source segment in each segment pair was identified by using a probabilistic model; - personal information identification (“biroamer-entities”): segments containing personal information are flagged, so final users of the corpus can decide whether to use these segments; - language variants: the language variant of English (British or American) was identified for every segment pair on document and domain level. Notice and take down: Should you consider that our data contains material that is owned by you and should therefore not be reproduced here, please: (1) Clearly identify yourself, with detailed contact data such as an address, telephone number or email address at which you can be contacted. (2) Clearly identify the copyrighted work claimed to be infringed. (3) Clearly identify the material that is claimed to be infringing and information reasonably sufficient in order to allow us to locate the material. (4) Please write to the contact person for this resource whose email is available in the full item record. We will comply with legitimate requests by removing the affected sources from the next release of the corpus. This action has received funding from the European Union's Connecting Europe Facility 2014-2020 - CEF Telecom, under Grant Agreement No. INEA/CEF/ICT/A2020/2278341. This communication reflects only the author’s view. The Agency is not responsible for any use that may be made of the information it contains

    The sentiment corpus of parliamentary debates ParlaSent-BCS v1.0

    No full text
    The dataset consists of mid-length sentences from the Bosnian, Croatian and Serbian parliamentary proceedings, annotated with a 6-level sentiment schema (defined below). The first 1,300 instances were annotated by two annotators, and a reconciliation procedure was performed if there was disagreement on the simplified 3-level schema (Positive, Negative, Neutral). The latter 1,300 instances were annotated by second annotator only. Besides having the annotations of the two annotators and potential reconciliation annotations, there is also a handy 3-level label available for all instances. Each sentence can be followed back to the original datasets (https://doi.org/10.5281/zenodo.6517697, https://doi.org/10.5281/zenodo.6521372, https://doi.org/10.5281/zenodo.6521648) via a document and sentence identifier. Date of the speech and the speaker name are given as well. If the speaker is MP, information on party, gender and year of birth are available as well. The dataset is split into a training (2,150 instances), development (150 instances) and testing subset (300 instances). The full 6-level annotation schema is the following: - Positive for sentences that are entirely or predominantly positive - Negative for sentences that are entirely or predominantly negative - M_Positive for sentences that convey an ambiguous sentiment or a mixture of sentiments, but lean more towards the positive sentiment in a strict binary classification - M_Negative for sentences that convey an ambiguous sentiment or a mixture of sentiments, but lean more towards the negative sentiment in a strict binary classification - P_Neutral for sentences that only contain non-sentiment-related statements, but still lean more towards the positive sentiment in a strict binary classification - N_Neutral for sentences that only contain non-sentiment-related statements, but still lean more towards the negative sentiment in a strict binary classificatio

    Slovenian datasets for contextual synonym and antonym detection

    No full text
    Slovenian datasets for contextual synonym and antonym detection can be used for training machine learning classifiers as described in the MSc thesis of Jasmina Pegan "Semantic detection of synonyms and antonyms with contextual embeddings" (https://repozitorij.uni-lj.si/IzpisGradiva.php?id=141456). Datasets contain example pairs of synonyms and antonyms in contexts together with additional information on a sense pair. Candidates for synonyms and antonyms were retrieved from the dataset created in the BSc thesis of Jasmina Pegan "Antonym detection with word embeddings" (https://repozitorij.uni-lj.si/IzpisGradiva.php?id=110533). Example sentences were retrieved from The comprehensive Slovenian-Hungarian dictionary (VSMS) (https://www.clarin.si/repository/xmlui/handle/11356/1453). Each dataset is class balanced and contains an equal amount of examples and counterexamples. An example is a pair of example sentences where the two words are synonyms/antonyms. A counterexample is a pair of example sentences where two words are not synonyms/antonyms. Note that a word pair can be synonymous or antonymous in some sense of the two words (but not in the given context). Datasets are divided into two categories, datasets for synonyms and datasets for antonyms. Each category is further divided into base and updated datasets. These contain three dataset files: train, validation and test dataset. Base datasets include only manually-reviewed sense pairs. These are generated from all pairs of VSMS sense examples for all confirmed pairs of antonym and synonym senses. Updated datasets include automatically generated sense pairs while constraining the maximal number of examples per word. In this way, the dataset is more balanced word-wise, but is not fully manually-reviewed and contains less accurate data. A single dataset entry contains the information on the base word, followed by data on synonym/antonym candidate. The last column discerns whether the sense pair is a pair of synonyms/antonyms or not. More details on this can be found inside the included README file

    Slovene Natural Language Inference Dataset SI-NLI

    No full text
    SI-NLI (Slovene Natural Language Inference Dataset) contains 5,937 human-created Slovene sentence pairs (premise and hypothesis) that are manually labeled with the labels "entailment", "contradiction", and "neutral". We created the dataset using sentences that appear in the Slovenian reference corpus ccKres (http://hdl.handle.net/11356/1034). Annotators were tasked to modify the hypothesis in a candidate pair in a way that reflects one of the labels. The dataset is balanced since the annotators created three modifications (entailment, contradiction, neutral) for each candidate sentence pair. The dataset is split into train, validation, and test sets, with sizes of 4,392, 547, and 998. We used Slovenian pre-trained language models to create splits, thereby ensuring that difficult and easy instances are evenly distributed in all three subsets. The dataset is released in a tabular TSV format. The README.txt file contains a description of the attributes. Only the hypothesis and premise are given in the test set (i.e. no annotations) since SI-NLI is integrated into the Slovene evaluation framework SloBENCH (https://slobench.cjvt.si/). If you use the dataset to train your models, please consider submitting the test set predictions to SloBENCH to get the evaluation score and see how it compares to others

    A Dictionary of the Welsh Language - GPC (ELEXIS)

    No full text
    Geiriadur Prifysgol Cymru. This is a subset of data from every entry in GPC online. GPC is the only standard historical dictionary of the Welsh language. It presents the vocabulary of the Welsh language from the earliest Old Welsh texts, through the abundant literature of the Medieval and Modern periods, to the huge expansion in vocabulary resulting from the wider use of Welsh in all aspects of life in the last half century. This vocabulary is defined in Welsh, and English equivalents are also given. Detailed attention is given to variant forms, collocations, and etymology. The work is based on an ever-expanding collection of over two million citation slips gathered from a range of texts over many years

    Serbian Dictionary - Karadžić (1818, 1852) (ELEXIS)

    No full text
    Српски рјечник (Serbian Dictionary) or Lexicon Lexicon serbico-germanico-latinum is the foundational dictionary of the modern Serbian language, compiled by Vuk Stefanović Karadžić and published in two editions (1818 and 1852). It is the first dictionary of the Serbian language that was based on the vernacular, as opposed to the hybrid, literary Slaveno-Serbian language of the time. It contains a wide range of dialectal and ethnographic material

    Q-CAT Corpus Annotation Tool 1.4

    No full text
    The Q-CAT (Querying-Supported Corpus Annotation Tool) is a tool for manual linguistic annotation of corpora, which also enables advanced queries on top of these annotations. The tool has been used in various annotation campaigns related to the ssj500k reference training corpus of Slovenian (http://hdl.handle.net/11356/1210), such as named entities, dependency syntax, semantic roles and multi-word expressions, but it can also be used for adding new annotation layers of various types to this or other language corpora. Q-CAT is a .NET application, which runs on Windows operating system. Version 1.1 enables the automatic attribution of token IDs and personalized font adjustments. Version 1.2 supports the CONLL-U format and working with UD POS tags. Version 1.3 supports adding new layers of annotation on top of CONLL-U (and then saving the corpus as XML TEI). Version 1.4 introduces new features in command line mode (filtering by sentence ID, multiple link type visualizations

    ASR model evaluator

    No full text
    Docker image with ASR evaluation tool that has support for WER calculation on punctuated and capitalised transcripts. The UI allows uploading the reference and predicted transcripts, and choice to perform WER calculation with or without consideration of punctuation and capitalisation. When punctuation and capitalisation is considered the resulting WER is augmented by the accuracy, recall and F1 statistics for punctuation (comma, full stop, question mark, exclamation point) and capital letters. Start the service by running "docker image load <evaluator.1.0.3.tar.gz" followed by "docker run -ti --rm eval-service:1.0.3". This will start a web service on http://localhost:8888. Use the tool by visiting the said address

    The EKI Combined Dictionary 2022 (ELEXIS)

    No full text
    Eesti Keele Ühendsõnastik 2022 (EKI Combined Dictionary 2022) displays information from different lexical databases: "The Dictionary of Estonian 2019", "Estonian Collocations Dictionary 2019", "Basic Estonian Dictionary" (2014), "The Estonian Morphological Database of the Institute of the Estonian Language 2022". It displays also information from bilingual lexical databases: "Estonian-Russian orthographic dictionary for students 2018" (1st edition 2011), "Estonian-Russian Dictionary 2018" (1st edition 1997–2009), "The Russian Morphological Database of the Institute of the Estonian Language 2022". The data is stored in Ekilex's PostgreSQL database and accessible through API. Ekilex is in-house DWS of the Institite of the Estonian Language. Ekilex is hosted in the Estonian Scientific Computing Infrastructure (ETAIS) cloud. See also: https://doi.org/10.15155/3-00-0000-0000-0000-08C0A

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇