Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Corpus of combined Slovenian corpora MetaFida 0.1

    No full text
    Slovenia has a large number of diverse corpora available for online analysis via the CLARIN.SI concordancers. However, if users are interested in the same queries across different corpora, they have to search for relevant information in each corpus separately, and then combine this information manually, which is time-consuming and also prone to analysis errors. An additional problem is that corpora typically have different metadata and may also be labeled at different linguistic levels, which further complicates identical searches across different corpora. For these reasons we combined a number of existing corpora of the Slovenian available through the CLARIN.SI concordances into the MetaFida corpus. Here it was first necessary to unify the metadata and harmonize the linguistic and structural annotations between the corpora, and to create conversions of individual corpora from their vertical formats, which are used as input by the CLARIN.SI concordances, into the MetaFida vertical format. As the source corpora are not completely distinct, MetaFida is deduplicated on the level of paragraphs. In the MetaFida corpus, we kept only that information that is common to most of the selected corpora. The structure is nested very shallowly, as it is easier to create subcorpus or limit the search to individual text types. All Metafida positional attributes are considered to have multiple values, separated by a space. More values ​​are needed because some corpora have normalized words (older Slovenian, user-generated content), where one original word can be mapped to several normalized ones or vice versa. There are 34 corpora included in this version of MetaFida: * classlawiki_sl, CLASSLAWiki-sl (Slovenian Wikipedia), 54,608,642 tokens * dgt15_sl, EU DGT 2015: Slovene, 62,303,744 tokens * dsi, DSI (informatics), 5,245,073 tokens * eltec_slv, ELTeC-slv (100 novels), 6,901,534 tokens * filmi, FILMI (film reviews), 936,446 tokens * gfida20_dedup, Gigafida v2.0 (reference, deduplicated), 1.333,360,653 tokens * gos_vl42, GosVL 4.2 (spoken, VideoLectures), 179,063 tokens * gos11, Gos 1.1.1 (reference, speech), 1,063,861 tokens * imp, IMP (older texts), 17,723,874 tokens * ispac_sl, ISPAC: Slovenian, 1,432,798 tokens * janes_blog, Janes Blog (blogs with comments), 34,534,431 tokens * janes_forum, Janes Forum (web forums), 47,066,575 tokens * janes_news, Janes News (news comments), 14,838,074 tokens * janes_tweet, Janes Tweet (tweets 2013-2017), 151,457,091 tokens * janes_wiki, Janes Wiki (Wikipedia comments), 5,008,067 tokens * jaslo_sl, jaSlo: Slovenian, 532,395 tokens * kas_dipl, KAS Dipl (diplomas), 1,101,796,659 tokens * kas_dr, KAS Dr (PhD theses), 101,473,395 tokens * kas_mag, KAS Mag (master theses), 495,827,656 tokens * konji, Konji (equestrianism), 469,894 tokens * korp, KoRP (public relations), 2,194,130 tokens * lemonde_sl, LeMonde: Slovenian, 615,617 tokens * maj68, Maj68 (May 1968 in literature), 794,382 tokens * maks, MAKS (youth literature), 12,072,273 tokens * prilit, PriLit (older narrative prose), 1,275,209 tokens * rsdo5, RSDO5 (term-annotated texts), 310,588 tokens * sbsj, SBSJ (school texts), 1,836,810 tokens * siparl20, siParl 2.0 (parliament 1990-2018), 239,749,733 tokens * slwac, slWaC (Slovene Web), 895,903,321 tokens * solar, Šolar v2 Clear (school essays), 1,907,731 tokens * suss, ŠUSS (FAQ on Slovenian language), 365,371 tokens * trans5_sl, TRANS5: Slovenian, 1,594,120 tokens * tweet_sl, Tweet-sl (older tweets), 6,291,820 tokens * vayna, VAYNA (attacks on the YNA), 300,666 tokens Σ 34 corpora, 4,601,971,696 token

    Valency lexicon extracted from the Gigafida 2.1 corpus

    No full text
    The valency lexicon was extracted from the Gigafida 2.1 Corpus of Written Standard Slovene (https://www.clarin.si/noske/run.cgi/corp_info?corpname=gfida21) using specialized scripts for extracting data from corpora containing syntactic and semantic role annotations. The lexicon contains valency patterns for 14,595 Slovene verbs based on the JOS syntactic dependency system (http://nl.ijs.si/jos/bib/jos-skladnja-navodila.pdf) and the semantic role labelling system for Slovene with 25 semantic role labels. The lexicon consists of separate XML files for each verb. Each file contains the verb's lemma, its aspect, and its frequency in the Gigafida 2.1 corpus, followed by a list of all the semantic role labels present in all the verb's valency patterns, along with two measures for each semantic role label (listed in ): (1) "valency_pattern_ratio", which indicates the percentage of the verb's valency patterns where the semantic role label is present; and (2) "valency_sentence_ratio", which indicates the percentage of all the corpus sentences containing both the verb and the semantic role label out of a total of all corpus sentences containing the verb. The valency patterns (listed in ) contain the following: - the valency pattern's ID-number (); - the number of corpus sentences in which the verb follows the valency pattern (); - the semantic roles occurring in the pattern (); - the syntactic structures occurring in the pattern (; if the syntactic structures contain prepositions, these are also included as additional information); - a human-readable representation of the valency pattern in Slovene (, e.g. KDO/KAJ abdicira); - the corpus examples in which the verb occurs with the valency pattern (). In the examples, the components of the valency patterns are annotated with their semantic role labels and syntactic structures. Each valency pattern contains at least one example from the Gigafida 2.1 corpus and all the relevant examples from the ssj500k 2.2 corpus (http://hdl.handle.net/11356/1210)

    Slovenian Twitter dataset 2018-2020 1.0

    No full text
    The dataset represents the Twitter production in Slovenian in the period from 2018 until 2020. It consists of tweet IDs, retweet IDs, pseudo-anonymized user IDs, publication dates, and automatically assigned hate labels (acceptable, inappropriate, offensive, violent) with https://huggingface.co/IMSyPP/hate_speech_slo. The dataset is the basis for the two following papers: - "Retweet communities reveal the main source of hate speech" - https://arxiv.org/pdf/2105.14898.pdf - "Community evolution in retweet networks" - https://arxiv.org/pdf/2105.06214.pd

    Spoken corpus Gos 1.1

    No full text
    Gos is a corpus of spoken Slovene that includes the transcripts of approximately 120 hours of speech recorded in various situations: radio and TV shows, school lessons and lectures, private conversations between friends or within the family, work meetings, consultations, conversations in buying and selling situations, etc. All speech is transcribed in two versions – with pronunciation-based spelling and with standardized spelling – and it comprises over one million words. The corpus can be searched by means of the web concordancer where it is also possible to listen to the corresponding recordings: http://www.korpus-gos.net. As opposed to the previous version, this one corrects some errors in the transcriptions and introduces various changes in the TEI and vertical encodings

    Offensive language dataset of Croatian, English and Slovenian comments FRENK 1.1

    No full text
    The FRENK dataset consists of comments to Facebook posts (news articles) of mainstream media outlets from Croatia, Great Britain, and Slovenia, on the topics of migrants and LGBT. The dataset contains whole discussion threads. Each comment is annotated by the type of socially unacceptable discourse (e.g., inappropriate, offensive, violent speech) and its target (e.g., migrants/LGBT, commenters, media). The annotation schema in its details is described in https://arxiv.org/pdf/1906.02045.pdf. Usernames in the metadata are pseudo-anonymised and removed from the comments. The data in each language (Croatian (hr), English (en), Slovenian (sl), and topic (migrants, LGBT) is divided into a training and a testing portion. The training and testing data consist of separate discussion threads, i.e., there is no cross-discussion-thread contamination between training and testing data. The sizes of the splits are the following: Croatian, migrants: 4356 training comments, 978 testing comments; Croatian LGBT: 4494 training comments, 1142 comments; English, migrants: 4540 training comments, 1285 testing comments; English, LGBT: 4819 training comments, 1017 testing comments; Slovenian, migrants: 5145 training comments, 1277 testing comments; Slovenian, LGBT: 2842 training comments, 900 testing comments. The difference to the first version of the dataset are the additions of 1. the annotation guidelines in English and 2. the link to the huggingface dataset

    Multilingual comparable corpora of parliamentary debates ParlaMint 2.0

    No full text
    ParlaMint is a multilingual set of comparable corpora containing parliamentary debates mostly starting in 2015 and extending to mid-2020, with each corpus being about 20 million words in size. The sessions in the corpora are marked as belonging to the COVID-19 period (after October 2019), or being "reference" (before that date). The corpora have extensive metadata, including aspects of the parliament; the speakers (name, gender, MP status, party affiliation, party coalition/opposition); are structured into time-stamped terms, sessions and meetings; with speeches being marked by the speaker and their role (e.g. chair, regular speaker). The speeches also contain marked-up transcriber comments, such as gaps in the transcription, interruptions, applause, etc. Note that some corpora have further information, e.g. the year of birth of the speakers, links to their Wikipedia articles, their membership in various committees, etc. The corpora are encoded according to the Parla-CLARIN TEI recommendation (https://clarin-eric.github.io/parla-clarin/), but have been validated against the compatible, but much stricter ParlaMint schemas. This entry contains the ParlaMint TEI-encoded corpora with the derived plain text version of the corpus along with TSV metadata on the speeches. Also included is the 2.0 release of the data and scripts available at the GitHub repository of the ParlaMint project. Note that there also exists the linguistically marked-up version of the corpus, which is available at http://hdl.handle.net/11356/1405

    Latvian user comment dataset 1.0

    No full text
    The dataset is an archive of reader comments from the Delfi news site from 2014-2019, containing approximately 12M comments, mostly in the Latvian language, with some in Russian. Description of the Datasets There are 6 CSV files: * ``lv-comments-2014.csv`` contains **2 753 655** comments from year 2014 * ``lv-comments-2015.csv`` contains **2 221 122** comments from year 2015 * ``lv-comments-2016.csv`` contains **1 897 669** comments from year 2016 * ``lv-comments-2017.csv`` contains **1 896 083** comments from year 2017 * ``lv-comments-2018.csv`` contains **2 222 051** comments from year 2018 * ``lv-comments-2019.csv`` contains **1 421 883** comments from year 2019 **In sum: 12 412 463 comments** Columns: * ``comment_id`` (string) - the ID of the written comment * ``article_id`` (string) - the ID of the article for which the comment was written * ``created_time`` (string) - the time and date of the comment * ``subject`` (string) - the title of the comment * ``reply_to_comment_id`` (string) - the parent comments ID * ``content`` (string) - the comment itself * ``is_anonymous`` (string) - * 1 if the comment was published anonymously * 0 if the comment was published by a registered user * ``is_enabled`` (string) - * 1 if the comment was published (online) * 0 if it wasn’t published * Questionable field: not all have been manually moderated * No additional information from the moderators * ``channel_language`` (string) - the language of the channel * 'nat' for Latvian * 'rus' for Russian * ``create_user_id`` (string) - the user ID of the commentator * ``modereted_by`` (string) - the ID of the moderato

    Ekspress user comment dataset 1.0

    No full text
    This dataset is an archive of reader comments on the Ekspress Meedia news site from 2009-2019, containing approximately 31M comments, mostly in the Estonian language, with some in Russian. Description of the Datasets. There are 11 CSV files: comments_2009.csv contains 2 898 438 comments from the year 2009 comments_2010.csv contains 2 377 591 comments from the year 2010 comments_2011.csv contains 2 729 389 comments from the year 2011 comments_2012.csv contains 3 372 776 comments from the year 2012 comments_2013.csv contains 3 289 393 comments from the year 2013 comments_2014.csv contains 3 195 502 comments from the year 2014 comments_2015.csv contains 3 202 592 comments from the year 2015 comments_2016.csv contains 2 848 624 comments from the year 2016 comments_2017.csv contains 2 838 075 comments from the year 2017 comments_2018.csv contains 3 194 597 comments from the year 2018 comments_2019.csv contains 1 526 755 comments from the year 2019 May In sum: 3 1473 732 comments Columns: comment_id (string) - the ID of the written comment article_id (string) - the ID of the article for which the comment was written created_time (string) - the time and date of the comment subject (string) - the title of the comment reply_to_comment_id (string) - the parent comments ID content (string) - the comment itself is_anonymous (string) - 1 if the comment was published anonymously 0 if the comment was published by a registered user is_enabled (string) - 1 if the comment was published (online) 0 if it wasn’t published Questionable field: not all have been manually moderated No additional information from the moderators channel_language (string) - the language of the channel: 'nat' for Estonian, 'rus' for Russian create_user_id (string) - the user ID of the commentator '0' for all blocked comments. moderated_by (string) - the ID of the moderato

    24sata news article archive 1.0

    No full text
    The 24sata news portal consists of a portal with daily news and several smaller portals covering news from specific topics, such as automotive news, health, culinary content, and lifestyle advice. The dataset contains over 650,000 articles in Croatian from 2007 to 2019, as well as assigned tags. Description of the Dataset The dataset consists of 11 columns and 657806 rows. Each row represents a single news article published on the 24sata news portals. Besides the 'www.24sata.hr', the biggest news portal, articles from other niche portals affiliated with 24sata are also included. Columns: 'article_id' - Public id of the article on the new site. The article can be accessed by concatenating the site URL and article_id. For example, to access the article with article_id 614684, you can access it on 'www.24sata.hr/--614684'. This id is, by itself, not unique across the dataset - articles from different portals can share the same article_id. 'site' - The location of the portals where the article came from. There are eight different portals covering topics of daily news, to the more focused portals about automotive technologies and trends, health and wellness, culinary trends and recipes, or lifestyle advices. 'title' - The title of the news article. 'lead' - Lead text, a short introduction to the content of an article. Can be empty. 'content' - The content of the news article, contains the bulk of the text. Can be empty if the whole article could fit in the lead text. 'tags' - Tags, zero or more, separated with a '|' character. Article tags are chosen by the author of the article. 'section' - The main section of the news portal where the article was posted (does not need to be set). The most frequent section is 'Vijesti' (News). 'subsection' - The subsection of the section where the article was posted (does not need to be set). Each section can have multiple subsections. 'authors' - Article authors, zero or more, separated with a '|' character. The author does not need to sign the article if he chooses not to so this can be empty. 'published_from' - A date when this article appeared on the portal. Journalists can write the article in advance and pick a future date and time when it will appear on the site. Due to this strategy, the 'published_from' can be much later than the 'date_created'. 'date_created' - A date when this article was originally written. For all articles published before 2nd Feb 2010 the 'date_created' is set to 2nd Feb 2010 - this is the date when the portal was redesigned and the database with news articles recreated

    Comparable corpora of South-Slavic Wikipedias CLASSLA-Wikipedia 1.0

    No full text
    This comparable corpus collection consists of Wikipedia dumps of the Bosnian, Croatian, Macedonian, Montenegrin, Serbian, Serbo-Croatian and Slovenian Wikipedia, harvested on October 17th 2020. The text was extracted from the dumps with the process documented at https://github.com/clarinsi/classla-wikipedia, and linguistic annotation was performed with the classla package (https://pypi.org/project/classla/), on all levels available for a specific language, with the Bosnian and Serbo-Croatian Wikipedias processed with the standard Croatian models

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇