Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
840 research outputs found
Sort by
GLOBAL Polish-French Dictionary - MLDS (ELEXIS)
A general language Polish to French dictionary
Latvian Delfi article archive (in Latvian and Russian) 1.0
This dataset is an archive of articles from the Delfi news site from 2015-2019, containing over 180,000 articles (c. 50% in Latvian and 50% in the Russian language). Keywords for articles are included.
There are 5 JSON files:
lv_2015.json contains 42 001 articles from the year 2015
lv_2016_.json contains 40 342 articles from the year 2016
lv_2017_.json contains 37 256 articles from the year 2017
lv_2018_.json contains 31 732 articles from the year 2018
lv_2019_.json contains 29 070 articles from the year 2019
In sum: 180 401 articles
Description of the dataset
This JSON file is a list of dictionaries, i.e. each article is represented as a dictionary. Each dictionary contains the following:
id (integer) - the ID of the article
title (string) - the title of the article
lead (string) - the lead of the article
tags [1] (list of dictionaries or None): each dictionary represents one tag. The tag dictionary contains the following:
domain_id (string) - the ID of the domain
id (string) - the ID of the tag
lang (string) - the language of the tag
tag (string) - the tag itself, e.g. Šokolāde
translitted_name (string) - a modified version of the tag, e.g. sokolade
rawBody (string) - the raw text of the article (contains HTML)
bodyText (string) - clean article text (stripped from HTML)
publishDate (string) - published date & time of the article
categoryPrimary (dictionary or empty list) - the dictionary contains the following information:
categoryId (integer) - the ID of the category
categoryName (string)- the name of the category (e.g. Futbols)
channelId (integer) - the ID of the channel
groupId - None
channelLanguage (string) - the language of the channel (nat - Latvian, rus - Russian)
categoryLanguage (integer) - ID of the channel language
relatedArticles (list of integers or None) - a list of related articles' ID's
relatedTags(string or None) -- related tags are comma-separate
Keyword extraction datasets for Croatian, Estonian, Latvian and Russian 1.0
EACL Hackashop Keyword Challenge Datasets
In this repository you can find ids of articles used for the keyword extraction challenge at
EACL Hackashop on News Media Content Analysis and Automated Report Generation (http://embeddia.eu/hackashop2021/). The article ids can be used to generate train-test split used in paper:
Koloski, B., Pollak, S., Škrlj, B., & Martinc, M. (2021). Extending Neural Keyword Extraction with TF-IDF tagset matching. In: Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation, Kiev, Ukraine, pages 22–29.
Train and test splits are provided for Latvian, Estonian, Russian and Croatian.
The articles with the corresponding ID-s can be extracted from the following datasets:
- Estonian and Russian (use the eearticles2015-2019 dataset): https://www.clarin.si/repository/xmlui/handle/11356/1408
- Latvian: https://www.clarin.si/repository/xmlui/handle/11356/1409
- Croatian: https://www.clarin.si/repository/xmlui/handle/11356/1410
dataset_ids folder is organized in the following way:
- latvian – containing latvian_train.json: a json file with ids from train articles to replicate the data used in Koloski et al. (2020), the latvian_test.json: a json file with ids from test articles to replicate the data
- estonian – containing estonian_train.json: a json file with ids from train articles to replicate the data used in Koloski et al. (2020), the estonian_test.json: a json file with ids from test articles to replicate the data
- russian – containing russian_train.json: a json file with ids from train articles to replicate the train data used in Koloski et al. (2020), the russian_test.json: a json file with ids from test articles to replicate the data
- croatian - containing croatian_id_train.tsv file with sites and ids (note that just ids are not unique across dataset, therefore site information also needs to be included to obtain a unique article identifier) of articles in the train set, and the croatian_id_test.tsv file with sites and ids of articles in the test set.
In addition, scripts are provided for extracting articles (see folder parse containing scripts parse.py and build_croatian_dataset.py, requirements for scripts are pandas and bs4 Python libraries):
parse.py is used for extraction of Estonian, Russian and Latvian train and test datasets:
Instructions:
ESTONIAN-RUSSIAN
1) Retrieve the data ee_articles_2015_2019.zip
2) Create a folder 'data' and subfolder 'ee'
3) Unzip them in the 'data/ee' folder
To extract train/test Estonian articles:
run function 'build_dataset(lang="ee", opt="nat")' in the parse.py script
To extract train/test Russian articles:
run function 'build_dataset(lang="ee", opt="rus")' in the parse.py script
LATVIAN:
1) Retrieve the latvian data
2) Unzip it in 'data/lv' folder
3) To extract train/test Latvian articles:
run function 'build_dataset(lang="lv", opt="nat")' in the parse.py script
build_croatian_dataset.py is used for extraction of Croatian train and test datasets:
Instructions:
CROATIAN:
1) Retrieve the Croatian data (file 'STY_24sata_articles_hr_PUB-01.csv')
2) put the script 'build_croatian_dataset.py' in the same folder as the extracted data and run it (e.g., python build_croatian_dataset.py).
For additional questions: {Boshko.Koloski,Matej.Martinc,Senja.Pollak}@ijs.s
Croatian corpus of non-professional written language by typical speakers and speakers with language disorders RAPUT 1.0
The corpus consists of texts produced by nonprofessional typical speakers and speakers with different language disorders (developmental language disorder, dyslexia, traumatic brain injury, aphasia, other). Roughly half of the corpus consists of texts of typical speakers, and the other half of speakers with language disorders. Language samples were elicited by six groups of tasks representing different writing styles (descriptive, expository, narrative, and letter) and different levels of formality.
The corpus has been manually annotated for normalized forms, lemmas, morphosyntactic information (by following the MULTEXT-East tagset), and type of error (phonological segmentation, orthography, non-standard spelling, typo, syntax, etc.). UD morphosyntactic description has been to the most part automatically generated from the MULTEXT-East morphosyntactic information
Franček portal dialect module
The Franček Portal Dialect Module contains data on dialect variation of select lexemes and/or their meanings including labels describing their respective relative frequencies and spatial distribution. The underlying data stems from the Slovenian Linguistic Atlas, volumes 1 and 2. The content has been transformed and adjusted for primary- and secondary-school users.
The dataset is linked to the "Franček Portal Headword List" (http://hdl.handle.net/11356/1445)
Comprehensive Slovenian-Hungarian Dictionary 1.0
The Comprehensive Slovenian-Hungarian dictionary is a general bilingual dictionary that is being compiled at the Centre for Language Resources and Technologies of the University of Ljubljana (CJVT UL). Version 1.0 contains 10,946 headwords, 33,298 translations, 15,265 collocations and other word combinations, and 2,416 examples. The file also contains links between synonymous entries or entry senses, and links between single-word headwords and compounds/phrases.
The Comprehensive Slovenian-Hungarian dictionary is a growing dictionary, which means that new headwords will be added in regular intervals. The Comprehensive Slovenian-Hungarian dictionary is based on a concept (Kosem et al. 2018) that was prepared in the targeted research project KOMASS (the Concept of Hungarian-Slovenian dictionary: from a language resource to its user), funded by the Slovenian Research Agency and the Ministry of Education, Science and Sport of the Republic of Slovenia. The dictionary concept follows the state-of-the-art international lexicographic practice, e.g. bilingual dictionaries compiled at established international publishers and institutes
eSSKJ: Dictionary of the Slovenian standard language (phraseological database): school edition
The eSSKJ: Dictionary of the Slovenian Standard Language (phraseological database): School Edition represents a part of phraseological entries in the educational language portal Franček. The underlying data stems from the eSSKJ: Dictionary of the Slovenian Standard Language. The content is adjusted for primary- and secondary-school users.
The dataset is linked to the "Franček Portal Headword List" (http://hdl.handle.net/11356/1445)
Glossary (EN-LT) of Terms with the Adjective 'Green' from EUR-Lex (ELEXIS)
Bilingual (EN-LT) glossary of the EU English terms with the adjective 'green' and their Lithuanian equivalents
Glossary (LT, EN, DA, NO) of Cybersecurity Terms (ELEXIS)
Quadrilingual (EN-LT-DA-NO) glossary of cybersecurity terms
Bulgarian Dictionary of Synonyms (ELEXIS)
Синонимен речник.
The Dictionary of Synonyms in Bulgarian contains about 29,998 words pertaining to four parts-of-speech, (10,704 nouns, 7,070 verbs, 8,763 adjectives, and 3,461 adverbs). distributed into 8,862 synonym sets