Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    GLOBAL Polish-French Dictionary - MLDS (ELEXIS)

    No full text
    A general language Polish to French dictionary

    Latvian Delfi article archive (in Latvian and Russian) 1.0

    No full text
    This dataset is an archive of articles from the Delfi news site from 2015-2019, containing over 180,000 articles (c. 50% in Latvian and 50% in the Russian language). Keywords for articles are included. There are 5 JSON files: lv_2015.json contains 42 001 articles from the year 2015 lv_2016_.json contains 40 342 articles from the year 2016 lv_2017_.json contains 37 256 articles from the year 2017 lv_2018_.json contains 31 732 articles from the year 2018 lv_2019_.json contains 29 070 articles from the year 2019 In sum: 180 401 articles Description of the dataset This JSON file is a list of dictionaries, i.e. each article is represented as a dictionary. Each dictionary contains the following: id (integer) - the ID of the article title (string) - the title of the article lead (string) - the lead of the article tags [1] (list of dictionaries or None): each dictionary represents one tag. The tag dictionary contains the following: domain_id (string) - the ID of the domain id (string) - the ID of the tag lang (string) - the language of the tag tag (string) - the tag itself, e.g. Šokolāde translitted_name (string) - a modified version of the tag, e.g. sokolade rawBody (string) - the raw text of the article (contains HTML) bodyText (string) - clean article text (stripped from HTML) publishDate (string) - published date & time of the article categoryPrimary (dictionary or empty list) - the dictionary contains the following information: categoryId (integer) - the ID of the category categoryName (string)- the name of the category (e.g. Futbols) channelId (integer) - the ID of the channel groupId - None channelLanguage (string) - the language of the channel (nat - Latvian, rus - Russian) categoryLanguage (integer) - ID of the channel language relatedArticles (list of integers or None) - a list of related articles' ID's relatedTags(string or None) -- related tags are comma-separate

    Keyword extraction datasets for Croatian, Estonian, Latvian and Russian 1.0

    No full text
    EACL Hackashop Keyword Challenge Datasets In this repository you can find ids of articles used for the keyword extraction challenge at EACL Hackashop on News Media Content Analysis and Automated Report Generation (http://embeddia.eu/hackashop2021/). The article ids can be used to generate train-test split used in paper: Koloski, B., Pollak, S., Škrlj, B., & Martinc, M. (2021). Extending Neural Keyword Extraction with TF-IDF tagset matching. In: Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation, Kiev, Ukraine, pages 22–29. Train and test splits are provided for Latvian, Estonian, Russian and Croatian. The articles with the corresponding ID-s can be extracted from the following datasets: - Estonian and Russian (use the eearticles2015-2019 dataset): https://www.clarin.si/repository/xmlui/handle/11356/1408 - Latvian: https://www.clarin.si/repository/xmlui/handle/11356/1409 - Croatian: https://www.clarin.si/repository/xmlui/handle/11356/1410 dataset_ids folder is organized in the following way: - latvian – containing latvian_train.json: a json file with ids from train articles to replicate the data used in Koloski et al. (2020), the latvian_test.json: a json file with ids from test articles to replicate the data - estonian – containing estonian_train.json: a json file with ids from train articles to replicate the data used in Koloski et al. (2020), the estonian_test.json: a json file with ids from test articles to replicate the data - russian – containing russian_train.json: a json file with ids from train articles to replicate the train data used in Koloski et al. (2020), the russian_test.json: a json file with ids from test articles to replicate the data - croatian - containing croatian_id_train.tsv file with sites and ids (note that just ids are not unique across dataset, therefore site information also needs to be included to obtain a unique article identifier) of articles in the train set, and the croatian_id_test.tsv file with sites and ids of articles in the test set. In addition, scripts are provided for extracting articles (see folder parse containing scripts parse.py and build_croatian_dataset.py, requirements for scripts are pandas and bs4 Python libraries): parse.py is used for extraction of Estonian, Russian and Latvian train and test datasets: Instructions: ESTONIAN-RUSSIAN 1) Retrieve the data ee_articles_2015_2019.zip 2) Create a folder 'data' and subfolder 'ee' 3) Unzip them in the 'data/ee' folder To extract train/test Estonian articles: run function 'build_dataset(lang="ee", opt="nat")' in the parse.py script To extract train/test Russian articles: run function 'build_dataset(lang="ee", opt="rus")' in the parse.py script LATVIAN: 1) Retrieve the latvian data 2) Unzip it in 'data/lv' folder 3) To extract train/test Latvian articles: run function 'build_dataset(lang="lv", opt="nat")' in the parse.py script build_croatian_dataset.py is used for extraction of Croatian train and test datasets: Instructions: CROATIAN: 1) Retrieve the Croatian data (file 'STY_24sata_articles_hr_PUB-01.csv') 2) put the script 'build_croatian_dataset.py' in the same folder as the extracted data and run it (e.g., python build_croatian_dataset.py). For additional questions: {Boshko.Koloski,Matej.Martinc,Senja.Pollak}@ijs.s

    Croatian corpus of non-professional written language by typical speakers and speakers with language disorders RAPUT 1.0

    No full text
    The corpus consists of texts produced by nonprofessional typical speakers and speakers with different language disorders (developmental language disorder, dyslexia, traumatic brain injury, aphasia, other). Roughly half of the corpus consists of texts of typical speakers, and the other half of speakers with language disorders. Language samples were elicited by six groups of tasks representing different writing styles (descriptive, expository, narrative, and letter) and different levels of formality. The corpus has been manually annotated for normalized forms, lemmas, morphosyntactic information (by following the MULTEXT-East tagset), and type of error (phonological segmentation, orthography, non-standard spelling, typo, syntax, etc.). UD morphosyntactic description has been to the most part automatically generated from the MULTEXT-East morphosyntactic information

    Franček portal dialect module

    No full text
    The Franček Portal Dialect Module contains data on dialect variation of select lexemes and/or their meanings including labels describing their respective relative frequencies and spatial distribution. The underlying data stems from the Slovenian Linguistic Atlas, volumes 1 and 2. The content has been transformed and adjusted for primary- and secondary-school users. The dataset is linked to the "Franček Portal Headword List" (http://hdl.handle.net/11356/1445)

    Comprehensive Slovenian-Hungarian Dictionary 1.0

    No full text
    The Comprehensive Slovenian-Hungarian dictionary is a general bilingual dictionary that is being compiled at the Centre for Language Resources and Technologies of the University of Ljubljana (CJVT UL). Version 1.0 contains 10,946 headwords, 33,298 translations, 15,265 collocations and other word combinations, and 2,416 examples. The file also contains links between synonymous entries or entry senses, and links between single-word headwords and compounds/phrases. The Comprehensive Slovenian-Hungarian dictionary is a growing dictionary, which means that new headwords will be added in regular intervals. The Comprehensive Slovenian-Hungarian dictionary is based on a concept (Kosem et al. 2018) that was prepared in the targeted research project KOMASS (the Concept of Hungarian-Slovenian dictionary: from a language resource to its user), funded by the Slovenian Research Agency and the Ministry of Education, Science and Sport of the Republic of Slovenia. The dictionary concept follows the state-of-the-art international lexicographic practice, e.g. bilingual dictionaries compiled at established international publishers and institutes

    eSSKJ: Dictionary of the Slovenian standard language (phraseological database): school edition

    No full text
    The eSSKJ: Dictionary of the Slovenian Standard Language (phraseological database): School Edition represents a part of phraseological entries in the educational language portal Franček. The underlying data stems from the eSSKJ: Dictionary of the Slovenian Standard Language. The content is adjusted for primary- and secondary-school users. The dataset is linked to the "Franček Portal Headword List" (http://hdl.handle.net/11356/1445)

    Glossary (EN-LT) of Terms with the Adjective 'Green' from EUR-Lex (ELEXIS)

    No full text
    Bilingual (EN-LT) glossary of the EU English terms with the adjective 'green' and their Lithuanian equivalents

    Glossary (LT, EN, DA, NO) of Cybersecurity Terms (ELEXIS)

    No full text
    Quadrilingual (EN-LT-DA-NO) glossary of cybersecurity terms

    Bulgarian Dictionary of Synonyms (ELEXIS)

    No full text
    Синонимен речник. The Dictionary of Synonyms in Bulgarian contains about 29,998 words pertaining to four parts-of-speech, (10,704 nouns, 7,070 verbs, 8,763 adjectives, and 3,461 adverbs). distributed into 8,862 synonym sets

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇