Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
840 research outputs found
Sort by
Croatian parliamentary corpus ParlaMeter-hr 1.0
The ParlaMeter-hr corpus contains minutes of the National Assembly of the Republic of Croatia and currently covers its VIth mandate (2016-11-15 - 2018-11-21). The corpus contains speaker metadata (gender, age, education, party affiliation), while the transcriptions of their speeches are MSD tagged, lemmatised, and marked with named entities
The Dictionary of the Clothing Terminology of the Zilja Dialect in Canale Valley (Kanalska dolina – Val Canale – Kanaltal – Valcjanâl): audio
The collection of sound clips for The Dictionary of the Clothing Terminology of the Zilja Dialect in Canale Valley (Kanalska dolina – Val Canale – Kanaltal – Valcjanâl) (dictionary: http://hdl.handle.net/11356/1217, photographs: http://hdl.handle.net/11356/1221) consists of two sets of files in the WAV format – the pronunciation sound clips for dialect lemmas (383 files) and text recordings associated with dictionary examples (698 files). The sound clips were cut out from around 16 hours of field recordings, serving as the main data source for the dictionary
Corpus extraction tool LIST 1.0
The LIST corpus extraction tool is a Java program for extracting lists from text corpora on the levels of characters, word parts, words, and word sets. It supports VERT and TEI P5 XML formats and outputs .CSV files that can be imported into Microsoft Excel or similar statistical processing software
Multilingual Culture-Independent Word Analogy Datasets
Word analogy task evaluates word embeddings, based on analagous word pairs (eg. "Paris - France" should be equivalent to "Rome - Italy", "son - daughter" should be equivalent to "brother - sister"). The dataset has been inspired by Mikolov's analogy test set in English (http://download.tensorflow.org/data/questions-words.txt). It was first written for Slovenian and then partially translated and partially done from scratch for the other languages (Croatian, Finnish, Estonian, Swedish, Latvian, Lithuanian, Russian and English).
The analogy dataset is composed of fifteen categories, five semantical and ten syntactical. Each dataset has about 19,000 entries.
In addition to nine monolingual datasets (one for each language), we also composed 72 cross-lingual datasets (one for each language pair), where one half of the entry (one analogy, eg. "mother-father") is in one language and the other half of the entry (eg. "sister-brother") is in another language
Corpus of Written Standard Slovene Gigafida 2.0
Gigafida 2.0, with about 1.1 billion words, is a reference corpus of written Slovene text published in the period 1990-2018. It is comprised of daily news, magazines, a selection of web texts (a certain portion of which covers news texts as well), and different types of publications (fiction, school books, and non-fiction). The texts have been selected and automatically processed with the aim of creating a corpus that represents a sample of modern standard Slovene and can be used for research in linguistics and other branches of the humanities, for compiling modern dictionaries, grammars, and learning materials, as well as for developing language technologies for Slovene.
Gigafida 2.0 is an upgraded version of the Gigafida corpus (Logar et al. 2012), which was made publicly available in 2012 at www.gigafida.net. Unlike its predecessor, Gigafida 2.0 is a corpus of standard Slovene, as the majority of texts containing non-standard language features (such as user comments from news forums, etc.) was removed during the upgrade. Other improvements include the removal of duplicate texts and text fragments, an improved automatic linguistic tagging, and several new features in the interface design. All of the above is described in more detail in the corpus specifications (https://www.cjvt.si/gigafida/wp-content/uploads/sites/10/2019/06/Gigafida2.0_specifikacije.pdf).
The Gigafida 2.0 corpus is comprised predominantly of newspapers, internet texts, and magazines. The texts published after the release of the first version of the corpus (2012–2018) represent approximately 27% of the new corpus.
For linguistic research, the entire corpus is freely accessible online in the NoSketchEngine and Kontext concordancers, as well as the commercial SketchEngine tool.
References:
Simon Krek, Špela Arhar Holdt, Tomaž Erjavec, Jaka Čibej, Andraz Repar, Polona Gantar, Nikola Ljubešić, Iztok Kosem and Kaja Dobrovoljc. Gigafida 2.0: The Reference Corpus of Written Standard Slovene. Proceedings of The 12th Language Resources and Evaluation Conference. Marseille, May 2020. https://www.aclweb.org/anthology/2020.lrec-1.409/
LOGAR BERGINC, Nataša, GRČAR, Miha, BRAKUS, Marko, ERJAVEC, Tomaž, ARHAR HOLDT, Špela and KREK, Simon. Korpusi slovenskega jezika Gigafida, KRES, ccGigafida in ccKRES: gradnja, vsebina, uporaba. Ljubljana: Trojina, zavod za uporabno slovenistiko; Fakulteta za družbene vede, 2012
Spoken corpus Gos VideoLectures 4.0 (audio)
Gos VideoLectures is an add-on to the Gos reference corpus of spoken Slovene (http://hdl.handle.net/11356/1040), and covers public academic speech. The Gos VideoLectures corpus contains a selection of public lectures available through the web portal Videolectures.net provided by the Jožef Stefan Institute, and covers 55 lectures with 22 hours of speech.
This resource contains only audio recordings of the corpus – annotated transcriptions are available at http://hdl.handle.net/11356/1444.
The recordings are available for each lecture separately, as well as split for utterance, segment, and words
Developmental corpus Šolar 2.0
The Developmental corpus Šolar 2.0 consists of 5,485 texts written by students in Slovene secondary schools (age 15-19) and pupils in the 7th-9th grade of primary school (13-15), with a small percentage also from the 6th grade. School essays form the majority of the corpus while other material includes texts created during lessons, such as text recapitulations or descriptions, examples of formal applications etc. Most of the texts were produced at the subject of the Slovenian language.
Part of the corpus (2,094 texts) is annotated with teachers' corrections using a system of labels described in the attached document (in Slovene). Teacher corrections were part of the original files and reflect real classroom situations of essay marking. Corrections were then inserted into texts by annotators, and subsequently categorized.
This corpus also exists in two derived versions, Šolar Clear (http://hdl.handle.net/11356/1219), which contains only the text of the students without the teacher corrections, and Šolar Error (http://hdl.handle.net/11356/1231), which contains only those sentecens that have teacher corrections
Developmental corpus (without language corrections) Šolar 2.0 Clear
Šolar 2.0 Clear is an adapted version of the Šolar 2.0 corpus, cf. http://hdl.handle.net/11356/1214.
The Šolar 2.0 Clear corpus consists of texts written by students in Slovene primary and secondary schools. School essays form the majority of the corpus while other material includes texts created during lessons, such as text recapitulations or descriptions, examples of formal applications etc. For each text, the information on school (elementary or secondary), subject, level (grade or year), type of text, region and date of production is provided.
Unlike the original Šolar 2.0 corpus (http://hdl.handle.net/11356/1214), Šolar 2.0 Clear includes student texts only: error annotations and other types of feedback from the teachers have been removed. The corpus can thus be used for processing tasks where the inclusion of corrections hinders or complicates the procedures (e.g. for comparative data extraction, training of language models etc)
Frequency lists of word parts from the Gigafida 2.0 corpus
Frequency lists of words split into word parts were extracted from the Gigafida 2.0 Corpus of Written Standard Slovene (https://viri.cjvt.si/gigafida/) using the LIST corpus extraction tool (http://hdl.handle.net/11356/1227). The lists contain all lemmas or lower-case word forms occurring in the corpus, split into their initial or final part (i.e. the initial or final string of 1, 2, 3, 4 or 5 characters in the word) and the rest of the word. In addition, the lists also contain absolute and relative frequencies, percentages, and distribution across the text-types included in the corpus taxonomy.
The lists were extracted for each part-of-speech category. For each part-of-speech, a total of 20 lists were extracted:
1) 10 lists for initial or final word parts extracted from lemmas,
2) 10 lists for initial or final word parts extracted from lower-case word forms.
In addition, 20 lists were extracted from all words (regardless of their part-of-speech category). For easier processing in statistical analysis software, shortened versions of longer lists were made containing the first 150,000 lines
Corpus of Informatics DSI 5.0
The DSI corpus is meant as a terminological resource for the field of informatics, esp. for the development of the on-line terminological dictionary of informatics, Islovar (www.islovar.org). The corpus contains articles from the proceedings of the conference series 2003-2019 "Dnevi slovenske informatike" (Days of Slovene Informatics), proceedings of the conference series 2015-2018 "Informatika v javni upravi" (Informatics in Public Administration), and volumes 2010-2019 of the journal "Uporabna informatika" (Applied Informatics)