Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
840 research outputs found
Sort by
Glossary of equestrianism
The glossary consists of terms related to bacis horse and rider equipment as well as five equestrian disciplines: dressage, jumping, driving, harness racing and outdoor riding. The glossary will aid users who alredy have some knowledge of equestrianism with finding the English term equivalent in Slovenian. A short explanation, meant for those who lack experience in equestrianism also accompanies each term.
The dictionary is distributed in XML using the TBX (TermBase eXchange) standard for representing and exchanging information from termbases
ŠUSS archive of questions and answers about the Slovenian language (1998-2010)
This corpus contains the Q&A archive of the ŠUSS language consultancy service. The ŠUSS internet forum was active 1998-2010. Questions posted by users were answered by a group of students of linguistics and related disciplines. Answers were often given with extensive commentary and references. After the end of its active life the articles were stored in an online acrhive, and this corpus is made from the XHTML files of the archive.
The corpus is encoded in TEI in two versions, the text one with structural annotations, and the linguistically annotated one with a simplified structure and automatic word-level annotations for lemmas and morphosyntactic descriptions. In cooperation with DARIAH-SI, the text version is also available as a digital library
Keywords and n-grams from a textbook corpus
Wordlists, keywords and n-grams were extracted from a corpus of textbooks for Slovenian elementary and secondary schools. The corpus contains 4,302,857 words (5,373,268 tokens), and consists of 127 textbooks from 16 different subjects:
- Biology (6 textbooks; 293,935 words),
- State, society and ethics (1 textbook; 21,881 words),
- Society (4 textbooks; 64,126),
- Physics (5 textbooks; 185,171),
- Geography (7 textbooks; 202,101 words),
- Music (8 textbooks; 224,034 words),
- Home Economics (3 textbooks; 33.803),
- Chemistry (7 textbooks; 282,543 words),
- Art (3 textbooks; 146,681),
- Mathematics (23 textbooks; 764,012),
- Science (5 textbooks; 226,191 words),
- Science and technology (6 textbooks; 183,749 words),
- Slovene language (37 textbooks; 1,437,945 words),
- Environmental Education (7 textbooks; 38,645 words),
- Technology (1 textbook; 24,733 words)
- History (4 textbooks; 173,307 words).
The lists were manually cleaned, most items not found in the reference morphological lexicon Sloleks (http://hdl.handle.net/11356/1039) were removed, which mainly consisted of conversion errors.
The lists include only those words, keywords or n-grams that were found in at least 8 different subjects. Keyword lists were extracted using the Sketch Engine tool, minimum frequency was set to 5, the statistics used was average relative frequency. Minimum frequency for n-grams was 10
The Dictionary of the Clothing Terminology of the Zilja Dialect in Canale Valley (Kanalska dolina – Val Canale – Kanaltal – Valcjanâl): photographs
The collection of illustrative photographs for The Dictionary of the Clothing Terminology of the Zilja Dialect in Canale Valley (Kanalska dolina – Val Canale – Kanaltal – Valcjanâl) (dictionary: http://hdl.handle.net/11356/1217, audio: http://hdl.handle.net/11356/1220) contains 76 photographs of traditional clothing items. They were selected from the archive of the ethnographic research of clothing culture in Canale Valley in 20th century, conducted by the Slovenian Cultural Centre (SKS) Planika Kanalska dolina in years 2003–2014. The same archive material is the source for the virtual clothing exhibition Glasovi Kanalske doline (The voices of Canale Valley) of the cultural heritage project Zborzbirk (https://as.parsis.si/zborzbirk/zbirka.a5w?zid=1040)
Inflectional lexicon srLex 1.3
srLex is a large inflectional lexicon of Serbian language where each entry consists of a (wordform, lemma, MSD, MSD features, UPOS, morphological features, frequency, per-million frequency) 8-tuple. The (wordform, lemma, MSD) triple frequencies are calculated on the srWaC v1.2 corpus. The MSD tagset follows the MULTEXT-East V6 tagset for the Serbo-Croatian macro-language available at http://nl.ijs.si/ME/V6/msd/html/msd-hbs.html. The UPOS + morphological features follow the UD v2 specifications available at http://universaldependencies.org/guidelines.html
Slovene corpus for aspect-based sentiment analysis - SentiCoref 1.0
SentiCoref 1.0 corpus consists of 837 documents selected from SentiNews 1.0 corpus (http://hdl.handle.net/11356/1110). The documents were selected based on the number of automatically detected named entities (using Polyglot, https://polyglot.readthedocs.io/) which contained between 50 and 73 named entities.
The corpus is provides an initial dataset for aspect-based sentiment analysis. The annotations consist of named entities (persons, organizations and locations), coreferences to the named entities, and 5-level sentiment annotation for each entity (coreference chain). Together there are 31,419 manually tagged named entities - 15,285 organizations, 8,606 persons and 7,528 locations. The dataset contains 14,572 coreference chains. Sentiment distribution for entities is as follows - 30 Very negative, 1801 Negative, 10869 Neutral, 1705 Positive and 24 Very positive.
Each document was annotated by two linguist students. In the preparation of the dataset, 8 students participated: Rednak Pia, Roblek Rebeka, Jelovšek Tjaša, Agović Haris, Vaupotič Jana, Grego Annamaria, Vidic Zala, Žvanut Kaja. The final curation was done by Neli Blagus and Slavko Žitnik.
The data is in WebAnno TSV 3 format (similar to CoNLL format) which is compatible with the WebAnno tool (https://webanno.github.io/webanno/)
The Dictionary of the Clothing Terminology of the Zilja Dialect in Canale Valley (Kanalska dolina – Val Canale – Kanaltal – Valcjanâl)
The Dictionary of the Clothing Terminology of the Zilja Dialect in Canale Valley (Kanalska dolina – Val Canale – Kanaltal – Valcjanâl) is the result of dialectological research team-work, conducted between 2003–2014. It was designed as a pilot project for the documentation of endangered variety of Gailtal dialect in the linguistically mixed area of Canale Valley. Considering the importance of preservation of the words for culturally relevant terms and concepts, the topic was chosen by the members of a speech community themselves. The sound material was gathered with various fieldwork techniques, from workshops on the old clothing dialect terms, classic dialect elicitations and guided conversations to monitoring spontaneous speech. As an additional source we used spoken texts, prepared for Slovenian Broadcasting Radio Trst (Trieste) A and the cultural heritage project Zborzbirk.
The dictionary contains 594 entries and is based on approximately 1,400 sound clips from around 16 hours of recordings. It is formatted as a trilingual (Slovene-German-Italian) concordance-based dictionary and of which the most common collocations are presented alongside clothing terminology. Each entry is equipped with sound clips for dialect lemma and examples of dialect use (http://hdl.handle.net/11356/1220). The encyclopaedic information is contained in illustrative photographs of traditional clothing items (http://hdl.handle.net/11356/1221) or added as occasional commentary subsection.
The documentary section presents the relations between the dialect and standard lexicon and documents the potential geographical area of the word. The multimedia lexical database is organised as a source for grammatical description of the local dialect varieties, especially phonology and morphology.
The morphological section of the dictionary thus contains all word forms, confirmed in the speech corpus, with special attention to individual phonetic and inflectional variation. For that purpose, the narrow phonetic transcription (Slovene dialectological transcription) was used for transcribing dialect material.
The previous editions were published as printed book: Kenda-Jež, Karmen, Shranli smo jih v bančah: slovarski prispevek k poznavanju oblačilne kulture v Kanalski dolini = contributo lessicale alla conoscenza dellʼabbigliamento in Val Canale, Ukve: S.K.S. Planika Kanalska dolina; [s. l.]: Slori, ATS Od me-je; Ljubljana: Inštitut za slovenski jezik Frana Ramovša ZRC SAZU = Istituto per la lingua slovena “Fran Ramovš” CRS ASSA, ¹2007; Ljubljana: Založba ZRC, ZRC SAZU = Casa editrice CRS, CRS ASSA, ²2015
Spoken corpus Gos VideoLectures 4.0 (transcription)
Gos VideoLectures is an add-on to the Gos reference corpus of spoken Slovene (http://hdl.handle.net/11356/1040), and covers public academic speech.
The Gos VideoLectures corpus contains a selection of public lectures available through the web portal Videolectures.net provided by the Jožef Stefan Institute, and covers 55 lectures and 22 hours of speech.
This resource contains only annotated transcriptions of the corpus – audio recordings are available at http://hdl.handle.net/11356/1222.
The transcriptions for Gos VideoLectures were done manually and carefully checked. The main guidelines for transcription were those of the Gos corpus (http://www.korpus-gos.net/Support/About). The transcription tool Transcriber 1.5.1 (http://trans.sourceforge.net/en/presentation.php) was used for making transcriptions. It can be also used for reading or exporting transcriptions (.trs files) to different formats.
The transcriptions comprise the TRS files with tabular metadata, their conversion to TEI and to vertical file format (as used e.g. by Sketch Engine). Each recording has two TRS files, one with pronunciation-based and the other with the standardised/normalised transcription. The TEI and CWB encodings join these two transcriptions at the token level, with the normalised words being also automatically PoS tagged and lemmatised. The TRS pack also contains files with automatically produced word and phone-level alignment with the speech signal.
The corpus can be used for training continuous speech recognition for Slovene language, for phonetic research or any other research of Slovene academic speech
Serbian Twitter training corpus ReLDI-NormTagNER-sr 2.1
ReLDI-NormTagNER-sr 2.1 is a manually annotated corpus of Serbian tweets. It is meant as a gold-standard training and testing dataset for tokenisation, sentence segmentation, word normalisation, morphosyntactic tagging, lemmatisation and named entity recognition of non-standard Serbian. Each tweet is also annotated for its automatically assigned standardness levels (T = technical standardness, L = linguistic standardness).
As an update to version 2.0, version 2.1 corrects some annotation errors and adds morphosyntactic annotations in the Universal Dependencies formalism in addition to the MULTEXT-East morphosyntactic descriptions. The corpus is now also available in CoNLL-U format
The CLASSLA-StanfordNLP model for lemmatisation of standard Croatian
The model for lemmatisation of standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1183) and using the hrLex inflectional lexicon (http://hdl.handle.net/11356/1232). The estimated F1 of the lemma annotations is ~97.6