Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
840 research outputs found
Sort by
GLOBAL Greek-French Dictionary - MLDS (ELEXIS)
A general language Greek to French dictionary
GLOBAL Japanese-French Dictionary - MLDS (ELEXIS)
A general language Japanese to French dictionary
GLOBAL Portuguese-French Dictionary - MLDS (ELEXIS)
A general language Portuguese to French dictionary
Dictionary of Old Norse Prose - ONP (ELEXIS)
Ordbog over det norrøne prosasprog.
The Dictionary of Old Norse Prose (ONP) is an online dictionary containing over 50,000 words. The headwords are in Old Norse, while the definitions are in Danish and/or English. The dictionary covers medieval Old Norse material found in prose texts from the oldest written documents to early modern times (Old Icelandic 1150-1540 and Old Norwegian 1150-1370). After the publication of the first four volumes of an intended thirteen volume printed edition, a decision was made to change the format of the dictionary to a digital edition available online. Published by the University of Copenhagen, the project is ongoing as of 2020 with new words added regularly. The dictionary is searchable in a number of ways: by headword, text, manuscript, type of word, and numerous others
Corpus of Serbian Forms of Address 1.0
The corpus consists of transcripts of audio-recorded biographical interviews with 19 participants. The interviews are about forms of address that speakers use in colloquial and in formal settings, and about their attitudes and evaluations concerning particular forms of address. We provide original transcripts (written according to GAT conventions) and TEI-XML transcripts. The corpus has been normalised, tagged with morphosyntactic and lemma information, and aligned with the respective turns in the audio files. The audio files are available on request
School dictionary of Slovenian language (semantic database)
The "School Dictionary of the Slovenian Language" ("Šolski slovar slovenskega jezika") includes 2045 dictionary entries and is aimed at pupils aged 6 to 10. It provides the most relevant language information for the target user group. The dictionary examines only standard-language lexical units and includes multi-word lexemes, both non-phraseological and phraseological.
The main source of materials for the dictionary is the "Corpus of Slovenian school texts" (http://hdl.handle.net/11356/1413). The dataset is linked to the "Franček Portal Headword List" (http://hdl.handle.net/11356/1445). The etymological part of the dictionary is available separately as "School Dictionary of Slovenian Language (etymological database)" (http://hdl.handle.net/11356/1452).
Note: some entries are copies of each other except for references to different entries in the (Franček Portal Headword List)
Parallel corpus EN-SL RSDO4 1.0
The RSDO4 parallel corpus of English-Slovene and Slovene-English translation pairs was collected as part of work package 4 of the Slovene in the Digital Environment project. It contains texts collected from public institutions and texts submitted by individual donors through the text collection portal created within the project. The corpus consists of 964433 translation pairs (extracted from standard translation formats (TMX, XLIFF) or manually aligned) in randomized order which can be used for machine translation training
Dictionary of Arghezian Terms (Letters A-F) (ELEXIS)
The dataset contains lexical items representing the poetic vocabulary encountered in the work of the Romanian poet Tudor Argezi. Each entry contains the number of occurences (entries with more than 10-20 occurences are considered thematic foci in the author's literary work), the grammatical category, etymology and literary interpretations (proper and figurative meaning assignements). Each lexeme is clarified by literary citations that exemplify occurence patterns
TED-ELH Parallel Corpus (ELEXIS)
The corpus contains parallelly aligned scripts of TED Talks in English, Lithuanian, and Hebrew. It contains spoken language data.
See also: http://hdl.handle.net/20.500.11821/3
Parallel Corpus (EN-LT-DA) of General Data Protection Regulation (ELEXIS)
Trilingual parallel corpus on general data protection regulation. The size of the corpus is 54,468 words in English, 42,566 words in Lithuanian, and 47,740 words in Danish.
The resource was available at https://hmi.mruni.eu/languageresources/index.php/apps/cms_pico/pico/bachelor_mikneviciute but the domain seems to no longer be active