Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
840 research outputs found
Sort by
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Serbian 1.0
The model for lemmatisation of non-standard Serbian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200), the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the hr500k training corpus (http://hdl.handle.net/11356/1210) and the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/), using the srLex inflectional lexicon (http://hdl.handle.net/11356/1233). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~97.62
The CLASSLA-StanfordNLP model for lemmatisation of standard Serbian 1.2
The model for lemmatisation of standard Serbian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200) and using the srLex inflectional lexicon (http://hdl.handle.net/11356/1233). The estimated F1 of the lemma annotations is ~97.9.
The difference to the previous version is that now it relies solely on XPOS annotations, and not on a combination of UPOS, FEATS (lexicon lookup) and XPOS (lemma prediction) annotations
Annotated corpus of Serbian language-related news articles MetaLangNEWS-Sr
A comprehensive corpus of news articles on the topic of language, published in major Serbian daily newspapers and news portals in the five-year period of January 1, 2015 - January 1, 2020. The corpus is designed to facilitate research on metalanguage (‘language about language’), linguistic ideologies, language policy and planning, as well as the specific contemporary debates on language defining, naming, and standardisation, ongoing in post-Yugoslav societies.
The corpus has been tagged using the CLASSLA-StanfordNLP models for morphosyntactic annotation and lemmatisation of standard Serbian. The corpus is available in plain text version, XML with full metadata, and tagged CONLL-U format.
MetaLangNEWS-Sr is complemented with a separate corpus of citizen metalanguage comments, i.e. online comments to the news articles, available as MetaLangNEWS-COMMENTS-Sr (http://hdl.handle.net/11356/1372). Parallel versions from Slovenia (http://hdl.handle.net/11356/1360) and Croatia (http://hdl.handle.net/11356/1369) are also available
A Machine-readable Dictionary of Damascus Arabic - dc-apc-eng (ELEXIS)
This dictionary has been prepared to support the Syrian Textbook prepared at the University of Vienna.
See also: https://hdl.handle.net/11022/0000-0007-C093-
A Machine-readable Dictionary of Dagaare - dc-dga-yue (ELEXIS)
The Dagaare - English Lexicon was compiled by Adams Bodomo and published in this form in 2015. This edition of the lexicon comprises more than 1250 entries. Headwords or entries contain other words in paradigmatic relation to the headword. Thus the dictionary actually comprises 3000-4000 words. Even though tone is not indicated in standard Dagaare orthography, tonal markings are indicated for each entry. This is followed by categorial, and, where necessary, subcategorial information, such as intransitive verb, question word, etc. of the entry. A salient feature of the lexicon is a comprehensive provision of verbal and nominal paradigms for each verbal and nominal headword. These paradigms serve as important indicators of the salient aspects of Dagaare grammar
New Idioticon Viennense - Loritza (ELEXIS)
Neues Idioticon Viennense.
Digitized version of a historic dialect dictionary of Viennese (1847)
Dictionary and Thesaurus of Latvian - Tezaurs.lv (ELEXIS)
Tēzaurs.lv: An extensive dictionary and thesaurus of Latvian, comprising more than 320,000 lexical entries, including multi-word units. Compiled and edited based on more than 300 sources. Provides detailed morphological information; being extented into a Latvian WordNet
Consonant-vowel structures in the GOS 1.0 corpus
The lists contain consonant-vowel structures of all lemmas, word forms, and normalized word forms in the GOS 1.0 Corpus of Spoken Slovene (http://hdl.handle.net/11356/1040). In each unit, its characters were converted as follows: C - consonant (in lists with finegrained character categorizations, consonants were divided into Z - sonorant, G - voiced obstruent, and K - voiceless obstruent), V - vowel, X - foreign consonant, Y - foreign vowel, S - symbol, P - punctuation, N - number, F - non-Latin-script character, ! - other.
Each consonant-vowel structure also contains its frequency in the corpus (i.e. the total sum of the frequencies of all units corresponding to the consonant-vowel structure), as well as the set of all units (in the lists labeled "entire") or the set of its 30 most frequent units (in the lists labeled as "short"), along with their part-of-speech categories and their individual frequencies). They also contain the number of all unique units within the consonant-vowel structure.
The lists were prepared based on frequency lists extracted from GOS 1.0 using LIST: http://hdl.handle.net/11356/1276
Note that there exists a related resource, "Consonant-vowel structures in the Gigafida 2.0 corpus", http://hdl.handle.net/11356/128
List of formulaic sequences in spoken Slovenian
This document contains 2,374 formulaic sequences in spoken Slovenian, i.e. frequently recurring strings of two to five words, manually annotated for syntactic structure, pragmatic function, and dictionary relevance. The list of sequences with a minimum frequency threshold of 20/million is based on the Frequency lists of word-level n-grams from normalized word forms in GOS 1.0 (http://hdl.handle.net/11356/1271) and contains the union of top-1,000 formulaic sequences ranked by frequency and five association measures (Dice, t-test, MI, MI3, simple-LL).
Note that there exists a related entry, "List of formulaic sequences in standard written Slovenian", http://hdl.handle.net/11356/1280
School Dictionary of the Croatian Language (ELEXIS)
Školski rječnik hrvatskoga jezika.
The printed edition of the School Dictionary of the Croatian Language was published in 2012 as a result of the project Croatian Normative One-Volume Dictionary. School Dictionary of the Croatian Language was not a born-digital dictionary but it was updated and adapted to web publishing and published online in 2020 (http://rjecnik.hr/). It is a normative dictionary of the contemporary standard Croatian language consisting of 30,000 entries. It is based on a corpus of elementary and high school textbooks from which the lexicographers manually extracted the alphabetical list of entries. It was written consulting the Croatian Language Repository corpus (http://riznica.ihjj.hr) as well as the Internet, i.e. all entry words were checked in the corpus for examples and collocations but examples and collocations were not directly taken from the corpus. It was written in the Softlex dictionary compilation program. While compiling the dictionary special attention was paid to semantic relations between entries and meanings and synonyms and antonyms are connected by links