Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
840 research outputs found
Sort by
The CLASSLA-StanfordNLP model for lemmatisation of standard Slovenian 1.1
The model for lemmatisation of standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and using the Sloleks inflectional lexicon (http://hdl.handle.net/11356/1230). The estimated F1 of the lemma annotations is ~99.0.
The difference to the previous version of the model is that it is trained with the lemmatiser padding bug removed, cf. https://github.com/stanfordnlp/stanfordnlp/issues/143
The CLASSLA-StanfordNLP model for lemmatisation of standard Bulgarian 1.0
The model for lemmatisation of standard Bulgarian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the BulTreeBank training corpus (http://hdl.handle.net/11495/D93F-C6E9-65D9-2) and using the Bulgarian inflectional lexicon (Popov, Simov, and Vidinska 1998). The estimated F1 of the lemma annotations is ~98.8
Frequency lists of word parts from the GOS 1.0 corpus 1.1
Frequency lists of words split into word parts were extracted from the GOS 1.0 Corpus of Spoken Slovene (http://hdl.handle.net/11356/1040) using the LIST corpus extraction tool (http://hdl.handle.net/11356/1227). The lists contain all lemmas, lower-case word forms or standardized word forms occurring in the corpus, split into their initial or final part (i.e. the initial or final string of 1, 2, 3, 4 or 5 characters in the word) and the rest. In addition, the lists also contain absolute and relative frequencies, percentages, and distribution across the text-types included in the corpus taxonomy.
The lists were extracted for each part-of-speech category. For each part-of-speech, a total of 30 lists were extracted:
1) 10 lists for initial or final word parts extracted from lemmas,
2) 10 lists for initial or final word parts extracted from lower-case word forms,
3) 10 lists for initial or final word parts extracted from standardized word forms.
In addition, 30 lists were extracted from all words (regardless of their part-of-speech category).
Compared to the previous version (http://hdl.handle.net/11356/1270), this one includes fixes of several typos and substitutes all instances of "normalized forms" with the more adequate term "standardized forms" (as used in the SSJ project)
Slovene translation of SuperGLUE
SuperGLUE is a benchmark styled after GLUE with a new set of more difficult language understanding tasks, improved resources, and a public leaderboard. It is comprised of 8 corpora (BoolQ, CB, COPA, MultiRC, ReCoRD, RTE, WiC, WSC), which cover 4 different types of tasks (QA, NLI, WSD, coref.). Slovene translation of SuperGLUE consists of machine and human translations of the benchmark. BoolQ, CB, COPA, MultiRC, and RTE are completely translated by the Google Machine Translation service. Human translators translated COPA and WSC completely, and BoolQ, CB, MultiRC, ReCoRD, and RTE in different ratios. Slovene translation of SuperGLUE is provided in three different file formats: jsonl, csv, and txt.
Tools for conversion between file formats are available on https://github.com/clarinsi/SuperGLU
The LiLaH Emotion Lexicon of Croatian, Dutch and Slovene
The lexicon contains manual translations of the NRC Emotion Lexicon (http://saifmohammad.com/WebPages/NRC-Emotion-Lexicon.htm) that encodes the sentiment of a word (positive, negative) and its emotion association (anger, anticipation, disgust, fear, joy, sadness, surprise, trust) for Croatian, Dutch and Slovene with a binary schema. Manual translations were produced by inspecting and correcting the automatic translations from English provided with the original lexicon. While translations to all 14,182 entries are provided for Slovene and Croatian, only translations for the 6,468 entries that have any sentiment or emotion associated with the word are given for Dutch. For English entries, please refer to the original NRC Emotion Lexicon
The CLASSLA-StanfordNLP model for lemmatisation of standard Bulgarian 1.1
The model for lemmatisation of standard Bulgarian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the BulTreeBank training corpus (http://hdl.handle.net/11495/D93F-C6E9-65D9-2) and using the Bulgarian inflectional lexicon (Popov, Simov, and Vidinska 1998). The estimated F1 of the lemma annotations is ~98.8.
The difference to the previous version of the lemmatizer is that now it relies solely on XPOS annotations, and not on a combination of UPOS, FEATS (lexicon lookup) and XPOS (lemma prediction) annotations
The CLASSLA-StanfordNLP model for lemmatisation of standard Croatian 1.2
The model for lemmatisation of standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1183) and using the hrLex inflectional lexicon (http://hdl.handle.net/11356/1232). The estimated F1 of the lemma annotations is ~97.6.
The difference to the previous version is that now it relies solely on XPOS annotations, and not on a combination of UPOS, FEATS (lexicon lookup) and XPOS (lemma prediction) annotations
Lemma list of the Danish Dictionary - DDO (ELEXIS)
Den Danske Ordbog (DDO), Lemma list.
Contents and format:
This list contains the headwords of the online version of DDO (ordnet.dk/ddo).
DDO describes Danish lemmas from 1950 til today.
The elements are:
headword (attributes: entryid, homno (if present)), POS.
If a word has several official spellings every form is listed in it’s own list entry. These forms are placed alphabetically, but share the ID (as they origin from the same DDO entry).
Remarks:
When new editions of the Retskrivningsordbogen (the official dictionary of Danish standard orthography) are published, the official spelling of specific words (or groups of words) might change. These changes are not shown in DDO until the following version is released. For this reason the DDO lemma list is date-stamped. We aim at updating the list after every new version of DDO.
The lemma list reflects the DDO headwords (and their POS, ID, homograph number etc.) at the time of the latest list update. This information is subject to change in later versions.
The POS inventory is as in DDO, see the dictionary website. In addition to the basic parts of speech there are a few combined or modified POS markers, e.g. "sb. pl. bf." and "sb. itk. og pl."
Dictionary of Russian Dialects - SRNG (ELEXIS)
Словарь русских народных говоров.
The Russian Dialect Dictionary is the largest dialect dictionary in Russia. It includes vocabulary of Russian dialects from all over the Russian Federation and former Soviet Union
The Orange workflow for observing collocation trends ColTrend 1.0
The Orange workflow for observing collocation trends ColTrend 1.0
ColTrend is a workflow (.OWS file) for Orange Data Mining (an open-source machine learning and data visualization software: https://orangedatamining.com/) that allows the user to observe temporal collocation trends in corpora. The workflow consists of a series of Python scripts, data filters, and visualizers.
As input, the workflow takes a .CSV file with data on collocations and their relative frequencies by year of publication extracted from a corpus. As output, it provides a .TSV file containing the same data (or a filtered selection thereof) enriched with four measures that indicate the collocation’s temporal trend in the corpus: (1) the slope (k) of a linear regression model fitted to the frequency data, which indicates whether the frequency of use of the collocation is increasing or declining; (2) the coefficient of determination (R2) of the linear regression model, indicating how linear the change in the collocation’s use is; (3) the ratio (m) of maximum relative frequency and average relative frequency, which indicates peaks in collocation usage; and (4) the coefficient of recent growth (t), which indicates an increased usage of the collocation in the last three years of the observed corpus data.
The entry also contains three .CSV files that can be used to test the workflow. The files contain collocation candidates (along with their relative frequencies per year of publication) extracted from the Gigafida 2.0 Corpus of Written Slovene (https://viri.cjvt.si/gigafida/) with three different syntactic structures (as defined in http://hdl.handle.net/11356/1415):
1) p0-s0 (adjective + noun, e.g. rezervni sklad),
2) s0-s2 (noun + noun in the genitive case, e.g. ukinitev lastnine), and
3) gg-s4 (verb + noun in the accusative case, e.g. pripraviti besedilo).
It should be noted that only collocation candidates with absolute frequency of 15 and above were extracted.
Please note that the ColTrend workflow requires the installation of the Text Mining add-on for Orange. For installation instructions as well as a more detailed description of the different phases of the workflow and the measures used to observe the collocation trends, please consult the README file