Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
840 research outputs found
Sort by
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Croatian 1.0
The model for lemmatisation of non-standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1210), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), using the hrLex inflectional lexicon (http://hdl.handle.net/11356/1232). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~97.54
The CLASSLA-StanfordNLP model for named entity recognition of non-standard Croatian 1.0
This model for named entity recognition of non-standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1183), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1205). The training corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed
Frequency lists of words from the GOS 1.0 corpus 1.1
Frequency lists of words were extracted from the GOS 1.0 Corpus of Spoken Slovene (http://hdl.handle.net/11356/1040) using the LIST corpus extraction tool (http://hdl.handle.net/11356/1227). The lists contain all words occurring in the corpus along with their absolute and relative frequencies, percentages, and distribution across the text-types included in the corpus taxonomy.
The lists were extracted for each part-of-speech category. For each part-of-speech, two lists were extracted:
1) one containing lemmas and their text-type distribution,
2) one containing lower-case word forms as well as their standardized forms, lemmas, and morphosyntactic tags along with their text-type distribution.
In addition, four lists were extracted from all words (regardless of their part-of-speech category):
1) a list of all lemmas along with their part-of-speech category and text-type distribution;
2) a list of all lower-case word forms with their lemmas, part-of-speech categories, and text-type distribution;
3) a list of all lower-case word forms with their standardized word forms, lemmas, part-of-speech categories, and text-type distribution;
4) a list of all morphosyntactic tags and their text-type distribution (the tags are also split into several columns).
Compared to the previous version (http://hdl.handle.net/11356/1269), this one includes fixes of several typos and substitutes all instances of "normalized forms" with the more adequate term "standardized forms" (as used in the SSJ project)
Slovenian RoBERTa contextual embeddings model: SloBERTa 1.0
The monolingual Slovene RoBERTa (A Robustly Optimized Bidirectional Encoder Representations from Transformers) model is a state-of-the-art model representing words/tokens as contextually dependent word embeddings, used for various NLP tasks. Word embeddings can be extracted for every word occurrence and then used in training a model for an end task, but typically the whole RoBERTa model is fine-tuned end-to-end.
SloBERTa model is closely related to French Camembert model https://camembert-model.fr/. The corpora used for training the model have 3.47 billion tokens in total. The subword vocabulary contains 32,000 tokens. The scripts and programs used for data preparation and training the model are available on https://github.com/clarinsi/Slovene-BERT-Tool
The released model here is a pytorch neural network model, intended for usage with the transformers library https://github.com/huggingface/transformers
A Machine-readable Persian-English Dictionary - dc-pes-eng (ELEXIS)
Dictionary of Contemporary Persian.
This dictionary project has grown out of a university language course. It is highly experimental in nature. The focus in compiling the dictionary has been on contemporary language. In the beginning, all lexical items needed in the classes were entered into the dictionary. Later we started to integrate data not available in other dictionaries. Particular attention has been paid to neologisms which can not be found in most of the usually older print dictionaries.
Usage examples are adapted from various sources, for the most part coming from the Internet. Of particular importance is a Wikipedia version which we have made available as a TEI corpus
Lemma list of the Beseda Corpus Lemmatisation Lexicon (ELEXIS)
Lematizacijski slovar (leksikon besednih oblik za Besedo). Beseda Corpus Lemmatisation Lexicon for Slovenian language was generated at the Fran Ramovš Institute of Slovenian Language, primarily through inflection of open class words from the Dictionary of Standard Slovenian (Slovar slovenskega knjižnega jezika), augmented by wordforms, their part of speech tags and their lemmas used during the PoS tagging and lemmatization of the Beseda corpus. It was initially (2000) composed of 1 million words from the following texts:
Ciril Kosmač Opus - 408,000 words
Tomo Križnar: O iskanju ljubezni / On Search for Love or Around the World by Bicycle - 132,000 words
George Orwell: 1984 / 1984 - 91,000 words
Plato: Država / Republic - 93,000 words
Sveto pismo Nove zaveze / The Bible - New Testament - 150,000 words
Gustave Flaubert: Bouvard in Pécuchet / Bouvard and Pécuchet - 86,000 words
Časopis DELO na internetu (vzorec iz 6.5.1997 - 17.6.1997) / Newspaper DELO on Internet (a sample from 5/6/1997 - 6/17/1997) - 52,000 words
After 2000 the following texts were added:
Marko Uršič: Štirje časi / Four Seasons - 171,000 words
Državni zbor RS 3. sklica - dobesedni zapisi sej: 29. redna seja, zasedanje 01.10.2003 / National Assembly of the Republic of Slovenia - session transcripts: 29th regular session, meeting of 10/1/2003 - 47,000 words
Časopis DELO za 3.1.2004 / Newspaper DELO for 1/3/2004 - 75,000 words
to round the corpus to 1,300,000 words.
Current lexicon was taken from the database of the online "Determination of Lemmas and PoS Tags for a List of Words" service at the Institute, available through the web page: http://bos.zrc-sazu.si/dol_lem1.html.
Wordform frequencies were compiled from the latest update of the abovementioned corpus (version 138, 1,300,626 words, August 2017) and are therefore approximate.
See also: http://hdl.handle.net/11356/114
The Dictionary of the Comparisons (ELEXIS)
Palyginimų žodynas.
The dictionary contains comparisons collected from the Dictionary of the Lithuanian Language (I–XX volumes), dialect dictionaries, printed and manuscript collections of folklore, Lithuanian literature. Comparisons in the dictionary are presented according to their key word
Dictionary of Slovenian Phrasemes - SSF (ELEXIS)
Slovar slovenskih frazemov.
The 3,002 entries of this dictionary cover the description and explanation of 13,125 Slovenian phrasemes. The use of phrasemes is represented by citations from lexical files built for the Dictionary of the Slovenian Standard Language as well as from the Nova beseda corpus and other resources. The entries also contain etymological information and equivalents in other languages.
See also: http://hdl.handle.net/11356/1129
This dictionary was published as a printed book:
Keber, Janez. Slovar slovenskih frazemov. Ljubljana : Založba ZRC, ZRC SAZU, 2011
The CLASSLA-StanfordNLP model for UD dependency parsing of standard Slovenian
The model for UD dependency parsing of standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the UD-parsed portion of the ssj500k training corpus (http://hdl.handle.net/11356/1210) and using the CLARIN.SI-embed.sl word embeddings (http://hdl.handle.net/11356/1204). The estimated LAS of the parser is ~92.7