Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Consonant-vowel structures in the Gigafida 2.0 corpus

    No full text
    The lists contain consonant-vowel structures of all lemmas and word forms in the Gigafida 2.0 corpus. In each unit, its characters were converted as follows: C - consonant (in lists with finegrained character categorizations, consonants were divided into Z - sonorant, G - voiced obstruent, and K - voiceless obstruent), V - vowel, X - foreign consonant, Y - foreign vowel, S - symbol, P - punctuation, N - number, F - non-Latin-script character, ! - other. Each consonant-vowel structure also contains its frequency in the corpus (i.e. the total sum of the frequencies of all units corresponding to the consonant-vowel structure), as well as the set of all units (in the lists labeled "entire") or the set of its 30 most frequent units (in the lists labeled as "short"), along with their part-of-speech categories and their individual frequencies). They also contain the number of all unique units within the consonant-vowel structure. The lists were prepared based on frequency lists extracted from Gigafida 2.0 using LIST: http://hdl.handle.net/11356/1276 Note that there exists a related resource, "Consonant-vowel structures in the GOS 1.0 corpus", http://hdl.handle.net/11356/1290

    New Words and Meanings (ELEXIS)

    No full text
    Новые слова и значения. Словарь-справочник по материалам прессы и литературы 90-х годов 20 в. New Words and Meanings. Dictionary on the printed-media materials of the 1990s records lexical units that entered the Russian language in a given decade and were included in the language use

    Norwegian Dictionary Norsk Ordbank (Bokmål) - NO (ELEXIS)

    No full text
    Norsk Ordbank is an inventory of Norwegian (Bokmål) lemma forms, their paradigms and inflections, including full forms with morphosyntactic features. Standardisation information details which forms/paradigms were standard, secondary forms, of substandard in which period of time

    Italian-Ukrainian Glossary of Tourism (ELEXIS)

    No full text
    Bilingual thematic glossary of tourism neologisms

    Word list of the collection Words of Slovenian Language - SBSJ (ELEXIS)

    No full text
    Seznam besed iz zbirke Besede slovenskega jezika. A list of 354.205 different words from the headwords of the collection Besede slovenskega jezika / Words of Slovenian Language. See also: http://hdl.handle.net/11356/103

    General Basque Dictionary - OEH (ELEXIS)

    No full text
    Orotariko Euskal Hiztegia. This is the largest Basque dictionary ever produced. It considers all Basque written production from the very beginnings, until the 1970ies. Authors are Mitxelena & Sarasola, IPR holder is the Basque Language Academy. This is a cc-licensed PDF version, released in 2008. We have had some superficial experiments with that using GROBID using this PDF (see https://digilex.hypotheses.org/250)

    A Dictionary of the Russian Language of the 18th Century (ELEXIS)

    No full text
    Словарь русского языка XVIII века. The dictionary describes vocabulary of the 18th century - the time of the intensive borrowing from the European languages and the first efforts to create the Russian literary language

    Mining Documentation Ontolоgy (ELEXIS)

    No full text
    Ontologija rudarske dokumentacije RuDokOnto is developed within PhD Thesis "Model development for managing mining project documentation" to support the system which could enable efficient managing of mining project documentation in electronic form, based on human language technology, for information retrieval and information extraction by using different language resources. The mining domain ontology RuDokOnto is developed for the purpose of collecting, describing, and systematization of mining project documentation throughout the phases of the mining project's life cycle in a way that links other related ontologies, such as for example EarthResource, MinExOnt etc.

    CroSloEngual BERT 1.1

    No full text
    Trilingual BERT (Bidirectional Encoder Representations from Transformers) model, trained on Croatian, Slovenian, and English data. State of the art tool representing words/tokens as contextually dependent word embeddings, used for various NLP classification tasks by finetuning the model end-to-end. CroSloEngual BERT are neural network weights and configuration files in pytorch format (i.e. to be used with pytorch library). Changes in version 1.1: fixed vocab.txt file, as previous verson had an error causing very bad results during fine-tuning and/or evaluation

    The CLASSLA-StanfordNLP model for lemmatisation of non-standard Slovenian 1.1

    No full text
    The model for lemmatisation of non-standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and the Janes-Tag corpus (http://hdl.handle.net/11356/1238), using the Sloleks inflectional lexicon (http://hdl.handle.net/11356/1230). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~98.86. The difference to the previous version of the lemmatizer is that now it relies solely on XPOS annotations, and not on a combination of UPOS, FEATS (lexicon lookup) and XPOS (lemma prediction) annotations

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇