Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Dictionary of the Slovenian Normative Guide (2001)

    No full text
    The dictionary part of Slovenian Normative Guide (first published in 2001) is a normative orthographic dictionary of Slovenian standard language. In 92,617 entries it contains 140,266 lemmas and sublemmas. The entries contain information on spelling, pronunciation, inflection, part of speech, normative information, synonyms, valency, while semantic identifications, labels and usage examples are also provided. This dictionary was published as a printed book: Toporišič, Jože; Jakopin, Franc; Moder, Janko; Dular, Janez; Suhadolnik, Stane; Menart, Janez; Pogorelec, Breda; Gantar, Kajetan; Ahlin, Martin; Hajnšek - Holz, Milena; Bokal, Ljudmila; Gložančev, Alenka; Keber, Janez; Lazar, Branka; Praznik, Zvonka; Snoj, Jerica. Slovenski pravopis. Ljubljana: Založba ZRC, ZRC SAZU, 2001. ISBN 961-6358-37-5

    Corpus of academic Slovene KAS 1.0

    No full text
    The KAS corpus of Slovene academic writing consists of almost 65,000 BSc/BA, 16,000 MSc/MA and 1,600 PhD theses (82 thousand texts, 5 million pages or 1,7 billion tokens) written 2000 - 2018 and gathered from the digital libraries of Slovene higher education institutions via the Slovene Open Science portal (http://openscience.si/). The theses have associated with them significant metadata, while each thesis in the corpus contains its textual body, i.e. without their front and back matter. The body is divided into pages, these into paragraphs, and then into sentences. The sentence tokens are morphosyntactically annotated, words are lemmatised and English-Slovene pairs of term candidates are marked up and linked. The PhD theses in the corpus also have marked-up Slovene monolingual term candidates. The corpus is distributed in the canonical TEI encoding, in the so-called vertical format used by the (no)Sketch Engine and CWB concordancers, and as plain text files. Each format distribution also contains a file with thesis metadata. This repository entry contains the complete corpus; separate entries are available that contain only the PhD theses (KAS-dr: http://hdl.handle.net/11356/1265), the MSc/MA theses (KAS-mag: http://hdl.handle.net/11356/1266) and BSc/BA theses (KAS-dipl: http://hdl.handle.net/11356/1267)

    Collocations Dictionary of Modern Slovene KSSS 1.0

    No full text
    The database of the Collocations Dictionary of Modern Slovene 1.0 contains entries for 35,862 headwords (18,043 nouns, 5,148 verbs, 10,259 adjectives and 2,412 adverbs) and 7,310,983 collocations that were automatically extracted from the Gigafida 1.0 corpus. For the automatic extraction via the Sketch Engine API we used a specially adapted Sketch grammar for Slovene, and, based on manual evaluation, a set of parameters that determined: maximum number of collocates per grammatical relation, minimum frequency of a collocate, minimum frequency of a grammatical relation, minimum salience (logDice) score of a collocate, and minimum salience of a grammatical relation. The procedure of automatic extraction, which produced a list of collocates (lemmas) in a particular relation, was followed by a set of post-processing steps: - removal of collocations that were represented by repetitions of the same sentence - preparation of full collocations by the addition of the headword, and, if needed, the third element in the grammatical relation (such as preposition). The headwords/collocates were also put in the correct case, depending on the grammatical relation. - addition of IDs from the Slovenian morphological lexicon Sloleks (http://hdl.handle.net/11356/1230) to every element in the collocation

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Croatian

    No full text
    The model for morphosyntactic annotation of standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1183) and using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1205). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~94.1

    Dictionary of Middle Dutch - MNW (ELEXIS)

    No full text
    Middelnederlandsch Woordenboek. Describes the vocabulary of the Dutch spoken from the thirteenth to the sixteenth century

    Training corpus ssj500k 2.2

    No full text
    The ssj500k training corpus contains about 500,000 tokens manually annotated on the levels of tokenisation, sentence segmentation, morphosyntactic tagging, and lemmatisation. About half of the corpus is also manually annotated with syntactic dependencies, named entities, and verbal multiword expressions. About a quarter of the corpus is annotated with semantic role labels. The morphosyntactic tags and syntactic dependencies are included both in the JOS/MULTEXT-East framework, as well as in the framework of Universal Dependencies. The annotations of the ssj500k corpus follow (1) the MULTEXT-East V6 morphosyntactic specifications for Slovene, http://nl.ijs.si/ME/V6/msd/, (2) the JOS dependency schema, http://nl.ijs.si/jos/bib/jos-skladnja-navodila.pdf, the Universal Dependencies morphosyntactic specifications and syntactic dependencies for Slovene-SSJ, https://universaldependencies.org/, (4) the Janes annotation guidelines for Slovenian named entities, http://nl.ijs.si/janes/wp-content/uploads/2017/09/SlovenianNER-eng-v1.1.pdf, and (5) the Guidelines of the PARSEME shared task on verbal multiword expressions, http://parsemefr.lif.univ-mrs.fr/parseme-st-guidelines/1.1/ The vocabulary of (1) and (2) is provided in the back element and (3), (4), and (5) in the teiHeader of the TEI encoded corpus. The semantic role labels are also documented in the teiHeader. In contrast to the previous version 2.1, this version corrects various errata in spacing and text metadata and adds UD morphological and (where it was possible to do so automatically) dependency annotations to the corpus. Note that the UD annotations are not included in the vertical file

    Error-annotated developmental corpus Šolar 2.0 Error

    No full text
    The corpus contains 2094 texts from the corpus Šolar 2.0 (http://hdl.handle.net/11356/1214), i.e. only those in which error annotations can be found. For each text, the information on school (elementary or secondary), subject, level (grade or year), type of text, region and date of production is provided. The original error annotations from Šolar 1.0 have been re-categorized according to a new system (the specifications in Slovene are attached). There are 36,671 error annotations in total, which also include corrections made by teachers. The corpus consists of 756,130 words from student texts (this word count does not include teacher corrections)

    Developmental corpus ccŠolar 1.0

    No full text
    The ccŠolar corpus contains 1693 texts collected during 2016-2018, as part of the upgrade of the corpus Šolar project. The project aims were to increase the size of the Šolar 1.0 corpus and to improve text balance across regions and education level. For each text, the information on school (elementary or secondary), subject, level (grade or year), type of text, region and date of production is provided. The ccŠolar 1.0 corpus is offered separately because the new texts were collected under CC BY 4.0 licence, a more open licence than the earlier texts

    Dependency tree extraction tool STARK 1.0

    No full text
    STARK is a python-based command-line tool for extraction of dependency trees from parsed corpora, aimed at corpus-driven linguistic investigations of syntactic phenomena of various kinds. It supports the CONLL-U format (https://universaldependencies.org/format.html) as input and returns a list of all relevant dependency trees, frequencies, and other associated information in the form of a tab-separated .tsv file. For installation, execution and the description of various user-defined parameter settings, see the official project page at: https://gitea.cjvt.si/lkrsnik/STARK. This entry corresponds to commit 421f12cac6 in the Git repository

    Dictionary of Old Dutch - ONW (ELEXIS)

    No full text
    Oudnederlands Woordenboek. A scientific dictionary of the oldest Dutch

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇