Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Slovene ontology of semantic types for nouns SLONEST-noun 1.0

    No full text
    SLONEST stands for Slovene Ontologies of Semantic Types. The first subset – SLONEST-noun 1.0 – represents an ontology developed for nouns. SLONEST-noun contains an XML file with a total of 271 categories of semantic types: 21 top-level categories, which are further divided into up to three levels of hierarchical subcategories. The ontology was developed and evaluated using the data from the Collocations Dictionary of Modern Slovene (Kosem et al. 2018; https://viri.cjvt.si/kolokacije; http://hdl.handle.net/11356/1250) and the Comprehensive Slovene-Hungarian Dictionary (https://www.cjvt.si/en/research/cjvt-projects/slovene-hungarian-dictionary), which are being compiled at the Centre for Language Resources and Technologies, University of Ljubljana. The semantic types in the SLONEST-noun ontology are accompanied with numerical ids (listed in the attribute SEMCODE; e.g. "1.1.1") and full ontology path (attribute SEMFULLNAME; e.g. "HUMAN-ACTIVITY-OTHER"). Every semantic type is provided with a definition (e.g. "Other denominations for humans related to activities."). Where relevant, especially at top-level semantic types, the corresponding semantic type (i.e. lexicographer file) from Wordnet (https://wordnet.princeton.edu/) is listed, along with the level of matching ("full" or "partial"). For most semantic types, examples of Slovene lemmas or multiword units are also provided. As the ontology was also developed for, and tested on, collocation data, a selection of collocations is also provided for most categories. For every collocation, noun headwords and collocates are clearly labelled, and the information on grammatical structure (id and name) is provided, based on the most recent database of Slovene collocations (http://hdl.handle.net/11356/1415). The ontology was developed as part of the KOLOS project. The authors acknowledge that the project titled Collocation as a basis for language description: semantic and temporal perspectives (J6-8255) was financially supported by the Slovenian Research Agency

    ISLEX Dictionary (audio) (ELEXIS)

    No full text
    The data contains audio files for the Icelandic lemmas of the ISLEX dictionary

    Croatian Language Resources for NooJ (ELEXIS)

    No full text
    Croatian Language Resources for NooJ is a set of files holding a list of Croatian words marked for POS, and depending on the POS with type, form, gender, category, case, number, tense, degree. Also, words found in the medical domain are marked with a semantic marker for a domain and subdomain. The resource can be helpful for any research performed on general, or health-related text; but also for any interlingual comparisons

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Serbian 1.0

    No full text
    This model for morphosyntactic annotation of non-standard Serbian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200), the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the hr500k training corpus (http://hdl.handle.net/11356/1210) and the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/), using the CLARIN.SI-embed.sr word embeddings (http://hdl.handle.net/11356/1206). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~94.91

    Dataset of Slovene idiomatic expressions SloIE

    No full text
    SloIE is a manually labelled dataset of Slovene idiomatic expressions. It contains 29,400 sentences with 75 different expressions that can occur with either a literal or an idiomatic meaning, with appropriate manual annotations for each token. The idiomatic expressions were selected from the Slovene Lexical Database (http://hdl.handle.net/11356/1030). We selected only expressions that can occur with both a literal and an idiomatic meaning. The sentences were extracted from the Gigafida corpus. For each sentence, the file first contains the text of the sentence prefixed by #. This is followed by a line of numbers indicating the positions of tokens that belong to the expression. The numbers also indicate the word order for expressions where the word order is flexible. They are ordered according to the dictionary form of the expression (e.g., the first number indicates the position where the first word of the expression - in its dictionary form - occurs). Each token is labelled with either 'DA', indicating tokens in an expression that have an idiomatic meaning, 'NE', indicating tokens in an expression that have a literal meaning, or '*', indicating tokens outside the expression. Additionally, 'NEJASEN ZGLED' indicates tokens where the annotators could not determine the meaning from the example sentence. Each token is also tagged with the dictionary form of the expression that is present in the sentence. Key reference: Škvorc, Tadej, Polona Gantar, and Marko Robnik-Šikonja. "MICE: Mining Idioms with Contextual Embeddings." arXiv preprint arXiv:2008.05759 (2020)

    Reference List of Slovene Frequent Common Words

    No full text
    The reference list of Slovene most frequent common words was prepared by selecting vocabulary at the intersection of the most frequent 10,000 lemmas of four Slovene text corpora: the balanced reference corpus of written Slovene Kres, the reference corpus of spoken Slovene GOS, the corpus of computer-mediated communication Janes and the corpus of school written production Šolar 2.0. The list was additionally manually cleaned and contains 4,768 common general lemmas. The file is in a tab separated format, containing lemma, part-of-speech (following the MULTEXT-East tagset for Slovene), relative average reduced frequency in each of the corpora, and the final average score computed from these values. The dataset is described in more detail in: Špela Arhar Holdt, Senja Pollak, Marko Robnik Šikonja, Simon Krek (2020). Referenčni seznam pogostih splošnih besed za slovenščino. In the Proceedings of the Conference on Language Technologies and Digital Humanities, pp. 10-15

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Serbian 1.1

    No full text
    The model for morphosyntactic annotation of standard Serbian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200) and using the CLARIN.SI-embed.sr word embeddings (http://hdl.handle.net/11356/1206). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~95.2. The difference to the previous version of the model is that now the whole XPOS tag is predicted and not specific characters, as was the case in stanfordnlp, which resulted in illegal XPOS tags (and slightly decreased performance)

    Word embeddings CLARIN.SI-embed.mk 0.1

    No full text
    CLARIN.SI-embed.mk contains word embeddings induced from a large collection of Macedonian texts crawled from the .mk top-level domain. The embeddings are based on the skip-gram model of fastText trained on 323,158,626 tokens of running text for 448,182 lowercased surface forms

    Annotated corpus of Serbian language-related news comments MetaLangNEWS-COMMENTS-Sr

    No full text
    A comprehensive corpus of user comments on online news articles on the topic of language from major Serbian daily newspapers and news portals, published in the five-year period of January 1, 2015 - January 1, 2020. The corpus is designed to facilitate research on metalanguage (‘language about language’), linguistic ideologies, language policy and planning, as well as the specific contemporary debates on language defining, naming, and standardisation, from the bottom-up perspective. The corpus has been tagged using the CLASSLA-StanfordNLP models for morphosyntactic annotation and lemmatisation of non-standard Serbian. The corpus is available in plain text version, XML with full metadata, and tagged CONLL-U format. This collection is complementary to the corpus of news articles MetaLangNEWS-Sr (http://hdl.handle.net/11356/1371). Parallel versions from Croatia (http://hdl.handle.net/11356/1370) and Slovenia (http://hdl.handle.net/11356/1362) are also available

    The "Arcticae horulae" dictionary of German borrowings in Slovenian

    No full text
    The "Arcticae horulae" dictionary of German borrowings in Slovenian was a project of continuous development, from a private amateur collection of German borrowings in the Slovenian language, via an art exhibition of the collection in the National and University Library of Slovenia in 1995, to the publication of the printed edition in 1997, which generated unexpected media interest and was sold out in a just few months. This development process connected the fields of art and science and highlighted the artist's responsibility: the artistic modes (e.g. travesty, performance) misdirected the public to treat the booklet as a reference dictionary. This is the reason that the author takes the end of her project to be the 8th of January 1998, when Marko Snoj, a well-known Slovenian etymologist, published a very negative review of the dictionary in the "Književni listi" supplement of the "Delo" newspaper. As the printed edition was sold out, and to document the Arcticae horulae project, the http://www.arcticae-horulae.si site was created in 2005. In addition to the PDF of the printed edition of the dictionary booklet, the site contains also the chronology of the project, and various publications about it. In 2017 CLARIN.SI supported the conversion of the PDF to a TEI XML encoding (using the module for dictionaries, https://www.tei-c.org/release/doc/tei-p5-doc/en/html/DI.html), in order to provide a structured digital edition of this interesting resource. In the process of correcting OCR errors in the XML, some mistakes in the original were also corrected. Note, however, that the conversion to the structured XML was performed automatically, so various structure errors most likely remain in the dictionary. This repository item archives the Arcticae horulae project, and contains the mirror of the www.arcticae-horulae.si site (as of November 2020), the PDF of the printed dictionary, the dictionary in a TEI XML encoding, and a simple HTML rendering of the dictionary, which attempts to preserve the formatting of the original

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇