Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    SimLex-999 Slovenian translation SimLex-999-sl 1.0

    No full text
    The resource contains English SimLex-999 (Hill et al. 2015) and their Slovene translations. In the translation process, the word pairs were first translated by two translators independently, and next, for the examples where the translations differed, the final translations were chosen in a consensus meeting. The translators had also access to Croatian Simlex-999 translations (Mrkšić et al. 2017) and received translation guidelines (see next sheet) inspired by guidelines of Multi-SimLex (Vulić et al. 2020). The resources was used for building the CoSimLex resource (Armendariz et al. 2020). The list contains English original pair of words (Word1 and Word2), their part-of-speech, followed by Slovene translations (Trans1 and Trans2). The last column Comment relates to special cases: - "multiword_translation" -> translators were asked to opt for single-word equivalents, in some cases the only appropriate translation was a multi-word expression (for example, "birthday" -> "rojstni dan"). - "no_translation" -> pairs without a proper translation, i.e. translation pair contains two identical words. Although the translators were asked to find two different translations for the words, in a few examples that was not possible. For example, for the English pair "taxi" and "cab", only "taksi" was considered a good Slovene equivalent. - "duplicated_translation" -> in cases where a pair of words is repeated for two different English original pairs, both occurrences are marked as duplicate translations. - "duplicated_original" -> in one case, the original word pair was a duplicate, which is also marked. Cite: If you use the dataset, please cite the Clarin handle and the following paper: Armendariz, Carlos Santos, Purver, Matthew, Ulčar, Matej, Pollak, Senja, Ljubešić, Nikola, Granroth-Wilding, Mark, and Vaik, Kristiina (2020). CoSimLex: A Resource for Evaluating Graded Word Similarity in Context. In Proceedings of the 12th Language Resources and Evaluation Conference, p. 5878--5886. https://www.aclweb.org/anthology/2020.lrec-1.720/ References: Armendariz, Carlos Santos, Purver, Matthew, Ulčar, Matej, Pollak, Senja, Ljubešić, Nikola, Granroth-Wilding, Mark, and Vaik, Kristiina (2020). CoSimLex: A Resource for Evaluating Graded Word Similarity in Context. In Proceedings of the 12th Language Resources and Evaluation Conference, p. 5878--5886. https://www.aclweb.org/anthology/2020.lrec-1.720/ Hill, F., Reichart, R., and Korhonen, A. (2015). Simlex-999: Evaluating semantic models with (genuine) similarity estimation. Computational Linguistics, 41(4):665–695. https://www.aclweb.org/anthology/J15-4004/ Mrkšić, Nikola, Ivan Vulić, Diarmuid Ó Séaghdha, Ira Leviant, Roi Reichart, Milica Gašić, Anna Korhonen, and Steve Young. (2017). Semantic specialisation of distributional word vector spaces using monolingual and cross-lingual constraints. Transactions of the ACL, 5:309–324. https://www.mitpressjournals.org/doi/abs/10.1162/tacl_a_00063 Vulić, Ivan, Baker, Simon, Ponti, Edoardo Maria, Petti, Ulla, Leviant, Ira, Wing, Kelly, Majewska, Olga, Bar, Eden, Malone, Matt, Poibeau, Thierry, Reichart, Roi and Anna Korhonen (2020). Multi-SimLex: A Large-Scale Evaluation of Multilingual and Cross-Lingual Lexical Semantic Similarity. Computational Linguistics. https://doi.org/10.1162/coli_a_0039

    Dictionary of New Slovenian Words - SNB (ELEXIS)

    No full text
    Slovar novejšega besedja slovenskega jezika. Dictionary of New Slovenian Words represents a basic new lexical supplement to the Slovar slovenskega knjižnega jezika (Dictionary of the Slovenian Standard Language). It contains 6399 new words and phrases that appeared in Slovenian or gained ground after 1991 as well as new meanings of previously standardised lexis. Two important new features of the dictionary are a corpus-driven analysis of new words that are in actual language use and etymological explanations of the included words. See also: http://hdl.handle.net/11356/1091 This dictionary was published as a printed book: Bizjak Končar, Aleksandra, Snoj, Marko, Gložančev, Alenka, Kern, Boris, Kostanjevec, Polona, Krvina, Domen, Ledinek, Nina, Michelizza, Mija, Perdih, Andrej, Petric, Špela, Šircelj-Žnidaršič, Ivanka, Žele, Andreja, Mirtič, Tanja, Gliha Komac, Nataša, Klemenčič, Simona. Slovar novejšega besedja slovenskega jezika. Ljubljana : Založba ZRC, ZRC SAZU, 2012

    ISLEX Dictionary (ELEXIS)

    No full text
    ISLEX-orðabókin. ISLEX is an online multilingual dictionary between modern Icelandic and six Scandinavian TLs: Danish, Norwegian (Bokmål and Nynorsk), Swedish, Faroese and Finnish. It is accessible on the web, free of charge (https://islex.arnastofnun.is/is/). The project is published by The Árni Magnússon Institute for Icelandic Studies in Reykjavík, Iceland. Its editorial staff is responsible for the description of the Icelandic source language and the development and maintenance of the database. Its partners, responsible for the TLs, are located in the other Nordic countries. Screen reader support enabled

    Slovenian parliamentary corpus (1990-2018) siParl 2.0

    No full text
    The siParl corpus contains minutes of the Assembly of the Republic of Slovenia for 11th legislative period 1990-1992, minutes of the National Assembly of the Republic of Slovenia from the 1st to the 7th legislative period 1992-2018, minutes of the working bodies of the National Assembly of the Republic of Slovenia from the 2nd to the 7th legislative period 1996-2018, and minutes of the Council of the President of the National Assembly from the 2nd to the 7th legislative period 1996-2018. The corpus comprises over 10 thousand sessions, one million speeches or 200 million words. The corpus contains meta-data about the speakers, a typology of sessions etc. and structural, editorial and linguistic annotations. The corpus is encoded according to the Parla-CLARIN schema (https://github.com/clarin-eric/parla-clarin). Each mandate is in one directory, and each session in one file. This item comprises the following datasets: 1. source DARAH-SI Parla-CLARIN encoded corpus; 2. linguistically annotatated Parla-CLARIN encoded corpus: tokenisation, MSD tagging, lemmatisation, Universal Dependencies features and syntactic parses, named entities; 3. linguisticaly annotated corpus in vertical format used by CWB and Sketch Engine concordancers; this format is simpler and smaller but does not contain all the information from the source TEI; 4. linguisticaly annotated corpus in CONLL-U format as used by Universal Dependencies 5. plain text of the corpus Note that each dataset also includes TSV meta-data files on sessions (files) and speakers. As opposed to the previous version 1.0, this version corrects many errors, has substantially better meta-data and the linguistic processing has more levels and less errors

    Multimodal corpus EVA 1.0

    No full text
    EVA Corpus 1.0 consists of one episode of an audio/video session plus corresponding orthographic transcriptions with a duration of 57 minutes. The multi-party spontaneous discourse in the recording is from an entertaining evening TV-talk show "A si ti tut not padu", broadcasted by the POP-TV Slovene commercial TV station in 2008, and represents a part of the Slovene spoken corpus GOS (http://hdl.handle.net/11356/1040). The show contains casual conversation about general, informal and personal aspects of interviewee's life. The transcribed session of this recording has been annotated using ELAN 4.9.4. In addition to the original transcription and morphosyntactic annotation from the GOS corpus, the following layers of information are added: - statement sentiment - phrase breaks within statements - prominence of statements - sentences within the statement - sentence sentiment - sentence type - speaker visibility on the scene - gesture units - gesture phrases - emotions - semiotic intent - dialogue rol

    Dictionary of the Slovenian Language in the Works of Janez Svetokriški - JSV (ELEXIS)

    No full text
    Slovar jezika Janeza Svetokriškega. The Dictionary of the Slovenian Language in the Works of Janez Svetokriški presents and explains the lexis, including proper nouns, from 233 sermons published by Janez Svetokriški in five volumes under the common title Sacrum promptuarium between 1691 and 1707. The dictionary contains 8,540 dictionary entries, which display and treat the entire Slovenian lexis, including proper nouns, used in the above-mentioned work. Each dictionary entry consists of 1. the headword, 2. the presentation of morphological characteristics, 3. the description of meaning and 4. examples of use. Entries containing loanwords additionally include etymologies. Some entries may here provide other philological or linguistic comments. Each entry describing a proper noun ends with the most basic encyclopaedic information. The Dictionary of the Slovenian Language in the Works of Janez Svetokriški is the first dictionary to treat the lexis of a Slovenian author from a period before the introduction of Gaj's Latin alphabet. The dictionary is distinguished by a modern, but not too complex display of material, by comprehensive citations of all attested variants and by the inclusion of encyclopaedic information about proper nouns, all of which in many respects facilitates the reading of the original baroque text or makes it possible in the first place. This dictionary was published as a printed book: Snoj, Marko. Slovar jezika Janeza Svetokriškega. Ljubljana: Založba ZRC, 2006. See also: http://hdl.handle.net/11356/109

    Basque Lexical Data in Wikidata (ELEXIS)

    No full text
    This dataset contains Basque lemma-sense pairs with a POS tag, and definitions, extracted from Wikidata using this query: https://w.wiki/qWH . Other RDF statements related to Basque lexemes can be retrieved, such as links from lexeme sense to Wikidata concept

    Dialogue act annotated spoken corpus GORDAN 1.0 (transcription)

    No full text
    The GORDAN 1.0 corpus contains authentic data of spoken communication, annotated for dialogue acts according to the GORDAN 1.0 dialogue act annotation scheme, included in the data. The corpus data were selected from existing Slovene speech corpora: GOS (http://hdl.handle.net/11356/1040), Gos Videolectures (http://hdl.handle.net/11356/1223) and BERTA. Four criteria were taken into account in the selection: public/non-public, interactive/monologic, channel and intention. The total length of the data is 1 hour of recordings (6,909 words). The selected data were annotated using the Transcriber 1.5.1 tool and its function Event. Annotation was done based on multimodal data, listening to the audio or watching the video recording, where available. This resource contains only annotated transcriptions of the corpus – audio and video recordings are available at http://hdl.handle.net/11356/1292

    Metaphor corpus KOMET 1.0

    No full text
    KOMET 1.0 is a hand-annotated corpus for metaphorical expressions which contains about 200,000 words from Slovene journalistic, fiction and on-line texts. To annotate metaphors in the corpus an adapted and modified procedure of the MIPVU protocol (Steen et al., 2010: A method for linguistic metaphor identification: From MIP to MIPVU, https://www.benjamins.com/catalog/celcr.14) was used. The lexical units (words) whose contextual meanings are opposed to their basic meanings are considered metaphor-related words. The basic and contextual meaning for each word in the corpus was identified using the Dictionary of the standard Slovene Language. The corpus was annotated for the metaphoric following relations: indirect metaphor, direct metaphor, borderline case and metaphor signal. In addition, the corpus introduces a new ‘frame’ tag, which gives information about the concept to which it refers

    The CLASSLA-StanfordNLP model for lemmatisation of standard Serbian 1.1

    No full text
    The model for lemmatisation of standard Serbian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200) and using the srLex inflectional lexicon (http://hdl.handle.net/11356/1233). The estimated F1 of the lemma annotations is ~97.9. The difference to the previous version of the model is that it is trained with the lemmatiser padding bug removed, cf. https://github.com/stanfordnlp/stanfordnlp/issues/143

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇