Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Offensive language dataset of Croatian, English and Slovenian comments FRENK 1.0

    No full text
    The FRENK dataset consists of comments to Facebook posts (news articles) of mainstream media outlets from Croatia, Great Britain, and Slovenia, on the topics of migrants and LGBT. The dataset contains whole discussion threads. Each comment is annotated by the type of socially unacceptable discourse (e.g., inappropriate, offensive, violent speech) and its target (e.g., migrants/LGBT, commenters, media). The annotation schema in its details is described in https://arxiv.org/pdf/1906.02045.pdf. Usernames in the metadata are pseudo-anonymised and removed from the comments. The data in each language (Croatian (hr), English (en), Slovenian (sl), and topic (migrants, LGBT) is divided into a training and a testing portion. The training and testing data consist of separate discussion threads, i.e., there is no cross-discussion-thread contamination between training and testing data. The sizes of the splits are the following: Croatian, migrants: 4356 training comments, 978 testing comments; Croatian LGBT: 4494 training comments, 1142 comments; English, migrants: 4540 training comments, 1285 testing comments; English, LGBT: 4819 training comments, 1017 testing comments; Slovenian, migrants: 5145 training comments, 1277 testing comments; Slovenian, LGBT: 2842 training comments, 900 testing comments

    TermFrame: Terms, definitions and semantic annotations for karstology

    No full text
    The resource contains several datasets containing domain-specific data in three languages, English, Slovenian and Croatian, which can be used for various knowledge extraction or knowledge modelling tasks. The resource represents knowledge for the domain of karstology, a subfield of geography studying karst and related phenomena. It contains: 1. Definitions Plain text files contain definitions of karst concepts from relevant glossaries and encyclopaedia, but also definitions which had been extracted from domain-specific corpora. 2. Annotated definitions Definitions were manually annotated and curated in the WebAnno tool. Annotations include several layers including definition elements, semantic relations following the frame-based theory of terminology (FBT), relation definitors which can be used for learning relation patterns, and semantic categories defined in the domain model. 3. Terms, definitions and sources The TermFrame knowledge base contains terms and their corresponding concept identifiers, definitions and definition sources

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Slovenian 1.2

    No full text
    This model for morphosyntactic annotation of standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and using the CLARIN.SI-embed.sl word embeddings (http://hdl.handle.net/11356/1204). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~97.06. The difference to the previous version of the model is that the pre-trained embeddings are limited to 250 thousand entries and adapted to the new code base

    Parallel corpus EN-SL RSDO4 2.0

    No full text
    The RSDO4 parallel corpus of English-Slovene and Slovene-English translation pairs was collected as part of work package 4 of the Slovene in the Digital Environment project. It contains texts collected from public institutions and texts submitted by individual donors through the text collection portal created within the project. The updated corpus consists of 3143624 (previously 964433) translation pairs (extracted from standard translation formats (TMX, XLIFF) or manually aligned) in randomized order which can be used for machine translation training

    Linguistically annotated multilingual comparable corpora of parliamentary debates ParlaMint.ana 2.0

    No full text
    ParlaMint is a multilingual set of comparable corpora containing parliamentary debates mostly starting in 2015 and extending to mid-2020, with each corpus being about 20 million words in size. The sessions in the corpora are marked as belonging to the COVID-19 period (after October 2019), or being "reference" (before that date). The corpora have extensive metadata, including aspects of the parliament; the speakers (name, gender, MP status, party affiliation, party coalition/opposition); are structured into time-stamped terms, sessions and meetings; with speeches being marked by the speaker and their role (e.g. chair, regular speaker). The speeches also contain marked-up transcriber comments, such as gaps in the transcription, interruptions, applause, etc. Note that some corpora have further information, e.g. the year of birth of the speakers, links to their Wikipedia articles, their membership in various committees, etc. The corpora are encoded according to the Parla-CLARIN TEI recommendation (https://clarin-eric.github.io/parla-clarin/), but have been validated against the compatible, but much stricter ParlaMint schemas. This entry contains the linguistically marked-up version of the corpus, while the text version is available at http://hdl.handle.net/11356/1388. The ParlaMint.ana linguistic annotation includes tokenization, sentence segmentation, lemmatisation, Universal Dependencies part-of-speech, morphological features, and syntactic dependencies, and the 4-class CoNLL-2003 named entities. Some corpora also have further linguistic annotations, such as PoS tagging or named entities according to language-specific schemes, with their corpus TEI headers giving further details on the annotation vocabularies and tools. The compressed files include the ParlaMint.ana XML TEI-encoded linguistically annotated corpus; the derived corpus in CoNLL-U with TSV speech metadata; and the vertical files (with registry file), suitable for use with CQP-based concordancers, such as CWB, noSketch Engine or KonText. Also included is the 2.0 release of the data and scripts available at the GitHub repository of the ParlaMint project

    Business English learner speech corpus SAPS

    No full text
    SAPS is a specialized speech corpus which contains business meeting simulations in English between undergraduate students of Languages for Business and Economics at the School of Economics and Business, University of Ljubljana. The corpus contains group business meeting simulations based on a fictional "Marbi" case study with max. 6 students in the group. The corpus is composed of the 32 audio files with corresponding transcriptions. The trascriptions are plain-text TSV files with code(s) of the the anonymised speaker(s) in the first column, and the transcription in the second column. Meta-textual notes are in square brackes. The recordings (about 4h 45' in length) were made in 2013, with half of the interviews recorded at the start (files with "START" in their name) and half at the end (END) of the semester. The compilaton and use of the corpus is described in: DOSTAL, Mateja. Učitelj kot posredovalec korektivne povratne informacije za razvijanje tujejezikovne sporazumevalne zmožnosti v simulacijah poslovnega sestanka. Šolsko polje : revija za teorijo in raziskave vzgoje in izobraževanja, ISSN 1581-6036. 2015, Vol. 26, no. 1/2, pp. 23-42, 144-146). http://www.dlib.si/details/URN:NBN:SI:doc-3IZTZY8H. DOSTAL, Mateja. Developing Foreign language communicative competence for English business meetings using business meeting simulations. Scripta manent : revija Slovenskega društva učiteljev tujega strokovnega jezika, ISSN 1854-2042, 2016, vol. 11, 1, pp. 2-20. http://scriptamanent.sdutsj.edus.si/ScriptaManent/article/view/153/138. DOSTAL, Mateja. Vloga učitelja pri razvijanju tujejezikovne sporazumevalne zmožnosti za poslovne sestanke v angleškem jeziku : primer simulacije poslovnih sestankov : doktorska disertacija. Maribor: 2015. https://dk.um.si/IzpisGradiva.php?id=55341 DOSTAL, Mateja. Simulacije poslovnega sestanka v angleškem jeziku za učinkovite mednarodne poslovne sestanke. In: REDEK, Tjaša (ed.). Izzivi podjetij, države in družbe v uresničevanju odgovornosti za trajnostni razvoj. Ljubljana: Ekonomska fakulteta, 2021. pp. 289-304. Zbirka Ekonomska fakulteta raziskuje. ISBN 978-961-240-374-4. http://www.ef.uni-lj.si/zaloznistvo/raziskovalne_publikacije. DOSTAL, Mateja. Gradnja govornega učnega korpusa simulacij poslovnih sestankov za raziskavo o vlogi učitelja pri razvijanju tujejezikovne sporazumevalne zmožnosti za poslovne sestanke v angleškem jeziku. In: JURKOVIČ, Violeta (ed.), ČEPON, Slavica (ed.). Raziskovanje tujega jezika stroke v Sloveniji. Ljubljana: Slovensko društvo učiteljev tujega strokovnega jezika. 2015, pp. 193-223

    SloBENCH evaluation framework

    No full text
    The evaluation framework contains public evaluation scripts. All the scripts contain additional Dockerfiles that allow for platform-independent evaluation and exact comparison of results. Pre-built Docker images are available in slobench/eval DockerHub repository. The evaluation framework is used and maintained by the SloBENCH leaderboard Web site team. SloBENCH submitters are able to check their compliance of submissions and evaluate theri model on training/validation data prior to submission. The initial version of SloBENCH contains evaluation scripts with examples of training and testing datasets for nine different tasks: named entity recognition, part-of-speech tagging, lemmatization, dependency parsing, semantic role labeling, translation (ENG-SLO, SLO-ENG), summarization and question answering

    School dictionary of Slovenian language (human audio recordings)

    No full text
    2,060 recordings in mp3 format were made for the School Dictionary of the Slovenian Language based on the original recordings in wav format (48 kHZ, 24-bit). Around 600 recordings were made at the Institute of Ethnomusicology, ZRC SAZU, in collaboration with Peter Vendramin, whereas the rest of them were recorded and edited at the phonetics laboratory of the Fran Ramovš Institute of the Slovenian Language by Tanja Mirtič. The text to be recorded was read by Marko Snoj. The recordings represent standard language pronunciation of a speaker with a pitch-accent system from the Upper Carniolan dialect. Recordings also include pronunciation variants that are equivalent from the normative standpoint. The XML provides pairing between audio files and the headwords of the "Franček Portal Headword List" (http://hdl.handle.net/11356/1445)

    Franček portal historical module

    No full text
    The Franček Portal Historical Module contains data on earliest usage of Slovenian words in literary language as attested in the texts of 16th-century Slovenian Protestant authors; it also combines linked entries from three Slovenian historical dictionaries describing lexis from the late 17th, early 18th, and late 19th centuries, respectively. The underlying data for the first part of the module stems from the Words of the 16th-Century Slovenian Literary Language [http://hdl.handle.net/11356/1127], while the second part consists of IDs linking each module entry and relevant entries in respective historical dictionaries as listed below. The dataset is linked to: * Franček Portal Headword List [http://hdl.handle.net/11356/1445] * Dictionary of Maks Pleteršnik (1894–1895) [http://hdl.handle.net/11356/1114] * Dictionary of the Slovenian Language in the Works of J. Svetokriški [http://hdl.handle.net/11356/1092] * Dictionary of Kastelec and Vorenc (1680–1710) (1997) [https://fran.si/iskanje?FilteredDictionaryIds=138&View=1&Query=*

    Parallel Corpus (EN-LT) of EUR-Lex Documents That Include Terms with the Adjective 'Green' (ELEXIS)

    No full text
    Bilingual parallel corpus of the EU English documents containing terms with the adjective 'green' and their Lithuanian translations. The size of the corpus is 4,447,683 words in English, and 3,649,812 words in Lithuanian. The corpus is composed of 502 documents: 251 of each language

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇