Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Slovenian 1.0

    No full text
    This model for morphosyntactic annotation of non-standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and the Janes-Tag corpus (http://hdl.handle.net/11356/1238), using the CLARIN.SI-embed.sl word embeddings (http://hdl.handle.net/11356/1204). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~96.14

    The CLASSLA-StanfordNLP model for lemmatisation of non-standard Slovenian 1.0

    No full text
    The model for lemmatisation of non-standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and the Janes-Tag corpus (http://hdl.handle.net/11356/1238), using the Sloleks inflectional lexicon (http://hdl.handle.net/11356/1230). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~98.86

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Croatian 1.1

    No full text
    The model for morphosyntactic annotation of standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1183) and using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1205). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~94.1. The difference to the previous version of the model is that now the whole XPOS tag is predicted and not specific characters, as was the case in stanfordnlp, which resulted in illegal XPOS tags (and slightly decreased performance)

    Annotated corpus of Croatian language-related news comments MetaLangNEWS-COMMENTS-Hr

    No full text
    A comprehensive corpus of user comments on online news articles on the topic of language from major Croatian daily newspapers and news portals, published in the five-year period of January 1, 2015 - January 1, 2020. The corpus is designed to facilitate research on metalanguage (‘language about language’), linguistic ideologies, language policy and planning, as well as the specific contemporary debates on language defining, naming, and standardisation, from the bottom-up perspective. The corpus has been tagged using the CLASSLA-StanfordNLP models for morphosyntactic annotation and lemmatisation of non-standard Croatian. The corpus is available in plain text version, XML with full metadata, and tagged CONLL-U format. This collection is complementary to the corpus of news articles MetaLangNEWS-Hr (http://hdl.handle.net/11356/1369). Parallel versions from Slovenia (http://hdl.handle.net/11356/1362) and Serbia (http://hdl.handle.net/11356/1372) are also available

    Epigraphic corpus of Medieval and Early Modern inscriptions in Slovenia MEMIS 1.0

    No full text
    The Epigraphic corpus of Mediaeval and Early Modern inscriptions in Slovenia collects carefully made transcriptions of Latin inscriptions that are found or have been discovered on the territory of Slovenian coastal towns, focused on Koper and Piran. The 51 collected inscriptions date from 1222 to the middle of the 17th century. The purpose of the corpus is to provide a methodological basis for the processing of inscriptions found or discovered in the Slovenian ethnic territory. The corpus brings together inscriptions that are either still located in their primary context or have been moved or even destroyed and are therefore accessible only in transcriptions. The material was collected through field research, i.e. the recording and documentation of inscriptions in situ. The corpus transcriptions have been carefully annotated, with the ligatures and abbreviations marked-up and expanded, and the texts translated into Slovenian. Various metadata (in Slovenian) have also been collected and included in the descriptions. The corpus is encoded in TEI XML, using the elements for manuscript descriptions for inscription metadata

    Morphological patterns from the Sloleks 2.0 lexicon 1.0

    No full text
    This entry consists of XML files with 96,290 lexical units (nouns, verbs, adjectives, and adverbs) from the Sloleks Morphological Lexicon of Slovene 2.0 (http://hdl.handle.net/11356/1230) that include codes for morphological patterns. The pattern codes were designed based on a manual analysis of automatically extracted paradigms and were obtained as follows: The lexical units from Sloleks 2.0 were first automatically clustered into groups through a rule-based approach based on (1) a number of predetermined grammatical features from the MULTEXT-East Version 6 morphosyntactic specifications for Slovenian (http://nl.ijs.si/ME/V6/), such as part of speech, gender and properness for nouns, aspect for verbs, and (2) the differentiating characteristics of their morphological paradigms (i.e. their mutable word parts, which are similar to but not always overlapping with the linguistic definition of word endings – for example: čas-Ø; čas-a; čas-om / prijatelj- Ø; prijatelj-a; prijatelj-em / odstot-ek; odstot-ka; odstot-kom). More than 1,000 automatically extracted pattern candidates were subsequently linguistically analyzed, combined into groups, and hierarchically organized. As a result, every lexical unit in the XML file features a code (listed as ) corresponding to the relevant morphological paradigm in the hierarchy (available in the accompanying file titled "nssss_morphological_pattern_hierarchy_1.0.tsv"). Because the patterns were extracted from Sloleks 2.0, they reflect the decisions that were implemented in its initial compilation, particularly in terms of the degree of morphological variation documented in the lexicon (e.g. not all morphological variants are necessarily included in the lexicon) and paradigm integrity (for instance, some nouns in Sloleks 2.0 only feature singular or plural forms). It should be noted that non-standard word forms were not included in the design of the patterns. In addition, the XML file does not contain lexical units from Sloleks 2.0 that consist of word forms from more than one morphological paradigm (e.g. lesketati – lesketam / leskečem; or lojen – lojenega / lojnega), or other problematic units (such as those with missing or erroneous data)

    List of single-word male and female occupations in Slovenian

    No full text
    The list of single-word occupations in Slovene is based on the Slovene Standard Classification of Occupations (https://www.uradni-list.si/glasilo-uradni-list-rs/vsebina?urlid=199728&stevilka=1641). The list includes 234 occupation pairs. For each occupation, it contains its masculine word form (e.g. fotograf), its possible synonym, its feminine equivalent (e.g. fotografka) and the corresponding synonym of the feminine form (e.g. fotografinja). The cases where no synonyms were added for a specific occupation are denoted with the label 0 (note that only synonyms with the same root are considered). Several conditions for inclusion or exclusion of an occupation to the list were applied: - Our list contains only single word occupation pairs, while the majority of the occupations in the aforementioned classification are multi-word expressions. - An occupation has to exist both in female and male grammatical gender (gender-neutral words such as pismonoša [en. postman] are not included in the list). - At least one of the variants of an occupation (masculine or feminine) occurs at least 500 times in the Corpus of Written Standard Slovene Gigafida 2.0. - The occupations that are also proper names in Slovene, e.g. kovač [en. blacksmith], were filtered out if in the Slovene Morphological Lexicon Sloleks 2.0 (Dobrovoljc et al., 2019) the proper name form exists. - Occupations that could be easily associated with a context unrelated to occupations (e.g. čarovnik/čarovnica [en. wizard/witch]) or where a male or female variant is a homograph of a common noun (e.g. detektivka [en. detective] also denotes a detective novel) were excluded from the final set of occupations. When a more established version of an occupation exists, we manually add a synonym with the same root (e.g. in the case of fotografka, an arguably more established fotografinja was added [en. photographer]). If the standard classification does not include the female (e.g. dramatik [en. playwright]) or the male version (e.g. prostitutka [en. prostitute]) of an occupation, the missing version is manually added if it exists and appears in Gigafida corpus (e.g. there are no established words for female and male versions of postrešček [en. porter] and hostesa [en. hostess]). The list of occupations can be used for different natural language processing tasks including evaluation of word embeddings models through analogies, which can point to bias in language use. If you use the dataset, please cite the following paper: SUPEJ, Anka, ULČAR, Matej, ROBNIK ŠIKONJA, Marko, POLLAK, Senja (2020). Primerjava slovenskih besednih vektorskih vložitev z vidika spola na analogijah poklicev. Zbornik konference Jezikovne tehnologije in digitalna humanistika / Proc. of the Conference on Language Technologies and Digital Humanities, p. 93-100

    Karelian Dictionary - KKS (ELEXIS)

    No full text
    Karjalan kielen sanakirja. The Dictionary of the Karelian Language was published in 1968–2005 by the Institute for the Languages of Finland and the Finno-Ugrian Society: Part 1 (A – J) 1968, Part 2 (K) 1974, Part 3 (L – N) 1983, Part 4 (O – P) 1993, part 5 (1997) and part 6 (T – Ö) 2005. The six-part work comprises a total of about 3,800 pages and almost 83,000 headwords. It cover the period from the 1880s to 1970. The dictionary is a dialect dictionary. It represents almost all the dialects of the actual Karelia - the Viennese Karelia and the South Karelia - and Aunus, or Livvi. Lydian dialects are not included in the dictionary. Headwords are usually in accordance with Viennese dialects. The metalanguage is Finnish. The dictionary has been transferred to electronic format as follows: The articles in the first three parts (alphabet A to N), printed using the old handwriting technique, have been scanned, corrected and converted to structured form. The articles in the last three sections (alphabet O – Ö) have been converted to a structured format from WordPerfect files. The online version was released in 2009. There have been few changes to the online dictionary compared to the printed dictionary. The errors have been corrected and some of the articles in the first part have been later adapted to established delivery principles. However, the articles have not been specifically harmonized. The dictionary is available for download in Kielipankki - the Language Bank of Finland, http://urn.fi/urn:nbn:fi:lb-2015110501, as well as on the website of The Institute for the Languages of Finland, http://kaino.kotus.fi/kks/lataa/kksxml.zip, it is also available at http://kaino.kotus.fi/cgi-bin/kks/karjala.cgi. The online dictionary (link above) is freely available and contains e.g. an introduction to the printed dictionary, which introduces the material used in the dictionary work and the stages of the dictionary work, and gives instructions to the user of the dictionary. The editors-in-chief of the Karelian Dictionary were Pertti Virtaranta (parts 1–3, alphabet a – n, 1955–1983), Jaakko Sivula (oto 1983–1985) and Raija Koponen (parts 4–6, alphabet o – ö, 1986–2005) . The editorial board of the dictionary has included Katariina Jeskanen, Matti Jeskanen, Leena Joki, Eero Kiviniemi, Pirkko Poutanen, Matti Punttila, Laila Rissanen, Tauno Särkkä, Marja Torikka (formerly Lehtinen) and Helmi Virtaranta. The online dictionary has been edited by Marja Torikka. The responsible reporter since September 8, 2010 is Leena Joki. The XML version has the same content as the online dictionary. Jari Vihtari is responsible for the web application and the creation of the XML file

    Dictionary of Finnish Dialects - SMS (ELEXIS)

    No full text
    Suomen murteiden sanakirja. The Dictionary of Finnish Dialects defines meanings and describes uses of words used in Finnish dialects from 1900s to 1970s. In addition, there are real maps describing the distribution of words and their meanings in Finnish dialects

    The Dictionary of the Modern Lithuanian Language - DLKZ (ELEXIS)

    No full text
    Dabartinės lietuvių kalbos žodynas. This is a monolingual explanatory dictionary of the Lithuanian language of 20th century

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇