Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Bulgarian Valency Dictionary - BVFL (ELEXIS)

    No full text
    БТБ-Валентен речник. The Valency Lexicon is a treebank-driven resource of extracted valency frames from BulTreeBank. It consists of 1000 most frequent verbs and their valence frames and it is based on a paper dictionary compiled by Balabanova and Ivanova. Each frame defines the number and the kind of the arguments and imposes morphosyntactic and semantic restrictions over them. The semantic restrictions over the arguments are extracted and matched against the SIMPLE core ontology. The frames of the most frequent verbs are compared to the corpus data and repaired if necessary (new frames are added, some of the existing frames are deleted or fine-grained)

    Dictionary of Slovenian Particles - SSC (ELEXIS)

    No full text
    Slovar slovenskih členkov. The dictionary describes the particles in the Slovenian language. It contains 429 entries with information on variants, dynamic and tonal accent, particle type, the meaning and etymology. See also: https://www.clarin.si/repository/xmlui/handle/11356/1128 This dictionary was published as a printed book: Žele, Andreja. Slovar slovenskih členkov. Ljubljana : Založba ZRC, ZRC SAZU, 2014

    English-Slovene term candidates KAS-biterm 1.0

    No full text
    KAS-biterm is an automatically generated glossary of English terms with their translations into Slovene. The pairs, possibly with their English and Slovene acronyms, were extracted from the Corpus of Academic Slovene KAS 1.0 (http://hdl.handle.net/11356/1244), where they have been annotated with the kas-biterm tool (https://github.com/clarinsi/kas-biterm) trained on the Bilingual terminology extraction dataset KAS-biterm 1.0 (http://hdl.handle.net/11356/1199). Note that only Query 1 was used for pre-selection of the sentences and for training the tool, and that the bi-lingual terms from the KAS corpus have been filtered to remove noise. The glossary is encoded in TEI-Lex0 (https://github.com/DARIAH-ERIC/lexicalresources) and gives, for each entry, also up to three examples of use, together with their bibliographic information. Various parts of the lexical entries also have links to the appropriate queries to CLARIN.SI noSketch Engine concordancer. The TEI encoded corpus is also available in a variant that is a much smaller document as it does not contain the examples of use and links

    Dictionary of living Slovenian "Razvezani Jezik" (The Unleashed Tongue)

    No full text
    Launched in December 2004 by the Domestic Research Society, Razvezani jezik (The Unleashed Tongue) is the first user-generated online dictionary of spoken Slovenian language. As a Wiki project, it allowed every visitor to add new entries freely. It quickly gained popularity and was labelled "the most entertaining Slovenian dictionary" by the media. Almost 3,000 authors belonging to different generations and subcultures contributed slang, neologisms and similar words and phrases from all Slovene regions. In 2019 the publisher decided to close outside edits to the dictionary and to archive its content in several ways, and this repository entry is one of them. The dictionary was exported and stored in XML, structured according to the accompanying DTD. Each dictionary entry includes the headword, a list of its keywords, a list of all revisions of the entry, which then contain the paragraphs with the text of the entry, and optional user comments. The authors of this resource would like to thank Luka Prinčič, Marko Mrđenović, Milan Erič, Damijan Kracina, Inge Pangos, Jani Pirnat, Polona Tavčar, Darja Vuga, Jaka Železnikar, Kaja Dolar, Anna Ehrlemark, Vasja Lebarič, and Ajdin Bašić for their invaluable help in making the Razvezani jezik project a success

    Spoken Torlak dialect corpus 1.0 (transcription)

    No full text
    Torlak corpus represents a spoken variety of the endangered Torlak dialect from the Timok area in Southeast Serbia. It comprises transcripts of interviews with the local population, collected in the field between 2015 and 2017. Semi-structured interviews were conducted eliciting spontaneous speech in the form of long narratives about traditional culture and history. The corpus is made up of semi-orthographic transcripts of 86.5 hours of recordings from locations evenly distributed across the Timok area of the Torlak dialect zone. The dialect is presently under the influence of a more prestigious Standard Serbian variety and expresses a great deal of variation in the use of non-standard features. The corpus contains samples of the typical representatives of the dialect with little influence of the standard, as well as a smaller portion of speakers who use both dialect and standard features. The corpus contains 489,021 tokens with accentuation, morphosyntacitc tags and lemmatisation. Accentuation was done manually by trained transcribers. Morphosyntactic annotation and lemmatisation (available in the TEI and vertical formats of the corpus) were done automatically, with minor manual corrections. The morphosyntactic tags follow the MULTEXT-East specificatins for Torlak, cf. https://github.com/clarinsi/mte-msd

    Frequency lists of character-level n-grams from the GOS 1.0 corpus 1.1

    No full text
    Frequency lists of character-level n-grams were extracted from the GOS 1.0 Corpus of Spoken Slovene (http://hdl.handle.net/11356/1040) using the LIST corpus extraction tool (http://hdl.handle.net/11356/1227). The lists contain 1-5-gram combinations of characters occurring in the corpus along with their absolute and relative frequencies, percentages, and distribution across the text-types included in the corpus taxonomy. Character-level n-grams were extracted from lemmas (5 files), lower-case word forms (5 files), and standardized word forms (5 files). Compared to the previous version (http://hdl.handle.net/11356/1268), this one includes fixes of several typos and substitutes all instances of "normalized forms" with the more adequate term "standardized forms" (as used in the SSJ project)

    The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Macedonian 1.0

    No full text
    This model for morphosyntactic annotation of standard Macedonian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the 1984 training corpus (to be published) and using the Macedonian CLARIN.SI word embeddings (http://hdl.handle.net/11356/1359). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~97.6

    Dictionary of Contemporary Dutch - ANW (ELEXIS)

    No full text
    Algemeen Nederlands Woordenboek (ANW). The ANW is a corpus-based, digital dictionary that describes contemporary Dutch in the Netherlands, Flanders, Suriname, and the Caribbean as comprehensively as possible. The language period covered by the ANW runs from 1970 to the present and more or less coincides with the post-war generations of adult language users. It is a synchronous dictionary with a focus on written language. The term 'Algemeen (general)' in the title should be understood as: not tied to a particular region, a particular group of people, or a particular field. In addition to words belonging to the core vocabulary, the ANW also describes neologisms (new words, new connections, new expressions, new meanings of already existing words). The ANW is an online dictionary and is not based on a printed version. The dictionary entries are designed for this purpose, and from the start, we have considered the different opportunities, demands and problems that come with the development of a new digital dictionary, with regard to both data collection, editing, and publication. For example, where relevant, images, videos or audio samples are added to the description of a word. It is an interactive dictionary. The ANW website is updated frequently to process additional information, corrections and revisions The ANW was first published online in 2009. See also: http://hdl.handle.net/10032/tm-a2-k

    Historical Finnish Dictionary - VKS (ELEXIS)

    No full text
    Vanhan kirjasuomen sanakirja. The Dictionary of Old Literary Finnish defines meanings and describes uses of all words used in Finnish literature from the 1540s to 1810

    Dictionary of Spanish Language 22 ed. (2001) - DLE22 (ELEXIS)

    No full text
    Diccionario de la lengua española 22 ed. (2001). The Diccionario de la lengua española is the standard dictionary of Spanish (a.k.a. Castilian) edited and produced by the Royal Spanish Academy (RAE). Its first edition dates from 1780, and its latest one is the 23rd edition published in 2014. The online version is comprised of the 22nd edition plus some of the work done for the 23rd edition. DLE is considered the most authoritative dictionary for the Spanish language. It includes commonly used words in any of the Spanish speaking countries. It also includes numerous archaic and unusual words with aims of understanding ancient Spanish literature

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇