Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
840 research outputs found
Sort by
Ladin-German Dictionary (Letter F) (ELEXIS)
Dizionar Ladin-Deutsch (Letter F).
This dataset contains the letter F of the Ladin-German dictionary by Giovanni Mischí
Lemma list of the Dictionary of Contemporary Portuguese - DLPC (ELEXIS)
Dicionário da Língua Portuguesa Contemporânea (DLPC) is a monolingual Portuguese dictionary published by Academia das Ciências de Lisboa (2001). This dictionary also represents the first complete edition of a Portuguese Academy dictionary, from A to Z. The DLPC contains around 70,000 entries. This dictionary is being updated.
The dictionary was accessible on-line at www.acad-ciencias.pt/academia/illlp-novo-dicionario. However, this URL no longer works
Dictionary of Bavarian Dialects of Austria - WBOE (ELEXIS)
Wörterbuch der Bairischen Mundarten in Österreich.
Automatically created XML/TEI versions of first 5 volumes of Dictionary of Bavarian Dialects in Austria, published between 1963 and 2015, covering A, B/P, C, D/T and E.
The dictionary of Bavarian dialects in Austria (WBÖ) is a large-scale dialect dictionary in which the Bavarian vocabulary in Austria and the neighboring South Tyrol is documented. The first 5 volumes also take into account the historical language of the office and other Bavarian-speaking areas of the former monarchy
Bilingual Corpus of Underground Mining (ELEXIS)
PodzemniRadovi-sr-en, dvojezični poravnati korpus radova iz oblasti rudarstva. Undeground-mining-sr-en: bilingual texts from the Underground Mining Engineering journal (55 papers from 8 issues), aligned at the sentence level (4831 translation units), represented in the TMX (Translation Memory eXchange) format
GeologyTerm (ELEXIS)
Podskup GeolISSTerm delimično poravnat sa GeoSciML. Subsets of GeolISSTerm - Dictionary of geological terms and terms used in Geological information system of Serbia (GeolISS) partially aligned with GeoSciML
Treq Translation Equivalents (ELEXIS)
Data for Treq interface 2.0 derived from the InterCorp parallel corpus release 12
Basque Dictionary (ELEXIS)
Hiztegi Batua. Standard Basque Dictionary 2010 TEI-XML release
List of formulaic sequences in standard written Slovenian
This document contains 1,891 formulaic sequences in standard written Slovenian, i.e. frequently recurring strings of two to five words, manually annotated for syntactic structure, pragmatic function, and dictionary relevance. The list of sequences with a minimum frequency threshold of 20/million is based on the Frequency lists of word-level n-grams from lowercase word forms in Gigafida 2.0 (http://hdl.handle.net/11356/1274) and contains the union of top-1,000 formulaic sequences ranked by frequency and five association measures (Dice, t-test, MI, MI3, simple-LL).
Note that there exists a related entry "List of formulaic sequences in spoken Slovenian", http://hdl.handle.net/11356/1279
The CLASSLA-StanfordNLP model for lemmatisation of standard Croatian 1.1
The model for lemmatisation of standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1183) and using the hrLex inflectional lexicon (http://hdl.handle.net/11356/1232). The estimated F1 of the lemma annotations is ~97.6.
The difference to the previous version of the model is that it is trained with the lemmatiser padding bug removed, cf. https://github.com/stanfordnlp/stanfordnlp/issues/143
The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Bulgarian 1.0
This model for morphosyntactic annotation of standard Bulgarian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the BulTreeBank training corpus (http://hdl.handle.net/11495/D93F-C6E9-65D9-2) and using the CoNLL2017 word embeddings (http://hdl.handle.net/11234/1-1989). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~96.8