Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Bulgarian Explanatory Dictionary (ELEXIS)

    No full text
    Речник на съвременния български език. The dictionary includes 59,623 headwords and attached grammatical and stylistic notes, word senses and examples. In addition to the general vocabulary, following a certain selection, the dictionary includes some obsolete words, words gradually moving to the passive vocabulary, and foreign words which are widely used in modern Bulgarian

    24sata news comment dataset 1.0

    No full text
    The dataset of user comments provided for research purposes for the EMBEDDIA, a Horizon 2020 project, extracted from the database of user comments from the 24sata.hr news portal. The 24sata.hr is the largest-circulation daily newspaper in Croatia, reaching on average 2 million readers daily. The dataset provides the comments metadata including the link to the relevant article, the ID of the comment author (anonymized), and timestamp. The comments are also labelled if they are blocked by human moderators. Description of the Datasets. The 24sata dataset consists of 11 columns and 21548192 rows. Each row represents one user comment on the 24sata news portal. Comments are added by registered users below the published news article. Columns: 'comment_id' - The internal id of the comment. Unique for each row. 'user_id' - The internal id of the user writing the comment. Unique for each user. '0' for all blocked comments. 'content' - The content (text) of the user comment. 'site' - The site the comment came from. 'reply_to_id' - The 'comment_id' of the parent comment - if this comment was intended as a reply. 'created_date' - The date the comment was created. 'last_change' - The date the comment was last edited. 'article_id' - A public id of the article where this comment was posted. The article itself can be accessed by appending article_id to the site. So an article with article_id 614684 and site 'www.24sata.hr' can be found on 'www.24sata.hr/a-614684'. (note the added 'a-' before the article name) 'infringed_on_rule' - If the user has infringed on rules with this comment, id of the rule is given. The description of the rules is given below. 'like_counts' - A number of times other users have voted in favour of this comment, similar to the Like button. 'dislike_counts' - A number of times other users have voted against this comment, opposite of the Like button

    Linguistically annotated multilingual comparable corpora of parliamentary debates ParlaMint.ana 2.1

    No full text
    ParlaMint 2.1 is a multilingual set of 17 comparable corpora containing parliamentary debates mostly starting in 2015 and extending to mid-2020, with each corpus being about 20 million words in size. The sessions in the corpora are marked as belonging to the COVID-19 period (from November 1st 2019), or being "reference" (before that date). The corpora have extensive metadata, including aspects of the parliament; the speakers (name, gender, MP status, party affiliation, party coalition/opposition); are structured into time-stamped terms, sessions and meetings; with speeches being marked by the speaker and their role (e.g. chair, regular speaker). The speeches also contain marked-up transcriber comments, such as gaps in the transcription, interruptions, applause, etc. Note that some corpora have further information, e.g. the year of birth of the speakers, links to their Wikipedia articles, their membership in various committees, etc. The corpora are encoded according to the Parla-CLARIN TEI recommendation (https://clarin-eric.github.io/parla-clarin/), but have been validated against the compatible, but much stricter ParlaMint schemas. This entry contains the linguistically marked-up version of the corpus, while the text version is available at http://hdl.handle.net/11356/1432. The ParlaMint.ana linguistic annotation includes tokenization, sentence segmentation, lemmatisation, Universal Dependencies part-of-speech, morphological features, and syntactic dependencies, and the 4-class CoNLL-2003 named entities. Some corpora also have further linguistic annotations, such as PoS tagging or named entities according to language-specific schemes, with their corpus TEI headers giving further details on the annotation vocabularies and tools. The compressed files include the ParlaMint.ana XML TEI-encoded linguistically annotated corpus; the derived corpus in CoNLL-U with TSV speech metadata; and the vertical files (with registry file), suitable for use with CQP-based concordancers, such as CWB, noSketch Engine or KonText. Also included is the 2.1 release of the data and scripts available at the GitHub repository of the ParlaMint project. As opposed to the previous version 2.0, this version corrects some errors in various corpora and adds the information on upper / lower house for bicameral parliaments. The vertical files have also been changed to make them easier to use in the concordancers

    Multilingual comparable corpora of parliamentary debates ParlaMint 2.1

    No full text
    ParlaMint 2.1 is a multilingual set of 17 comparable corpora containing parliamentary debates mostly starting in 2015 and extending to mid-2020, with each corpus being about 20 million words in size. The sessions in the corpora are marked as belonging to the COVID-19 period (after November 1st 2019), or being "reference" (before that date). The corpora have extensive metadata, including aspects of the parliament; the speakers (name, gender, MP status, party affiliation, party coalition/opposition); are structured into time-stamped terms, sessions and meetings; with speeches being marked by the speaker and their role (e.g. chair, regular speaker). The speeches also contain marked-up transcriber comments, such as gaps in the transcription, interruptions, applause, etc. Note that some corpora have further information, e.g. the year of birth of the speakers, links to their Wikipedia articles, their membership in various committees, etc. The corpora are encoded according to the Parla-CLARIN TEI recommendation (https://clarin-eric.github.io/parla-clarin/), but have been validated against the compatible, but much stricter ParlaMint schemas. This entry contains the ParlaMint TEI-encoded corpora with the derived plain text version of the corpus along with TSV metadata on the speeches. Also included is the 2.0 release of the data and scripts available at the GitHub repository of the ParlaMint project. Note that there also exists the linguistically marked-up version of the corpus, which is available at http://hdl.handle.net/11356/1431

    Abstracts from the KAS corpus KAS-Abs 1.0

    No full text
    The KAS-abs corpus contains 108,254 automatically identified Slovenian and/or English abstracts (30 million words) from 62,000 BSc/BA, MSc/MA, and PhD theses included in the KAS Corpus of Academic Slovene. This corpus is made available because the public version of KAS (http://hdl.handle.net/11356/1244) does not contain the front matter, and hence the abstracts. The abstracts were identified on a per-page basis, and are either in Slovenian (*-abs-sl.txt, 47,273 files), English (*-abs-en.tx, 49,261 files) or, for cases where the abstracts in both languages were on the same page, in both languages (*-abs-slen.txt, 11,720 files). The files contain the plain text of the abstracts, one paragraph per line. Note that as the cleaning of source PDF files and identification of the abstracts was done automatically, this corpus contains various types of errors. The files are stored in the same manner as for the complete KAS corpus, i.e. in 1,000 directories with the same filename prefix as in KAS. The file with the metadata for the corpus texts is also included. The abstracts can be useful for research in e.g. machine translations and terminology extraction, and, using also the full texts from the KAS corpus, for studies in automatic summarisation

    Multiword Expressions lexicon extracted from the Gigafida 2.1 corpus

    No full text
    The MWE lexicon was extracted from the Gigafida 2.1 Corpus of Written Standard Slovene (https://www.clarin.si/noske/run.cgi/corp_info?corpname=gfida21) using specialized scripts for extracting data from corpora containing syntactic dependency annotations. The lexicon contains 5,242 Multiword Expressions with 12,358 examples from Gigafida 2.1. Each MWE entry (or sense) contains at least one and up to three extracted examples. MWEs were analysed using the JOS dependency parser system (http://nl.ijs.si/jos/bib/jos-skladnja-navodila.pdf) and were assigned matching syntactic structure IDs. The corpus sentences containing the MWE components and matching syntactic structure features were identified in the corpus and assigned to the corresponding headword or sense. MWEs variants (or variant senses) are linked with the "senseKey" attribute values, forming a MWE cluster of related variants or variant senses. A sample of MWE headwords also contains manually created sense division with descriptions of meaning for each sense

    Basic vocabulary of the Dictionary of the Danish Language - ODS (ELEXIS)

    No full text
    Ordbog over det Danske Sprog (ODS), Basic vocabulary. ODS is a historic dictionary. It was published 1918-46 in 28 volumes and has since been expanded with 5 supplementary volumes. The dictionary covers the Danish lexicon from 1700 till 1950 approximately. An online version has been available since 2005. Contents: This resource contains the basic content of the basic vocabulary of the online version of ODS (ordnet.dk/ods). The data has been processed by Elexifier. Entries included: 4.515 entries 3.959 entries corresponding to DDO entries holding senses linked to WordNet base concepts in connection with the DanNet project. 556 entries corresponding to the most comprehensive DDO entries not included in the group mentioned above. Information types included. Elements are listed under their source names = the element names in m:e="...” in the uploaded, Elexified file: headword: 1st occurring headword in ODS. headword2: Secondary hedword(s) in ODS. Typically spelling variants, but also closely related words (e.g. “derivational siblings”), that are treated in the same entry. Words from this last group typically have an independant entry in DDO. POS sense (1 or more) definition (1 or more (#2ff = subsenses)) “Definition” is merely the raw text following the sense delimiters (as the detailed markup of definitions, examples etc. is still far from complete)

    GLOBAL Turkish-French Dictionary - MLDS (ELEXIS)

    No full text
    A general language Turkish to French dictionary

    High-German Idiom - Adelung (1811) (ELEXIS)

    No full text
    The Grammatical-Critical Dictionary of the High-German Idiom by Johann Christoph Adelung (1811) describes the vocabulary of the second half of the 18th century (1750-1799). It was the first scientific dictionary of the German language

    The Croatian Web Dictionary Mrežnik (ELEXIS)

    No full text
    Hrvatski mrežni rječnik ­– Mrežnik. Croatian Web Dictionary – Mrežnik is a free, corpus-based, born-digital, monolingual, easily searchable, hypertext, normative online dictionary of the Croatian standard language. It consists of three modules: for adult native speakers of Croatian, schoolchildren, and non-native speakers of Croatian. It is the central meeting point of the existing language resources of the Institute of Croatian Language and Linguistics, but also of the language resources created within the project. From 1st March 2017 to 31st August 2021 the work on Mrežnik has been financed by the Croatian Science Foundation. See also: http://hdl.handle.net/11356/147

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇