Common Language Resources and Technology Infrastructure - Slovenia
Not a member yet
    840 research outputs found

    Ekspress news article archive (in Estonian and Russian) 1.0

    No full text
    The dataset is an archive of articles from the Ekspress Meedia news site from 2009-2019, containing over 1.4M articles, mostly in Estonian language (1,115,120 articles) with some in Russian (325,952 articles). Keywords are included for articles after 2015. The main archive is in file ee_articles_2009_2019. Other files contain derived versions and subsets - please see README files inside those zip files. The main archive contains JSON files of all the Estonian articles from the year 2009 to 2019 May. These datasets are intended for usage in EMBEDDIA, a H2020 project. Articles are in Estonian language with some in Russian. The main archive is in file ee_*articles_*2009_2019. Other files contain derived versions and subsets (please see README files inside those zip files), in short: - eearticles2015-2019: This dataset contains Estonian and Russian articles - 5 years, with tags, that were missing in the previous versions. - files eearticles20152019lemmatized and eearticles20092014lemmatized are the files preprocessed by TEXTA (contact [email protected]) - in file eeandsttarticlelemmasembeddingsand_usage you can find w2v embeddings trained by TEXTA (contact [email protected]) Description of the Main Dataset (eearticles_2009_2019) There are 12 JSON files: articles_2009_ver2.json contains 161394 articles from the year 2009 articles_2010_ver2.json contains 151033 articles from the year 2010 articles_2011_ver2.json contains 168273 articles from the year 2011 articles_2012_ver2.json contains 152772 articles from the year 2012 articles_2013_ver2.json contains 141012 articles from the year 2013 articles_2014_ver2.json contains 128388 articles from the year 2014 articles_2015_ver2.json contains 127425 articles from the year 2015 articles_2016_ver2.json contains 130704 articles from the year 2016 articles_2017_ver2.json contains 119318 articles from the year 2017 articles_2018_ver2.json contains 117388 articles from the year 2018 articles_2019_Jan-Apr_ver2.json contains 35076 articles from the year 2019 January to April articles_2019_May_ver2.json contains 8329 articles from the year 2019 May In sum: 1 441 112 articles Each JSON file is a list of dictionaries, i.e. each article is represented as a dictionary. Each dictionary contains the following: id (integer) - the ID of the article title (string) - the title of the article lead (string) - the lead of the article (can contain HTML, e.g. tag) url (string) - the URL of the article tags (list of dictionaries or None) [1]: each dictionary represents one tag. The tag dictionary contains the following: domain_id (string) [2] - the ID of the domain id (string) - the ID of the tag lang (string) - the language of the tag tag (string) - the tag itself, e.g. Kert Kingo (a name) translitted_name (string) - a modified version of the tag, e.g. kert-kingo rawBody (string) - the raw text of the article (contains HTML) bodyText (string) - clean article text (stripped from HTML) publishDate (string) - published date & time of the article categoryPrimary (dictionary or empty list) - the dictionary contains the following information: categoryId (integer) - the ID of the category categoryName (string)- the name of the category (e.g. World) channelId (integer) - the ID of the channel OR articleId (integer) - the ID of the article categoryId (integer) - the ID of the category categoryName (string)- the name of the category (e.g. World) categoryPrimary (boolean) - unknown categorySort (integer) - unknown categoryUrl (string) - the URL of the category categoryVisible (boolean) - unknown channelId (integer) - the ID of the channel channelUrl (string) - the URL of the channel (e.g. 'https://sport.delfi.ee') directoryName (string) - unknown parentId (integer) - unknown channelLanguage (string or None) [3] - the language of the channel categoryLanguage (int or None) [4] -unknown commentCount (int) [5] - the number of comments relatedArticles (list of integers) - a list of related articles' ID'

    Text collection for training the BERTić transformer model BERTić-data

    No full text
    The BERTić-data text collection contains more than 8 billion tokens of mostly web-crawled text written in Bosnian, Croatian, Montenegrin or Serbian. The collection was used to train the BERTić transformer model (https://huggingface.co/classla/bcms-bertic). The data consists of web crawls before 2015, i.e. bsWaC (http://hdl.handle.net/11356/1062), hrWaC (http://hdl.handle.net/11356/1064), and srWaC (http://hdl.handle.net/11356/1063); previously unpublished 2019-2020 crawls, i.e. cnrWaC, CLASSLA-bs, CLASSLA-hr, and CLASSLA-sr; the cc100-hr and cc100-sr parts of CommonCrawl (https://commoncrawl.org/); and the Riznica corpus (http://hdl.handle.net/11356/1180). All texts were transliterated to the Latin script. The format of the text collection is one-sentence-per-line, empty-line-as-document-boundary. More details, especially on the applied near-deduplication procedure, can be found in the BERTić paper (https://arxiv.org/pdf/2104.09243.pdf)

    Montenegrin web corpus meWaC 1.0

    No full text
    The Montenegrin web corpus meWaC was built by crawling the .me top-level domain in 2019. The corpus was near-deduplicated on paragraph level, normalised via transliteration into the Latin script, and morphosyntactically annotated, lemmatised and dependency-parsed with a prototype version of the classla pipeline (https://pypi.org/project/classla/). Each document is accompanied by the URL and title metadata. The corpus is available in CoNLL-U format and as vertical file (wilth included registry) for mounting on CQP-compatible concordancers

    GLOBAL Arabic-French Dictionary - MLDS (ELEXIS)

    No full text
    A general language Arabic to French dictionary

    GLOBAL French-Polish Dictionary - MLDS (ELEXIS)

    No full text
    A general language French to Polish dictionary

    GLOBAL Russian-French Dictionary - MLDS (ELEXIS)

    No full text
    A general language Russian to French dictionary

    GLOBAL Chinese-French Dictionary - MLDS (ELEXIS)

    No full text
    A general language Chinese to French dictionary

    The Explanatory Dictionary of the Hungarian Language - ÉrtSz (ELEXIS)

    No full text
    A magyar nyelv értelmező szótára. The seven volumes of The Explanatory Dictionary of the Hungarian Language were published between 1959 and 1962 by the Research Institute for Linguistics, Hungarian Academy of Sciences. It covers Hungarian literary language of the 19th century as well as the written and spoken standard Hungarian of the first half of the 20th century. The main source of input was a corpus of about 6 million examples collected since the end of the 19th century. Each sense unit is illustrated by a few examples: citations from the classical Hungarian literature and example sentences created by the lexicographers

    Portal of Contemporary Croatian Personal Names (ELEXIS)

    No full text
    Portal suvremenih hrvatskih osobnih imena is based on the Dictionary of Contemporary Croatian Personal Names published in 2018 by Institute of Croatian Language and Linguistics. It covers pesonal names from IX. century to modern times

    Parallel Corpus (EN-LT-FR) of EUR-Lex Document Extracts That Include Terms with the Colour Names Black, White and Grey

    No full text
    Trilingual parallel corpus of EUR-Lex Document Extracts that include terms with colour names (black, white and grey). The size of the corpus is 23,198 words in English, 19,262 words in Lithuanian, and 27,274 words in French. The corpus is composed of 1326 document extracts: 442 of each language

    5

    full texts

    840

    metadata records
    Updated in last 30 days.
    Common Language Resources and Technology Infrastructure - Slovenia
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇