504 research outputs found

    Corpus2MWE

    No full text
    A CCL reader (Corpus2) with MWE detection

    Polish-Ukrainian Parallel Corpus

    No full text
    Polish-Ukrainian Parallel Corpu

    plWordNet 4.0

    No full text
    PLWordNet ver. 4.0 is a lexico-semantic network which reflects the lexical system of the Polish language with projection to the English language. Słowosieć, Princeton Wordnet, EnWordnet together the whole resource currently contains 506815 senses, 347564 synsets and over 1.5M relations and 361177 inter-lingual relations between lexical units. It is now the largest wordnet in the world and is still growing

    enWordNet 1.0

    No full text
    The extension of Princeton WordNet built within the CLARIN-PL project. The attached file also contains the mapping to Open Multilingual Wordnet

    Word embeddings for Polish (KGR10, Fasttext binary) kgr10_fasttext_bin_v1

    No full text
    Distributional language model (binary) for Polish trained on KGR10 using Fasttext (vector dimension: 100)

    Lilia

    No full text
    sample of historical text

    CorpoGrabber

    No full text
    CorpoGrabber: The Toolchain to Automatic Acquiring and Extraction of the Website Content Jan Kocoń, Wroclaw University of Technology CorpoGrabber is a pipeline of tools to get the most relevant content of the website, including all subsites (up to the user-defined depth). The proposed toolchain can be used to build a big Web corpora of text documents. It requires only the list of the root websites as the input. Tools composing CorpoGrabber are adapted to Polish, but most subtasks are language independent. The whole process can be run in parallel on a single machine and includes the following tasks: downloading of the HTML subpages of each input page URL [1], extracting of plain text from each subpage by removing boilerplate content (such as navigation links, headers, footers, advertisements from HTML pages) [2], deduplication of plain text [2], removing of bad quality documents utilizing Morphological Analysis Converter and Aggregator (MACA) [3], tagging of documents using Wrocław CRF Tagger (WCRFT) [4]. Last two steps are available only for Polish. The result is a corpora as a set of tagged documents for each website. References [1] https://www.httrack.com/html/faq.html [2] J. Pomikalek. 2011. Removing Boilerplate and Duplicate Content from Web Corpora. Ph.D. Thesis. Masaryk University, Faculcy of Informatics. Brno. [3] A. Radziszewski, T. Sniatowski. 2011. Maca – a configurable tool to integrate Polish morphological data. Proceedings of the Second International Workshop on Free/Open-Source Rule-Based Machine Translation. Barcelona, Spain. [4] A. Radziszewski. 2013. A tiered CRF tagger for Polish. Intelligent Tools for Building a Scientific Information Platform: Advanced Architectures and Solutions. Springer Verlag

    Genology

    No full text
    Corpu

    Diachrono - sample

    No full text
    Sample of diachronic corpu

    Linguistic

    No full text
    Corpu

    40

    full texts

    504

    metadata records
    Updated in last 30 days.
    CLARIN-PL
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇