504 research outputs found

    BlogReader

    No full text
    BlogReader - corpus acquisition from structured web source

    Wnuk

    No full text
    opi

    Transcriptions of the Polish Film Chronicles (Polska Kronika Filmowa) - years 1945-1962

    No full text
    This is the orthographic transcription of the audio of the Polish Film Chronicles (Polska Kronika Filmowa - PKF) between the years 1945-1962. The transcription is mostly hand-checked and should match the audio to a high degree. Only the narrator is transcribed (which is the vast majority of all speech in the recordings)

    Liner2.6 model NER NKJP

    No full text
    Liner2.6 NER NKJP model The package contains a pre-trained Liner2 (https://github.com/CLARIN-PL/Liner2) model for recognition named entities according to NKJP guidelines. The model was trained on the NKJP corpus (http://nkjp.pl/) and evaluated in the PolEval 2018 Task 2 (http://poleval.pl/tasks/). The model won third place with the following results: Exact — 0.778, Overlap — 0.818, Final — 0.810. References: * NKJP corpus in TEI format — http://clip.ipipan.waw.pl/NationalCorpusOfPolish?action=AttachFile&do=view&target=NKJP-PodkorpusMilionowy-1.2.tar.gz * PolEval 2018 Task 2 evaluation corpus — http://mozart.ipipan.waw.pl/~axw/poleval2018

    Tagger SentiOne - version 1

    No full text
    The SentiOne tagger is a tagger for the Polish language adapted to processing of user-generated content. It was trained on the Polish UGC-corpus (prepared within the same research project and soon to become available in the CLARIN repository)

    Multisłownik: Linking plWordNet-based Lexical Data for Lexicography and Educational Purposes

    Get PDF
    Multisłownik is an automated integrator of Polish lexical data retrieved from multiple available online sources intended to be used in various scenarios requiring access to such data, most prominently dictionary creation, linguistic studies and education. In contrast to many available internet dictionaries Multisłownik is WordNet-centric, capturing the core definitions from Słowosieć synsets. The paper provides details of construction of the resource, discussed the difficulties related to linking different logical structures of underlying data and investigates two sample scenarios for using the resulting platform

    Mahabharata_sample

    No full text
    lemmatised Sanskrit text with lemma IDs, white-space delimite

    Blogi_zip 02

    No full text
    blogi zi

    Testowy MPW

    No full text
    Wypowiedzi europosłó

    Liner2.5

    No full text
    Generic framework for information extraction tasks, including recognition of named entities, temporal expressions, spatial expressions and events

    40

    full texts

    504

    metadata records
    Updated in last 30 days.
    CLARIN-PL
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇