CLARIN-PL
Not a member yet
504 research outputs found
Sort by
Transcriptions of the Polish Film Chronicles (Polska Kronika Filmowa) - years 1945-1962
This is the orthographic transcription of the audio of the Polish Film Chronicles (Polska Kronika Filmowa - PKF) between the years 1945-1962. The transcription is mostly hand-checked and should match the audio to a high degree. Only the narrator is transcribed (which is the vast majority of all speech in the recordings)
Liner2.6 model NER NKJP
Liner2.6 NER NKJP model
The package contains a pre-trained Liner2 (https://github.com/CLARIN-PL/Liner2) model for recognition named entities according to NKJP guidelines. The model was trained on the NKJP corpus (http://nkjp.pl/) and evaluated in the PolEval 2018 Task 2 (http://poleval.pl/tasks/).
The model won third place with the following results: Exact — 0.778, Overlap — 0.818, Final — 0.810.
References:
* NKJP corpus in TEI format — http://clip.ipipan.waw.pl/NationalCorpusOfPolish?action=AttachFile&do=view&target=NKJP-PodkorpusMilionowy-1.2.tar.gz
* PolEval 2018 Task 2 evaluation corpus — http://mozart.ipipan.waw.pl/~axw/poleval2018
Tagger SentiOne - version 1
The SentiOne tagger is a tagger for the Polish language adapted to processing of user-generated content. It was trained on the Polish UGC-corpus (prepared within the same research project and soon to become available in the CLARIN repository)
Multisłownik: Linking plWordNet-based Lexical Data for Lexicography and Educational Purposes
Multisłownik is an automated integrator of Polish lexical data retrieved from multiple available online sources intended to be used in various scenarios requiring access to such data, most prominently dictionary creation, linguistic studies and education. In contrast to many available internet dictionaries Multisłownik is WordNet-centric, capturing the core definitions from Słowosieć synsets. The paper provides details of construction of the resource, discussed the difficulties related to linking different logical structures of underlying data and investigates two sample scenarios for using the resulting platform
Liner2.5
Generic framework for information extraction tasks, including recognition of named entities, temporal expressions, spatial expressions and events