CLARIN-PL
Not a member yet
504 research outputs found
Sort by
Vector representations of polish words (Word2Vec method)
Model skip gram with vectors of length 100. Trained on kgr 10, a corpora with over 4 billion tokens. Data preprocessing involved segmentation, lemmatization and mophosyntactic disambiguation with MWE annotation
Morfeusz 2
Morfeusz 2 is a dictionary based morphological analyser and generator for Polish. This version of the program is decoupled from the dictionary. Two dictionaries of Polish developed within other projects are distributed with Morfeusz 2, namely SGJP and Polimorf
Cinderella - tool for Clustering and Classifications of Texts in Polish
System for clustering and classifications of Texts in Polish. Source code
Bilingual Cascade Dictionary
Bilingual Cascade Dictionary is a collection of dictionaries organised in a cascade with the top-most dictionaries having the highest priority in applications
Część komentarzy internetowych dłuższych niż 500 znaków do filmu YT: Mazurek Kapeli - Polacy witają uchodźców
Testowy, próbny korpus komentarzy internetowych opublikowanych do filmu "Mazurek Kapeli - Polacy witają uchodźców! - YouTube"
https://www.youtube.com/watch?v=dAX4vJiO9Aw
komentarze dłuższe niż 500 znakó
plWordNet 3.0 – Almost There
It took us nearly ten years to get from no wordnet for Polish to the largest wordnet ever built. We started small but quickly learned to dream big. Now we are about to release plWordNet 3.0-emo – complete with sentiment and emotions annotated – and a domestic version of PrincetonWordNet, larger thanWordNet 3.1 by nearly ten thousand newly added words. The paper retraces the road we travelled and talks a little about the future
PELCRA for National Corpus of Polish Search Engine 2
The PELCRA for NKJP search engine 2 provides access to the full National Corpus of Polish dataset (over 1.5 billion word tokens). In addition to linguistically motivated corpus queries, it supports a number of data exploration and visualisation features. Most of the functionality of the search engine is available through a REST web service. Access to the API is available upon request
Slowal
Slowal is a web tool designed for creating, editing and browsing valence dictionaries. So far, it has mainly been used for creating The Polish Valence Dictionary (Walenty).
Slowal supports the process of creating the dictionary; it also facilitates access by making it possible to browse the dictionary using an advanced built-in filtering system, covering both syntactic and semantic phenomena. Slowal also gives control over the work of lexicographers involved in creating dictionary, for instance by using predefined lists of values, which prevents spelling errors and enforces consistency, as well as by imposing strict validation rules.
Last but not least, the created dictionary can be exported from Slowal in various formats: plain text, TeX, PDF, and TEI XML