CLARIN-PL
Not a member yet
504 research outputs found
Sort by
Teksty reklam TVP ABC
teksty reklam emitowane na kanale TVP ABC miedzy lipcem 2014 a styczniem 201
NE_SUMO_PLWN_mapping
Mapping between named entities types, SUMO catagories and plWordNet synset
WebStylo
Web based, open stylometry system based on Multilevel Text Analysis. Runs cluto and stylo (R system) clusterisation methods. Based on Natural Language Processing Workflow engine (included in the distribution)
Poliqarp2
Poliqarp2 is a linguistic search engine, capable of searching through large corpora annotated on multiple levels. It is not an upgraded version of Poliqarp, it is a completely new software developed from scratch
Słowosieć 3.0
plWordNet is a lexico-semantic network which reflects the lexical system of the Polish language. plWN currently contains 178 000 nouns, verbs, adjectives, and adverbs, 259 000 word senses, and over 600 000 relations and 240 000 inter-lingual relations between lexical units. It is now the largest wordnet in the world and is still growing.
Senses in plWordNet are interconnected by relations. In the resulting network, each word is defined implicitly in reference to other words. For example, samochód 'car' is a kind of pojazd drogowy 'road vehicle'; it is a whole consisting of silnik 'engine', spryskiwacz 'windscreen washer', podwozie 'chassis' and so on; its close counterpart is the colloquial fura 'wheels'.
Among plWordNet's numerous applications there is its use as a Polish-English and English-Polish dictionary -- the effect of mapping onto Princeton WordNet (the first and for many years the largest wordnet in the world). plWordNet is also an important resource in natural language processing and in artificial intelligence research. For example, it is used by Google Translate for the purpose of machine translation.
The University has made plWordNet available free of charge for all applications, including commercial ones, on a licence modelled on the Princeton WordNet licence. Users may browse plWordNet via mobile version and via WordNetLoom-Viewer (application enabling display of plWN entries), as well as download source files. Programmers may access plWordNet via Web service.
We provide (currently only in download version) 31 000 lexical units marked with their sentiment values: positive, negative, ambiguous or neutral
Polish Grapheme-to-phoneme tool and service
This archive contains the source code of the Polish grapheme-to-phoneme conversion tool and the webservice located at http://mowa.clarin-pl.eu/transcriber
ENIAMtoolkit
ENIAMtoolkit is a collection of libraries that:
- perform tokenization, lemmatization, part of speech tagging;
- detect MWE and abbreviations;
- split text into sentences
ChronoPress -- Chronologica Corpus
ChronoPress is a unique resource containing samples of Polish press texts from the period 1945-1954. The corpus was designed as a representative set of samples for Polish public discourse
Late 19th- and Early 20th-Century Polish Novels
Corpus of late 19th- and early 20th-century literary texts intended as benchmark collection for text categorization. It contains 100 Polish novels written by various authors. Each text is stored as separate .txt file