CLARIN-PL
Not a member yet
504 research outputs found
Sort by
MACA
Utilities are simple programs referencing the corresponding API functions, hence similar functionality may be easily obtained by using the libraries
Disaster
Disaster (DISAmbiguator and STatistical chunkER) is a Python module for chunking and morphosyntactic disambiguation
ChunkRel WS
ChunkRel-WS is a prototype service for recognition of three syntactic relations between chunks. The service may be run against plain text (input format: text), then the necessary processing steps will be run automatically (tagger and chunker). You can also process already tagged and chunked input (input format: ccl). The output will be enriched with inter-chunk relations. The service is based on a prototype implementation, hence it works slowly. The configuration used here operates on shallow syntactic annotation scheme from the KPWr corpus
Toki
Toki is a configurable tokeniser, i.e. a module for segmentation of running text into tokens (word-like units) and sentences
Inforex
Inforex is a web-based system designed for managing and annotating text corpora on the semantic level including annotation of Named Entities (NE), anaphora, Word Sense Disambiguation (WSD) and relations between named entities. The system also supports manual text clean-up and automatic text pre-processing including text segmentation, morphosyntactic analysis and word selection for WSD annotation
Corpus-SUCK
Proces przetwarzania umożliwia pobranie zawartości serwisów internetowych. Wejściem dla procesu jest lista adresów URL, na wyjściu uzyskuje się zbiór plików zawierających najbardziej istotną zawartość (tylko tekst, np. treść artykułu, bez dodatkowych informacji na stronie) najbardziej istotnych podstron (tylko podstrony zawierające tekst w odpowiedniej ilości, bez zawartości typu obrazy, filmy, itp.). Pliki pogrupowane są według źródła - dla każdego linku z wejściowej listy tworzony jest osobny katalog, w którym znajdują się pliki. Każdy plik jest osobną podstroną. Najbardziej istotna zawartość jest poddana filtrowaniu (domyślnie dokument powinien mieć min. 300 znaków istotnych (należących do tokenów) oraz min. 20% słów musi być znanych (znajdować się w słowniku Morfeusz). Dokumenty po filtrowaniu są tagowane przy pomocy narzędzia WCRF
WordNet
plWordNet is a lexico-semantic network which reflects the lexical system of the Polish language. There are at present ca. 144,000 nouns, verbs and adjectives in plWordNet, ca. 203,000 word senses and ca. 500,000 relations. It is already the second-largest wordnet in the world, and it keeps growing
Tagger WS
Tagger-WS is a web service that reads Polish text and outputs sentences divided into tokens where each token is labelled with a morphosyntactic tag and a lemma. The tagger uses NKJP tagset and Morfeusz SGJP analyser. The service is based on WCRFT
CEN
oai:clarin-pl.eu:11321/6Corpus of Economic News (CEN) contains 797 documents from Polish Wikipedia annotated with 65 categories of proper names in ccl format.
http://nlp.pwr.edu.pl/inforex/?corpus=5&page=brows
Enriched corpus of [Polish] frequency dictionary
Wzbogacony korpus slownika frekwencyjnego, cf. http://clip.ipipan.waw.pl/PL196x?action=AttachFile&do=view&target=wksf.pd