504 research outputs found

    MACA

    No full text
    Utilities are simple programs referencing the corresponding API functions, hence similar functionality may be easily obtained by using the libraries

    Disaster

    No full text
    Disaster (DISAmbiguator and STatistical chunkER) is a Python module for chunking and morphosyntactic disambiguation

    ChunkRel WS

    No full text
    ChunkRel-WS is a prototype service for recognition of three syntactic relations between chunks. The service may be run against plain text (input format: text), then the necessary processing steps will be run automatically (tagger and chunker). You can also process already tagged and chunked input (input format: ccl). The output will be enriched with inter-chunk relations. The service is based on a prototype implementation, hence it works slowly. The configuration used here operates on shallow syntactic annotation scheme from the KPWr corpus

    Toki

    No full text
    Toki is a configurable tokeniser, i.e. a module for segmentation of running text into tokens (word-like units) and sentences

    Inforex

    No full text
    Inforex is a web-based system designed for managing and annotating text corpora on the semantic level including annotation of Named Entities (NE), anaphora, Word Sense Disambiguation (WSD) and relations between named entities. The system also supports manual text clean-up and automatic text pre-processing including text segmentation, morphosyntactic analysis and word selection for WSD annotation

    Corpus-SUCK

    No full text
    Proces przetwarzania umożliwia pobranie zawartości serwisów internetowych. Wejściem dla procesu jest lista adresów URL, na wyjściu uzyskuje się zbiór plików zawierających najbardziej istotną zawartość (tylko tekst, np. treść artykułu, bez dodatkowych informacji na stronie) najbardziej istotnych podstron (tylko podstrony zawierające tekst w odpowiedniej ilości, bez zawartości typu obrazy, filmy, itp.). Pliki pogrupowane są według źródła - dla każdego linku z wejściowej listy tworzony jest osobny katalog, w którym znajdują się pliki. Każdy plik jest osobną podstroną. Najbardziej istotna zawartość jest poddana filtrowaniu (domyślnie dokument powinien mieć min. 300 znaków istotnych (należących do tokenów) oraz min. 20% słów musi być znanych (znajdować się w słowniku Morfeusz). Dokumenty po filtrowaniu są tagowane przy pomocy narzędzia WCRF

    WordNet

    No full text
    plWordNet is a lexico-semantic network which reflects the lexical system of the Polish language. There are at present ca. 144,000 nouns, verbs and adjectives in plWordNet, ca. 203,000 word senses and ca. 500,000 relations. It is already the second-largest wordnet in the world, and it keeps growing

    Tagger WS

    No full text
    Tagger-WS is a web service that reads Polish text and outputs sentences divided into tokens where each token is labelled with a morphosyntactic tag and a lemma. The tagger uses NKJP tagset and Morfeusz SGJP analyser. The service is based on WCRFT

    CEN

    No full text
    oai:clarin-pl.eu:11321/6Corpus of Economic News (CEN) contains 797 documents from Polish Wikipedia annotated with 65 categories of proper names in ccl format. http://nlp.pwr.edu.pl/inforex/?corpus=5&page=brows

    Enriched corpus of [Polish] frequency dictionary

    No full text
    Wzbogacony korpus slownika frekwencyjnego, cf. http://clip.ipipan.waw.pl/PL196x?action=AttachFile&do=view&target=wksf.pd

    40

    full texts

    504

    metadata records
    Updated in last 30 days.
    CLARIN-PL
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇