CLARIN-PL
Not a member yet
504 research outputs found
Sort by
Description of nominal lexico-semantic relations in plWordNet 4.0 (Guidelines)
The pdf document contains guidelines of decription of Nouns in the Polish part of plWordNet
Świgra
Świgra is a parser of Polish generating constituency trees using a DCG style grammar stemming from Marek Świdziński’s grammar “Gramatyka formalna języka polskiego” (1992). The grammar was heavily rewritten for the purpose of annotating the Składnica treebank. The structure of trees was simplified with respect to Świdziński’s version, many new types of constructions were included (in particular various forms of coordinated structures), a statistical disambiguating component was added. Moreover, the Clarin version of Świgra uses the valency dictionary Walenty developed within Clarin
Krokodyl: A hybrid depencency parser of Polish
Krokodyl is an experimental hybrid deep depencency parser of Polish.
Krokodyl has been developed at the Institute of Computer Science, Polish Academy
of Sciences (IPI PAN) within the CLARIN-PL project. It was create to evaluate
a hybrid approach to parsing: combining syntactic, lexical and semantic
features for dependency parsing.
It uses a number of tools as components of the feature generation chain, namely
the Spejd Grammar, MALT parsing engine, MATE, the Polish Wordnet, the Skladnica
treebank
MWELexicon
Lexicon of 55k multi-word lexical units linked to plWordNet, together with description of their syntactic bahaviour obtained in constraint language (WCCL)
Paralela corpus and search engine
Paralela is as an open-ended, opportunistic parallel corpus of Polish-English and English-Polish translations. It currently contains 262 million words in 10,877,000 translation segments. The Paralela online search engine supports the SlopeQ query syntax for bilingual Polish-English corpus queries for the full dataset. Both the full texts and query results can be accessed and exported through the online application at http://paralela.clarin-pl.eu
WiKNN Text Classifier
WiKNN is an online text classifier service for Polish and English texts. It supports hierarchical labelled classification of user-submitted texts with Wikipedia categories. WiKNN is available through a web-based interface (http://pelcra.clarin-pl.eu/tools/classifier/) and as a REST service with interactive documentation available at http://clarin.pelcra.pl/apidocs/wiknn
KPWr annotation guidelines - spatial expressions (1.0)
Spatial expressions annotation guidelines describing the process of manual annotation of documents in Polish Corpus of Wrocław University of Technology (KPWr
SlopeQ for BNC Search Engine
The SlopeQ for BNC Search Engine provides access to the British National Corpus dataset.
In addition to linguistically motivated corpus queries, it supports a number of data exploration and
visualisation features. Most of the functionality of the search engine is available through a
REST web service