CLARIN-PL
Not a member yet
504 research outputs found
Sort by
KPWr EVENTS (Attributes and Relations)
Documents from Polish Corpus of Wrocław University of Technology manually annotated with Attributes for EVENT instances and relations between EVENTS instance
Polish Dependency Bank
Polish Dependency Bank (PDB) is the largest set of manually annotated dependency trees. PDB consists of more than 22K trees with 15.8 tokens per sentence on the average
TimeAssign
TimeAssign is a program which recognizes temporal expressions and assigns TimeML labels to words in Polish text using a Bi-LSTM based neural net and wordform embeddings
Polish Parliamentary Corpus
The Polish Parliamentary Corpus (PPC) is a large collection of linguistically analysed documents from the proceedings of Polish Parliament, Sejm and Senate. The corpus files are made available in TEI P5 format compatible with the annotation used by the National Corpus of Polish
Protestant Architecture Bohemia
Research Project : Protestant Building in Europe in Baroqu
KPWr annotation guidelines - events (attributes and relations)
KPWr annotation guidelines - events instances attributes and relations between events instance
Dependency parsing models for Polish
PDB-based parsing models are trained on the current version of Polish Depedency Bank with the publicly available parsing systems: MaltParser, MateParser, and UDPipe
Wordnet-based Evaluation of Large Distributional Models for Polish
The paper presents construction of large scale test datasets for word embeddings on the basis of a very large wordnet. They were next applied for evaluation of word embedding models and used to assess and compare the usefulness of different word embeddings extracted from a very large corpus of Polish. We analysed also and compared several publicly available models described in literature. In addition, several large word embeddings models built on the basis of a very large Polish corpus are presented
Periphraser
Periphraser is a tool for storing and presenting knowledge base of conventionalized periphrastic nominal expressions (i.e. phrases headed by a noun) together with their textually attested realizations. For instance, the database entry for the phrase ,,Robert Lewandowski'' in the demo for Polish will include the phrase ,,the Polish international'' while ,,pediatrics'' will be featured as ,,medical care for children''. It allows contacting with database using REST API as well as exporting it to XML or CSV format. For Polish language, it also provides some more complex mechanisms like: automatic semantic and syntactic normalization, errors autodetection (also based on NKJP frequency and amount of results returned by the web browser), and simple interface for commenting and marking possibly wrong entries or ones needing improvement
Świgra — a parser of Polish
Świgra is a parser of Polish generating constituency trees using a DCG style grammar stemming from Marek Świdziński’s grammar “Gramatyka formalna języka polskiego” (1992). The grammar was heavily rewritten for the purpose of annotating the Składnica treebank. The structure of trees was simplified with respect to Świdziński’s version, many new types of constructions were included (in particular various forms of coordinated structures), a statistical disambiguating component was added. Moreover, the Clarin version of Świgra uses the valency dictionary Walenty developed within Clarin