504 research outputs found

    Toposław

    No full text
    Toposław is an editor of multi-word unit inflection lexicons

    1000 Novels Corpus

    No full text
    Corpus of literary texts intended as benchmark collection for text categorization. It contains 1000 novels written in polish or translated to polish by various authors. Each text is stored as separate .txt file

    Cyfry

    No full text
    A small spoken digits corpus in polish. Contains 488 recordings of 25 speakers reading 20 digits (0-9) each. Amounts to around 76 minutes of recordings. Split into train (~72%), valid (~8%) and test (~20%) sets

    Walenty (2016-04-28)

    No full text
    Walenty is a valence dictionary of Polish developed at the Institute of Computer Science, Polish Academy of Sciences (IPI PAN). The original formalism of Walenty was established by Filip Skwarski, Elżbieta Hajnicz, Agnieszka Patejuk, Adam Przepiórkowski, Marcin Woliński, Marek Świdziński, and Magdalena Zawisławska. It has been further developed by Elżbieta Hajnicz, Agnieszka Patejuk, Adam Przepiórkowski, and Marcin Woliński. The semantic layer has been developed by Elżbieta Hajnicz and Anna Andrzejczuk. The original seed of Walenty was provided by the automatic conversion, manually reviewed by Filip Skwarski, of the verbal valence dictionary used by the Świgra2 parser (6396 schemata for 1462 lemmata), which was in turn based on SDPV, the Syntactic Dictionary of Polish Verbs by Marek Świdziński (4148 schemata for 1064 lemmata). Afterwards, Walenty has been developed independently by adding new entries, syntactic schemata, in particular phraseological ones, and semantic frames. Walenty has been edited and compiled using the Slowal tool (http://zil.ipipan.waw.pl/Slowal) created by Bartłomiej Nitoń and Tomasz Bartosiak

    POLFIE-OT: an LFG grammar of Polish with OT marks

    No full text
    POLFIE-OT is a version of POLFIE, an LFG grammar of Polish implemented in the XLE system (Xerox Linguistic Environment), enriched with OT (Optimality Theory) constraints for the purpose of automatic disambiguation of resulting parses – according to defined criteria, certain parses are considered optimal, while the remaining ones are considered unoptimal (it is possible, however, to view the unoptimal parses). POLFIE has been developed at the Institute of Computer Science, Polish Academy of Sciences (IPI PAN) within two projects: NEKST and CLARIN-PL. It provides a two-layer representation: constituent structure (c-structure, tree representation) and functional structure (f-structure, AVM representation). It is based on two previous implemented grammars of Polish: its c-structure is based on GFJP2, a DCG grammar used by the parser Świgra, while its f-structure is inspired by FOJP, an HPSG grammar of Polish. Lexical entries used by the grammar are created with the help of two state-of-the-art resources for Polish: Morfeusz2, a morphological analyser, and Walenty, a valence dictionary. POLFIE-OT is available via XLE-Web (a part of INESS; it does not require a local installation of XLE): • go to http://iness.mozart.ipipan.waw.pl/iness/xle-web or http://clarino.uib.no/iness/xle-web • choose "POLFIE-OT" grammar from the "Grammar" menu • write a sentence in the relevant field • click the "Parse sentence" button – only optimal solutions will be presented • however, if you want to see unoptimal solutions, check the "Show unoptimal" checkbox and click "Parse sentence" butto

    Toposław 2 (2016-05-31)

    No full text
    Toposław 2 is an editor of multi-world unit inflection lexicons

    Defender

    No full text
    Deepened lexical parser into nominal phrase

    KPWr annotation guidelines - named entities

    No full text
    Named entities annotation guidelines describing the process of manual annotation of documents in Polish Corpus of Wrocław University of Technology (KPWr

    NPSemRel

    No full text
    NPSemrel is a tool for recognizing semantic roles into nominal Phrases

    Słownik kolokacji rzeczownikowo-przymiotnikowych bez uzgodnienia

    No full text
    Słownik nieuzgodnionych kolokacji rzeczownikowo-przymiotnikowych z korpusu Słowosieci. W nagłówku mamy base 0 base 1 relacja czestość AgrAdjSubstP0 AgrAdjSubstH1P0 AgrSubstAdjP0 AgrSubstAdjH1P0 gdzie base 0 - to lemat pierwszego wyrazu, base 1 - lemat drugiego wyrazu, relacja - rodzaj operatora, który wykrył połączenie, częstość - frekwencja połączenia w korpusie, AgrAdjSubstP0 - szyk AN (przymiotnik - rzeczownik) bez uzgodnienia, nieprzedzielony przez 3 wyraz, AgrAdjSubstH1P0 - szyk A_N (z przerwą na 1 separujący połączenie wyraz), AgrSubstAdjP0 - szyk NA bez uzgodnienia, bez separującego wyrazu, AgrSubstAdjH1P0 - szyk NA z przerwą na 1 separujący wyraz, bez uzgodnienia. Przykład: base 0 base 1 relacja czestość AgrAdjSubstP0 AgrAdjSubstH1P0 AgrSubstAdjP0 AgrSubstAdjH1P0 ger:uzyskać adj:bezcłowy AgrSubstAdjH1P0 3 0 0 1 2 subst:misja adj:kotlarski AgrSubstAdjH1P0 1 0 0 0

    40

    full texts

    504

    metadata records
    Updated in last 30 days.
    CLARIN-PL
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇