504 research outputs found

    Introduction to the special issue: On wordnets and relations

    Get PDF
    The present paper is concerned with the issues of wordnets and relations

    Pytania i odpowiedzi z serwisu wikipedyjnego "Czy wiesz", wersja 1.1

    No full text
    Czy wiesz” (pol. “Did you know”) is a set of 4721 questions, each linked to a Wikipedia article that contains the answer. For 250 questions a detailed manual analysis has been performed. What results is attachment of manually-checked answer-bearing fragments to each of those selected questions. Some questions are assigned multiple fragments. The data set has been obtained from the Polish “Did you know” wikiproject. The dataset is made to facilitate evaluation and development of Polish QA systems

    The chicken-and-egg problem in wordnet design: synonymy, synsets and constitutive relations

    Get PDF
    Wordnets are built of synsets, not of words. A synset consists of words. Synonymy is a relation between words. Words go into a synset because they are synonyms. Later, a wordnet treats words as synonymous because they belong in the same synset. . . Such circularity, a well-known problem, poses a practical difficulty in wordnet construction, notably when it comes to maintaining consistency. We propose to make a wordnet a net of words or, to be more precise, lexical units. We discuss our assumptions and present their implementation in a steadily growing Polish wordnet. A small set of constitutive relations allows us to construct synsets automatically out of groups of lexical units with the same connectivity. Our analysis includes a thorough comparative overview of systems of relations in several influential wordnets. The additional synset-forming mechanisms include stylistic registers and verb aspect

    Serel (WS)

    No full text
    Serel is a Python framework for recognition relations between annotations in text

    HaskPL

    No full text
    HaskPL is a Polish phraseological database designed for language professionals including linguists, language teachers, lexicographers, language materials developers and translators. Query results can be visualised and exported as spreadsheets. A complementary tool is HaskProof (http://pelcra.clarin-pl.eu:9894/#/lang/pl) identifying potential collocations in any text inserted by the user

    Wikipedia articles extracted from Polish Corpus of Wrocław University of Technology

    No full text
    The resource is the part of the Polish Corpus of Wrocław University of Technology (fully available on the website http://nlp.pwr.wroc.pl/kpwr). The documents within this collection are the samples of the Polish Wikipedia articles manually annotated on the level of chunks and selected predicate-argument relations, named entities, relations between named entities, anaphora relations and word senses

    Chunker WS

    No full text
    Chunker-WS provides shallow parsing of Polish. The parser may be run against plain text (input format: text, then it runs WCRFT for tagging) or already tagged input (other input formats). Service output will contain tokenisation and tagging, but also boundaries of syntactic phrases as well as phrases' syntactic heads. The service is based on IOBBER, a chunker for Polish. The configuration used here operates on chunk definitions from the KPWr corpus

    IOBBER

    No full text
    IOBBER is a chunker for Polish. Its job is to recognise syntactic phrases (chunks) in Polish text. The name comes from IOB tags that are assigned to tokens to represent chunks (strictly speaking, we use IOB2 representation). Here is an example sentence annotated with NP and VP chunks

    NELexicon

    No full text
    NELexicon to gazetteer nazw własnych, który zawiera ponad 1.4 miliona unikalnych nazw własnych przypisanych do kategorii (par kategoria; nazwa), w tym ponad 1.37 miliona unikalnych napisów (z pominięciem powtórzeń nazw własnych przypisanych do kilku kategorii)

    Fextor

    No full text
    Fextor is a tool for extracting features from the collections of texts. It is characterized by high flexibility, while maintaining the performance and simplicity. Features are extracted from text snippets, defined according to the type of pointer (token, annotation or pair annotations). This allows the simultaneous generation of multiple features for a single document. Defining new types of features can be done by implementing in python or using a description in wccl language. Fextor supports two formats of corpora - Poliqarp and CCL. The extracted features are saved in CSV format, with the possibility of converting to a matrix format, for use in LexCSD package

    40

    full texts

    504

    metadata records
    Updated in last 30 days.
    CLARIN-PL
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇