504 research outputs found

    Blogs_2018

    No full text
    Teksty z blogów książkowyc

    Mowa Wrocławia lat 80-tych - corpus

    No full text
    The corpus comprises spoken data collected in the 1980s in Wrocław. The data were retrieved from tapes and digitalised

    Knowledge base of Polish conventionalized periphrastic nominal expressions

    No full text
    The resource includes free Periphraser export with a knowledge base of Polish conventionalized periphrastic nominal expressions (i.e. phrases headed by a noun) together with their textually attested realizations. For instance, the database entry for the phrase ,,Robert Lewandowski'' in the referred resource will include the phrase ,,the Polish international'' while ,,pediatrics'' will be featured as ,,medical care for children''. Export is available in two formats provided by Periphraser - XML and CSV. Associated files include a free version of data (available on the CC BY-SA 4.0 License). In case you are interested in full (authenticated) export of Polish periphrastic data please contact resource contact person

    KGR10 FastText Polish word embeddings

    No full text
    Distributional language model (both textual and binary) for Polish (word embeddings) trained on KGR10 corpus (over 4 billion of words) using Fasttext with the following variants (all possible combinations): - dimension: 100, 300 - method: skipgram, cbow - tool: FastText, Magnitude - source text: plain, plain.lower, plain.lemma, plain.lemma.lower The link below leads to the NextCloud directory with all variants of embeddings. If you use it, please cite the following article: @article{kocon2018embeddings, author = {Koco\'{n}, Jan and Gawor, Micha{\l}}, title = {Evaluating {KGR10} {P}olish word embeddings in the recognition of temporal expressions using {BiLSTM-CRF}}, journal = {Schedae Informaticae}, volume = {27}, year = {2018}, url = {http://www.ejournals.eu/Schedae-Informaticae/2018/Volume-27/art/13931/}, doi = {10.4467/20838476SI.18.008.10413}

    Context-sensitive Sentiment Propagation inWordNet

    Get PDF
    In this paper we present a comprehensive overview of recent methods of the sentiment propagation in a wordnet. Next, we propose a fully automated method called Classifier-based Polarity Propagation, which utilises a very rich set of features, where most of them are based on wordnet relation types, multi-level bag-ofsynsets and bag-of-polarities. We have evaluated our solution using manually annotated part of plWordNet 3.1 emo, which contains more than 83k manual sentiment annotations, covering more than 41k synsets. We have demonstrated that in comparison to existing rule-based methods using a specific narrow set of semantic relations our method has achieved statistically significant and better results starting with the same seed synsets

    Acoustic Data Building Toolset

    No full text
    This folder contains data and software tools (in python) that can be used in experiments with phoneme recognition in speech samples recorder in Polish. Acoustic data used here were extracted from CLARIN-PL speech corpus after rejecting speech samples, where recorded sequence of words does not correspond strictly to the word sequence declared as the sample orthographic transcription. In order to use the python programs and data published here, the appropriate folder structure should be created. Follow the steps below: 1) create the root folder and set the environment variable ASR_DATASET_ROOT that points to this folder (let's call it ROOT), 2) create the subfolders in the ROOT folder: train, devel, test, doc and src in the root folder, 3) download ar files: test.tar.gz, devel.tar.gz, train.tar.gz, doc.tar.gz, src.tar.gz and unpack then in the corresponding folders, 4)download aux.tar.gz and unpack it directly to ROOT folder. More information can be found in doc/README.pdf. If you find this dataset useful, please make reference in your related papers to the paper: " Acoustic Data Building Toolset for Easy Experimentation with Neural Network-based Speech Recognition in Polish and English" (https://ieeexplore.ieee.org/document/8431366/) Bibtex: @INPROCEEDINGS{8431366, author={J. Sas}, booktitle={2018 11th International Conference on Human System Interaction (HSI)}, title={Acoustic Data Building Toolset for Easy Experimentation with Neural Network-based Speech Recognition in Polish and English}, year={2018}, volume={}, number={}, pages={93-99}, doi={10.1109/HSI.2018.8431366}, ISSN={}, month={July},

    C1_essays

    No full text
    C1 essay

    Speech Recognition System for Polish: Polish Film Chronicles

    No full text
    This resource contains dockerized models and scripts of an automatic speech recognition system for Polish trained on recording of the Polish Film Chronicles. The system is based on the Kaldi toolkit. The scripts include methods for performing speech recognition, forced alignment and a lenient alignment of audio. The Github repository contains information on how to use the tool

    PELCRA EMI corpus

    No full text
    The corpus comprises open interviews with Polish people residing in Scotland

    Polimorf

    No full text
    PoliMorf is a morphological dictionary for Polish resulting from the standardization and merger of Morfeusz SGJP and Morfologik. The present version includes extended information on proper names

    40

    full texts

    504

    metadata records
    Updated in last 30 days.
    CLARIN-PL
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇