SADiLaR Language Resource Repository
Not a member yet
    536 research outputs found

    African Speech Technology Coloured-Afrikaans Speech Corpus

    No full text
    African Speech Technology speech and transcription data for the Coloured-Afrikaans database. The "speech" directory contains Afrikaans speech as spoken by South African Coloured speakers

    NCHLT Sesotho Speech Corpus

    No full text
    Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers

    NCHLT Sesotho Annotated Text Corpora

    No full text
    Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project

    NCHLT isiNdebele Morphological Decomposer

    No full text
    Morphological decomposer developed during the NCHLT Text project

    NCHLT Afrikaans Speech Corpus

    No full text
    Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers

    Autshumato Xitsonga Monolingual Corpora

    No full text
    Xitsonga monolingual corpus as deliverable of the Autshumato project. The data is given as a UTF-8 text file; with each sentence on a newline. NOTE: There is a newer version for an English-Xitsonga Monolingual Corpus. See https://hdl.handle.net/20.500.12185/57

    NCHLT Siswati Annotated Text Corpora

    No full text
    Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project

    NCHLT Sepedi Text Corpora

    No full text
    Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project

    NCHLT isiNdebele Annotated Text Corpora

    No full text
    Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project

    NCHLT isiNdebele Text Corpora

    No full text
    Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project

    8

    full texts

    536

    metadata records
    Updated in last 30 days.
    SADiLaR Language Resource Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇