SADiLaR Language Resource Repository
Not a member yet
    536 research outputs found

    Woefzela

    No full text
    The primary purpose of the Woefzela software application is to record a list of prompts by a number of different speakers. The resultant output is then intended to be used as training and/or testing data (audio and associated transcriptions) for building Automatic Speech Recognition (ASR) Systems in the selected language

    NCHLT isiZulu Morphological Decomposer

    No full text
    Morphological decomposer developed during the NCHLT Text project

    NCHLT Tshivenda Text Corpora

    No full text
    Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project

    African Speech Technology Black-Afrikaans Speech Corpus

    No full text
    African Speech Technology speech and transcription data for the Black-Afrikaans database. The "speech" directory contains Afrikaans speech as spoken by South African speakers whose mother tongue language is one of the indigenous Southern African languages (e.g. isiXhosa, isiZulu, Sesotho)

    NCHLT Xitsonga Morphological Decomposer

    No full text
    Morphological decomposer developed during the NCHLT Text project

    African Speech Technology English-English Speech Corpus

    No full text
    African Speech Technology speech and transcription data for the English-English database. The "speech" directory contains English speech as spoken by South African English mother tongue speakers

    African Speech Technology Indian-English Speech Corpus

    No full text
    African Speech Technology speech and transcription data for the Indian-English database. The "speech" directory contains English speech as spoken by South African people of Indian descent

    NCHLT Tshivenda Annotated Text Corpora

    No full text
    Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project

    NCHLT Sesotho Lemmatiser

    No full text
    Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line. Output format: "Token tab Lemma"

    NCHLT Setswana Annotated Text Corpora

    No full text
    Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project

    8

    full texts

    536

    metadata records
    Updated in last 30 days.
    SADiLaR Language Resource Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇