SADiLaR Language Resource Repository
Not a member yet
536 research outputs found
Sort by
African Speech Technology Coloured-Afrikaans Speech Corpus
African Speech Technology speech and transcription data for the Coloured-Afrikaans database. The "speech" directory contains Afrikaans speech as spoken by South African Coloured speakers
NCHLT Sesotho Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
NCHLT Sesotho Annotated Text Corpora
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project
NCHLT isiNdebele Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project
NCHLT Afrikaans Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
Autshumato Xitsonga Monolingual Corpora
Xitsonga monolingual corpus as deliverable of the Autshumato project. The data is given as a UTF-8 text file; with each sentence on a newline.
NOTE: There is a newer version for an English-Xitsonga Monolingual Corpus. See https://hdl.handle.net/20.500.12185/57
NCHLT Siswati Annotated Text Corpora
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project
NCHLT Sepedi Text Corpora
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project
NCHLT isiNdebele Annotated Text Corpora
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project
NCHLT isiNdebele Text Corpora
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project