SADiLaR Language Resource Repository
Not a member yet
536 research outputs found
Sort by
NCHLT Siswati Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
NCHLT isiZulu Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
NCHLT Sesotho Text Corpora
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project
NCHLT Setswana Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
NCHLT Sepedi Annotated Text Corpora
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project
African Speech Technology isiXhosa Speech Corpus
African Speech Technology speech and transcription data for the isiXhosa database. The "speech" directory contains isiXhosa speech as spoken by isiXhosa mother tongue speakers
NCHLT Siswati Text Corpora
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project
African Speech Technology isiZulu Speech Corpus
African Speech Technology speech and transcription data for the isiZulu database. The "speech" directory contains isiZulu speech as spoken by isiZulu mother tongue speakers
NCHLT English Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
NCHLT Sesotho Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project