SADiLaR Language Resource Repository
Not a member yet
536 research outputs found
Sort by
NCHLT isiNdebele Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
NCHLT Part of Speech Taggers
Part of speech taggers developed during the NCHLT Text project.
Available for the following languages: Afrikaans, English, isiNdebele, isiXhosa, isiZulu, Sesotho sa Leboa (Sepedi), Setswana, Sesotho (Southern Sotho),
Siswati, Tshivenda, and Xitsonga
Autshumato Xitsonga Frequency Word List
A list of the most frequent Xitsonga words as deliverable of the Autshumato project
NCHLT Tshivenda Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project
Lara2
Tool for annotating texts with lemma, part of speech and morphological analysis informatio
Autshumato English-Xitsonga Parallel Corpora
Aligned English-Xitsonga parallel corpus. The data is given as two seperate UTF-8 text files; with each segment on a newline
NCHLT Setswana Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project
NCHLT isiXhosa Annotated Text Corpora
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project
NCHLT isiZulu Annotated Text Corpora
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project
African Speech Technology Black-English Speech Corpus
African Speech Technology speech and transcription data for the Black-English database. The "speech" directory contains English speech as spoken by South African speakers whose mother tongue language is one of the indigenous Southern African languages (e.g. isiXhosa, isiZulu, Sesotho)