SADiLaR Language Resource Repository
Not a member yet
536 research outputs found
Sort by
Woefzela
The primary purpose of the Woefzela software application is to record a list of prompts by a number of different speakers. The resultant output is then intended to be used as training and/or testing data (audio and associated transcriptions) for building Automatic Speech Recognition (ASR) Systems in the selected language
NCHLT isiZulu Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project
NCHLT Tshivenda Text Corpora
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project
African Speech Technology Black-Afrikaans Speech Corpus
African Speech Technology speech and transcription data for the Black-Afrikaans database. The "speech" directory contains Afrikaans speech as spoken by South African speakers whose mother tongue language is one of the indigenous Southern African languages (e.g. isiXhosa, isiZulu, Sesotho)
NCHLT Xitsonga Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project
African Speech Technology English-English Speech Corpus
African Speech Technology speech and transcription data for the English-English database. The "speech" directory contains English speech as spoken by South African English mother tongue speakers
African Speech Technology Indian-English Speech Corpus
African Speech Technology speech and transcription data for the Indian-English database. The "speech" directory contains English speech as spoken by South African people of Indian descent
NCHLT Tshivenda Annotated Text Corpora
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project
NCHLT Sesotho Lemmatiser
Lemmatiser developed during the NCHLT Text project.
\n\n
Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line.
Output format: "Token tab Lemma"
NCHLT Setswana Annotated Text Corpora
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project