SADiLaR Language Resource Repository
Not a member yet
536 research outputs found
Sort by
NCHLT Xitsonga Lemmatiser
Lemmatiser developed during the NCHLT Text project. \n\n
Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line.
Output format: "Token tab Lemma"
NCHLT Tshivenda Lemmatiser
Lemmatiser developed during the NCHLT Text project.
\n\n
Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line.
Output format: "Token tab Lemma"
NCHLT Tshivenda Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
NCHLT Afrikaans Lemmatiser
Lemmatiser developed during the NCHLT Text project. \n\n
Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line. Output format: "Token tab Lemma"
Automated multilingual telephone access to financial services
A working prototype automated telephone based enquiry and payment system functioning in three African languages, i.e. Zulu, Xhosa and Southern Sotho, in a financially related application, making provision for code-mixing between these languages and English
Lwazi Tshivenda ASR corpus
Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems
isiZulu Spelling Checker 1.1
Spelling checkers and hyphenators for South African languages compatible with Microsoft® Office 2000, XP, 2003, 2007, 2010 or Microsoft® Office 2013. Extended research by CTexT ensures that spell checking and hyphenation occur effortlessly within the parameters of the standard written variant of each language. It also includes a comprehensive word list and hyphenation
Lwazi isiXhosa Pronunciation Dictionary
General phonemic pronunciations for frequently occurring words in SA languages. Dictionaries were developed to be practically usable for speech technology systems, rather than phonetically accurate. Audio samples of all phonemes included. A letter-to-sound rule set for predicting the pronunciations of generic words included. (Separate entry describes rule sets.
SpeechGREP
Text-initiated search for words and phrases in recorded speech data bases. Searches in the phonetic database Speech mining (more info ??) GUI-base
Lwazi isiNdebele Pronunciation Dictionary
General phonemic pronunciations for frequently occurring words in SA languages. Dictionaries were developed to be practically usable for speech technology systems, rather than phonetically accurate. Audio samples of all phonemes included. A letter-to-sound rule set for predicting the pronunciations of generic words included. (Separate entry describes rule sets.