SADiLaR Language Resource Repository
Not a member yet
536 research outputs found
Sort by
Asterisk Nuance 1.4.
Integration of commercial Nuance speech-recognition engine to the Asterisk open-source platform
Calomo
Calomo is a hyphenator for Afrikaans, which can be implemented in any NLP system. It takes as input a string, and produces as output an analysed string, without any tags. For example, the string "hondehokdak" ('dog house roof') will be analysed as "hon-de-hok-dak", where the hyphen indicates syllable boundaries on an orthographic level. Calomo is a C5 classifier, trained on data consisting of circa 40,000 words. The resulting decision tree and cases can be converted to C code by means of a script written by M.M. van Zaanen. This C code can then be implemented in any other system
Pretoria Tshivenda Corpus
Collection of texts for general linguistic research, in particular for lexicograph
Afrikaans TnT-Tagger
The Afrikaans TnT-tagger is a part of speech tagger that can be used to add part of speech tags to Afrikaans texts.The tagger is an Afrikaans version of TnT (Brants, 2000). It was trained on 20,000 annotated Afrikaans text units. The tagger takes as input a tokenised Afrikaans text without any tags. It gives as output a text containing a text unit with its appropriate pos-tag per line. The tagset that the Afrikaans TnT-tagger uses was especially designed for Afrikaans and consists of 139 pos-tags
CorpusCatcher
Corpus Catcher is a tool that is designed to crawl the web to retrieve data for inclusion in a corpus. It makes use of seed documents/wordlists to construct queries to retrieve document
Sepedi Pronunciation Dictionary
A pronunciation dictionary for Sepedi ASR system was created from transcriptions of speech data collected
Spelt
Spelt allows a linguist to classify surface forms of words. The word can be associated with a root form and with a word classification. The primary use of Spelt is to create classified wordlists for use in spell checkin
SADE Municipality Hotline IVR Prompts
Audio and corresponding transcriptions for the SADE Municipality Hotline IVR prompts in English, Sesotho and isiZulu. The English SADE municipality hotline's prompts were translated (and recorded) in Sesotho and isiZulu. While the underlying recognition system is able to recognise Afrikaans, English, Sesotho and isiZulu speakers, pronouncing any of the South African municipality names, the interface language can also be customised to be Sesotho or isiZulu
Afribooms Afrikaans Dependency Treebank
This is the annotated corpus developed for Afrikaans for the Afribooms project. The corpus includes annotations for lemma, part-of-speech (POS) and dependency relations. Lemma and POS information originates from the source corpus used, in this project only the dependency tags and relations were added