SADiLaR Language Resource Repository
Not a member yet
    536 research outputs found

    Lwazi II Xitsonga TTS Corpus

    No full text
    Orthographic and phonemically aligned transcriptions

    Asterisk Nuance 1.4.

    No full text
    Integration of commercial Nuance speech-recognition engine to the Asterisk open-source platform

    Calomo

    No full text
    Calomo is a hyphenator for Afrikaans, which can be implemented in any NLP system. It takes as input a string, and produces as output an analysed string, without any tags. For example, the string "hondehokdak" ('dog house roof') will be analysed as "hon-de-hok-dak", where the hyphen indicates syllable boundaries on an orthographic level. Calomo is a C5 classifier, trained on data consisting of circa 40,000 words. The resulting decision tree and cases can be converted to C code by means of a script written by M.M. van Zaanen. This C code can then be implemented in any other system

    Pretoria Tshivenda Corpus

    No full text
    Collection of texts for general linguistic research, in particular for lexicograph

    Afrikaans TnT-Tagger

    No full text
    The Afrikaans TnT-tagger is a part of speech tagger that can be used to add part of speech tags to Afrikaans texts.The tagger is an Afrikaans version of TnT (Brants, 2000). It was trained on 20,000 annotated Afrikaans text units. The tagger takes as input a tokenised Afrikaans text without any tags. It gives as output a text containing a text unit with its appropriate pos-tag per line. The tagset that the Afrikaans TnT-tagger uses was especially designed for Afrikaans and consists of 139 pos-tags

    CorpusCatcher

    No full text
    Corpus Catcher is a tool that is designed to crawl the web to retrieve data for inclusion in a corpus. It makes use of seed documents/wordlists to construct queries to retrieve document

    Sepedi Pronunciation Dictionary

    No full text
    A pronunciation dictionary for Sepedi ASR system was created from transcriptions of speech data collected

    Spelt

    No full text
    Spelt allows a linguist to classify surface forms of words. The word can be associated with a root form and with a word classification. The primary use of Spelt is to create classified wordlists for use in spell checkin

    SADE Municipality Hotline IVR Prompts

    No full text
    Audio and corresponding transcriptions for the SADE Municipality Hotline IVR prompts in English, Sesotho and isiZulu. The English SADE municipality hotline's prompts were translated (and recorded) in Sesotho and isiZulu. While the underlying recognition system is able to recognise Afrikaans, English, Sesotho and isiZulu speakers, pronouncing any of the South African municipality names, the interface language can also be customised to be Sesotho or isiZulu

    Afribooms Afrikaans Dependency Treebank

    No full text
    This is the annotated corpus developed for Afrikaans for the Afribooms project. The corpus includes annotations for lemma, part-of-speech (POS) and dependency relations. Lemma and POS information originates from the source corpus used, in this project only the dependency tags and relations were added

    8

    full texts

    536

    metadata records
    Updated in last 30 days.
    SADiLaR Language Resource Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇