SADiLaR Language Resource Repository
Not a member yet
536 research outputs found
Sort by
KALAS
KALAS is a rule-based compound analyser for Afrikaans, to be used for the detection of word boundaries within compounds. It takes as input a string, and produces as output an analysed string, without any tags. For example, the string "hondehokdak" ('dog house roof') will be analysed as "hond _ e + hok + dak", where the plus sign indicates the beginning of an independent constituent, and the underscore the beginning of a dependent constituent (i.e. a valence morpheme). The algorithm is based on a basic longest-string matching algorithm, with certain restrictions build into it. It has been implemented in both C and Perl
Format Normaliser 1.0.
Normalises input files to txt, utf8, replaces smart quotes with straight quotes, removes empty lines, etc
African Speech Technology isiZulu Text Corpus
Monolingual text corpus developed during the African Speech Technology project
NCHLT Afrikaans Text Corpora
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project
NCHLT Xitsonga Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
NCHLT Sepedi Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project
NCHLT Afrikaans Annotated Text Corpora
Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project
NCHLT isiXhosa Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
Autshumato English-Xitsonga Manually Translated Parallel Corpora
Aligned English-Xitsonga parallel corpus. The data is given as two seperate UTF-8 text files; with each segment on a newline
NCHLT Xitsonga Text Corpora
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project