SADiLaR Language Resource Repository
Not a member yet
    536 research outputs found

    KALAS

    No full text
    KALAS is a rule-based compound analyser for Afrikaans, to be used for the detection of word boundaries within compounds. It takes as input a string, and produces as output an analysed string, without any tags. For example, the string "hondehokdak" ('dog house roof') will be analysed as "hond _ e + hok + dak", where the plus sign indicates the beginning of an independent constituent, and the underscore the beginning of a dependent constituent (i.e. a valence morpheme). The algorithm is based on a basic longest-string matching algorithm, with certain restrictions build into it. It has been implemented in both C and Perl

    Format Normaliser 1.0.

    No full text
    Normalises input files to txt, utf8, replaces smart quotes with straight quotes, removes empty lines, etc

    African Speech Technology isiZulu Text Corpus

    No full text
    Monolingual text corpus developed during the African Speech Technology project

    NCHLT Afrikaans Text Corpora

    No full text
    Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project

    NCHLT Xitsonga Speech Corpus

    No full text
    Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers

    NCHLT Sepedi Morphological Decomposer

    No full text
    Morphological decomposer developed during the NCHLT Text project

    NCHLT Afrikaans Annotated Text Corpora

    No full text
    Lemmatised, part of speech tagged and morphologically analysed corpora developed during the NCHLT Text project

    NCHLT isiXhosa Speech Corpus

    No full text
    Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers

    Autshumato English-Xitsonga Manually Translated Parallel Corpora

    No full text
    Aligned English-Xitsonga parallel corpus. The data is given as two seperate UTF-8 text files; with each segment on a newline

    NCHLT Xitsonga Text Corpora

    No full text
    Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project

    8

    full texts

    536

    metadata records
    Updated in last 30 days.
    SADiLaR Language Resource Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇