SADiLaR Language Resource Repository
Not a member yet
    536 research outputs found

    NCHLT isiZulu Text Corpora

    No full text
    Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project

    NCHLT isiXhosa Lemmatiser

    No full text
    Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line. Output format: "Token tab Lemma"

    NCHLT isiZulu Lemmatiser

    No full text
    Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line. Output format: "Token tab Lemma"

    NCHLT Sepedi Speech Corpus

    No full text
    Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers

    NCHLT Siswati Morphological Decomposer

    No full text
    Morphological decomposer developed during the NCHLT Text project

    NCHLT Afrikaans Morphological Decomposer

    No full text
    Morphological decomposer developed during the NCHLT Text project

    NCHLT Setswana Lemmatiser

    No full text
    Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line. Output format: "Token tab Lemma"

    NCHLT isiXhosa Morphological Decomposer

    No full text
    Morphological decomposer developed during the NCHLT Text project

    NCHLT isiXhosa Text Corpora

    No full text
    Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project

    NCHLT Siswati Lemmatiser

    No full text
    Lemmatiser developed during the NCHLT Text project. \n\n Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line. Output format: "Token tab Lemma"

    8

    full texts

    536

    metadata records
    Updated in last 30 days.
    SADiLaR Language Resource Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇