SADiLaR Language Resource Repository
Not a member yet
536 research outputs found
Sort by
NCHLT isiZulu Text Corpora
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project
NCHLT isiXhosa Lemmatiser
Lemmatiser developed during the NCHLT Text project.
\n\n
Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line.
Output format: "Token tab Lemma"
NCHLT isiZulu Lemmatiser
Lemmatiser developed during the NCHLT Text project.
\n\n
Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line.
Output format: "Token tab Lemma"
NCHLT Sepedi Speech Corpus
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test suite of 8 speakers
NCHLT Siswati Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project
NCHLT Afrikaans Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project
NCHLT Setswana Lemmatiser
Lemmatiser developed during the NCHLT Text project.
\n\n
Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line.
Output format: "Token tab Lemma"
NCHLT isiXhosa Morphological Decomposer
Morphological decomposer developed during the NCHLT Text project
NCHLT isiXhosa Text Corpora
Collection of source text documents, genre classified text documents, raw corpus, clean corpus, lexicon, frequency list and named-entity lists developed during the NCHLT Text project
NCHLT Siswati Lemmatiser
Lemmatiser developed during the NCHLT Text project.
\n\n
Available in the Readme.txt - Input format: Text data (encoding: UTF8 without BOM), one lowercase token per line.
Output format: "Token tab Lemma"