1,720,966 research outputs found
Dakhil wordlist for Persian vocabulary
The Dakhil Wordlist consists of roughly 260K words from Persian (though there may be some duplicates) ending in "ان" "ات" "ون" "ین" but it is not part of their morphological boundary. the wordlist has been uploaded in both the affixes in one txt file and separated by their ending. for more information read the paper: Rahimi, Adel. "A new hybrid stemming algorithm for Persian." arXiv preprint arXiv:1507.03077 (2015)
Dakhil wordlist for Persian vocabulary
The Dakhil Wordlist consists of roughly 260K words from Persian (though there may be some duplicates) ending in "ان" "ات" "ون" "ین" but it is not part of their morphological boundary. the wordlist has been uploaded in both the affixes in one txt file and separated by their ending. for more information read the paper: Rahimi, Adel. "A new hybrid stemming algorithm for Persian." arXiv preprint arXiv:1507.03077 (2015)
Corpus of Computational Linguistics' Academic Literature
The corpus of Computational linguistics is an 8 million corpus of Journal publications, books, and theses. these include interdisciplinary topics such as Speech Recognition, Experimental Phonology, Language Models, Machine Learning, Semantics, Syntactic Theory, and Information Retrieval
Replication Data for: Rahimi, A., Critical discourse analysis of Adolf hitler's speeches
this is the corpus used in the research "Critical discourse analysis of Adolf hitler's speeches
Corpus of World National Anthems
This is the first corpus of world National Anthems consisting of ~264 national anthems from ~194 countries across the globe.
The password for the zip file is: kurmanj
Corpus of World National Anthems
This is the first corpus of world National Anthems consisting of ~264 national anthems from ~194 countries across the globe.
The password for the zip file is: kurmanj
Kurmanji Wikipedia Corpus
Corpus of Wikipedia in Kurmanji language. txt format ~1 million word
Corpus of Computational Linguistics' Academic Literature
The corpus of Computational linguistics is an 8 million corpus of Journal publications, books, and theses. these include interdisciplinary topics such as Speech Recognition, Experimental Phonology, Language Models, Machine Learning, Semantics, Syntactic Theory, and Information Retrieval
Replication Data for: Rahimi, A., Critical discourse analysis of Adolf hitler's speeches
this is the corpus used in the research "Critical discourse analysis of Adolf hitler's speeches
Kurmanji Wikipedia Corpus
Corpus of Wikipedia in Kurmanji language. txt format ~1 million word
- …
