1,720,964 research outputs found
Resource Repositories and linking resources: An exploratory study
In this article the existence, use and importance of repositories are explored. An introduction into language resources (LRs) is given as well as a discussion of two platforms for the distribution of language resources, namely, the repository of the South African Centre for Digital Language Resources (SADiLaR) and Lanfrica, a site that links resources. In this article, types of repositories, such as institutional and language resource repositories, will be distinguished and compared. Language preservation is proposed as an important aspect which can be strengthened by the presence and use of repositories. The view expressed in this article is that the availability of language resources and repositories are pivotal for the development, preservation and advancement of languages.
Having a host site that links available resources and a repository where resources could be uploaded is a positive attribute of the mentioned online platforms, however as it will be discussed, the fact that information is available online is not a guarantee that the resources are or will be used by researchers or other interested persons, especially if they are not aware of their existence.
The article is concluded with suggestions for future work, for example measuring the influence of inaccurate metadata of language resources on linguistic research
A Critical Evaluation of Three Sesotho Dictionaries: 'n Kritiese evaluering van drie Sesotho woordeboeke
This article gives a perspective on Sesotho lexicography and a critical analysis of the macrostructures and microstructures of three selected Sesotho dictionaries. The monolingual paper dictionary Sethantšo sa Sesotho, the bilingual paper dictionary Southern Sotho–English Dictionary and the Sesotho online Bukantswe v.3 are evaluated. Their virtues and shortcomings as reference works will be viewed against dictionaries of high lexicographic achievement in order to establish to what extent they fulfil the most basic requirements of macrostructures and microstructures. The inconsistencies addressed in this article reflect the need for Sesotho lexicographers to use corpora in dictionary compilation in order to enhance the quality of entries on both microstructural and macrostructural levels. It will be argued that much more research and description of lexicographic issues is required to bring Sesotho lexicography on a par with its sister languages, Sepedi and Setswana and with good dictionaries for major languages of the world. After decades in existence, currently available Sesotho dictionaries are in dire need for revision and new dictionaries aimed at specific target users should be compiled.
Hierdieartikel gee 'n perspektief op Sesotho-leksikografie en 'n kritiese ontleding van die makrostrukture en mikrostrukture van drie geselekteerde Sesotho woordeboeke. Die eentalige papierwoordeboek Sethantšo sa Sesotho, die tweetalige papierwoordeboek Southern Sotho–English Dictionary en die Sesotho online Bukantswe v.3 word geëvalueer. Hulle deugde en tekortkominge as naslaanwerke sal beskou word teenoor woordeboeke van hoë leksikografiese gehalte om vas te stel in watter mate hulle aan die mees basiese vereistes van makrostrukture en mikrostrukture voldoen. Die teenstrydighede wat in hierdie artikel aangespreek word, weerspieël die noodsaaklikheid dat Sesotho leksikograwe korpora in woordeboeksamestelling gebruik om die gehalte van inskrywings op mikrostrukturele sowel as makrostrukturele vlak te verhoog. Daar sal geargumenteer word dat baie meer navorsing en beskrywing van leksikografiese kwessies nodig is om die leksikografie van Sesotho op gelyke voet te bring met die sustertale Sepedi en Setswana asook met goeie woordeboeke van wêreldtale. Na dekades van gebruik, moet die Sesotho woordeboeke wat tans beskikbaar is dringend hersien word en nuwe woordeboeke saamgestel word wat op spesifieke teikengebruikers gerig is
An overview of Sesotho BLARK content
This article overviews digital language resources available for Sesotho, an official language of South Africa. The South African Center for Digital Language Resources (SADiLaR) repository is used as a reference as it is the official host of various language resources for South African languages. A total of 18 written resources are identified from the repository, and a further 16 spoken resources are identified. Finally, a total of 45 applications and modules were identified. Findings indicate that the majority of applications and modules available for Sesotho are in fact general resources aimed at all eleven official South African languages. Furthermore, the available resources indicate an inclination to the development of entry level, basic language resources and an absence of middle and higher resources with functionalities such as semantic analyses for written resources and prosody prediction for spoken resources. The study is hindered by the dearth of resource specific evaluations and related research and exacerbated by the absence of some of the resources on the repository
Early Child Language Resources and Corpora Developed in Nine African Languages by the SADiLaR Child Language Development Node
Prior to the initiation of the project reported on in this paper, there were no instruments available with which to measure the language skills of young speakers of nine official African languages of South Africa. This limited the kind of research that could be conducted, and the rate at which knowledge creation on child language development could progress. Not only does this result in a dearth of knowledge needed to inform child language interventions but it also hinders the development of child language theories that would have good predictive power across languages. This paper reports on (i) the development of a questionnaire that caregivers complete about their infant’s communicative gestures and vocabulary or about their toddler’s vocabulary and grammar skills, in isiNdebele, isiXhosa, isiZulu, Sesotho, Sesotho sa Leboa, Setswana, Siswati, Tshivenda, and Xitsonga; and (ii) the 24 child language corpora thus far developed with these instruments. The potential research avenues opened by the 18 instruments and 24 corpora are discussed
Corpus-based Lexicography for Sesotho
Dissertation (MA)--University of Pretoria, 2018.For centuries, dictionaries were compiled based upon the knowledge of the lexicographer and information retrieved from manually consulted sources, mainly through a process of reading and marking. This approach meant that much of the information used in the dictionary relied upon the knowledge of the lexicographer. It is vital to rely on the lexicographer’s knowledge of the language but this has its shortcomings, since there is no single individual who knows all the words or terms, their meanings and usage, the words they combine with, and so on, in a specific language. The utilization of this method left room for errors and omissions because the lexicographer could easily overlook some words due to factors like time, fatigue, limited knowledge of the lexicographer, etc. Important words, for example words likely to be looked for by the target users of the dictionary, could accidentally be omitted. In the 1980s, the corpus era was born and the lexicography field changed forever. Collins COBUILD in Birmingham spearheaded this era with the publication of the first corpus-based dictionary, the Collins COBUILD Dictionary in 1987. Since the corpus era began, lexicographers no longer rely solely on their knowledge of the language, intuition, or the limited information gathered from available written sources, which are very limited for African languages. The corpus allows the lexicographer to have access to huge volumes of authentic data from written texts and transcribed oral data. This research will therefore critically discuss dictionary compilation for Sesotho and spearhead the use of corpora in the compilation of Sesotho dictionaries, so that lexicographers do not compile dictionaries as if they are compiling the first dictionary for the language. In addition, they should take into account tasks like lexicographic planning, amongst other factors required to compile a good user-friendly dictionary.
Key words
Corpora, collocations, concordances, lexicography, lexicographical planning, microstructure, macrostructure, lemmatisation.African LanguagesMAUnrestricte
The annotators agree to not agree on the fine-grained annotation of hate-speech against women in Algerian dialect comments
A significant number of research studies have been presented for detecting hate speech in social media during the last few years. However, the majority of these studies are in English. Only a few studies focus on Arabic and its dialects (especially the Algerian dialect) with a smaller number of them targeting sexism detection (or hate speech against women). Even the works that have been proposed on Arabic sexism detection consider two classes only (hateful and non-hateful), and three classes(adding the neutral class) in the best scenario. This paper aims to propose the first fine-grained corpus focusing on 13 classes. However, given the challenges related to hate speech and fine-grained annotation, the Kappa metric is relatively low among the annotators (i.e. 35%). This work in progress proposes three main contributions: 1) Annotation of different categories related to hate speech such as insults, vulgar words or hate in general. 2) Annotation of 10,000 comments, in Arabic and Algerian dialects, automatically extracted from Youtube. 3) Highlighting the challenges related to manual annotation such as subjectivity, risk of bias, lack of annotation guidelines, etc.</p
A corpus-based list of frequently used words in Sesotho
This article describes the development of a list of frequently used words in written Sesotho. The list has been created with the aim of incorporating it into frequency-based text readability metrics. The list was derived using a corpus-based approach. By leveraging three existing Sesotho corpora, frequency lists could be derived, which were subsequently merged and qualitatively analysed and fine-tuned by an experienced speaker of Sesotho. The main challenges in compiling the list included reconciling the spelling variations, the treatment of abbreviations, and the presence of unexpected words in the preliminary lists. The final list comprises 3037 entries and is made publicly available to the research community
An exploration of the computational identification of English loan words in Sesotho
South Africa, with its twelve official languages, is an inherently multilingual country. As such, speakers of many of the languages have been in direct contact. This has led to a cross-over of words and phrases between languages. In this article, we provide a methodology to identify words that are (potentially) borrowed from another language. We test our approach by trying to identify words that moved from English into Sesotho (or potentially the other way around). To do this, we start with a bilingual Sesotho-English dictionary (Bukantswe). We then develop a lexicographic comparison method that takes a pair of lexical items (English and Sesotho) and computes a range of distance metrics. These distance metrics are applied to the raw words (i.e., comparing orthography), but using the Soundex algorithm, an approximate phonological comparison can be made as well. Unfortunately, Bukantswe does not contain complete annotation of loan words, so a quantitative evaluation is not currently possible. We provide a qualitative analysis of the results, which shows that many loan words can be found, but in some cases lexical items that have a high similarity are not loan words. We discuss different situations related to the influence of orthography, phonology, syllable structure, and morphology. The approach itself is language independent, so it can also be applied to other language pairs, e.g., Afrikaans and Sesotho, or more related languages, suchas isiXhosa and isiZulu
The annotators agree to not agree on the fine-grained annotation of hate-speech against women in Algerian dialect comments
A significant number of research studies have been presented for detecting hate speech in social media during the last few years. However, the majority of these studies are in English. Only a few studies focus on Arabic and its dialects (especially the Algerian dialect) with a smaller number of them targeting sexism detection (or hate speech against women). Even the works that have been proposed on Arabic sexism detection consider two classes only (hateful and non-hateful), and three classes(adding the neutral class) in the best scenario. This paper aims to propose the first fine-grained corpus focusing on 13 classes. However, given the challenges related to hate speech and fine-grained annotation, the Kappa metric is relatively low among the annotators (i.e. 35%). This work in progress proposes three main contributions: 1) Annotation of different categories related to hate speech such as insults, vulgar words or hate in general. 2) Annotation of 10,000 comments, in Arabic and Algerian dialects, automatically extracted from Youtube. 3) Highlighting the challenges related to manual annotation such as subjectivity, risk of bias, lack of annotation guidelines, etc.</p
IsiXhosa Intellectual Traditions Digital Archive: Digitizing isiXhosa texts from 1870-1914
This article offers an overview of the IsiXhosa Intellectual Traditions Digital Archive, which hosts digitized texts and images of early isiXhosa newspapers and books from 1870-1914. The archive offers new opportunities for a range of research across multiple fields, and responds to debates around the importance of African intellectual traditions and their indigenous language sources in generating African social sciences which is contextually relevant. We outline the content and context of these materials and offer qualitative and quantitative details with the aim of providing an overview for interested scholars and a reference for those using the archive
- …
