1,720,961 research outputs found
Detecting news bias in the emerging market using transformed model
Thesis (M.Sc. (e-Science)) -- University of Limpopo, 2023The news readers from western countries are exposed to news reports from the global news media outlets that focus fundamentally on the conflicts happening in third-world countries such as those in South Africa and Nigeria. Many financial institutions from western countries provide financial opportunities to emerging markets in Africa. For this purpose, an important source of information used by those institutions is the global media outlets reporting the financial news. Therefore, there is a need to assess the credibility of global news outlets in covering financial news for African countries. This study focuses on detecting financial news bias in the emerging markets countries such as South Africa and Nigeria using a transformer model called BERT. The transformer model for sentiment classification was developed using the pre-trained model to achieve this goal. The financial news titles from the local and global news media outlets about South Africa and Nigeria were downloaded from the online news database called Global Database of Events, Language, and Tone project, and they were injected into the BERT model to compute the sentiment scores and labels for each news title. The pretrained BERT model achieved a good accuracy of 89.76% after being fine-tuned with a sample of 5000 movie review dataset. We found that both the local and global news media outlets report more positive financial news than negative news for both countries based on the results from the sentiment scores and labels. To assess news bias between the local and global news media outlets for South Africa and Nigeria, the average monthly sentiment scores were compared with the average monthly exchange rates for each country to determine the correlation between them and test if that correlation is significant or not. It was found that only the Nigerian global news coverage correlates with the exchange rates. Therefore, it was concluded that there is no evidence to suggest that the global news media outlets are biased in reporting the financial news in emerging markets countries since it was seen that the sentiment scores from the local news outlets for both countries do not correlate with their respective exchanges rates. It is evident that the transformer model can be used to accurately compute the sentiment in the financial news articles for the purpose of detecting news bias
Prediction of South African crime rate using supervised machine learning techniques
Thesis (M.Sc. (E-Science)) -- University of Limpopo, 2023Kaggle crime statistics for South Africa were used to create machine learning categorization models. Although the techniques used in the experiments that came before this one differed, the dataset that was used was. The accuracy of other previous studies conducted on different datasets from this one and compared during the experiment stage were utilized to identify the three classification algorithms employed in this study. The study chose to use the random forest, K-nearest neighbor, and Naive Bayes classifier models. The Python-based algorithms were trained on a pre-processed crime dataset. Data preparation and processing, missing value analysis, exploratory analysis, and finally model construction and evaluation made up the analytical process. The best model should be chosen in accordance with the results. In both approaches, RF is outperforming the other models. According to the study's evaluation of both metrics and logloss, RF appears to be doing better.The National e-Science Postgraduate Teaching and Training Platform (NEPTTP
Development of a text-independent automatic speaker recognition system
Thesis (M. Sc. (Computer Science)) -- University of Limpopo, 2021The task of automatic speaker recognition, wherein a system verifies or identifies
speakers from a recording of their voices, has been researched for several decades.
However, research in this area has been carried out largely on freely accessible
speaker datasets built on languages that are well-resourced like English. This study
undertakes automatic speaker recognition research focused on a low-resourced
language, Sepedi. As one of the 11 official languages in South Africa, Sepedi is
spoken by at least 2.8 million people. Pre-recorded voices were acquired from a
speech and language national repository, namely, the National Centre for Human
Language Technology (NCHLT), were we selected the Sepedi NCHLT Speech
Corpus. The open-source pyAudioAnalysis python library was used to extract three
types of acoustic features of speech namely, time, frequency and cepstral domain
features, from the acquired speech data. The effects and compatibility of these
acoustic features was investigated. It was observed that combining the three acoustic
features of speech had a more significant effect than using individual features as far
as speaker recognition accuracy is concerned. The study also investigated the
performance of machine learning algorithms on low-resourced languages such as
Sepedi. Five machine learning (ML) algorithms implemented on Scikit-learn namely,
K-nearest neighbours (KNN), support vector machines (SVM), random forest (RF),
logistic regression (LR), and multi-layer perceptrons (MLP) were used to train different
classifier models. The GridSearchCV algorithm, also implemented on Scikit-learn, was
used to deduce ideal hyper-parameters for each of the five ML algorithms. The
classifier models were evaluated on recognition accuracy and the results show that
the MLP classifier, with a recognition accuracy of 98%, outperforms KNN, RF, LR and
SVM classifiers. A graphical user interface (GUI) is developed and the best performing
classifier model, MLP, is deployed on the developed GUI intended to be used for real time speaker identification and verification tasks. Participants were recruited to the
GUI performance and acceptable results were obtaine
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
The automatic recognition of emotions in speech
Thesis(M.Sc.(Computer Science)) -- University of Limpopo, 2020Speech emotion recognition (SER) refers to a technology that enables machines to detect and recognise human emotions from spoken phrases. In the literature, numerous attempts have been made to develop systems that can recognise human emotions from their voice, however, not much work has been done in the context of South African indigenous languages. The aim of this study was to develop an SER system that can classify and recognise six basic human emotions (i.e., sadness, fear, anger, disgust, happiness, and neutral) from speech spoken in Sepedi language (one of South Africa’s official languages). One of the major challenges encountered, in this study, was the lack of a proper corpus of emotional speech. Therefore, three different Sepedi emotional speech corpora consisting of acted speech data have been developed. These include a RecordedSepedi corpus collected from recruited native speakers (9 participants), a TV broadcast corpus collected from professional Sepedi actors, and an Extended-Sepedi corpus which is a combination of Recorded-Sepedi and TV broadcast emotional speech corpora. Features were extracted from the speech corpora and a data file was constructed. This file was used to train four machine learning (ML) algorithms (i.e., SVM, KNN, MLP and Auto-WEKA) based on 10 folds validation method. Three experiments were then performed on the developed speech corpora and the performance of the algorithms was compared. The best results were achieved when Auto-WEKA was applied in all the experiments. We may have expected good results for the TV broadcast speech corpus since it was collected from professional actors, however, the results showed differently. From the findings of this study, one can conclude that there are no precise or exact techniques for the development of SER systems, it is a matter of experimenting and finding the best technique for the study at hand. The study has also highlighted the scarcity of SER resources for South African indigenous languages. The quality of the dataset plays a vital role in the performance of SER systems.National research foundation (NRF) and
Telkom Center of Excellence (CoE
The development of an automatic pronunciation assistant
Thesis (M. Sc. (Computer Science)) -- University of Limpopo, 2019The pronunciation of words and phrases in any language involves careful manipulation of linguistic features. Factors such as age, motivation, accent, phonetics, stress and intonation sometimes cause a problem of inappropriate or incorrect pronunciation of words from non-native languages. Pronunciation of words using different phonological rules has a tendency of changing the meaning of those words. This study presents the development of an automatic pronunciation assistant system for under-resourced languages of Limpopo Province, namely, Sepedi, Xitsonga, Tshivenda and isiNdebele.
The aim of the proposed system is to help non-native speakers to learn appropriate and correct pronunciation of words/phrases in these under-resourced languages. The system is composed of a language identification module on the front-end side and a speech synthesis module on the back-end side. A support vector machine was compared to the baseline multinomial naive Bayes to build the language identification module. The language identification phase performs supervised multiclass text classification to predict a person’s first language based on input text before the speech synthesis phase continues with pronunciation issues using the identified language. The speech synthesis on the back-end phase is composed of four baseline text-to-speech synthesis systems in selected target languages. These text-to-speech synthesis systems were based on the hidden Markov model method of development. Subjective listening tests were conducted to evaluate the performance of the quality of the synthesised speech using a mean opinion score test. The mean opinion score test obtained good performance results on all targeted languages for naturalness, pronunciation, pleasantness, understandability, intelligibility, overall quality of the system and user acceptance. The developed system has been implemented on a “real-live” production web-server for performance evaluation and stability testing using live data
The development of a text generation model for Sepedi language using transformer-based machine-learning techniques
Thesis (M.Sc. (Computer Science)) -- University of Limpopo, 2025The transformer-based machine learning technique is a deep learning model that
processes the sequential input data using an encoder-decoder process. Transformers process the input data simultaneously using a parallelism approach while paying attention to each word at the time by applying an attention mechanism to each unit text being processed. The transformer-based model has been known to provide more state-of-the-art performance in natural language processing (NLP) tasks than a recurrent neural network (RNN) such as Long Short-Term Memory (LSTM). RNNs have the drawback of suffering from the problem of vanishing gradients and exploding gradients in implementation. The GPT-Sepedi transformer-based model has shown great success in dealing with the process of text generation for the Sepedi language. This has led to a limited text generation system developed using a transformer-based model for the under resourced African language, namely, the Sepedi language. This research project aimed to develop a text generation model for the Sepedi language using transformer based machine learning techniques. The LSTM-Sepedi Attention-based model and the GPT-Sepedi Transformer-based model were developed and trained using a National Centre for Human Language Technology (NCHLT) Sepedi text corpus. The models were compared based on the results that they generated. A GPT-Sepedi Transformer-based model was used to generate the text. The generated text was then compared with the Sepedi language vocabulary to
determine the validity of the text. It was found that 61% of the text within the generated texts is found in the Sepedi language vocabulary. The Recall-Oriented Understudy for Gisting Evaluation (ROUGE) score was used to compare the model generated text to human-written text. The ROUGE score result indicates that the GPT-Sepedi Transformer-based text generation model was able to generate words that humans can write with 83% precision. Even though the precision results indicated a better percentage, the text generated cannot be comprehensible with the recall percentage of 0.05% and 0.1% F1-score results.NRF (National Research Foundation
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
- …
