1,720,996 research outputs found
BERT-based multi-task learning for aspect-based sentiment analysis
The Aspect Based Sentiment Analysis (ABSA) systems aims to extract the aspect terms (e.g., pizza, staff member), Opinion terms (e.g., good, delicious), and their polarities (e.g., Positive, Negative, and Neutral), which can help the customers and companies to identify product weaknesses. By solving these product weaknesses, companies can enhance customer satisfaction, increase sales, and boost revenues. There are several approaches to perform the ABSA tasks, such as classification, clustering, and association rule mining. In this research we have used a neural network-based classification approach. The most prominent neural network-based methods to perform ABSA tasks include BERT-based approaches, such as BERT-PT and BAT. These approaches build separate models to complete each ABSA subtasks, such as aspect term extraction (e.g., pizza, staff member) and aspect sentiment classification. Furthermore, both approaches use different training algorithms, such as Post-Training and Adversarial Training. Moreover, they do not consider the subtask of Opinion Term Extraction. This thesis proposes a new system for ABSA, called BERT-ABSA, which uses MultiTask Learning (MTL) approach and differentiates from these previous approaches by solving all three tasks such as aspect terms, opinion terms extraction, and aspect term related sentiment detection simultaneously by taking advantage of similarities between tasks and enhancing the model’s accuracy as well as reduce the training time. To evaluate our model’s performance, we have used the SemEval-14 task 4 restaurant datasets. Our model outperforms previous models in several ABOM tasks, and the experimental results support its validit
Diagnosis of pleural mesothelioma using machine learning
Mesothelioma is cancer that develops in the pleura. The most common cause of this disease is contact with asbestos. Patients with mesothelioma have a better chance of surviving if they are diagnosed quickly. This study utilizes a variety of machine learning to enhance pleural mesothelioma diagnosis. The possibility of misclassification was decreased by extracting features from a preexisting dataset. SVM, Decision Trees, and Random Forests are only a few machine learning classifiers trained using essential and foundational features. Accuracy, precision, recall, and F1-score were just a few measures used to evaluate these classifiers' performance in cross- validation. SVM demonstrated excellent accuracy, precision, recall, and F1-score when classifying individuals as either healthy or having mesothelioma. The results show the potential of machine learning techniques for early diagnosis of pleural mesothelioma. Machine learning algorithms improve diagnosis accuracy and turnaround time, improving patient outcomes. Using the results of this research, a fully automated technique for diagnosing mesothelioma might be developed, allowing clinicians more time to provide better care for their patients
Enhancing neural mean teacher learning-based emotion-centric model for image captioning
Image captioning is a task in computer vision and natural language processing that involves
generating a textual description of the content of an image. The goal of image captioning is to
create a system that can accurately recognize the objects, attributes, and relationships depicted
in an image, and generate a meaningful description of it in natural language, typically in the
form of a sentence or short paragraph. One of the state-of-the-art methods that we can use for
image captioning is Nemesis: Neural Mean Teacher Learning-based Emotion-centric Speaker.
Nemesis is a neural mean teacher learning-based emotion-centric speaker. It is a proposed
neural speaker capable of leveraging emotional supervision signals in the caption generation
process. Nemesis has been applied to the recently introduced ArtEmis dataset, which is the first
large-scale dataset for emotion-centric image captioning, containing 455K emotional
descriptions of 80K artworks from WikiArt. In this study, I employed a straightforward but
improved version of Self-Critical Sequence Training. By modifying the baseline function
choice in the REINFORCE algorithm, I introduced a simple alteration. The updated baseline
offers enhanced performance without any additional expenses, when compared to the baseline
that utilizes greedy decoding
A hydrid deep neural network for electroencephalogram (EEG)-based screening of depression
Technological development is a major contributor to improve people's quality of life. In recent
times happy life has been considered one of the major requirements as people live under stress and
face several mental disorders like depression, anxiety, and loneliness. In the mental disorder space,
depression is a major and common disease. According to the World Health Organization (WHO),
it is estimated that 5% of adults suffer from depression. Diagnosis of depression has several
challenges, for example, patient counseling is time consuming, over-dependence on doctors and
accuracy of diagnosis. To resolve these diagnosis issues, computer aided system is required with
the use of machine learning tools. The objective of this research to develop hybrid deep learning
model by using CNN and LSTM. The dataset used in this study contains 945 subjects of mental
disorders and healthy control subjects. Three hybrid models were developed and compared with
different sets of extracted features. Raw data was pre-processed and applied in hybrid model, and
at the end the model was validated with the unknown EEG dataset. The hybrid model with entire
features of dataset reported an accuracy of 98.0% and performed better in comparison with other
two models which were trained with extracted features by using decision tree classifier. The results
show that the developed hybrid CNN and LSTM model is accurate, less complex, and useful in
detecting mental disorders including depression using EEG signals
Survival analysis and prediction of lung cancer in patients based on clinical and image features using machine learning
Lung cancer develops in lung tissues, most commonly in the cells that line the airways. It is the leading cause of death from cancer in both men and women. To estimate the prevalence of lung cancer in the coming years, it is necessary to diagnose it in the early stages. This thesis work proposes
to perform a reliable diagnosis of patients with lung cancer. The goal of this research is to analyze the important variables impacting lung cancer based on p-value using image features as well as clinical data and is focused on quality analysis. Further, to enable early diagnosis of cancer with high efficiency, this work proposes to classify the patient’s images into cancer using a Convolutional Neural Network (CNN) to enable its early diagnosis. The thesis discusses the dataset, data pre-processing steps, survival rate risk analysis, classification, and performance evaluation of the process. This study used two kinds of data, clinical and image data. The Genomic Data Commons (GDC) Data Portal and The Cancer Imaging Archive (TCIA) were used as the data source. The Random Forest regression estimation method was used to fill in the missing values. It first imputes all missing data with the mean/mode, then fits a random forest on the observed part and predicts the missing part for each variable with missing values. Three models are used to test the significance of variables on cancer survival rates: Kaplan Meier (KM), Cox Proportional Hazards (CPH), and Accelerated Failure Time (AFT). The analysis took into account three types of data: clinical only, image only, and combined clinical and image data. All three models have been effectively applied and the outcome revealed the most robust data and the crucial variable to be focused upon for further experimentation. For classification, a Convolutional Neural Network (CNN), with low
computational cost and time overhead is used. The output of statistical models demonstrates the robustness of image data among all types, as it has the fewest chances of producing false results.
Image data, which is common in clinical data collection is less prone to human error. As a result of the data's robustness, only image features data was preferred over clinical data and combined in the next step to perform the classification of images for cancer prediction. Based on the accuracy, the CNN results were compared to the two other ensemble approaches, Random forest (RF) and XgBoost. CNN achieved an accuracy of 99% in image classification, which was higher than the accuracy rates of Random forest (RF) and XgBoost, which were 95.83% and 95.83%, respectively.
As a result, the CNN model can be applied to new Computerized Tomography (CT) scan images for lung cancer diagnosis to conduct additional research and to assist clinicians
Increased Prediction Accuracy in the Game of Cricket Using Machine Learning
ABSTRACT
Player selection is one the most important tasks for any sport and cricket is no exception. The performance of the players depends on various factors such as the opposition team, the venue, his current form etc. The team management, the coach and the captain select 11 players for each match from a squad of 15 to 20 players. They analyze different characteristics and the statistics of the players to select the best playing 11 for each match. Each batsman contributes by scoring maximum runs possible and each bowler contributes by taking maximum wickets and conceding minimum runs. This paper attempts to predict the performance of players as how many runs will each batsman score and how many wickets will each bowler take for both the teams. Both the problems are targeted as classification problems where number of runs and number of wickets are classified in different ranges. We used naïve bayes, random forest, multiclass SVM and decision tree classifiers to generate the prediction models for both the problems. Random Forest classifier was found to be the most accurate for both the problems.
KEYWORDS
Naïve Bayes, Random Forest, Multiclass SVM, Decision Trees, Cricke
Early risk prediction in acute aortic syndrome on clinical data using machine learning
Advancements in machine learning present novel opportunities for early prediction of Acute
Aortic Syndrome (AAS) as a critical and life-threatening clinical condition and the identification
of critical features influencing this prediction. This study concentrates on integrating, cleaning,
and handling missing data from extensive clinical datasets sourced from 150 emergency
departments across Canada and the USA. Covering medical histories of nearly 150,000 patients
from 2021 to 2022, the dataset comprises categorical clinical variables. Additionally, the research
focuses on constructing predictive machine learning models utilizing various data-splitting
strategies and classifiers to optimize AAS prediction. Methodologically, the study encompasses
data identification, acquisition, exploration, processing, and feature extraction, followed by
dimensionality reduction using Principal Component Analysis (PCA) and other feature selection
methods such as Correlation-based (CFS) and Relief. The multiple imputations method and the
SMOTE method are utilized for handling missing and imbalanced data, respectively. The findings
demonstrate that employing the Relief-feature method with an 80-10-10 split ratio alongside the
Random Forest classifier yields an exceptional accuracy of 99.3%, surpassing alternative models..
Furthermore, this research addresses a prevalent challenge encountered by many researchers
regarding dataset size limitations, thereby facilitating the utilization of the integrated and prepared
dataset for research on AAS and other cardiovascular diseases
Evaluating the impact of an educational intervention on the management of BPPV: an interrupted time series analysis
This study examines the impact of an educational intervention on the use of key diagnostic and treatment maneuvers for Benign Paroxysmal Positional Vertigo (BPPV) and the reduction of unnecessary CT scans in clinical practice. Data were sourced from three healthcare institutions in Ontario, encompassing over 2000 patient records from the pre- and post-intervention periods. The primary diagnostic and treatment outcomes analyzed include Dix-Hallpike, Supine Roll, and Canalith Repositioning Maneuver (CRM), while the secondary focus is on the utilization of CT scans. The analysis employs Interrupted Time Series (ITS) methodology, with segmented regression used to assess changes in the frequency of these procedures. Adjustments were made for potential confounders such as history of BPPV, vomiting, and vertigo. The intervention led to an immediate increase in the appropriate use of Dix-Hallpike and Supine Roll maneuvers, followed by sustained positive trends in their application. Similarly, CRM usage significantly increased post-intervention. However, the intervention had no significant effect on reducing CT scan usage, highlighting the complex challenges of minimizing unnecessary imaging in clinical settings. Nevertheless, after accounting for confounding variables, the post-intervention increases observed previously were no longer statistically significant. This research underscores the importance of targeted educational programs in promoting evidence-based diagnostic and treatment maneuvers for BPPV. While the intervention effectively increased the use of appropriate procedures, its impact on reducing CT scan usage was limited. Future work will focus on extending the follow-up period to assess long-term trends and exploring alternative analytical methods to uncover complex patterns and validate intervention effectiveness
Detecting Image Forgery over Social Media Using U-NET with Grasshopper Optimization
Currently, video and digital images possess extensive utility, ranging from recreational and social media purposes to verification, military operations, legal proceedings, and penalization. The enhancement mechanisms of this medium have undergone significant advancements, rendering them more accessible and widely available to a larger population. Consequently, this has facilitated the ease with which counterfeiters can manipulate images. Convolutional neural network (CNN)-based feature extraction and detection techniques were used to carry out this task, which aims to identify the variations in image features between modified and non-manipulated areas. However, the effectiveness of the existing detection methods could be more efficient. The contributions of this paper include the introduction of a segmentation method to identify the forgery region in images with the U-Net model’s improved structure. The suggested model connects the encoder and decoder pipeline by improving the convolution module and increasing the set of weights in the U-Net contraction and expansion path. In addition, the parameters of the U-Net network are optimized by using the grasshopper optimization algorithm (GOA). Experiments were carried out on the publicly accessible image tempering detection evaluation dataset from the Chinese Academy of Sciences Institute of Automation (CASIA) to assess the efficacy of the suggested strategy. The results show that the U-Net modifications significantly improve the overall segmentation results compared to other models. The effectiveness of this method was evaluated on CASIA, and the quantitative results obtained based on accuracy, precision, recall, and the F1 score demonstrate the superiority of the U-Net modifications over other models
- …
