Journal of Computer Networks, Architecture and High Performance Computing
Not a member yet
473 research outputs found
Sort by
Optimizing SMS Spam Detection Using Machine Learning: A Comparative Analysis of Ensemble and Traditional Classifiers
With the rapid rise of mobile communication, Short Message Service (SMS) has become an essential platform for transmitting information. However, the growing volume of unsolicited and harmful spam messages presents significant challenges for both users and mobile network operators. This study explores the effectiveness of various machine learning models, including Random Forest, Gradient Boosting, AdaBoost, Support Vector Machine (SVM), Logistic Regression, and an Ensemble Voting Classifier, in detecting SMS spam. A dataset containing 5,572 SMS messages, labeled as either spam or ham (legitimate), was used to evaluate these models. Hyperparameter tuning was performed on each model to optimize accuracy, and the models were assessed using metrics such as precision, recall, F1-score, and accuracy. The results indicated that the SVM and Ensemble Voting Classifier achieved the highest performance, with accuracies of 0.9857 and 0.9848, respectively. Both models demonstrated superior recall for spam messages, making them highly effective for real-world spam detection systems. While Random Forest, Gradient Boosting, and AdaBoost also performed well, their slightly lower recall for spam suggests that they may misclassify some spam as legitimate messages. The study highlights the effectiveness of machine learning models in addressing the SMS spam problem, particularly when using ensemble methods. Future research should focus on addressing class imbalance and exploring deep learning approaches to further enhance model performance. These findings offer valuable insights for developing more accurate and scalable SMS spam detection systems
Menu Sales Prediction at Kiyo Café Using Machine Learning
This research evaluates the performance of the K-Nearest Neighbors (KNN) and Naïve Bayes algorithms in predicting raw material stock for Café Kiyo. The study encompasses six key stages, including preparation, literature review, data collection, data mining processing, results and discussion, and conclusion with recommendations. The data mining process adheres to the Knowledge Discovery in Databases (KDD) framework, involving data selection, preprocessing, transformation, data mining, and interpretation and evaluation. The evaluation metrics reveal that KNN boasts a marginally higher accuracy of 98.71% compared to Naïve Bayes with 98.21%. KNN also demonstrates superior precision (81.25%) in identifying true positives, outperforming Naïve Bayes (72.59%). However, Naïve Bayes excels in recall, achieving 95.15% compared to KNN's 50.00%. The Area Under the Curve (AUC) analysis further confirms Naïve Bayes' superiority, with an AUC value of 0.995, indicating better performance in distinguishing between positive and negative classes
Analysis of Logistic Regression Regularization in Wild Elephant Classification with VGG-16 Feature Extraction
The research article explores the intersection of image-based wildlife classification and logistic regression regularization, focusing on the classification of wild elephant species. It begins by highlighting the significance of ecological research in biodiversity monitoring and conservation and introduces Convolutional Neural Networks (CNNs) as potent tools for feature extraction from images. The VGG-16 model is particularly emphasized for its ability to capture hierarchical representations of visual features crucial for classification tasks. The integration of VGG-16 feature extraction with logistic regression regularization is proposed as a compelling approach, offering a balance between sophisticated feature representation and efficient classification algorithms. The literature review delves into image-based wildlife classification, emphasizing the role of CNNs, especially VGG-16, in extracting discriminative features. It discusses the fusion of VGG-16 features with logistic regression and the challenges in this field, such as dataset annotation and environmental variability. The method section outlines the dataset acquisition, feature extraction using the VGG-16 architecture, and model configuration using logistic regression with lasso and ridge regularization. The process of finding the optimal regularization parameter (lambda) and model evaluation through cross-validation is detailed. Results showcase the optimal lambda values for lasso and ridge regularization and compare the performance of logistic lasso and logistic ridge models. Misclassification analysis reveals factors influencing classification accuracy, including feature variability and contextual complexity. The discussion reflects on the implications of the findings, emphasizing the importance of lambda selection and addressing challenges in wildlife classification. It suggests avenues for further research, such as advanced modeling techniques and feature engineering approaches. In conclusion, the study contributes to advancing wildlife classification efforts by leveraging state-of-the-art techniques and sheds light on opportunities to enhance classification accuracy in wildlife conservation
Usability Evaluation Of The Academic Information System Using The Concurrent Think-Aloud, Webuse, And Sus Methods
This research was motivated by the situation that the academic information system (SIAKAD) Universitas Teknologi Indonesia (UTI) had never been tested for usability based on user responses. Apart from that, the flow that occurs in the UTI academic information system is not yet coherent. From the user's side, the information guide feature for using SIAKAD UTI is not yet available. Based on this, the research aimed to: 1) describe the results of Usability evaluation using the Concurrent Think-Aloud and Webuse methods on the Academic Information System of Indonesia Technology University; 2) describe recommendations for usability evaluation of the Concurrent Think-Aloud and Webuse methods on the Academic Information System of the Indonesia Technology University; and 3) create a simulation of SIAKAD UTI development and describe the test results using the System Usability Scale (SUS) method. The research method used a mix of qualitative and quantitative methods, so that data obtained is more comprehensive, valid, reliable, and objective. Usability evaluation data through the Concurrent Think-Aloud method was collected using interviews, usability evaluation through the Webuse method was collected using a questionnaire, and satisfaction data through the SUS method was collected using a questionnaire. The results of the development carried out in this research can be concluded as follows. (1) The SIAKAD UTI evaluation results through the Concurrent Think-Aloud method showed that many navigation buttons did not function, inconsistent button features, disproportionate location/buttons position, and inappropriate color selection. Therefore, the evaluation results through the Concurrent Think-Aloud Method showed that SIAKAD UTI needs to be improved. (2) The SIAKAD UTI evaluation results through the Webuse method obtained an overall mean value of usability points, namely 0.23, in the range 0.2 < x ?0.4 with a usability level category of Poor, which means that it is necessary to improve the development of SIAKAD UTI. (3) Based on the evaluation results, a SIAKAD UTI development simulation was carried out, followed by a satisfaction test using SUS. The average score obtained after processing the assessment scores from respondents with an average value of 97.35, which showed that respondents were very satisfied using the results of the SIAKAD UTI development simulation
User Satisfaction Analysis of the PT Dikstra Cipta Solusi Website Using the Webqual 4.0 Method
The abundance of IT consulting firms in this digital era necessitates continuous improvement for these companies. From services to products, everything must be constantly updated. Websites have become a crucial platform for reaching customers connected to the internet. Customer satisfaction is equally important in building trust with IT consulting firms. This makes companies more aware of the need for customer satisfaction. The Webqual 4.0 method can be a solution to help IT consulting firms understand satisfaction from the end user's perspective by measuring website quality in terms of usability, information quality, and service interaction. The use of the Webqual 4.0 method also requires IT consulting firms to pose a series of questions to existing customers to gauge how attractive, informative, useful, and secure a website is when accessed. This makes the method highly effective if the questions asked are accurate and align with the views of the customers or end users. From the research conducted on PT Diksta Cipta Solusi, it was found that the dominant characteristics of the website users are males aged 35-39 years, with a usage frequency of about twice a month. User evaluations of the website fall into the good category, with the majority of users feeling satisfied or very satisfied, as evidenced by the high validity and reliability of the assessments. The most significant factor influencing user satisfaction is usability, particularly in terms of the ease of operating the website. It was discovered that the three independent variables—usability, information quality, and service interaction—collectively influence user satisfaction. This study makes an important contribution by demonstrating the effectiveness of the Webqual 4.0 method in measuring website quality and providing insights for IT consulting firms to focus on improving usability, information quality, and service interaction to enhance user satisfaction
Prediction of Obesity Categories Based on Physical Activity Using Machine Learning Algorithms
Obesity is a global health issue with rising prevalence, marked by excessive fat accumulation that poses health risks. Contributing factors include poor eating habits, lack of physical activity, and genetics, which elevate the risk of chronic diseases like type 2 diabetes, heart disease, stroke, and cancer. This study examines an obesity dataset with seven variables: Age, Gender, Height, Weight, BMI, Physical Activity Level, and Obesity Category. The analysis reveals strong correlations between Body Weight, BMI, and the Obesity Category, while Body Height shows a moderate negative correlation. Various machine learning algorithms were tested, including XGBoost, AdaBoost, Gradient Boosting, and Extra Trees Classification. XGBoost emerged as the top performer, achieving the highest accuracy (0.9961) and an almost perfect AUC (0.9992), making it highly effective for obesity prediction. The study's significance lies in its ability to elucidate the key factors contributing to obesity and their interactions. By recognizing the strong links between Body Weight, BMI, and Obesity Category, healthcare professionals can craft more targeted interventions. Furthermore, the successful application of advanced machine learning algorithms underscores the potential for technology to enhance predictive accuracy and support healthcare decision-making. The findings highlight XGBoost's superior performance, demonstrating its value in predicting obesity and aiding in early diagnosis and prevention strategies. This research emphasizes the critical role of data and technology in tackling obesity and improving public health outcomes
The Prediction of Thyroid Cancer Recurrence with the XGBoost Method: The Clinicopathological Feature-Based Approach
This research aims to develop a thyroid cancer recurrence prediction model using the XGBoost method with a clinicopathological feature-based approach. Thyroid cancer is one of the cancers that have a significant recurrence rate after initial treatment. Therefore, thyroid cancer recurrence prediction is important in determining treatment plans and patient management. In this study, we used a dataset containing 383 records of clinicopathological information on thyroid cancer patients who had undergone treatment. The features include various clinical and pathological parameters that are considered important in recurrence prediction. We used the XGBoost algorithm, which has proven effective in various classification tasks, to build a prediction model. The model evaluation results show good consistency in predicting the thyroid cancer recurrence with an average accuracy value of around 97.74% and an average F1-score value of around 95.94%. The results show that the XGBoost model can provide thyroid cancer recurrence prediction with good accuracy, with the ability to effectively detect both classes (recurrence and non-recurrence). The model is expected to be a valuable tool in supporting clinical decision-making related to the management of thyroid cancer patients
Apriori Algorithm to Predict Availability of Beauty Products
This study introduces the Apriori algorithm in beauty product availability prediction system as a solution to enhance stock prediction accuracy and mitigate inventory risks in the beauty industry. By applying data mining technology, specifically the Apriori algorithm, Kazana Kosmetik aims to gain insights into consumer purchasing patterns to optimize operations. The research analyzes transaction data to identify key buying patterns and improve stock management strategies. The results reveal seven main purchasing patterns with an average confidence value of 0.414, offering valuable guidance for Kazana Kosmetik in inventory control and marketing tactics. By leveraging data mining techniques, companies like Kazana Kosmetik can streamline sales strategies and enhance customer satisfaction. This research underscores the effectiveness of the Apriori algorithm in predicting beauty product availability and its potential to revolutionize operational efficiency in the cosmetics market
Literature Review Application of YOLO Algorithm for Detection and Tracking
A vehicle tracking system is a computer program that utilizes devices to monitor the position, movement and condition of a vehicle or fleet of vehicles. Multi-vehicle tracking on highways has significant research interest and practical value in building intelligent transportation systems. Nevertheless, traffic road video frames consist of various complex backgrounds and objects. Detection and tracking are very challenging because foreground to background switching occurs frequently. One-stage algorithm approaches such as YOLO and its various variants have been proven to be accurate for detecting vehicles. Meanwhile, the SORT, DeepSORT, ByteTrack and other algorithms can be combined in YOLO. The aim of this study is to highlight existing research on the application of YOLO and its variants in detecting and tracking vehicles, especially in traffic management. The journals used are limited to 2019 – 2024 and the journal sources consist of Hindawi, IEEE, MDPI, Research Gate, Science Direct, and Springer. Based on the research that has been reviewed, the YOLO variant algorithm approach has been successfully applied in the field of vehicle monitoring to support smart cities. In addition, many new model combinations and improvements have been proposed, proving that this algorithm has a big influence in the field of computer vision
Sentiment Analysis of Oppenheimer Movie Reviews: Naïve Bayes Algorithm for Public Opinion
The development of information and communication technology has revolutionized the way people consume and engage with media, particularly in the realm of film. Online platforms such as Netflix, Amazon Prime Video, and YouTube have transformed movie consumption habits, providing a vast array of options for viewers to explore and enjoy. A crucial aspect of this digital landscape is the proliferation of movie reviews, which serve as valuable guides for users seeking to discover films aligned with their preferences. However, the abundance of reviews, often varying in quality and objectivity, necessitates tools capable of effectively processing and understanding these textual data. This research delves into sentiment classification of Oppenheimer movie reviews, utilizing the Naive Bayes algorithm to categorize reviews into positive, negative, and neutral sentiments. The dataset comprising audience reviews and numerical ratings undergoes preprocessing using the TF-IDF method to facilitate numerical representation. Subsequently, the Naïve Bayes algorithm is trained on this processed data to accurately classify sentiments. The model demonstrates exceptional performance, achieving an accuracy rate of 97.45% in distinguishing between positive, negative, and neutral sentiments within Oppenheimer movie reviews. This study underscores the efficacy of the Naive Bayes algorithm in sentiment classification and emphasizes the significance of employing techniques like TF-IDF for enhancing sentiment analysis in the domain of movie reviews