1,720,959 research outputs found

    Handwritten Digit Recognition using Machine Learning

    Get PDF
    Handwritten Digit Recognition (HDR) remains a fundamental benchmark in pattern recognition and machine learning due to its practical applications and inherent classification challenges posed by diverse handwriting styles. This study investigates and compares two classical statistical classifiers—Gaussian Naive Bayes (GNB) and Linear Discriminant Analysis (LDA)—to recognize the digits from the MNIST dataset. Both models assume underlying normality in feature distributions and offer computational efficiency, making them suitable for high-dimensional input such as image pixels. Using 60,000 training and 10,000 test samples, we evaluate model performance through accuracy, precision, recall, F1 score, and confusion matrices. The results reveal that while GNB achieves moderate accuracy (55.58\%), LDA significantly outperforms it with an accuracy of 87.30\%, demonstrating superior capability in distinguishing visually similar digits. Our analysis further highlights the limitations of GNB’s independence assumption and underscores LDA’s strength in capturing shared variance across classes. These findings reinforce the effectiveness of LDA as a robust baseline for HDR tasks, especially when interpretability and computational simplicity are desired

    Application and Analysis of Machine Learning and Deep Learning Algorithms in Detection of DDoS Cyberattacks

    Get PDF
    A Distributed Denial-of-Service (DDoS) attack involves overwhelming a target system\u27s data bandwidth or computational resources, often using multiple attack systems, aiming to slow down or disable the targeted system. Detecting and mitigating DDoS attacks effectively remains challenging due to their varying characteristics. One of the promising approaches involves developing an AI based Intrusion Detection System (IDS) against cyberattacks. In this study, we aim to develop an AI based Intrusion Detection System (IDS) for DDoS threat detection using Machine Learning, Deep Learning, or hybrid techniques. Different Machine Learning (ML) and Deep Learning (DL) algorithms like Random Forest (RF), Naïve Bayes (NB), Logistic Regression (LR), K-Nearest Neighborhood (KNN), Deep Neural Network (DNN), Long Short-Term Memory (LSTM) have been used to build the AI based intrusion detection system. CIC-DDoS-2019 and CIC-IoT-2023 datasets were utilized in this work for training and testing the performance of the AI models. In this work, we also concentrated on addressing data imbalance issues, which arose from the presence of high volume of attack data compared to benign data. Four different data balancing techniques have been used to solve the data imbalance problem. The performance of ML and DL models was assessed using metrics such as accuracy, precision, recall, F1 score, balanced accuracy, and Area Under the ROC-Curve (AUC) score under four different balancing techniques. Lastly, we compared the performance of these ML and DL models with different balancing techniques to obtain a better solution

    Detecting Physical Activity Using Wearable Sensor Data

    No full text
    This study focuses on detecting physical activity using wearable sensor data, specifically distinguishing between walking and running. A dataset comprising accelerometer and gyroscope readings is used to train and evaluate various machine learning models, including logistic regression, random forest, k-nearest neighbors, naïve Bayes, and XGBoost. Extensive preprocessing, such as creating lag features and rolling statistics, is performed to enhance temporal data representation. The models are evaluated using metrics like accuracy, precision, recall, and F1 score. Incorporating lag and rolling features significantly improves model performance, with logistic regression achieving perfect scores across all metrics. These findings demonstrate the effectiveness of enhanced feature engineering for time-series data in human activity recognition and highlight the potential of wearable sensors in monitoring physical activities

    Clustering Dataset Using K-Mean Clustering

    Get PDF
    Clustering is a fundamental technique in unsupervised machine learning, widely applied in various domains such as pattern recognition, data segmentation, and anomaly detection. This study evaluates the performance of the K-Means clustering algorithm on multiple benchmark datasets, including low-dimensional, high-dimensional, and imbalanced datasets. The clustering results are assessed using four key evaluation metrics: Mean Squared Error (MSE), Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), and Silhouette Score. Experimental results demonstrate that K-Means performs effectively on datasets with well-separated clusters, particularly in high-dimensional spaces, where it achieves near-perfect clustering accuracy. However, its performance deteriorates in datasets with overlapping clusters and varying cluster densities, highlighting its sensitivity to initialization and cluster structure

    Data Science Job Salary Prediction Using Linear Regression

    Get PDF
    In the evolving landscape of data science, accurate salary prediction plays a crucial role in shaping career expectations, informing educational strategies, and guiding organizational hiring decisions. This study investigates the key factors influencing entry-level data science salaries in the United States by applying a multiple linear regression model to a recent dataset spanning from 2020 to 2024. Through data preprocessing, transformation, and diagnostic evaluation, we identify how job roles, experience levels, employment types, work arrangements, residency status, and company size impact compensation. Despite challenges such as outliers, heteroscedasticity, and non-normal residuals, model refinements like the Box-Cox transformation and variable selection enhance predictive performance. The final model, while modest in explanatory power, offers actionable insights into salary determinants and lays the groundwork for future predictive modeling improvements in the domain

    Performance of LASSO and Ridge Regression for Variable Selection in Genome-Wide Association Studies of Maize Flowering Time

    Get PDF
    Genome-Wide Association Studies (GWAS) are instrumental in identifying genetic variants linked to complex traits, providing valuable insights into trait heritability and biological mechanisms. This study applies GWAS to investigate flowering time in maize, a critical adaptive trait, using a diverse dataset of 5,000 recombinant inbred lines across eight environments. Traditional GWAS methods often encounter challenges in high-dimensional datasets due to the presence of multiple small-effect genetic loci. To address this, we compared two penalized regression methods—LASSO and Ridge regression—to perform variable selection and regression analysis within a GWAS framework. LASSO effectively reduced the number of predictors by selecting the most impactful variables, while Ridge regression retained more features, offering a broader genetic context for predicting flowering time. Results demonstrated that Ridge regression yielded slightly better predictive performance, achieving a lower Mean Squared Error (MSE) and Root Mean Squared Error (RMSE) than LASSO

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
    corecore