San Jose State University

SJSU ScholarWorks
Not a member yet
    32584 research outputs found

    Enhancing Environmental Health and Safety: Fine-Tuning Large Language Models for Domain-Specific Applications

    Get PDF
    This study aims to simplify Environmental Health and Safety (EHS) by leveraging the power of Large Language Models (LLMs). In this research, we focus on fine-tuning three LLMs — LLaMA, Mistral, and Falcon — using PEFT techniques such as QLoRA and SFT, to address domain-specific needs such as safety compliance, incident reporting, and knowledge dissemination. Our research methodology involves fine-tuning each LLM model on a custom dataset compiled from various regulatory agencies, supplemented by targeted web scraping and manual collection of questionnaires to capture and enrich the models with the latest regulations and guidelines. This study aims to compare the effectiveness of these fine-tuned models to identify the most effective model and fine-tuning techniques for specific EHS applications. We aim to integrate LLMs to support EHS practices and demonstrate a practical way of improving and automating EHS queries and reports to foster safer and more compliant workplace environments

    A framework for scientific data indexing, searching and sharing

    Get PDF
    Scientific data continues to grow. Wildfire simulation experiments performed by the WIRC team at SJSU have generated over 138 TB of data so far and it is expected to keep growing. It becomes difficult for researchers to search through that data to find the data of their interest. This data is stored on an HPC cluster that external users do not have access to. The WIRC team also conducts experiments and publishes their research, but the size of data makes it difficult to share these datasets. This project introduces a novel solution to indexing scientific data, searching through the data and sharing it

    Explaining the Maliciousness of URLs using SHAP and LIME

    Get PDF
    No system has ever reached the levels of proliferation that the Internet now enjoys. It stands as the most widely spread distributed system across the globe; yet this evolution has given rise to an ever-growing wave of malintent that challenges every user and entity on the vast expanse of cyberspace. Malicious URLs loom large as vulnerabilities leaving users naked as they traverse online landscapes, but cybersecurity experts craft models with esoteric algorithms in a bid to stem this tide and shield users from cybercrime. However, peering into the decision-making corridors of these models holds key importance, it’s through understanding such cognitive landscapes that robust forward protectors for users and platforms can be erected. Machine learning models are often dubbed black boxes; but not in our case. The term black boxes is often used to describe machine learning models as their workings are hidden from view; however, this paper delves into this very topic by investigating the interpretability of machine learning models with a specific focus on their use in detecting malicious URLs. Among the various models considered, attention is paid to identifying the most effective one out of a pool of five: MLP, deep models, RandomForest Classifier, SVM, and XGBoost through SHAP and LIME techniques - hoping that this dual approach will illuminate differing aspects regarding each model’s operation. By conducting an in-depth analysis and juxtaposition between these methodologies (SHAP and LIME), it is hoped that more light will be shed on how these models work differently and where exactly one can draw precise cybersecurity decisions from

    A Laplace-based model with flexible tail behavior

    Get PDF
    The proposed multiple scaled contaminated asymmetric Laplace (MSCAL) distribution is an extension of the multivariate asymmetric Laplace distribution to allow for a different excess kurtosis on each dimension and for more flexible shapes of the hyper-contours. These peculiarities are obtained by working on the principal component (PC) space. The structure of the MSCAL distribution has the further advantage of allowing for automatic PC-wise outlier detection – i.e., detection of outliers separately on each PC – when convenient constraints on the parameters are imposed. The MSCAL is fitted using a Monte Carlo expectation-maximization (MCEM) algorithm that uses a Monte Carlo method to estimate the orthogonal matrix of eigenvectors. A simulation study is used to assess the proposed MCEM in terms of computational efficiency and parameter recovery. In a real data application, the MSCAL is fitted to a real data set containing the anthropometric measurements of monozygotic/dizygotic twins. Both a skewed bivariate subset of the full data, perturbed by some outlying points, and the full data are considered

    Family building and pregnancy experiences of cisgender sexual minority women

    Get PDF
    BACKGROUND: Although 10% to 20% of cisgender women aged 18 to 40 years have a sexual minority identity (eg, bisexual, lesbian, and queer), there is limited research on the family building and pregnancy experiences of sexual minority cisgender women. Improving our understanding of the family building and pregnancy experiences of cisgender sexual minority women is critical for improving the perinatal health of this population. OBJECTIVE: This study aimed to compare the mode of family building, past pregnancy experiences, and future pregnancy intentions among cisgender sexual minority women by sexual orientation. STUDY DESIGN: This is an observational study which was conducted using cross-sectional data collected in 2019 from a national sample of 1369 cisgender sexual minority women aged 18 to 45 years. RESULTS: Most participants (n=794, 58%) endorsed multiple sexual orientations, most commonly queer (n=641, 47%), lesbian (n=640, 47%), and/or bisexual (n=583, 43%). There were 243 (18%) cisgender sexual minority women who were parents. Pregnancy was used by 74% (181/243) of women to build their families. Among participants who used pregnancy, 60% (108/181) became pregnant through sexual activity with another parent of the child, whereas 27% (64/243) of women used donor sperm. An additional 10% (n=24) became parents through second-parent adoption, 10% (n=25) through adoption, and 14% (n=35) through step-parenting. Bisexual women more often used sexual activity to become parents (61/100, 61%) compared with queer (40/89, 45%) and lesbian women (40/130, 31%). In contrast, lesbian (50/130, 39%) and queer (25/89, 27%) women more often used donor sperm to become parents compared with bisexual women (11/100, 11%). Among the 266 (19%) cisgender sexual minority women who had ever been pregnant, there were 545 pregnancies (mean, 2.05 pregnancies per woman). Among those pregnancies, 59% (n=327) resulted in live birth, 23% (n=126) resulted in miscarriage, 15% (n=83) resulted in abortion, and 2% (n=9) resulted in ectopic pregnancy. A quarter of women had future pregnancy intentions, with no differences by sexual orientation. Overall, few participants (16%) reported that all of their healthcare providers were aware of their sexual orientation. CONCLUSION: Cisgender sexual minority women primarily built their families through pregnancy and a quarter have future pregnancy desires. In addition, there were important differences in family building methods used by sexual orientation. Providers should be aware of the pregnancy and family-building patterns, plans, and needs of cisgender sexual minority women

    Classification and online clustering of zero-day malware

    Get PDF
    A large amount of new malware is constantly being generated, which must not only be distinguished from benign samples, but also classified into malware families. For this purpose, investigating how existing malware families are developed and examining emerging families need to be explored. This paper focuses on the online processing of incoming malicious samples to assign them to existing families or, in the case of samples from new families, to cluster them. We experimented with seven prevalent malware families from the EMBER dataset, four in the training set and three additional new families in the test set. The features were extracted by static analysis of portable executable files for the Windows operating system. Based on the classification score of the multilayer perceptron, we determined which samples would be classified and which would be clustered into new malware families. We classified 97.21% of streaming data with a balanced accuracy of 95.33%. Then, we clustered the remaining data using a self-organizing map, achieving a purity from 47.61% for four clusters to 77.68% for ten clusters. These results indicate that our approach has the potential to be applied to the classification and clustering of zero-day malware into malware families

    DEEP-LEARNING APPROACHES TO PREDICT REMAINING USEFUL LIFE OF HARD DISKS

    Get PDF
    On a daily basis, data centers process huge volumes of data using inexpensive hard disks. Data stored in these disks serve a range of critical functional needs from financial, and healthcare to aerospace. As such, premature disk failure and consequent loss of data can be catastrophic. To mitigate the risk of failures, cloud storage providers perform condition-based monitoring and replace hard disks before they fail. By estimating the remaining useful life (RUL) of hard disk drives, one can predict the time-to-failure of a particular device and replace it at the right time, ensuring maximum utilization whilst reducing operational costs. We analyze 10-years worth of data across manufacturers and understand failure trends. In this work, we aim to look at several deep learning architectures ranging for LSTMs to Transformers to predict the RUL of hard disk for better predictive maintenance

    Explaining Misinformation Detection Using Large Language Models

    Get PDF
    Large language models (LLMs) are a compressed repository of a vast corpus of valuable information on which they are trained. Therefore, this work hypothesizes that LLMs such as Llama, Orca, Falcon, and Mistral can be used for misinformation detection by making them cross-check new information with the repository on which they are trained. Accordingly, this paper describes the findings from the investigation of the abilities of LLMs in detecting misinformation on multiple datasets. The results are interpreted using explainable AI techniques such as Local Interpretable Model-Agnostic Explanations (LIME), SHapley Additive exPlanations (SHAP), and Integrated Gradients. The LLMs themselves are also asked to explain their classification. These complementary approaches aid in better understanding the inner workings of misinformation detection using LLMs and lead to conclusions about their effectiveness at the task. The methodology is generic and nothing specific is assumed for any of the LLMs, so the conclusions apply generally. Primarily, when it comes to misinformation detection, the experiments show that the LLMs are limited by the data on which they are trained

    Remote Sensing with Airborne Infrared Thermography for Assessment of Landscape Scale Wildfire Spread and Intensity

    Get PDF
    Wildland fire is one of the most complex environmental physical processes to quantify. The way in which we observe and measure wildland fire is critical in understanding how these processes drive fire dynamics. Fire behavior can be observed in several ways including the use of ground sensors and remote sensing packages aboard aircraft and satellites. Airborne sensors provide high spatial resolution and can provide high temporal resolution. Infrared thermography takes advantage of radiant heat transfer allowing for the study of fire characteristics such as fire spread and intensity. Aircraft observations utilizing infrared cameras have been a well-established method of fire observation, yet there is still a significant lack of comprehensive data available. Presented in this research are several advances in the use of infrared remote sensing techniques to evaluate fire behavior. First, a synthesis of knowledge and methods regarding the use of infrared camera systems to measure fire behavior is discussed and analyzed. Second, is the use of airborne infrared data collected during tactical firefighting operations to develop an automated method of extracting active fire edges for fire spread analysis. Third, is an analysis of high-resolution fire behavior data collected during several wildfires in the 2022 fire season and the new methods used to process the images, calculate fire radiative power, and evaluate fire spread. These methods are a contribution to the advancement of being able to use operational firefighting data for research applications, furthering automated methods of processing large data quantities to evaluate fire spread, and evaluating high-resolution landscape scale wildfire behavior

    27,861

    full texts

    32,584

    metadata records
    Updated in last 30 days.
    SJSU ScholarWorks
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇