SHAREOK Repository
Not a member yet
49261 research outputs found
Sort by
Sensitive Attribute Association Bias in Latent Factor Recommendation Algorithms: Theory and In Practice
This dissertation presents methods for evaluating and mitigating a relatively unexplored bias topic in recommendation systems, which we refer to as attribute association bias. Attribute association bias (AAB) can be introduced when leveraging latent factor recommendation models due to their ability to entangle model and implicit attributes into the trained latent space. This type of bias occurs when entity embeddings showcase significant levels of association with specific types of explicit or implicit entity attributes, thus having the potential to introduce representative harms for both consumer and provider stakeholders. We present a novel analysis method framework to help practitioners evaluate their latent factor recommendation models for AAB. This framework consists of three main techniques for gaining insight into sensitive AAB in the recommendation latent space: bias direction creation, bias evaluation metrics, and multi-group evaluation. Methods within our evaluation framework were inspired by techniques presented by the natural language processing research community for measuring gender bias in learned language representations. Additionally, we explore how this bias can be reinforced and produce feedback loops via retraining. Finally, we explore possible mitigation techniques for addressing said bias. Primarily, we demonstrate our methodology with two case studies that evaluate user gender association bias in latent factor recommendation. With our methods, we uncover the existence of user gender association bias and compare the various methods we propose to help guide practitioners in how best to use our techniques for their systems. In addition to exploring user gender, we experiment with measuring user age association bias as a means for evaluating non-binary AAB
Exploration of previous science experiences and sociocultural factors that influence science literacy of first-year college students
Science literacy is an important element for individuals in any society. How a person engages with scientific information can help or hinder the progress of society. Science literacy can advance society through evolution of technology, advancements in healthcare, and comprehensive ways to protect the planet. While a scient literate society itself should be a global goal, modern science is rooted in Eurocentric ways of knowing. It is important to understand that the differences in social and cultural experiences play a part in how people see and interact with the world around them. The purpose of this study was to examine relationships between previous science experiences, sociocultural experiences, and science literacy of first-year college students.
To gain a better understanding of factors that contribute to science literacy, an anonymous survey was distributed to first year college students at a single university. The survey was designed using the embedded method which included both quantitative and qualitative questions which pertained to participants’ demographics, previous science experiences, sociocultural experiences, as well as a science literacy assessment. Correlation analysis was utilized to assess quantitative responses. Reflexive thematic and sentiment analysis was used to analyze qualitative responses. Lastly, critical discourse analysis was employed to investigate instances of identity, agency, and power, identified within open-ended responses.
Results indicated that there were a greater number of relationships between sociocultural experiences and science literacy than between previous science experiences and science literacy. Furthermore, critical analysis identified several instances of identity, agency, and power in participant responses
Single-probe Single Cell Mass Spectrometry Studies: Investigation of Cell Heterogeneity and Quantification of Intracellular Small Molecules
Studying cell heterogeneity can provide a deeper understanding of biological activities, but corresponding studies cannot be performed using traditional bulk analysis methods. The development of diverse single cell bioanalysis methods is in urgent need and of great significance. Mass spectrometry (MS) has been recognized as a powerful technique for bioanalysis for its high sensitivity, wide applicability, label-free detection, and capability for quantitative analysis. The paramount significance of single cell mass spectrometry (SCMS) techniques have been recognized, and they are becoming indispensable tools in fundamental research and studies of human diseases such as cancers and infectious disease. My studies consist of two major parts: (1) the development novel method to quantify nitric oxide (NO) using combined chemical reactions and SCMS techniques and (2) the investigation of cell heterogeneity using integrated bioinformatics tools and SCMS methods.
In Chapter one, we reviewed the development of single cell mass spectrometry (SCMS) field and summarized multiple existing SCMS techniques. We also included the methods that have been used for quantitative studies of small molecules in single cells. In particular, we further developed the Single-probe, a microscale device that is ideally suited for SCMS study of live single cells under ambient environment, for molecular quantification in single cells. In Chapter two, the single-probe SCMS was coupled with chemical reactions to detect and quantify nitric oxide (NO) in single cells. We then performed detailed data analysis to study the subpopulations of cells based on their NO expression levels. In Chapter three, cellular heterogeneity in infectious disease was revealed using the Single-probe SCMS, and we discovered the bystander effect of cells, which are uninfected cells adjacent to infected cells. In Chapter four, we developed a novel data analysis method for assessing the global metabolomic profiles from the SCMS experiments, allowing us to identify subpopulations and determine the number of subpopulations without prior knowledge. Finally, in Chapter five, a new machine learning method was applied to classify cells with different drug resistant levels
Uncovering the Potential of Federated Learning: Addressing Algorithmic and Data-driven Challenges under Privacy Restrictions
Federated learning is a groundbreaking distributed machine learning paradigm that allows for the collaborative training of models across various entities without directly sharing sensitive data, ensuring privacy and robustness. This Ph.D. dissertation delves into the intricacies of federated learning, investigating the algorithmic and data-driven challenges of deep learning models in the presence of additive noise in this framework. The main objective is to provide strategies to measure the generalization, stability, and privacy-preserving capabilities of these models and further improve them. To this end, five noise infusion mechanisms at varying noise levels within centralized and federated learning settings are explored. As model complexity is a key component of the generalization and stability of deep learning models during training and evaluation, a comparative analysis of three Convolutional Neural Network (CNN) architectures is provided. A key contribution of this study is introducing specific metrics for training with noise. Signal-to-Noise Ratio (SNR) is introduced as a quantitative measure of the trade-off between privacy and training accuracy of noise-infused models, aiming to find the noise level that yields optimal privacy and accuracy.
Moreover, the Price of Stability and Price of Anarchy are defined in the context of privacy-preserving deep learning, contributing to the systematic investigation of the noise infusion mechanisms to enhance privacy without compromising performance. This research sheds light on the delicate balance between these critical factors, fostering a deeper understanding of the implications of noise-based regularization in machine learning. The present study also explores a real-world application of federated learning in weather prediction applications that suffer from the issue of imbalanced datasets. Utilizing data from multiple sources combined with advanced data augmentation techniques improves the accuracy and generalization of weather prediction models, even when dealing with imbalanced datasets. Overall, federated learning is pivotal in harnessing decentralized datasets for real-world applications while safeguarding privacy. By leveraging noise as a tool for regularization and privacy enhancement, this research study aims to contribute to the development of robust, privacy-aware algorithms, ensuring that AI-driven solutions prioritize both utility and privacy
Examining the Predictability of Tornadic and Nontornadic Non-Supercellular MCS Storms using GridRad-Severe Radar Data and Machine Learning Techniques
Many studies have aimed to identify novel storm characteristics that are indicative of current or future severe weather potential using a combination of ground-based radar observations and severe reports. However, this is often done on a small scale using limited case studies on the order of tens to hundreds of storms due to how time-intensive this process is. Herein, we introduce the GridRad-Severe dataset, a database including ~100 severe weather days per year and upwards of 1.3 million objectively tracked storms from 2010-2019. Composite radar volumes spanning objectively determined, report-centered domains are created for each selected day using the GridRad compositing technique, with dates objectively determined using report thresholds defined to capture the highest-end severe weather days from each year, evenly distributed across all severe report types (tornadoes, severe hail, and severe wind). Spatiotemporal domain bounds for each event are objectively determined to encompass both the majority of reports as well as the time of convection initiation. Severe weather reports are matched to storms that are objectively tracked using the radar data, so the evolution of the storm cells and their severe weather production can be evaluated. Herein, we apply storm mode (single cell, multicell, or mesoscale convective system) and right-moving supercell classification techniques to the dataset, and revisit various questions about severe storms and their bulk characteristics posed and evaluated in past work. Additional applications of this dataset are reviewed for possible future studies.
Given this large dataset of severe storms, questions about storm structure of very specific storm types can be investigated using what is still a large subsample of the total GridRad-Severe dataset. This study compares populations of tornadic non-supercellular MCS storm cells to their nontornadic counterparts, focusing on nontornadic storms that have similar radar characteristics to tornadic storms. Comparison of single-polarization radar variables during storm lifetimes show that median values of low-level, mid-level, and column-maximum azimuthal shear, as well as low-level radial divergence, enable the highest degree of separation between tornadic and nontornadic storms. Focusing on low-level azimuthal shear values, null storms were randomly selected such that the distribution of null low-level azimuthal shear values matches the distribution of tornadic values. After isolating the null cases from the nontornadic population, signatures emerge in single-polarization data that enable discrimination between nontornadic and tornadic storms. In comparison, dual-polarization variables show little deviation between storm types. Tornadic storms both at tornadogenesis and at 20-minute lead time show collocation of the primary storm updraft with enhanced near-surface rotation and convergence, facilitating the non-mesocyclonic tornadogenesis processes.
With this additional knowledge about the structure of tornadic vs. nontornadic storms and which radar variables best differentiate the two, machine learning methods can be used to learn the differences between these storm type at various lead times and improve tornado predictability. A convolutional neural network was trained on tornadic and nontornadic data where the nontornadic data were either sampled from storms that have similar radar characteristics to tornadic storms as in the PMM analyses or sampled from the entire population of non-supercellular MCS storms. These models were then tested on independent data from 2020-2021, again either including all tornadic storms and sampling nontornadic cases as in the PMM analyses or including all tornadic and nontornadic storms. Models that were tested on all tornadic and nontornadic storms, whether they were trained and validated on datasets including sampled strong nontornadic storms or a sample of all nontornadic storms, both performed well below the baseline performance metrics from the NWS. However, when the model was trained, validated, and tested using samples of all tornadic storms and only strong nontornadic storms, model test performance far exceeded the baseline NWS metrics. Performance metrics include a probability of detection (POD) of 79%, a false alarm ratio (FAR) of 58%, and a CSI of 0.38. Compared to the NWS metrics of 49%, 75%, and 0.2, respectively, this model shows clear promise as a supplemental forecasting tool for scenarios where a storm is identified as (at least) borderline tornadic. However, further analyses of the model performance scaled to account for the true proportion of tornadic vs. nontornadic storms shows that it was the unnatural ratio of tornadic to nontornadic storms, and not the focus on strong nontornadic storms, that was the cause for the improved model performance.
Finally, a brief analysis of the underlying populations and their demographic characteristics in the vicinity of tornadoes are examined. Special attention is given to non-supercellular MCS storms, as well as discrete supercells, whose tornadoes are often a main focus of tornado research in the U.S. Analyses show that groups making up ~3% or less of the CONUS mean population typically have lower relative population densities in the vicinity of storms. The Black or African American Alone demographic has higher relative populations in the vicinity of all tornadoes compared to their CONUS mean population density, as do all Non-Hispanic categories (Not Hispanic, Non-Hispanic White and Non-Hispanic Black). Comparing population densities near specific types of tornadoes (i.e., mode and combination of mode and human impact) to their densities near all tornadoes, the White Alone demographic has population densities near the CONUS mean for supercellular tornadoes, but that density jumps 6-7 percentage points in the vicinity of deadly supercellular tornadoes when examining underlying population density by deadly event and by death, suggesting that the deadliest supercellular tornadoes occur in predominantly White areas. On average, populations in the vicinity of all tornadoes have ~75-80% higher Black or African American Alone and Non-Hispanic Black densities when compared to the CONUS mean, with those demographics' relative densities only increasing when isolating MCS tornadoes and deadly MCS tornadoes, suggesting that the deadliest MCS tornadoes preferentially occur in areas with relatively higher Black or African American Alone and Non-Hispanic Black populations. One particularly striking result is that the mean Social Vulnerability Index (SVI) of populations near all tornadoes is just barely above the CONUS mean (0.52 vs. CONUS mean of 0.51), but is slightly lower for supercellular tornadoes (0.49) and higher for MCS tornadoes (0.57). Therefore, MCS tornadoes tend to occur in areas that are less resilient to natural disasters than both the CONUS mean and areas in the vicinity of supercellular tornadoes. For both MCS and supercellular tornadoes that were associated with deaths or injuries, the local SVI is higher, likely pointing to the applicability of SVI in identifying areas less resilient to natural disasters
Exploring cell heterogeneity: applications of single-probe single cell mass spectrometry for subpopulation analysis
Mass spectrometry (MS) has become an indispensable tool for transformed metabolomics studies, whereas exploring single-cell metabolomic profiles remains a challenge due to limited techniques and suitable algorithms. This dissertation delves into cell heterogeneity, including method development and applications to infectious disease, with a focus on Chagas disease caused by Trypanosoma cruzi. Using the effective Single-probe SCMS technique combined with a fixation method that can safely decontaminate samples, we examined individual cell responses during parasite infection. Our findings unveiled significant differences in cell metabolism, even in neighboring uninfected cells, shedding light on the broader impact of infection. This pioneering study, utilizing bioanalytical SCMS, offers versatile tools to understand infectious diseases and the complexities of cell behavior in diseases.
Additionally, this dissertation tackles these challenges by merging the Single-probe single-cell MS (SCMS) technique with SinCHet-MS, a specialized bioinformatics software package. This combination allowed us to understand cell diversity, quantify cell subgroups, and identify key metabolites representing cells in subpopulations. Testing this approach with melanoma cancer cell lines revealed new subgroups after drug treatment, showcasing the potential for in-depth exploration of cell diversity and marker identification. This label-free method enhances our comprehension of cell metabolism in diseases and therapeutic responses.
Furthermore, this research pioneers a novel approach by integrating CRISPR-Cas9 gene editing with the Single-probe SCMS metabolomics, focusing on FASN-knockout cells in the human cell model HEK293T. This innovative strategy provides valuable insights into gene-metabolite interactions at the single-cell level. By combining advanced SCMS techniques with gene editing, this dissertation opens new avenues for understanding gene editing efficiency and the complex relationship between genes and cell metabolomics. These integrated methods advance our understanding of cell diversity in cancer, infectious diseases, and gene therapy, offering a fresh perspective for future research and therapeutic intervention
The Concern about Losing Face and Social Anxiety: The Mediating Roles of Self-Compassion and Autonomy
Social anxiety is a prevalent mental health challenge among college students. Prior research has documented various antecedents of social anxiety, with one of them being the concern about losing face. Yet, less is known about the factors that could explain the link between concern about losing face and social anxiety. This study explores the mediating roles of self-compassion, and autonomy. A sample of 180 college students completed self-report measures of the variables of interest. The serial mediation model of concern about losing face on social anxiety, mediated by self-compassion and autonomy, explained 46% of the variance in social anxiety. Results suggested that individuals who reported having concerns about losing face were more likely to report experiencing social anxiety. This relationship was mediated by a lack of self-compassion and the failure to satisfy the fundamental psychological need for autonomy, which the model suggests leads to the experience of social anxiety. The current study provides an explanation for the link between concern about losing face and social anxiety in American college students, and it offers empirical support for the Basic Psychological Need and the Self-Compassion theories
Object detection in dual-band infrared
Dual-Band Infrared (DBIR) offers the advantage of combining Mid-Wave Infrared (MWIR) and Long-Wave Infrared (LWIR) within a single field-of-view (FoV). This provides additional information for each spectral band. DBIR camera systems find applications in both military and civilian contexts. This work introduces a novel labeled DBIR dataset that includes civilian vehicles,
aircraft, birds, and people. The dataset is designed for utilization in object detection and tracking algorithms. It comprises 233 objects with tracks spanning up to 1,300 frames, encompassing images in both MW and LW.
This research reviews pertinent literature related to object detection, object detection in the infrared spectrum, and data fusion. Two sets of experiments were conducted using this DBIR dataset: Motion Detection and CNNbased object detection. For motion detection, a parallel implementation of the Visual Background Extractor (ViBe) was developed, employing ConnectedComponents analysis to generate bounding boxes. To assess these bounding boxes, Intersection-over-Union (IoU) calculations were performed. The results demonstrate that DBIR enhances the IoU of bounding boxes in 6.11% of cases within sequences where the camera’s field of view remains stationary. A size analysis reveals ViBe’s effectiveness in detecting small and dim objects within
this dataset.
A subsequent experiment employed You Only Look Once (YOLO) versions 4 and 7 to conduct inference on this dataset, following image preprocessing. The inference models were trained using visible spectrum MS COCO data. The
findings confirm that YOLOv4/7 effectively detect objects within the infrared spectrum in this dataset. An assessment of these CNNs’ performance relative to the size of the detected object highlights the significance of object size in detection capabilities. Notably, DBIR substantially enhances detection capabilities in both YOLOv4 and YOLOv7; however, in the latter case, the number
of False Positive detections increases. Consequently, while DBIR improves the recall of YOLOv4/7, the introduction of DBIR information reduces the precision of YOLOv7.
This study also demonstrates the complementary nature of ViBe and YOLO in their detection capabilities based on object size in this data set. Though this is known prior art, an approach using these two approaches in a hybridized configuration is discussed. ViBe excels in detecting small, distant objects, while YOLO excels in detecting larger, closer objects. The research underscores that DBIR offers multiple advantages over MW or LW alone in modern computer vision algorithms, warranting further research investment
General survey of Native American participation in the Vietnam War
Native American participation during the Vietnam War is a subject woefully understudied, as whole, with a majority of the historiography consisting of one man’s work. This general survey of Native American/Indian/Indigenous American participation in the Vietnam War seeks to expand the current understanding of the experiences of Native Americans servicemen and women during the Vietnam era, with special focus paid to Native Servicemen stationed in West Germany and the views of non-Native servicemen towards their Native counterparts. This research shows that some of the more ubiquitous features in prior scholarship on the topic, such as ceremonies meant to send off and welcome back returning warriors, may be less common than previously thought. It also shows that discrimination faced by soldiers varied based on where they were stationed, opening up a new area of study for future research