95038 research outputs found
Sort by
Explaining Unintuitive Feature Importance Explanations
Explainable AI (XAI) research has proposed a variety of approaches for explaining machine learning model predictions to end-users. Feature importance explanation, which highlights input features that are most influential to the output, is a popular and effective approach. Studies have developed different algorithms to identity important features, and also empirically shown that such explanations can improve end-users' decision-making performance and understanding of the AI system in certain tasks. Despite these promising advancements, a critical gap remains open: features deemed predictive by algorithms can appear unintuitive to humans. For instance, prior studies found that the word ``problems'' predicts positive sentiment in product reviews, and the word ``Chicago'' is strongly associated with deceptive hotel reviews. However, most XAI research on feature importance explanations does not further explain why a feature is predictive. Therefore, this dissertation explores an underexplored area in XAI research: explaining unintuitive feature importance explanations. Using deception detection as a case study, this dissertation investigated two primary research questions. The first research question focused on developing a computational approach to explaining predictive but unintuitive words in deception detection. While previous research has used contextual information to explain unintuitive words in sentiment analysis, those in deception detection often represent some underlying phenomena that are too complex to be captured by local context. To this end, I leveraged a large language model (LLM) to conjecture the phenomena associated with words predictive of a review being genuine or deceptive, and then conducted algorithmic evaluations to validate the technical reliability of this LLM-based approach. The evaluation results confirmed that the LLM used in this dissertation (GPT-4o) generated non-hallucinated and generalizable phenomena that effectively explained why a word is predictive of genuine or deceptive reviews. The second research question focused on evaluating how unintuitive words and LLM-generated explanations influence participants in a decision-making task. To address this research question, I conducted a crowdsourced user study (N=220). In the study, participants first interacted with an AI system and judged a sequence of eight Chicago hotel reviews. Participants were randomly assigned to one of five interface conditions (i.e., a between-subjects design). The interface conditions varied in terms of the AI assistance features made available to participants. Then, participants independently completed a task of detecting deceptive hotel reviews from other cities. The study results found that showing predictive words without explaining why they are predictive was no better than not showing them at all, while explaining why these words are predictive significantly enhanced participants' learning of the task, appropriate reliance on AI assistance, and perceptions of the AI system. In summary, this dissertation makes three key contributions: (1) it sheds light on the critical issue of unintuitive feature importance explanations and the urgent need for making them understandable to end-users; (2) it introduces a conjecture-then-validate pipeline for explaining unintuitive words through LLMs; (3) it provides empirical insights into how unintuitive words and LLM-generated explanations influence end-users during decision-making. This dissertation offers an important design implication for future XAI research: machine-generated explanations should be aligned with human intuition and common sense to facilitate effective human-AI interaction.Doctor of Philosoph
DEFINING THE IMPACT OF CHRONIC ATF4 SIGNALING ON CD8+ T CELLS IN THE TUMOR MICROENVRIOMENT
Coral del Mar Alicea Pauneto: Defining the impact of chronic ATF4 signaling on CD8+ T cells in The Tumor Microenvrioment (Under the direction of Jessica Thaxton) While cancer immunotherapy has revolutionized oncology treatment, the hostile tumor microenvironment (TME) often hinders its efficacy. Characterized by metabolic dysregulation and oxidative stress, the TME impairs CD8+ tumor-infiltrating lymphocyte (TIL) function, limiting immunotherapy efficacy. Despite experiencing profound cellular stress in the TME, the key cell-intrinsic mechanisms driving CD8+ TIL dysfunction remain poorly understood. Here, we investigate how CD8+ TILs respond to oxidative stress driven by the hypoxic TME through the central transcription node of the integrated stress response (ISR), activating transcription factor 4 (ATF4). We explore the paradoxical nature of ATF4, which initially possesses a protective response that transcribes antioxidant genes to restore homeostasis and chaperones to restore cellular homeostasis. However, in chronic and persistent stress ATF4 response is linked with poor prognosis and cellular death. Our findings demonstrate that sustained ATF4 activity, driven by the hypoxic TME, results in metabolic polarization that reduces CD8+ TIL antitumor function. Specifically, chronic ATF4 activation drives persistent citric acid cycle activity, resulting in mitochondrial defects and accumulation of mitochondrial reactive oxygen species (ROS), culminating in T cell death. We identified that attenuation of ATF4 resulted in the maintenance of functional CD8+ T cell responses, enabling enhanced response to checkpoint therapy. This work identifies persistent ATF4 activity as a critical barrier to effective T cell-based immunotherapies, highlighting the therapeutic potential of targeting the ISR to improve the efficacy of immunotherapy combinations.Doctor of Philosoph
Scalable Statistical Analysis and Improved Missing Data Imputation for High-Dimensional Omics Data
Due to the special structure of large p and small n, with covariance structures between subjects in thecurrent generation omics dataset, many statistical and computational algorithms face challenges of scalabilityin higher dimension.We develop novel approaches to solve three particular problems in high-dimensional settings: i) scale-upthe computational efficiency of association analysis in twin studies; 2) improve the imputation performancefor omics data by combining existing algorithm; 3) develop novel imputation algorithm via transfer-learningand fine-tuning approaches to address the block-missing problem in cross-omics cohort studies.In the first project, we address a common problem in twin studies that have been widely used in researchesof inheritable diseases and traits. GWAS of multiple traits, such as eQTL studies in twins, requires associationtests between thousands of transcripts and millions of single nucleotide polymorphisms (SNPs). Standardmethods such as mixed-effects models are extremely computationally inefficient and impractical in the currentgeneration of computers. We introduce TwinEQTL, a computationally efficient alternative to commonly usedmixed effects models in the analysis of eQTL and GWAS in twin studies.In the second project, we move to missing data problem. Missing data is a common problem in variousrandomized and observational studies. Several approaches have been developed to handle missing data, andrecently multiple imputation has become increasingly popular. Relying on the sophisticated associationsamong the variables in high-dimensional setting, we propose a novel two-phase approach to impute themissing entries by combining existing imputation algorithms. The proposed method can impute continuousdata in high-dimensional space and has robust performance in different levels of missing rate.In the final project, we aim to address the persistent issue of block-missingness in omics data byleveraging pre-trained machine learning models for imputation, drawing on knowledge from large, existingmulti-omics datasets. Our approach utilizes existing large-scale cross-omics studies and advanced machinelearning algorithms to build robust and efficient models that can be pre-trained and fine-tuned for applicationto new datasets. Using transfer learning, we adapt and apply intricate patterns and dependencies learned fromexternal data sources, enabling more accurate and context-aware imputation in target studies.Doctor of Philosoph
BIOSOCIAL MECHANISMS OF ACCELERATED AGING: PERSISTENT INFECTIONS AND THE IMMUNE SYSTEM
Persistent infections and immune system aging are increasingly recognized as critical mechanisms underlying the development of age-related disease and health disparities. Yet, limited research has examined how these processes unfold prior to midlife or how socioeconomic disadvantage spanning early life and young adulthood may shape infection susceptibility, burden, and immune aging trajectories. This dissertation investigates the associations between life-course socioeconomic status (SES), persistent infections, and biomarkers of immunosenescence and biological aging in a nationally representative cohort of U.S. adults in early midlife, aiming to advance understanding of biosocial pathways that contribute to accelerated aging.Using data from the National Longitudinal Study of Adolescent to Adult Health (Add Health, we investigated these relationships in young adulthood and early midlife. We examined persistent infection outcomes (seropositivity and IgG antibody concentrations for CMV, EBV, HSV-1, and H. Pylori) in Wave IV (median age 28), and immune cell distributions and epigenetic markers of biological age acceleration at Wave V (median age 38). We estimated the associations between life-course socioeconomic disadvantage and both persistent infection and cellular immunosenescence outcomes, as well as associations between infection measures and cellular immunosenescence and biological age acceleration. Multivariable linear and logistic regression models were used to examine these associations. Survey weights and multiple sensitivity analyses were applied to assess generalizability and robustness of results.Results showed that persistent SES disadvantage from adolescence to young adulthood was strongly associated with higher odds of CMV, HSV-1, and H. Pylori infection, elevated CMV and HSV-1 IgG antibody concentrations, and more aged immune cell profiles as measured by increased CD4+ and CD8+ memory-to-naïve T cell ratios. Further, CMV emerged as the infection most robustly associated with both immunosenescence and epigenetic age acceleration. While other infections (HSV-1, EBV, H. Pylori) demonstrated weaker and more variable associations, they were still linked to epigenetic aging outcomes.These findings underscore that biological aging processes begin manifesting well before midlife and are shaped by interacting features of social environment and biological determinants. The study highlights the importance of integrating social and biomedical frameworks in aging research and public health, and points to early-life intervention as a promising strategy to reduce infection-related immune decline and promote healthy aging trajectories across the life course.Doctor of Philosoph
The Development and Pilot Test of a Culturally Relevant and Bilingual Math Game - I Am (Apply Math) In My World.
Stereotype threat experienced by African American and Hispanic students contributes to elevated levels of math anxiety and low math self-efficacy. Elementary school is a critical period to cultivate quantitative literacy and social-emotional skills that form the foundation for future success in mathematics and STEM career pathways. This dissertation uses Human-Centered Design (HCD) to develop and evaluate a culturally relevant, bilingual, web-based math game aimed at reducing math anxiety and enhancing math self-efficacy among African American and Hispanic students in 3rd-5th grade. Informed by the principles of Culturally Relevant Pedagogy (CRP), Transformative Social Emotional Learning (TSEL), and alignment with North Carolina Math Standards, this project involved collaboration with students, parents, teachers, mental health providers, and interdisciplinary experts in computer science, math education, and psychology. This dissertation comprises three manuscripts: (1) a content analysis of existing free online math games to evaluate their incorporation of CRP, TSEL, and math standards; (2) a description of the HCD-driven development of the intervention through co-designer input and iterative refinement; and (3) a pilot feasibility study testing the game’s usability, acceptability, and preliminary impact on math anxiety and self-efficacy. Leveraging the educational potential of serious games and cultural assets, this work aims to produce a scalable, sustainable, and culturally grounded intervention to address systemic barriers to math engagement.Doctor of Philosoph
Divided We Stand? Stories of Polarization
This dissertation traces the origins and development of “polarization” discourses from their ancient Greek roots to their contemporary articulations with fears of a second civil war in the United States. Drawing on keywords and discourse analyses, I argue that “polarization” discourses have been part of how Americans make sense of social problems that range from antidemocratic attitudes to partisan conduct, racialized conflict, and a pervasive pessimism regarding the nation’s future. When “polarization” first became applicable to human societies in the 1930s, its discourses were shaped by the underlying metaphor of an unbroken line with two “poles,” and this metaphor became increasingly salient during the Cold War. Yet as the new millennium dawned, the underlying metaphor changed as journalists and commentators increasingly wrote of the missing center in American life. A line with two “poles” became two entirely disconnected worlds or realities, enabling the articulation of “polarization” and civil war. Since the mid-20th century, “polarization” discourses have become common sense, and their naturalization has been shaped by a shifting discursive formation that has, in the last century, taken on sociostatistical and computational forms. Narratives of “polarization” are difficult to disrupt because they are sustained by practices of knowledge production with a monopoly on truth. And while political and social scientists have contested the truth of a “polarized” society, their efforts are hampered by the immense amount of commonsense work that “polarization” discourses do.Doctor of Philosoph
Working memory in Search and Sensemaking: An Exploratory Study
Working memory (WM) is involved in cognitive tasks such as reading comprehension, reasoning, and learning. Searching and sensemaking is not an exception---a range of cognitive activities are involved in the process of finding information and making sense of information. Prior work in interactive information retrieval (IIR) has found that WM impacts searchers' perceptions, behaviors, and outcomes. However, little work has been done to gather insights into how WM might affect the search and sensemaking process. This dissertation investigates the influence of WM on search and sensemaking along three dimensions: perceptions, behaviors, and learning outcomes. 60 participants were recruited for a working memory span task. Twenty-two individuals were selected from the upper and lower bounds of the score range to form the high- and low-WM groups, respectively. The influence of WM was explored by comparing these two groups across the aforementioned three dimensions. A total of 44 participants completed a learning-oriented search task, followed by a summary writing session. To examine search and sensemaking (SSM) activities, as well as cognitive processes, a combination of participants' screen activities and think-aloud comments were collected and analyzed through qualitative methods. The major findings of the study are as follows. First, in terms of search behaviors, as indicated by the search log data, low-WM participants exhibited more abandoned queries, spent more time on the SERP, but less time on landing pages and the note-taking page. Second, in terms of SSM activities, low-WM participants issued more data-driven queries, and high-WM participants engaged more with specific sensemaking activities, including accretion (adding information to their notes), instantiation (integrating new information into existing notes to refine, expand, or exemplify), and semantic fit (assessing how new information fits within the existing notes or goals). In terms of cognitive activities, high-WM participants exhibited greater levels of monitoring and active maintenance (keeping certain pieces of information active while searching and reading). Finally, in terms of learning outcome, high-WM participants demonstrated greater learning outcomes. Knowledge summaries were scored based on the number of correct statements recalled, as well as the breadth and depth of correct statements recalled. Interestingly, despite the observed differences in behaviors and learning outcomes, there was no difference in participants' perceptions of their task performance or perceived change in knowledge. This research has a few contributions. First, it goes beyond the question of whether WM affects search and sensemaking. It asks how by examining the cognitive and behavioral activities participants engage in during a learning-oriented search task. By leveraging qualitative techniques, this study reveals the specific ways in which WM might influence the search and sensemaking process, from query formulation to semantic fit. Second, this study employs a diverse set of data types, including search logs, survey data, observations, and learning outcomes. The combination of these data types enabled a more comprehensive, holistic understanding of the role of WM in search and sensemaking. Third, the findings of this study offer insights into how individuals with conditions related to WM capacity, such as those with ADHD, dyslexia, or age-related cognitive decline engage with complex information tasks. This understanding of the cognitive, metacognitve, and sensemaking activities involved in a learning-oriented task can inform the design of more accessible and effective information systems, digital learning tools, information literacy education materials, and interventions.Doctor of Philosoph
GOT TODDLER MILK? DESCRIBING AND EXAMINING THE MARKETING, PURCHASING AND PROVISION OF TODDLER MILK PRODUCTS
ABSTRACTIntroduction: Toddler milk products are the fastest growing category of commercial milk formulas globally. Marketed to children aged 12 to 36 months, these products typically contain added sugars and less protein than cow’s milk. Despite public health recommendations discouraging their use, toddler milk products are frequently promoted as essential for growth and development. This dissertation examines the marketing, consumption, nutritional composition, and purchasing patterns of toddler milk products.Methods: This dissertation uses systematic review, content analysis, and observational methods to examine how toddler milk products are marketed and purchased. Study 1 is a scoping review synthesizing international evidence on caregivers’ beliefs, feeding practices, and the influence of toddler milk marketing. Study 2 uses content analysis to examine front-of-package marketing strategies and nutritional content of products sold across the Americas. Study 3 is a repeated cross-sectional analysis of a nationally representative US consumer panel, assessing trends in toddler milk purchasing, volume, and expenditures from 2013 to 2023.Results: Study 1 found that toddler milk sales are increasing globally. Packaging often resembles infant formula and includes claims and imagery that may contribute to caregivers’ misperceptions. Young caregivers with higher income and education levels and from minority populations are more likely to provide toddler milk. Study 2 found that nearly all toddler milk products featured at least one front-of-package claim, with nutrition claims most common (on 97% of packages), followed by health claims (56%), and other marketing claims (39%). Visual marketing appeared on 92% of packages, with 21% depicting humans, and many using symbols such as shields (32%) and hearts (27%). Nutritionally, toddler milk products contained more sugar and less protein than cow’s milk. Study 3 found that toddler milk purchasing in the US remained relatively stable over the past decade. In 2023, Black or African American-headed households were more likely to purchase toddler milk products than White-headed households, while divorced, separated, or widowed households were less likely than married ones to purchase them.Conclusions: Toddler milk marketing is widespread and can be misleading. Although US purchasing remains stable, clear labeling regulations and public health messaging are needed to address inappropriate marketing practices.Doctor of Philosoph
MARINE MICROBIAL COMMUNITY STRUCTURE AND METABOLIC POTENTIAL VIA QUANTITATIVE METAGENOMICS
Marine microbes play central roles in global biogeochemical cycles, yet much about their ecology remains poorly understood due to challenges in studying diverse and uncultured microbial populations. While culture-independent approaches have improved our understanding of microbial diversity and metabolic capacities, they often lack the resolution and quantitative power needed to track individual populations and assess their ecological roles. This dissertation addresses this gap by developing a novel quantitative genome-resolved metagenomic framework that combines metagenome-assembled genomes (MAGs) with internal standard normalization to quantify the absolute abundance of microbial populations in environmental samples. In Chapter 1, I applied this approach to study the microbial communities at the Galápagos Archipelago, where strong environmental gradients were created by interacting oceanic features as well as the 2015/16 El Niño event. Microbial populations exhibited distinct spatial and temporal patterns linked to environmental variability and their genomic traits reflected niche differentiation. In Chapter 2, I applied this framework to a temperate estuarine-to-coastal continuum in North Carolina, USA, and resolved the genomic content for half of the microbial cells, gaining unprecedented insight into the structure and function of the microbial communities. Pronounced seasonal shifts among closely related populations were observed along the transect, suggesting temporal and spatial niche partitioning. In Chapter 3, I combined MAGs with quantitative metatranscriptomics to investigate vitamin B1 metabolism in coastal microbial communities. B1 auxotrophy was widespread within the communities but B1 synthesis gene abundances remained stable, indicating a balanced supply and demand in the environment. By integrating MAG B1 genotypes and transcription, active B1 producers and consumers were identified. Nutrient amendment experiments suggested that B1 addition did not stimulate overall bacterial growth as B1 was not a limiting nutrient, but it might have promoted grazing pressure on the bacteria. Overall, this dissertation presents a framework for integrating microbial community structure, metabolic potential, and activity across time and space. By enabling the quantification of population-specific abundance and transcription, this work enhances our understanding of population-level dynamics, environmental drivers of microbial communities, and microbial metabolism, advancing our ability to link microbial ecology to biogeochemical impact.Doctor of Philosoph
Connecting Probabilistic and Mechanistic Frameworks for Causality to the Dynamic Intervention Response Problem
This work focuses on the general problem of predicting the response of a dynamic system to a previously unobserved intervention. We examine the philosophical foundations of this problem arguing that it relates to the interventionist definition of causality and may therefore benefit from a causal approach. We then survey probabilistic and mechanistic frameworks for causality through the lens of intervention response prediction, contrasting their various strengths and weaknesses and identifying potential synergies. We describe the usage of a new software package, Interfere, for simulating and predicting intervention response scenarios and summarize the creation of a diverse, extensible benchmark for evaluating the relative ability of predictive methods. We proceed to test five methods on said benchmark, taking state of the art applied mathematics tools for learning equations from data and extending their use to the causally motivated intervention response context. Our experiments study deterministic, stochastic, and confounded scenarios providing evidence that, (1) intervention response is markedly different from standard forecasting problems (2) that the sparse identification of non-linear dynamics algorithm, (or SINDy) performs better than all other methods tested on the benchmark sets and (3) that deep learning methods, while capable forecasters, do not appear to generalize well about intervention response.Doctor of Philosoph