DASH

Harvard University

DASH
Not a member yet
    71446 research outputs found

    Chosen: A story based on true events

    No full text
    This is a story that is neither memoir nor fiction, but a work of autofiction. Memoir tells a true story as the author remembers it—real life experiences shared with intention. A novel is a tale spun from the author’s imagination—invented characters and plot taking place in an invented world. But somewhere between the two is a middle ground—a story that blends reality with fiction—a story of things that happened and things that didn’t, people who were real and people who weren’t. Therein lies autofiction and my book, Chosen. Chosen is the story of a boy who grew up adopted. When young Chase Engel’s well-intentioned, but tough-minded adoptive parents cut off his college funding, it was the push he needed to search for information about his birth and the people who gave him away. Longing for an emotional connection and concrete answers about who he is, he finds both, but is thrust into a reality of betrayal, desperation and desire. Chase learns his conception was a mistake made by two reckless searchers—she was a shy, doe-eyed art student, and he a local television personality and married father of four. Like Chase, they each carried a heavy emotional burden without resolution, for decades, until the day all three meet as adults. The circumstances of that day—a meeting of strangers who shared both a profound attachment and a fateful void—are tangled with deeply intense and confusing emotions that changes each of their lives, and the lives of those closest to them, forever.Extension Studie

    Women in Fire: The Lived Experiences of Women in the Professional Fire Service in Northern Virginia

    No full text
    This thesis is a work of applied anthropology that explores the lived experiences of women in the professional fire service in Northern Virginia. Utilizing a qualitative phenomenological and autoethnographic methodology, this research is grounded in in-depth interviews with women firefighters, whose narratives illuminate the professional challenges and personal resilience in a field historically dominated by men. This study seeks to understand how entrenched gender biases shape women’s experiences and the implications for women firefighters’ integration within the fire service by addressing the questions, what values do women firefighters emphasize when trying to succeed at their jobs and why? Key findings reveal that women seek respect, trust, and acceptance to achieve personal and professional fulfillment. The critical influence of firehouse culture and leadership dynamics largely plays into women’s experiences on the job. Participants detailed the complexities of navigating respect, trust, and acceptance in environments often steeped in gender bias. Participants highlighted how career progression, leadership dynamics, and the psychological toll of gendered challenges illustrate the systemic barriers that hinder women from achieving respect, trust, and acceptance in their career. The study underscores how respect, trust, and acceptance are often contingent on women exceeding performance expectations while navigating discriminatory practices, cultural stereotypes, and inconsistent support systems. As an applied anthropological study, the findings aim to contribute to understanding how informal and formal leadership can either cultivate an inclusive and equitable culture or perpetuate exclusionary practices. The findings also highlight the need for further research into leadership practices, career trajectories, department policy, and the nuanced impacts of firehouse culture on the professional and personal lives of women in this vital field.Extension Studie

    Streaming Algorithms for Code Sparsification

    No full text
    The rapid growth of data collection and analysis in the 21st century has prompted the development of several new techniques in theoretical computer science. The streaming model, in which an algorithm has sequential access to a stream of input data and can only store a small fraction of the input in random access memory, has emerged as a powerful approach for processing datasets that are too large to store on a single system. From the perspective of algorithmic techniques, on the other hand, sparsification algorithms allow for the compression of large structures like graphs while preserving important global properties. These two ideas have a rich and interconnected history, and streaming algorithms for graph and hypergraph sparsification have become essential tools in the processing of massive datasets like social networks. Recently, a new framework called code sparsification has been developed as a method of generalizing existing sparsification algorithms through the lens of compressing matrices over finite fields. However, existing works only describe how to implement code sparsification in the classical computational setting, limiting its power. In this thesis, we develop the first algorithms for code sparsification in the insertion-only and dynamic streaming models. This provides a unified theory of graph and hypergraph sparsification over streams, and extends these techniques to allow for the space-efficient sparsification of Cayley graphs and a subset of constraint satisfaction problems.Computer Scienc

    Theoretical and Observational Developments of Inflation in Cosmology

    No full text
    This thesis takes modest steps towards the goal of understanding the earliest period in the history of the universe—inflation. In the first leg, we focus on loop effects in de Sitter space. Therein we explain the physical meaning of anomalous dimensions and demonstrate how to compute them systematically for a broad class of scalar theories. Ultimately focusing on phenomenology, we highlight how otherwise undetectable light scalars can be visible in the cosmological collider signal of sufficiently heavy particles via its anomalous dimension. Moreover, we demonstrate how strong infrared effects in de Sitter lead to an enhancement of this quantity. Then, we focus on a special class of light scalars called axions, or compact scalar fields. Such particles are very theoretically well motivated as they can resolve several problems within the standard model of particle physics, one of which is the strong CP problem. Being field theoretic analogues of the quantum mechanical particle on a circle, we show that the quantization of these scalars crucially depends on their gauge constraint, or periodicity, unlike their behavior in flat-space. We further illustrate that taking care for their quantization, in addition, has phenomenological consequences which could aid their detection through a cosmological collider channel as well. We also point out that strong infrared effects qualitatively change their primordial signal from the naive expectation for a light scalar in inflation. In the process, we organize the primordial bispectrum generated by a broad class of higher-loop processes into a novel and useful form. In the second leg, we turn our focus towards the late universe. We specifically utilize a careful compression of the three-point correlation function, known as the skew-spectrum, in order to efficiently extract the non- Gaussian information in late-universe tracers of the dark matter density. It is well known that tracers of the matter density in the late universe are highly non-Gaussian, and extracting this information using traditional means can quickly become highly expensive. The skew-spectra due to its compressed nature enables access to this non-Gaussian information with a dramatic improvement in speed. We first test the skew-spectra on simulations of lensing and galaxy cross-correlation analyses in order to estimate the galaxy bias parameters. Finally, we apply the skew-spectrum to the SDSS-BOSS galaxy catalog in order to measure fNLequilf_{\rm NL}^{\rm equil}, finding no evidence of primordial non-Gaussianity.Physic

    Woven into War: Digital Humanities and Sephardi Women in the Late Ottoman Empire and Its Diaspora, 1908 - 1945

    No full text
    This thesis presents OverText_{history}, a novel piece of text visualization software, that enables historians to enhance their research by making sentence-level comparisons of primary source material. It also presents a use case of OverText_{history} by way of a novel historical study of Sephardi Jewish women's experiences of wars during the 1908 - 1945 period in territories of the Ottoman Empire. Through a hybrid digital and analog investigation, I tell a complicated story of changing attitudes, communities, and social roles surrounding Sephardi women and argue that OverText_{history} is helpful for historical research in both technical and creative ways.Computer Scienc

    Causal Inference in Complex Observational Settings with Applications to Electronic Health Record Data

    No full text
    Electronic health record (EHR) data serve as a valuable source of real-world evidence for assessing treatment effects, offering rich, longitudinal patient data that can enhance clinical research, inform healthcare decisions, and ultimately improve patient outcomes. However, leveraging these complex data sources for reliable statistical inference remains challenging. For example, one common difficulty is the limited availability of readily validated clinical outcomes in EHR data, which often requires labor-intensive manual annotation through chart review. A related challenge arises when studies span multiple institutions, where differences in patient populations can introduce heterogeneity and potential biases. These complications further exacerbate the fundamental challenge of using observational data to draw valid causal conclusions. In this dissertation, we propose novel statistical methods for causal inference that address certain complexities relevant to clinical studies leveraging EHR data. In particular, we focus on two settings of interest: (1) semi-supervised learning, where outcome labels are scarce, and (2) transfer learning, where data must be integrated across diverse populations. We hope that the contributions of this dissertation offer meaningful steps toward unlocking the full potential of EHR data for clinical and translational research. In Chapter 1, we address the problem of estimating treatment effects in a semi-supervised setting where labeled outcomes are sparse but unlabeled data are abundant. We develop a semi-supervised calibration method that leverages a subsample of labeled outcomes to calibrate inferred outcomes, ensuring that the downstream treatment effect estimator remains consistent despite potential errors in outcome imputation. Unlike most traditional semi-supervised methods, we allow the labeling mechanism to depend on the observed data, rather than assuming it is completely random. This problem is analogous to estimating mean outcomes in longitudinal studies with monotone missingness, and we show that our proposed estimator is asymptotically equivalent to the augmented inverse probability weighting (AIPW) estimator when a consistent estimate of the labeling propensity score is available. The estimator is multiply robust and locally semiparametric efficient. We also demonstrate improved finite-sample efficiency in semi-supervised settings, owing to an effective normalization of an implicit augmentation term. The finite-sample performance is evaluated through simulations, and we illustrate the method in a case study comparing the effectiveness of two anti-TNF therapies on remission outcomes in patients with rheumatoid arthritis. In Chapter 2, we consider the problem of causal mediation analysis in a semi-supervised setting. Causal mediation analysis is a fundamental tool for understanding how treatments or exposures affect outcomes through intermediate variables. However, existing methods lack theoretical and practical guarantees in the presence of substantial missingness in the outcome variable. To address this, we propose a robust and efficient method for the semi-supervised estimation of natural direct and indirect effects, accommodating settings both with and without surrogate outcomes. Our approach extends the double machine learning framework by constructing multiply robust one-step estimators, which enable the use of flexible machine learning algorithms for estimating nuisance functions. We show that the proposed estimators are asymptotically normal and locally minimax optimal under semi-supervised models that do not assume the distribution of the non-missing variables to be known. Through extensive simulations, we demonstrate that the estimators remain unbiased, yield valid inference, and achieve substantial efficiency gains by leveraging predictive surrogates, even under complex data-generating mechanisms. We illustrate the method using EHR data to examine whether racial disparities in Alzheimer’s disease–related outcomes are mediated through cardiovascular comorbidities. In Chapter 3, we explore the problem of estimating low-dimensional functionals in a target population using data from both source and target domains, where the two distributions are only weakly aligned. This setting arises frequently in practice when performing transfer learning, yet remains underexplored in the context of functional estimation. We focus on three canonical functionals relevant to statistics and causal inference: the quadratic regression functional, the expected conditional covariance, and the mean response in missing data models. Rather than making strong parametric assumptions on the source and target distributions, we model weak alignment by assuming that the differences in conditional mean functions across domains are bounded in L2 norm. For each functional, we propose two estimation strategies: (1) an influence function–based estimator that employs existing minimax-optimal transfer learning methods for nuisance estimation, and (2) a functional-level confidence thresholding estimator that selects between source- and target-based estimators. We derive non-asymptotic upper bounds on the estimation risk in mean absolute error and show that both strategies achieve faster rates than target-only estimators, provided that posterior drift between domains is sufficiently small. When functionals involve two nuisances, we demonstrate that some of the proposed estimators are able to outperform the minimax-optimal target-only estimator as long as the degree of posterior drift in either of the two nuisance functions is sufficiently small, which we refer to as distributional shift double robustness. Furthermore, we demonstrate that under additional assumptions -- for example, by directly bounding the separation between the functional evaluated at source and target distributions -- the functional-level confidence thresholding estimator can attain minimax-optimal rates up to log factors.Biostatistic

    Theoretical and Computational aspects of Polynomial Neural Networks: Training Stability, Algorithmic Complexity and Expressivity

    No full text
    This thesis is focused on the study of Polynomial Neural Networks (PNN) and their properties. PNNs are neural networks where the activation functions are themselves polynomials. The main result of the thesis focuses on the maximum learning rate for the stable training of polynomial neural networks in the ultra-rich regime; recent empirical bounds were observed showing a root-like (γ1/d) behavior of the maximum learning rate as a function of the richness factor γ. This thesis provides a fundamental and theoretical explanation of this observed phenomena in the cases of PNNs as well as transformer networks. Two more topics are investigated as part of this thesis: the establishment of a quasi-polynomial time algorithm for the training problem for PNNs, as well as a computational framework for the study of the expressivity of PNNs.Extension Studie

    Genetic and Proteomic Factors Underlying Complex Human Diseases

    No full text
    The rapid growth of genomic data over the past two decades has brought the biomedical field closer to precision medicine, which aims to develop data-driven and tailored approaches for disease prevention and treatment. Population-scale biobanks have played an instrumental role in generating vast amounts of human genetic sequencing data, empowering extensive investigations into the genetics underlying complex diseases. Notably, biobanks have facilitated thousands of genome-wide association studies (GWAS). These studies use genotype information from large cohorts to identify common genetic variants associated with phenotypes of interest, and have provided foundational insights into the genetic basis of numerous complex diseases and medically-relevant traits. However, fulfilling the vision of precision medicine requires deeper understanding of genetics in diverse populations and within broader biological contexts. First, most GWAS have been conducted in populations of European ancestry, which has not only limited GWAS discoveries but also the clinical translation of GWAS findings. Associations identified in GWAS are frequently utilized to construct polygenic risk scores (PRS), estimates of individuals’ genetic risk that have been incorporated into clinical risk models for some complex diseases. Imbalances in the representation of different populations in biobanks and GWAS have precipitated the development of many PRS with lower predictive accuracy in populations of non-European ancestries. Second, while genes provide the blueprint for biology, they are several steps removed from the molecular mechanisms driving disease. Biomarkers from other -omic data, such as proteomics, may help bridge the gap between the genome and disease etiology, providing a more comprehensive view of underlying causal pathways as well as the effects of non-genetic factors on disease processes. Emerging datasets of plasma proteomics from population biobanks have highlighted the great potential of this data for precision medicine, but challenges remain in delineating the interplay between the genome, proteome, environment, and disease. In Chapter 1, I conduct a multi-ancestry GWAS meta-analysis using data from 18 biobanks to characterize the genetic architecture of asthma, a heterogeneous and multifactorial disease with variable prevalence rates. I identify 49 novel associations, and demonstrate the value of integrating data from many diverse cohorts to assess shared genetic risk across ancestries, biobanks, and disease subtypes, as well as improve polygenic risk prediction. In Chapter 2, I develop PRS trained on multi-ancestry and multi-biobank data with up to 750,000 participants for 32 complex diseases and traits. By evaluating the prediction performance of these models across diverse ancestry groups, I elucidate strategies for constructing the optimal PRS depending on ancestry, method, and genetic architecture. In Chapter 3, I dissect the mechanisms driving thousands of associations between 2,935 plasma proteins and onset of 23 age-related diseases. By integrating disease and protein GWAS data in a causal inference approach, I identify a subset of proteomic biomarkers that play causal roles in disease development. I also demonstrate that a large proportion of the circulating proteome is associated with smoking, and develop a proteomic score that captures smoking behavior and history. Together, these studies leverage multi-omic data to expand our understanding of the factors underlying disease risk, with the ultimate goal of advancing precision medicine.Medical Science

    On the Quantification of Aging

    No full text
    This dissertation explores the complex biological process of aging through multiple methodological lenses, from molecular mechanisms to population-level analyses. As global demographics shift toward an increasingly older population, understanding aging mechanisms and developing interventions to extend healthy lifespan have become critical scientific priorities. Despite chronological age being the strongest risk factor for many diseases, individuals age at different rates, suggesting that chronological age alone is an insufficient measure of biological aging. This research addresses key challenges in aging research through a multifaceted approach combining traditional genetic and epidemiological analyses with advanced computational methods. The dissertation investigates causal relationships between aging and disease, develops causality-enriched epigenetic clocks, examines the role of germline mutations in exceptional longevity, constructs high-dimensional representations of aging, develops foundation models for analyzing complex aging-related data, establishes standardized frameworks for biomarker evaluation, and creates a unified theoretical definition of biological age. By pursuing these objectives, this work aims to contribute significantly to our understanding of the aging process and provide tools and frameworks that can accelerate research in this field.Biological Sciences in Public Healt

    Gary, Indiana...From Slag to the Sublime

    No full text
    “Gary, Indiana…” explores how slag, one of the main byproducts of steel production, can be used to connect a city to a waterfront. Rather than closing Gary Works – the largest steel mill in the United States – this thesis imagines a near future where this site will accommodate increased domestic steel production within the pressures of climate change and ecological deterioration. This situation necessitates a new relationship between production, industrial waste and human-environmental experience. By analyzing the history, techniques, and socio-ecological impact of slag use, this project investigates how slag can be expressed as a land-making material and how it can foster a primary successional ecosystem. In doing so, the manufacturing which sustains Northwest Indiana can act as the mechanism which dissipates the barrier between Gary and Lake Michigan, proposing a new typology of public space which allows heavy industry, ecological restoration, and recreation to coexist in the American Rust Belt.Department of Landscape Architectur

    26,027

    full texts

    71,446

    metadata records
    Updated in last 30 days.
    DASH is based in United States
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇