71446 research outputs found
Sort by
Harnessing the Power of a Novel Synthetic U.S. Population to Uncover Public Health Insights
Marginalized communities in the U.S. must contend with structural discrimination that has systematically placed them in under-resourced neighborhoods whose inherent social, economic, infrastructural, and environmental conditions adversely affect health and their wellbeing. Today, researchers’ understanding of these neighborhoods—characterized by substandard housing, hazardous workplaces, degraded infrastructures, close proximity to toxic sites and polluting facilities, food deserts and fast-food jungles, limited opportunities for quality education and employment, and restricted access to healthcare services and facilities—and their inhabitants rely on incomplete population data fraught with assumptions and biases symptomatic of a long history of residential segregation. Broadly, the two options for available population data are: (1) geographically-aggregated data that operates at a high spatial resolution, but only provides marginal distributions of single characteristics, or (2) individual-level microdata that provides joint distributions of multiple characteristics, but only operates at a low spatial resolution. Thus, researchers aiming to leverage available population data to devise and disperse public health interventions must forgo either an understanding of compounding risks or the ability to pinpoint localized risks.
This dissertation aims to construct, validate, and demonstrate the utility of a novel synthetic U.S. population which models 330 million agents that collectively reflect the real population with respect to key social, economic, infrastructural, and environmental characteristics at the census tract level. Generated through spatial microsimulation methods that avail the strengths of both geographically-aggregated data and individual-level microdata, it affords researchers the ability to answer previously unanswerable questions about large-scale, complex public health issues across the U.S. First, we constructed the model using iterative proportional fitting (IPF)—a widely-favored static deterministic reweighting method for spatial microsimulation. We then validated it against the 2018-2022 American Community Survey (ACS), the 2020 Decennial Census, and the 2020 Residential Energy Consumption Survey (RECS) with a suite of metrics including Pearson’s R2 (mean = 0.98), standard absolute error (SAE) (mean = 0.03), and the Kullback-Leibler (KL) divergence statistic (mean 0.08), finding that our synthetic U.S. population exhibited strong overall performance in simulating the real U.S. population with variation due to the size and diversity of census tracts and the complexity of linking variables. Second, we examined the well-established compounding risk faced by children under the age of five who we know—based on our synthetic U.S. population—live in homes built before 1980 to map residential lead exposure hotspots at the census tract level. Our novel approach for identifying at-risk areas therefore employed the joint distribution of these two characteristics in direct contrast to traditional approaches that assess risk by overlapping separate marginal distributions. This enabled us to uncover up to 73% more at-risk census tracts where an estimated 3.28 million additional at-risk children reside. Additionally, our approach revealed statistically significant systematic inequities in how traditional approaches detect risk, where newly identified at-risk children were disproportionately located in census tracts with higher concentrations of uninsured, White, and Native populations. It also highlighted that relying on housing age as a proxy for residential lead-based paint exposure may neglect important disparities in housing quality. Third, we coupled our synthetic U.S. population with the 2020 Residential Energy Consumption Survey (RECS) to estimate individual probabilities of residential air conditioning (AC) access—an essential cooling solution to protect against the adverse effects of heat stress—using a machine learning model to downscale data that was previously only available with spatial information at the state level to the census tract level. Unlike traditional approaches which focus solely on AC ownership as a proxy for AC access, our measure additionally accounted for the ability to install, maintain, and operate AC, resulting in probability estimates that were 17% lower and 9% more variable. By accounting for these additional dimensions of residential AC access, we also uncovered interstate and intrastate disparities shaped by fine-scale differences in state policies, infrastructures, and practices as well as climate conditions, housing conditions, and demographic characteristics.
In summary, our novel synthetic U.S. population addresses critical gaps in existing population data, uncovering limitations, assumptions, and biases that may otherwise persist as researchers aim to devise and disperse public health interventions. Beyond its utility in mapping residential lead exposure risk and residential AC access, this model can be adapted to examine a wide range of public health challenges, enabling researchers to explore how intersecting social, economic, infrastructural, and environmental forces shape outcomes at fine spatial resolutions and providing them with a powerful tool to guide more effective and equitable public health interventions across the U.S.Population Health Science
The Western Village
Thomas Fur, an amateur writer going through a midlife crisis, decides to quit his job and move to the Western Village, a small town on the Chesapeake Bay. He rents the basement of the house of Mrs. Vivian, a sixty-five-year-old widow. Given that only elderly people live in the village, Thomas feels that this place away from the busy towns appeals to him. The village, where only two things alternate, the dying old residents and the eternal sea, becomes the final destination, a reinvented frontier for a man who wants to overcome his limitations.
After arriving in the village, Thomas decides to write a novel. He uses the aging residents as characters for his literary project. For the very reason, he makes several acquaintances there, including a former criminal who works as a gravedigger's assistant, a veteran turned hunter and fisherman, and a bartender and her sister.
Thomas, unable to express his love to Mrs. Vivian, hires an unknown man to write love letters to her. He even helps her respond to these letters. What has started as a game soon turns into a strange love relationship.
Thomas later confronts the only physician in the town, Doctor Schroeder, who also works as a gravedigger. Doctor Schroeder makes money off the old people either way: if he cures them, he bills them while they are alive, and if he fails, he bills the deceased people's relatives for their death.
Later, Thomas discovers that Doctor Schroeder is a souls' collector. He locks up the dead's souls in glass bottles. Thomas's confrontation with the doctor turns into the final tragedy in his life.
At the end of the novel, we are left perplexed, not knowing whether Thomas invented the village and its residents to write his novel or whether it was a real chapter in his life, or whether Thomas himself was a fictional character in real people's life.Extension Studie
Impact of Different Policies on US EV Sales
Electric vehicles (EVs) are a key component in reducing greenhouse gas
emissions to contribute to the United Nations’ 2030 Agenda for Sustainable
Development. EV adoption varies widely across the United States due to differing
policies and incentives for individual states. The primary objective of this research was
to identify which policies or combination of policies were the most effective in increasing
EV sales in the United Sates from 2011-2021. This study utilized data from the US
National Conference of State Legislatures and the Alliance for Automotive Innovation.
Policies were categorized into tax incentives, financial incentives, subsidies, and
convenience benefits. Regression and correlation analyses were performed to assess the
impact of these incentives on EV growth from 2011-2021. Additionally, a sales forecast
for 2024-2034 was conducted using logistic growth models, from which future research
can study the validity of the forecast by comparing that data to the actual sales data for
2024. Tax incentives showed the strongest and only statistically significant correlation
with EV sales growth (r = .318, p = .024), with states like Maryland and California
leading with significant increases. Conversely, financial incentives and subsidies showed
insignificant correlations. While some states achieved over 26% annual growth,
particularly those with a mix of state and private incentive types, correlation results
suggest that individual policy types alone do not fully explain state-level sales
performance. Recommendations based on these findings are to enhance tax incentives,
develop subsidy programs, and promote private-public partnerships.Extension Studie
Genetic Risk and Phenotypic Diversity in Primary Open-Angle Glaucoma: Polygenic Risk Scores Association with Laser Trabeculoplasty Outcomes and Data-Driven Phenotyping through Unsupervised Clustering
ABSTRACT
Objective
To examine the relationship between a primary open-angle glaucoma (POAG) polygenic risk score (PRS) and laser trabeculoplasty (LTP) outcomes.
Design
Retrospective observational follow-up study in a Biobank-linked clinical sample and three case groups identified from population-based cohorts.
Participants
We included data from patients aged 40 and above with POAG who underwent LTP in the Mass General Biobank, the Nurses’ Health Study, Nurses’ Health Study 2, and the Health Professionals Follow-up Study.
Methods
A PRS was calculated using prior genome-wide association study (GWAS) summary statistics, and participants with LTP were categorized into either a high- (top 10%) or low-score (remaining 90%) category. Statistical analyses included Kaplan-Meier survival analysis to assess time to LTP failure, Cox proportional hazards models to evaluate the impact of PRS on failure risk, and Area Under the Curve (AUC) to measure predictive accuracy.
Main outcome measure
LTP failure, defined as less than a 20% reduction in intraocular pressure (IOP) from baseline or the need for additional glaucoma surgical or laser procedures, assessed six weeks after LTP and within two years of follow-up.
Results
The study included 109 participants, grouped into low PRS (bottom 90th percentile; 72 participants) and high PRS (top 10th percentile; 37 participants). The median time to LTP failure was significantly longer for low PRS patients (12 months, 95%CI: 8.94–16.79) than in those with high PRS (6 months, 95%CI: 5.62–7.36, p.0001). Participants in the high PRS group had a 2.03 times higher risk of LTP failure (95%CI: 1.25–3.32, p=0.004) after adjusting for age, sex, genetic ancestry, LTP type, baseline IOP, IOP-lowering medication number, and visual field mean deviation (MD). A model incorporating PRS in addition to clinical characteristics had higher predictive accuracy over time (AUC = 0.73 vs. 0.66, p=0.03 at 18 months after LTP).
Conclusions
A higher PRS was significantly associated with increased risk of LTP failure in this sample of POAG patients. These findings suggest PRS may help identify patients at higher risk for LTP failure, but further research with a prospective, population-based sample is needed to confirm this association.
ABSTRACT
Objective:
To compare three unsupervised clustering algorithms: k-means, Fuzzy c-means (FCM), and hierarchical clustering analysis (HCA) for primary open-angle glaucoma (POAG) phenotyping and to identify clinical POAG subtypes using multimodal structural and functional data from Electronic Health Records.
Design:
Retrospective cohort study.
Participants:
4,274 eyes from patients with POAG seen between 2016 and 2023 at a tertiary ophthalmology center.
Methods:
We extracted 21 continuous clinical features per eye, including baseline and longitudinal visual field indices, intraocular pressure (IOP) parameters, OCT-derived retinal nerve fiber layer (RNFL) metrics, and optic nerve morphology. Dimensionality reduction was performed using Sparse Principal Component Analysis (Sparse PCA), retaining seven components. The optimal number of clusters (K=5) was determined based on silhouette score and HCA dendrogram inspection. We then compared agreement between methods and evaluated cluster performance. Post hoc comparisons across clusters used Kruskal-Wallis and chi-square tests.
Main Outcome Measures:
Clustering performance was assessed using internal validation metrics (Calinski-Harabasz index, Davies-Bouldin index, Dunn index, and silhouette score) and cluster stability metrics (mean Jaccard similarity for k-means and HCA; fuzzy partition coefficient for FCM). Agreement across methods was quantified using unweighted Cohen’s kappa. Distinct POAG phenotypes were characterized based on demographic, clinical, and treatment features.
Results:
HCA showed the weakest internal validity (Calinski-Harabasz index = 308.5; Dunn index = 0.02; silhouette score = 0.102) and the lowest cluster stability (mean Jaccard similarity = 0.27). K-means and FCM achieved higher validity (Calinski-Harabasz = 731; Dunn index = 0.013; silhouette scores = 0.133 and 0.132, respectively). FCM yielded the highest overall stability (fuzzy partition coefficient = 0.89). Agreement between FCM and k-means was excellent (Cohen’s kappa = 0.97). Based on FCM-derived clustering, we identified five phenotypes: (1) Small-disc with minimal progression, (2) Early-stage with stable structure-function, (3) RNFL-thinned with preserved function, (4) Rapidly progressing with high IOP variability, and (5) End-stage with apparent functional plateau.
Conclusions:
In this large real-world cohort of POAG patients, FCM and k-means outperformed HCA in both internal validity and cluster stability. FCM yielded the highest fuzzy partition coefficient and excellent agreement with k-means, supporting its use for POAG phenotypic segmentation. The five derived POAG subtypes exhibited distinct structural-functional profiles and IOP dynamics. Future work will incorporate archetypal analysis to model spatial patterns of visual field loss and further refine individualized disease trajectories.Graduate Educatio
Computational Models for Algorithm and Data Structure Design in Systems
Computational models are powerful tools that can be applied in a number of ways. We present three examples of how they can be applied in the context of algorithms and data structure design for systems, both as components within a design and as tools for navigating design tradeoffs. We start with an example of the former, where we demonstrate how best to utilize a machine learning model to improve the false-positive rate of the classic Bloom Filter. We then demonstrate how a computational model can be used to navigate the design space of prefix-based range filters and how such a model can be used for automated online optimization. Finally, we examine the history of computational models used for evaluating storage algorithms and demonstrate how tailoring these models to the critical aspects of modern storage hardware can provide valuable new insights.Engineering and Applied Sciences - Computer Scienc
Investigating Pediatric Diffuse Midline Glioma Invasion: Chloride Modulation and Tumor Co-Culture Dynamics
This thesis investigates the invasive nature of diffuse midline glioma (DMG), a rare yet lethal brain tumor most common in children, by exploring chloride ion modulation and tumor-tumor interactions. The primary objectives were to examine the regulation of potassium-chloride cotransporter 2 (KCC2) expression in DMG cells through transcription factor knockouts, modulate intracellular chloride levels through genetic engineering for KCC2 and sodium-potassium-chloride cotransporter 1, and assess tumor-tumor spheroid interactions under various conditions.
The study used molecular biology techniques, including short guide RNA plasmid cloning and CRISPR-Cas9 gene editing, to manipulate chloride cotransporter expression in DMG XIII cells. Immunofluorescence imaging and analysis were used to evaluate KCC2 expression following transcription factor knockouts. Additionally, tumor spheroid co-culture experiments were conducted to investigate infiltration patterns between different DMG cell lines under varying media conditions.
Results showed that knockout of the transcription factor STAT5A significantly increased KCC2 expression in DMG XIII cells. Genetic engineering efforts to overexpress KCC2 and sodium-potassium-chloride cotransporter 1 (NKCC1) were successful, though further selection and growth were required for comprehensive analysis. Tumor spheroid co-culture experiments revealed distinct infiltration patterns between different DMG cell lines, with DMG XIII and DMG XXIV showing the highest level of interaction. However, media conditions did not significantly affect infiltration levels.
This work establishes a foundation for understanding chloride ion homeostasis in DMG progression and invasion. The findings suggest potential therapeutic targets for managing DMGs and provide insights into tumor-tumor interactions that may guide future strategies to limit tumor invasion into healthy brain tissue.Biomedical Engineering A
Improving Profiling Techniques for Pipeline Parallelism in Distributed AI Model Training
As the capabilities of AI models continue to grow exponentially, so does their scale and
complexity. Training such large models is very costly in both memory consumption and
runtime, so optimizing over GPU clusters only becomes more important as size increases.
With the number of GPUs required for such large models, the financial cost is a significant
burden. Differences between models mean that it is essential to choose the most efficient
algorithm to minimize this cost. This problem is so complex that modern day training still
involves numerous costly experiments to determine the optimal training parameters to balance
memory consumption, runtime, and financial cost. For large models even just parameter
tuning can result in hundreds of thousands of dollars in financial cost, not to mention the
time to devise and run the experiments.
Here, we provide an efficient, accurate, and unintrusive memory and runtime estimation
tool that works with the most advanced training algorithms we have today. As the cost
of parameter experiments comes from the many partial profiling runs of the model that
are required, an accurate memory and runtime estimation tool that does not need to run
the model is extremely valuable. This allows us to choose the optimal training algorithm
for a given model while eliminating the costly parameter search. Prior works have only
implemented this for a limited subset of models, specifically for limited types of parallelism.
Thus our solution provides an extended memory and runtime profiling tool that supports the
use of pipeline parallelism, a technique used in many state-of-the-art training algorithms.Computer Scienc
Development of Inhibitors of ADAR1
Adenosine Deaminase Acting on RNA 1 (ADAR1) is an RNA-editing enzyme that
deaminates adenosine residues within double stranded RNA, transforming adenosine into
inosine. The deregulation of ADAR1 activity has been implicated in a variety of human diseases,
including cancer, through mechanisms involving both its catalytic and non-catalytic activity.
Drug discovery campaigns against ADAR1 have been stymied by the lack of small molecular
scaffolds that bind to ADAR1 and the lack of assays with enough sensitivity to discover these
scaffolds. Three methods were attempted to address these obstacles.
1. C75 as a covalent modifier of ADAR1: The fatty acid synthase inhibitor C75 was
shown to covalently modify ADAR1. The mechanism was confirmed to be a non-specific
Michael addition by nucleophilic amino acid residues on ADAR1. Analogues were
designed and synthesized that tempered the electrophilicity of C75 and curbed the
promiscuity of the ligand, but they were ultimately unsuccessful in retaining covalent
binding to ADAR1.
iv
2. Imidazotriazines as competitive inhibitors of ADAR1: Contrary to previous literature
reports, a commercial sample of ADAR1 was observed to slowly deaminate monomeric
adenosine into inosine, leading to an assay that quantified the amount of inosine produced
as a proxy for ADAR1 activity. Small molecule probes using the imidazotriazine scaffold
were found to competitively inhibit the inosine generating activity. However, further
observations with a more potent analogue revealed that the deaminase activity was due to
a small impurity of Adenosine Deaminase in the sample of ADAR1 used.
3. DNA-Encoded Library (DEL) Screen to Identify Degraders of ADAR1: Several
ligands were identified as possible binding partners to ADAR1 in a DEL screen, but none
showed inhibitory activity towards ADAR1. The ligand with the highest binding affinity
towards ADAR1 was elaborated into a PROTAC molecule that was shown to decrease
the level of ADAR1 in a cellular assay.Chemistry and Chemical Biolog
UNDERSTANDING DISEASE PROGRESSION IN PULMONARY FIBROSIS: THE ROLE OF AUTOANTIBODIES AND FUNCTIONAL IMPAIRMENT
Overview of the thesis papers
Interstitial lung disease (ILD) is characterized by progressive fibrosis of the lung parenchyma, ultimately leading to significant morbidity and mortality. The advent of antifibrotic therapies has made it increasingly important to identify and monitor disease progression, as timely intervention may improve patient outcomes. In this context, I conducted two studies to better understand disease progression in ILD. One focused on anti-Ro52 autoantibodies as a prognostic marker; the other on pulmonary functional impairment, focusing on FVC % predicted in patients without evident progression.
The first paper addressed the clinical question: Does anti-Ro52 positivity predict poorer outcomes in ILD? In a retrospective cohort study, I identified patients tested for anti-Ro52 antibodies and compared those who were anti-Ro52 positive with those who were negative. The primary outcome was ILD progression or death, and results revealed that the anti-Ro52-positive group exhibited significantly higher rates of disease progression, lung transplantation, and all-cause mortality. These findings suggest that anti-Ro52 seropositivity is an important biomarker for prognostication and underscores the need for vigilant monitoring in affected patients.
The second paper addressed another critical clinical question: Among patients with non-idiopathic pulmonary fibrosis ILD who do not meet the commonly used ≥5% absolute FVC decline criterion, is a lower FVC % predicted at one year nonetheless associated with a worse prognosis? A retrospective cohort study demonstrated that patients maintaining “stable” disease yet exhibiting a lower one-year FVC % predicted had significantly worse transplant-free survival. This indicates that even without measurable pulmonary function decline, functional impairment can still have prognostic significance, highlighting the need for close clinical monitoring in these patients.Graduate Educatio
Male Authorship and Female Agency: Examining Khaled Hosseini’s Validation in Writing from the Female Perspective in A Thousand Splendid Suns
This thesis explores the validity of male authorship in depicting female suffering
and agency in Khaled Hosseini’s A Thousand Splendid Suns. By examining how Hosseini
portrays Afghan women’s oppression, resilience, agency, and desires, this study
questions whether male authorship enhances or diminishes the accuracy of female
experiences. Through the “solidarity-based sisterhood” between Mariam and Laila, the
novel illustrates how female bonds serve as both a refuge and a catalyst for resistance
against patriarchal oppression. Additionally, Hosseini provides his female characters
with the space to express desires that Afghan women were historically denied, such as
love, romance, bodily autonomy, and human connection. This romanticization of their
longing for agency highlights the emotional depth of their struggles. Furthermore, the
novel’s male characters reflect the spectrum of male responses to female empowerment,
with some experiencing a “feminine epiphany” that acknowledges women’s strength
despite Taliban-imposed restrictions. Through analyzing these themes, this thesis argues
that Hosseini’s male perspective does not invalidate female suffering, and offers insight
into male insecurity, which further highlights the necessity of female agency. While
female voices are essential in narratives of oppression, male authors, like Hosseini, can
contribute meaningfully by exposing the mechanisms of patriarchal control and
reinforcing the urgency of women’s resilience and autonomy.Extension Studie