80264 research outputs found
Sort by
Single-cell somatic copy number variants in brain using different amplification methods and reference genomes
The presence of somatic mutations, including copy number variants (CNVs), in the brain is well recognized. Comprehensive study requires single-cell whole genome amplification, with several methods available, prior to sequencing. Here we compare PicoPLEX with two recent adaptations of multiple displacement amplification (MDA): primary template-directed amplification (PTA) and droplet MDA, across 93 human brain cortical nuclei. We demonstrate different properties for each, with PTA providing the broadest amplification, PicoPLEX the most even, and distinct chimeric profiles. Furthermore, we perform CNV calling on two brains with multiple system atrophy and one control brain using different reference genomes. We find that 20.6% of brain cells have at least one Mb-scale CNV, with some supported by bulk sequencing or single-cells from other brain regions. Our study highlights the importance of selecting whole genome amplification method and reference genome for CNV calling, while supporting the existence of somatic CNVs in healthy and diseased human brain
Photopatterning of conductive hydrogels which exhibit tissue-like properties
Hydrogels are three-dimensional, highly tunable material systems that can match the properties of extracellular matrices. In addition to being widely used to grow and modulate cell behavior, hydrogels can be made conductive to further modulate electrically active cells, such as neurons, and even incorporated into multielectrode arrays to interface with tissues. To enable conductive hydrogels, graphene flakes can be mechanically suspended into a hydrogel precursor. The conductivity of the hydrogel can be increased by increasing the weight percentage of graphene flakes in the precursor while maintaining the mechanical properties of the formed gel similar to the properties of neural tissue. By using a photocrosslinkable hydrogel matrix, such as gelatin methacrylate, with a photoabsorber, the conductive precursor solutions can be crosslinked into predefined complex patterns. Finally, the formulations can be used to support the growth of sensory neurons, derived from human induced pluripotent stem cells, for more than 7 weeks while the neurons remain viable. These scaffolds can be patterned into components of multielectrode arrays, to enable ultrasoft electrodes with tissue-matched properties for further interactions, both in vitro and in vivo, with the nervous systems
Technologies to Better Study Cerebrospinal Fluid Movement in The Human Brain
This work describes the development of data processing pipelines and machine learning algorithms to better study the movement of cerebrospinal fluid (CSF) in the human brain. This was done by approaching two limitations of current technology: an inability to measure CSF flow without bulky and expensive magnetic resonance imaging (MRI) and a lack of image processing algorithms to measure CSF in perivascular spaces of the brain from MRI.
First, we developed a method to estimate CSF flow in the cerebral aqueduct of healthy control subjects using a set of simpler, portable, and non-invasive modalities. To do this, we extracted novel feature sets from signals acquired during a sleep study, including electrocardiogram (ECG), photoplethysmography (PPG), respiration, electroencephalogram (EEG), hypnograms, and body position as input into three machine learning models to estimate CSF flow. The estimates are compared to the ground truth CSF flow as determined by 7-Tesla phase contrast MRI (PC-MRI). We achieve a mean Pearson correlation coefficient of , , and for the three models (partial least squares regression, least absolute shrinkage and selection operator regression, and ridge regression) across all input feature sets for the development cohort. We also perform CSF estimation on a separate independent validation set with some models performing over and up to Pearson correlation coefficient. Thus, the regressor achieves high accuracy and is generalizeable making it more than adequate for clinical use. To the best of our knowledge, this is the first instance of CSF estimation without using MRI based signals. This lays the foundation towards a method to assess CSF movement in the brain with a portable and non-invasive device.
Second, we developed a novel image processing pipeline for extracting the velocity of cerebrospinal fluid (CSF) in perivascular spaces of humans using Tesla (7T) Phase Contrast Magnetic Resonance Imaging (PC-MRI). We applied the processing pipeline to a group of subjects ( healthy controls, normal pressure hydrocephalus, subarachnoid hemorrhage, stroke, and Alzheimer's disease) acquired by Houston Methodist Hospital. Using the healthy control subjects, we investigate if there exists differences in perivascular CSF movement between age groups and sex. We found that there were no significant differences amongst these groups, which is in alignment with findings for CSF velocity in the cerebral aqueduct. Next we investigated the relationship between CSF velocity in the cerebral aqueduct and the perivascular spaces and found a mean Pearson correlation coefficient between these two curves of across the subjects, indicating a real association between the ventricular and perivascular system of CSF movement. Lastly, we investigated the differences in perivascular CSF movement across the disease cases and found significant differences in the standard deviation of perivascular velocity between the healthy controls and normal pressure hydrocephalus. To our knowledge, this is the first example of quantitatively measuring CSF velocity in the perivascular spaces of humans and the first time anyone has compared the movement of CSF between the ventricular and perivascular systems. Our pipeline can be applied to both healthy and diseased subjects for accelerated discovery of differences in CSF and glymphatic function across populations of subjects.
These two bodies of work demonstrate the application and development of data processing and machine learning algorithms to studying CSF movement in the human brain to pave the way towards new biomarkers for brain health and disease etiology
Predicting Liver Segmentation Model Failure with Feature-Based Out-of-Distribution Detection and Generative Adversarial Networks
Advanced liver cancer is often treated with radiotherapy, which requires precise liver segmentation. Deep learning models excel at segmentation but struggle on unseen data, a problem exacerbated by the difficulty of amassing large datasets in medical imaging. Clinicians manually correct these errors, but as models improve, the risk of clinicians overlooking mistakes due to automation bias increases. To ensure quality care for all patients, this thesis aims to offer automated, scalable, and interpretable solutions for detecting liver segmentation model failures.
My first approach prioritized performance and scalability. It applied the Mahalanobis distance (MD) to the features of four Swin UNETR and nnU-net liver segmentation models. I proposed reducing the dimensionality of these features with either principal component analysis (PCA) or uniform manifold approximation and projection (UMAP), resulting in improved performance and efficiency. Additionally, I proposed a k-th nearest neighbors distance (KNN) as a non-parametric alternative to the MD for medical imaging. KNN drastically improved scalability and performance on raw and average-pooled bottleneck features.
My second approach emphasized interpretability by introducing generative modeling for the localization of novel information that a model will fail on. It employed a StyleGAN2 network to model a distribution of 3,234 abdominal computed tomography exams (CTs). It then localized metal artifacts and abnormal fluid buildup, two prevalent causes of liver segmentation model failure, in 55 CTs by reconstructing the scans with backpropagation on the StyleGAN’s input space and focusing on the regions with the highest reconstruction errors.
The computational cost, data requirements, and training complexity of generative adversarial networks, along with a lack of reliable evaluation measures, have impeded their application to medical imaging. Accordingly, a significant portion of this thesis is dedicated to evaluating the applications of StyleGAN2 and the Fréchet Inception Distance (FID), a common measure of synthetic image quality, to medical imaging.
The principal contributions of this thesis are integrating PCA and UMAP with MD, utilizing KNN for out-of-distribution detection in medical imaging, leveraging generative modeling to localize novel information at inference, providing a comprehensive application study of StyleGAN2 to medical imaging, and challenging prevailing assumptions about the FID in medical imaging
Computational Methods for Analyses of Single-cell DNA Sequencing Data in Cancer
The study of cancer using single-cell sequencing technology has opened up exciting new avenues for understanding the genomic complexity and heterogeneity of this disease. However, the analysis of such data presents computational challenges both in terms of designing novel mathematical models for biological discovery as well as devising new methods that are scalable to the newly emerged large-scale single-cell sequencing data. Throughout my Ph.D. studies, I focused on multiple research projects, each of which aimed to address such computational challenges in analyzing single-cell sequencing data in the context of cancer. In this thesis, I present my contributions to three studies and their corresponding methods, including Phylovar for phylogeny-aware detection of single-nucleotide variations (SNVs), MoTERNN for classifying the mode of cancer evolution, and MaCroDNA for integrating high-throughput single-cell DNA and RNA sequencing data.
In Phylovar, I improved the joint inference of cancer cells' SNVs (a common type of mutation in cancer) and their phylogeny, an approach known as phylogeny-aware SNV detection. Although this approach is highly accurate, its scalability to large-scale single-cell sequencing datasets was limited. To address this, I introduced a novel vectorized formulation for computing the likelihood function of this model, achieving very good improvement in calculation speed, enabling us to scale up accurate SNV detection from hundreds to millions of genomic loci suitable for the fast-expanding datasets from single-cell whole-genome and whole-exome sequencing technologies.
MoTERNN is aimed at determining modes of cancer evolution—linear, branching, neutral, or punctuated—each indicative of specific evolution patterns critical for diagnosis, prognosis, and treatment strategies. I treated this as a graph classification problem, using phylogenetic trees as graphs and evolution modes as classes, and employed Recursive Neural Networks (RvNNs) for classification. As the first application of RvNNs to phylogenetics, MoTERNN demonstrated very high accuracy in both the training and testing phases, showcasing the potential of RvNNs for learning on phylogenetic trees.
In the MaCroDNA project, I aimed to link DNA mutations to their impacts on RNA changes by pairing the cells that have been sequenced for either DNA or RNA data alone. In this work, I employed a maximum weighted bipartite matching algorithm for assigning the cells from the two data domains so that the sum of the Pearson correlation between all pairs is maximized. MaCroDNA achieved very good accuracy and outperformed the state-of-the-art method by a large margin
Accurate and Efficient Computational Approaches for Long-read Alignment and Genome Phasing of Human Genomes
The arrival of long-read sequencing technologies has enabled analysis of human genomes at unprecedented resolution. Long-read technologies have facilitated telomere-to-telomere assembly of the human genome and shed light on difficult to resolve structural variations, single nucleotide variations and epigenetic modifications, which all play a critical role in disease etiology and individual genetic diversity. Despite the technological advancement, novel computational methods are still needed to fully leverage long reads. In this dissertation, I tackle three key computational questions by leveraging long-read sequences of human genomes: 1. I improve on the efficiency and precision of long-read alignment, 2. I develop a novel variant phasing techniques based on methylation signal, and 3. I provide a novel method for clinical analysis specific to cancer samples and tumor purity estimation. These accomplishments are represented by three software tools I have developed: Vulcan, MethPhaser and MethPhaser-Cancer, respectively.
Vulcan is a read mapping pipeline that uses two distinct gap penalty modes, which is referred to as dual-mode alignment. Read aligners before Vulcan only use one type of scoring scheme during the pairwise alignment stage, which can struggle due to the variable diversity across the human genome. With Vulcan’s dual-mode alignment algorithm, the read-to-reference mapping quality and efficiency for Oxford Nanopore Technology (ONT) long-reads are improved for both simulated and real datasets. Notably, we also show Vulcan provides improvement in structural variation detection. Vulcan increased the SV detection F1 score of 30X human ONT reads from 82.66% (minimap2) to 84.94%.
MethPhaser is the first method that utilizes methylation, an epigenetic marker, from Oxford Nanopore Technologies to extend SNV-based phasing. Long-read human genomic variant phasing is limited by read length and stretches of homozygosity along the genome. The key innovation of MethPhaser is the utilization of the haplotype-specific long-read methylation signals. In benchmarking against human samples, MethPhaser nearly triples the phase length N50 while incurring a minimal increase in switch error from 0.06% to 0.07% using ONT R10 reads at 60X coverage. As an extension method to existing long-read SNV-based phasing workflows, MethPhaser offers substantial enhancements with a negligible rise in switch error rates.
Building upon MethPhaser, I have also innovated an algorithmic extension named MethPhaser-Cancer that uses methylation signals for the assessment of tumor purity and for categorizing reads. The tumor purity estimation is an important step in clinical treatment that is related to tailoring patient-specific therapeutic strategies and in the broader context of personalized medicine. MethPhaser-Cancer adeptly identifies hypomethylated areas within human tumor samples and utilizes the k-means algorithm to sort the reads into two distinct groups. This represents a pioneering approach in the long-read sequencing field to consider whole-genome methylation profiles in simulated clinical samples, capable of automatically estimating the tumor purity and distinguishing long-reads within specific regions between two samples.
To conclude, this dissertation represents a set of novel and efficient approaches that enhances the long-read human genomic analysis. The real-life usage of Vulcan, MethPhaser and MethPhaser-Cancer includes long-read alignment, human genome variant phasing and tumor purity estimation
Prompting Strategy Use and Beyond: Examining the Relationships between Elaboration, Quantity, and Diversity of Learning Strategies on Performance
Elaboration is a generative learning strategy wherein learners link prior knowledge and experiences with to-be-remembered information. It is positively related to an array of learning outcomes. However, most students do not independently use generative learning strategies. We explored whether prompting elaboration learning strategies when reading an academic passage influenced knowledge test performance. Participants were randomly assigned to two conditions: receiving a prompt (i.e., experimental; n = 94) and no prompt (i.e., control; n = 112). The results revealed that participants who received the elaboration prompt (M = 13.88, SD = 2.20) did not outperform learners who did not receive the prompt (M = 13.67, SD = 2.43) on the knowledge test. However, we did find a positive relationship between the extent of elaboration strategy use and knowledge test performance across conditions (r = 0.17, p < 0.05). Twelve themes emerged from an exploratory thematic analysis, wherein participants were asked about the learning strategies they used when reading the passage. Students used a variety of learning strategies unprompted, although 42.15% reported not using any additional learning strategies outside of the prompt or using low-utility learning strategies (e.g., relying on memory, skimming). Further exploratory analyses found that the quantity and diversity of learning strategies used individually influenced knowledge test performance. ANCOVA results revealed, however, that when controlling for quantity, the diversity of learning strategies used did not significantly influence knowledge test performance. Our findings contribute to prior literature by (1) demonstrating a relationship between elaboration strategy use and test performance, (2) highlighting learning strategies students use to retain information, and (3) exploring additional factors regarding learning strategy use that influence performance
AI and algorithms for ultra large scale information retrieval systems
Information Retrieval (IR) entails the task of matching a specific query to a single item or a group of items within a database full of potential matches. This pivotal operation serves numerous applications, such as genome sequence and document search, near-neighbor search, and classification. The exponential increase in data generation - encompassing text, audio, images, DNA sequences, and time series - from sources like sensors and the web has resulted in data volumes exceeding standard hardware's storage capabilities. Additionally, modern AI applications rely on substantial data quantities, ranging from Gigabytes to Terabytes, and a key-label pair set spanning billions to trillions.
Hashing-based algorithms have shown promise in enabling efficient large-scale retrieval due to their low latency and minimal hardware storage requirements. This research is set to examine two search modalities: exact-match search and similarity search. We propose RAMBO, a Bloom filter-based sub-linear search index for exact-match search. Particularly potent in genome sequence searches, this efficient algorithm indexes terabytes of DNA data for E-coli species, dramatically reducing indexing times from weeks to hours and search times from hours to seconds. This progress empowers any laboratory to scour through extensive genome archives using standard computers. We also extend this work by proposing a hash function IDL (Identity with Locality) that preserves the spatial locality and identity of keys for cache-efficient exact-match queries. Conversely, recommendation engines often lean on similarity matching when handling real-world data such as text, images, and audio. In this context, we introduce BLISS (Balanced Index for Scalable Search), a mechanism capable of learning and associating any two real-world data entities within an acceptable approximation. BLISS can index a billion items through an iterative learning algorithm, and its high accuracy, coupled with a small memory footprint, speeds up retrieval by a factor of 5 compared to current benchmarks. Moreover, BLISS offers the dual functionality of near-neighbor search and extreme classification. In practical retrieval engines, similarity match and exact match form the core components, with embedding and filters serving as the query. Current techniques involve a cascade of vector similarity match and exact match on filters, which often leads to slower and less precise results. To address this, we introduce CAPS (Constrained Approximate Partition Search), a unified, single-stage index designed for filter-based near-neighbor searches. Our findings demonstrate that CAPS not only streamlines the search process but also significantly improves the efficiency of Amazon’s search system
Population coding of strategic variables during foraging in freely moving macaques
Until now, it has been difficult to examine the neural bases of foraging in naturalistic environments because previous approaches have relied on restrained animals performing trial-based foraging tasks. Here we allowed unrestrained monkeys to freely interact with concurrent reward options while we wirelessly recorded population activity in the dorsolateral prefrontal cortex. The animals decided when and where to forage based on whether their prediction of reward was fulfilled or violated. This prediction was not solely based on a history of reward delivery, but also on the understanding that waiting longer improves the chance of reward. The task variables were continuously represented in a subspace of the high-dimensional population activity, and this compressed representation predicted the animal’s subsequent choices better than the true task variables and as well as the raw neural activity. Our results indicate that monkeys’ foraging strategies are based on a cortical model of reward dynamics as animals freely explore their environment