1,721,005 research outputs found
Recommended from our members
Parameter Estimation Procedure of Reaction Diffusion Equation with Application on Cell Polarity Growth
Cell polarity is a fundamental feature of almost all cells. An excellent example for studying cell polarity is pollen tube tip growth, which is a specialized form of cell growth, in which growth is limited to the apical end of the cell, allowing the cell to rapidly elongate and penetrate tissues. The key of pollen tube tip growth is polarized distribution on plasma membrane of a particle named ROP1 and Calcium. The oscillated distribution is the result of a feedback loop between ROP1 and Calcium. This dissertation takes a multidisciplinary approach by combining knowledge of cell biology, mathematics, and statistics in order to quantitatively study the full system for the interaction between ROP1 and Calcium in pollen tube tip growth. In the first part of the dissertation, we propose a mechanistic reaction-diffusion equation derived model of cell polarity, and analytically study the spatiotemporal dynamic of proposed model. In the second part, we first introduce parameter estimation procedure of linear reaction-diffusion equation, then extend it to nonlinear case. Our approach can be seen as an extension of gradient matching procedure. Moreover, consistency and asymptotic normality of proposed method will be discussed in linear case. And simulation study will be conducted in both cases. In the end, we will apply proposed method on the pollen tube tip growth model to predict ROP1 and Calcium concentration on plasma membrane. And guidance of refining proposed model is given based on predication of ROP1 and Calcium
Recommended from our members
Parameter Estimation in Differential Equation Based Models
There is a long history for differential equations to be utilized to model dynamic pro- cesses in many disciplines such as physics, engineering, computer science, Finance, Biology, etc.. Most original efforts have been devoted to simulating the dynamic process for given parame- ters that characterize the differential equations. In recent years, more and more attention has been given by scientists, especially statisticians who considered the problem inversely, that is, using the experimental data to recover the values of parameters that specifically describe the experimental process trajectory.In this dissertation, we introduced a general Integro-ordinary differential equation that describes a reaction diffusion process for the tip growth of pollen tubes and proposed a constrained nonlinear mixed effects model for the dynamic response. Accordingly, we developed two estimation procedures to estimate this model, Constrained Method of Moments (CMM) and Constrained Restricted Maximum Likelihood (CREML). The advantages and disadvantages of the two procedures were investigated. Simulation studies and a real data analysis demonstrated that both estimation procedures could provide accurate estimates for the parameters.As an extension, the parameter estimation for Integro-partial differential equation of multiple dimension would also be discussed
Recommended from our members
SNP Calling Using Genotype Model Selection on High-Throughput Sequencing Data
Recent advances in high-throughput sequencing (HTS) promise revolutionary impacts in science and technology, including the areas of disease diagnosis, pharmacogenomics, and mitigating antibiotic resistance. An important way to analyze the increasingly abundant HTS data, is through the use of single nucleotide polymorphism (SNP) callers. Considering a selection of popular HTS SNP calling procedures, it becomes clear that many rely mainly on base-calling and read mapping quality values. Thus there is a need to consider other sources of error when calling SNPs, such as those occurring during genomic sample preparation. Genotype Model Selection (GeMS), a novel method of consensus and SNP calling which accounts for genomic sample preparation errors, is thus given. Simulation studies demonstrate that GeMS has the best balance of sensitivity and positive predictive value (PPV) among a selection of popular SNP callers. Real data analyses also support this conclusion.As an extension to the aforementioned single sample GeMS, the multiple sample Genotype Model Selection (multiGeMS) method is also given. A simulation study and a real data analysis demonstrate that multiGeMS has a good balance of sensitivity and PPV when compared to a selection of popular multiple sample SNP callers
Recommended from our members
Estimation and Clustering on Longitudinal Data Using Penalized Spline Models
It is important to identify and characterize associations between microorganisms and human processes because they suggest possible involvements of health or disease processes and the interaction of organisms. My dissertation is motivated by a type of longitudinaldata in which the number of microorganisms in several mice are measured over time under different operational taxonomic units. This study focuses on such type of data using clustering on linear mixed models and generalized linear mixed models with the effects of operational taxonomic units and subject
Recommended from our members
Identifying the Best Predictive Biomarker in Pharmacogenomics
The traditional prescribing approaches in clinical therapies, such as “one drug fits all”, have limits due to consideration of drug effectiveness and safety. It is now well recognized that the Single Nucleotide Polymorphism (SNP) plays an important role in pharmacogenomics and personalized medicine by serving as the predictive biomarker for patient stratification and dose selection. Considering the economic efficiency of drug development, searching for the single best predictive SNP has just begun and has great potential. Current statistical methods for the best predictive biomarker selection rely on a variety of ranking procedures. However, these ordering approaches face three main issues: (1) they can potentially fail to distinguish predictive biomarkers from prognostic biomarkers; (2) the ordering is not necessarily correlated with the true significance, especially when the signal-to-noise is small; (3) such ranking approaches provide no control on the incorrect assertion in terms of the best predictive SNP detection.In this paper, we propose to overcome the first issue by quantifying the predictive ability of each candidate SNP in two parallel approaches: (1) adjusting the model with Least Square mean (LSmean); (2) adjusting the data with Virtual Matching (VM) before doing estimation. For the selection of the best predictive SNP, we propose to apply the Multiple Comparisons with the Best (MCB) to the ranking and selection procedures. Specifically, we will do the indifference zone selection with MCB lower bounds; and the subset selection with MCB upper bounds (suppose a larger response is better). The probability that both inferences are correct is controlled. Simulation studies show that our proposed method is more reliable in distinguishing between predictive and prognostic effects, and has a larger chance to detect the true best predictive SNP than varieties of methods in both multiple testing and machine learning approaches. Furthermore, the adjusting-model approach is generally better than the adjusting-data approach when we assume there is one single predictive SNP; however, the latter one is easier to be extended to the scenario with multiple predictive SNPs
Recommended from our members
Metagenomic Binning Algorithms
Metagenomics is the study of DNAs of microorganisms that are taken directly from environmental samples without cultivation and isolation. Recently, the emerging field of metagenome sequencing, facilitated by the high-throughput capability of NGS technology, allows the simultaneous sequencing all genomes in an environmental sample while also results in high complexity datasets. Although the NGS technology significantly improve the sequencing efficiency and cost, assembly of metagenomic sequences into genomes is extremely difficult since the reads are very short and sampled are from multiple genomes. Several computational methods have been developed to group metagenomic sequence reads into different bins, which can be categorized into two classes: supervised methods and unsupervised methods. Supervised methods may leave a large fraction of reads unclassified due to low rate of known reference genome in the database, while the unsupervised methods are still undergoing active development. The performance of existing unsupervised methods rely heavily on the length of reads, the number of species in the sample and the evenness of species abundance. It is also challenging for some algorithms to operate without a pre-specified number of species, which is not a trivial assumption to make.In this work, we present a novel algorithm, the DirichletCluster, based on Markovian assumption and sequential Monte Carlo (SMC) technique that has shown high binning accuracy under various scenarios with data-driven approach to estimate the number of species systematically. Specifically, we looked at the Markovian structure of the nucleotide reads, and implemented a mixture Dirichlet process model with the Markov chain structure. The Dirichlet process is a stochastic processs describing distribution over probability measures, which indicates draws from this process can be interpreted as random distributions. By using the mixture Dirichlet process model, we are able to characterize the individual genome sequence, as well as the clusters of sequences. Sequential Monte Carlo, together with GC content ordering, is implemented to cluster reads into species using a simulation based approach. We show through some simulation studies and a real data application that the proposed DirichletCluster binning algorithm to be robust to the evenness of abundance ratio and to be able to correctly identify the most number of species from the metagenomic data among alternatives. Moreover, it uses a complete data-driven approach to estimate the total number of species in the metagenomic sample. Therefore, we believe that DirichletCluster is a performant binning algorithm that is beneficial to the advancement of Metagenomics research
Recommended from our members
Adjusting for Population Differences Using Applied Machine Learning Methods
Clinical treatment evaluation based on real-world data often requires adjusting for population differences in order to draw meaningful inference. This problem is considered in the context of estimating mean outcomes and treatment effects in a well-defined target population using clinical data from a study population that differs from but overlaps with the target population in terms of patient characteristics. The current literature includes a variety of statistical models which generally require the correct specification of at least one parametric regression model. In this work, we propose the use of machine learning methods to estimate nuisance functions, incorporating these methods into existing doubly robust estimators. The resulting nonparametric estimators are -consistent, asymptotically normal, and asymptotically efficient under general conditions. Simulation results demonstrate that the proposed methods perform well in reasonable settings. These methods are also illustrated with a concrete cardiology example concerning standard of care for aortic stenosis. Finally, the ignorability assumption is examined through the development of global sensitivity analysis methods for two of the commonly used parametric approaches
Recommended from our members
Base-Calling of High-Throughput Sequencing Data Using a Random Effects Mixture Model
The emergence of high-throughput sequencing (HTS) technology has greatly influenced research in biological sciences including clinical applications such as in the understanding of disease etiology and pharmacogenomics. One widely used sequencing machine is the Illumina platform which uses a novel sequencing-by-synthesis method that involves chemical and optical imaging processes. The conversion of fluorescence intensity measures resulting from image processing to nucleotide bases is what is known as base-calling. The complex nature of sequencing-by-synthesis generates biases that affect accuracy of the sequenced DNA. Consequently, further analysis of sequences such as in genome assembly and variant detection may be directly influenced. Considering recently published methods to perform base-calling, it is evident that many methods perform transformations to the intensity data to reduce and/or eliminate biases. Thus, there is a need to model the original intensity data to maintain the information inherent within the data. Our novel method based on a Random Effects Mixture model, REMix, aims to capture the sequencing process while using the original data provided by the sequencing machine. Real data results demonstrate that REMix has the best balance of performance with respect to the validation metrics that are considered
An integrative framework for Bayesian variable selection with informative priors for identifying genes and pathways.
The discovery of genetic or genomic markers plays a central role in the development of personalized medicine. A notable challenge exists when dealing with the high dimensionality of the data sets, as thousands of genes or millions of genetic variants are collected on a relatively small number of subjects. Traditional gene-wise selection methods using univariate analyses face difficulty to incorporate correlational, structural, or functional structures amongst the molecular measures. For microarray gene expression data, we first summarize solutions in dealing with 'large p, small n' problems, and then propose an integrative Bayesian variable selection (iBVS) framework for simultaneously identifying causal or marker genes and regulatory pathways. A novel partial least squares (PLS) g-prior for iBVS is developed to allow the incorporation of prior knowledge on gene-gene interactions or functional relationships. From the point view of systems biology, iBVS enables user to directly target the joint effects of multiple genes and pathways in a hierarchical modeling diagram to predict disease status or phenotype. The estimated posterior selection probabilities offer probabilitic and biological interpretations. Both simulated data and a set of microarray data in predicting stroke status are used in validating the performance of iBVS in a Probit model with binary outcomes. iBVS offers a general framework for effective discovery of various molecular biomarkers by combining data-based statistics and knowledge-based priors. Guidelines on making posterior inferences, determining Bayesian significance levels, and improving computational efficiencies are also discussed
Recommended from our members
Sequential Multiple Testing for Variable Selection in High Dimensional Linear Model
Covariance test is proposed for testing the significance of the predictor variable that enters the current lasso model along the lasso solution path. In this paper, we propose the sequential multiple testing structure using covariance test p-values, which has good power properties with error rate controlled at a desired level. Specifically, we consider the full underlying hypotheses and the error rate control within each step as well as across all steps along the lasso solution path. Our sequential multiple hypothesis structure becomes valid because of the asymptotic distribution of the covariance test. And we prove that the minimum composite p-values under null hypothesis get larger along the steps, which is desirable when we apply step-down procedure for error rate control. The benefit of making use of the multivariate structure is, for some scenarios, to increase power while other procedures stop the selection too early. Also, our Hybrid procedures show higher power through simulations for weak signals and for high-dimensional data.To control FWER, we propose Hybrid Bonferroni-Holm Step-Down procedure along with Hybrid Hochberg-Holm and Hybrid Simes-Holm Step-Down procedures and compare them with StrongStop. To control FDR, we propose Hybrid Bonferroni-Benjamini-Liu Step-Down procedure and compare it with the ForwardStop, StrongStop, TailStop procedures and SLOPE . Simulation studies show that our proposed procedures have higher power with both FWER and FDR controlled at the desired level, especially for large scale and high-dimensional data, and very stable to use as well as in correlated design matrix.In our work, we first review the variable selection in statistics and most popular used variable selection methods. One of the most widely used and well developed method lasso is introduced in Chapter 2. The covariance test proposed for the significance test for lasso will be given in Chapter 3 , as well as its properties for orthogonal matrix X, and the asymptotic distribution. We then propose our sequential hypotheses multiple testing structure built on the lasso covariance test in Chapter 4. In Chapter 5, we develop the Hybrid Bonferroni-Holm Step-Down procedure to control FWER at alpha, and compare with StrongStop. In Chapter 6, we propose Hybrid Bonferroni-Benjamini-Liu Step-Down method and prove the FDR can be controlled at q. We also compare it with ForwardStop, StrongStop, TailStop and SLOPE through simulation studies. In Chapter 7, we show two applications of our proposed procedure on a diabetes data and a framingham heart study data and compare it with the other procedures. The conclusion and discussion are given in the last. The proofs for the main theorems and lemmas are included in the appendix
- …
