Freie Universität Berlin
Repository: Freie Universität Berlin (FU), Math Department (fu_mi_publications)Not a member yet
2251 research outputs found
Sort by
Analysis of Antigen Receptor Repertoires Captured by High Throughput Sequencing
In vertebrate species, the main mechanisms of defence against various types of pathogens are divided into the innate and the adaptive immune system. While the former relies on generic mechanisms, for example to detect the presence of bacterial cells, the latter features mechanisms that allow the individual to acquire defenses against specific, potentially novel features of pathogens and to maintain them throughout life. In a simplified sense, the adaptive immune system continuously generates new defenses against all kinds of structures randomly, carefully selecting them not to be reactive against the hosts own cells. The underlying generative mechanism is a unique somatic recombination process modifying the genes encoding the proteins responsible for the recognition of such foreign structures, the so-called antigen receptors. With the advances of high throughput DNA sequencing, we have gained the ability to capture the repertoire of different antigen receptor genes that an individual has acquired by selectively sequencing the recombined loci from a cell sample. This enables us to examine and explore the development and behaviour of the adaptive immune system in a new way, with a variety of potential medical applications. The main focus of this thesis is on two computational problems related to immune repertoire sequencing. Firstly, we developed a method to properly annotate the raw sequencing data that is generated in such experiments, taking into account various sources of biases and errors that either generally occur in the context of DNA sequencing or are specific for immune repertoire sequencing experiments. We will describe the algorithmic details of this method and then demonstrate its superiority in comparison with previously published methods on various datasets. Secondly, we developed a machine learning based workflow to interpret this data in the sense that we attempted to classify such recombined genes functionally using a previously trained model. We implemented alternative models within this workflow, which we will first describe formally and then assess their performances on real data in the context of a binary functional feature in T cells, namely whether they have differentiated into cytotoxic or helper T cells
Optimal data-driven estimation of generalized Markov state models for non-equilibrium dynamics
There are multiple ways in which a stochastic system can be out of statistical equilibrium. It might be subject to time-varying forcing; or be in a transient phase on its way towards equilibrium; it might even be in equilibrium without us noticing it, due to insufficient observations; and it even might be a system failing to admit an equilibrium distribution at all. We review some of the approaches that model the effective statistical behavior of equilibrium and non-equilibrium dynamical systems, and show that both cases can be considered under the unified framework of optimal low-rank approximation of so-called transfer operators. Particular attention is given to the connection between these methods, Markov state models, and the concept of metastability, further to the estimation of such reduced order models from finite simulation data. We illustrate our considerations by numerical examples
OpenPathSampling: A Python Framework for Path Sampling Simulations. 2. Building and Customizing Path Ensembles and Sample Schemes
The OpenPathSampling (OPS) package provides an easy-to-use framework to apply transition path sampling methodologies to complex molecular systems with a minimum of effort. Yet, the extensibility of OPS allows for the exploration of new path sampling algorithms by building on a variety of basic operations. In a companion paper [Swenson et al 2018] we introduced the basic concepts and the structure of the OPS package, and how it can be employed to perform standard transition path sampling and (replica exchange) transition interface sampling. In this paper, we elaborate on two theoretical developments that went into the design of OPS. The first development relates to the construction of path ensembles, the what is being sampled. We introduce a novel set-based notation for the path ensemble, which provides an alternative paradigm for constructing path ensembles, and allows building arbitrarily complex path ensembles from fundamental ones. The second fundamental development is the structure for the customisation of Monte Carlo procedures; how path ensembles are being sampled. We describe in detail the OPS objects that implement this approach to customization, the MoveScheme and the PathMover, and provide tools to create and manipulate these objects. We illustrate both the path ensemble building and sampling scheme customization with several examples. OPS thus facilitates both standard path sampling application in complex systems as well as the development of new path sampling methodology, beyond the default
Knock Control in Shockless Explosion Combustion by Extension of Excitation Time
Shockless Explosion Combustion is a novel constant volume combustion concept with an expected efficiency increase compared to conventional gas turbines. However, Shockless Explosion Combustion is prone to knocking because it is based on autoignition. This study investigates the potential of prolonging the excitation time of the combustible mixture by dilution with exhaust gas and steam to suppress detonation formation and mitigate knocking. Analyses of the characteristic chemical time scales by zero-dimensional reactor simulations show that the excitation time can be prolonged by dilution such that it exceeds the ignition delay time perturbation caused by a difference in initial temperature. This may suppress the formation of a detonation because less energy is fed into the pressure wave running ahead of the reaction front. One-dimensional simulations are performed to investigate reaction front propagation from a hot spot with various amounts of dilution. They demonstrate that dilution with exhaust gas or steam suppresses the formation of a detonation compared to the undiluted case, where a detonation ensues from the hot spot
Hybrid stochastic framework predicts efficacy of prophylaxis against HIV: An example with different dolutegravir regimen
A Bayesian approach to parameter identification in gas networks
The inverse problem of identifying the friction coefficient in an isothermal semilinear Euler system is considered. Adopting a Bayesian approach, the goal is to identify the distribution of the quantity of interest based on a finite number of noisy measurements of the pressure at the
boundaries of the domain. First well-posedness of the underlying non-linear PDE system is shown using semigroup theory, and then Lipschitz continuity of the solution operator with respect to the friction coefficient is established. Based on the Lipschitz property, well-posedness of the resulting
Bayesian inverse problem for the identification of the friction coefficient is inferred. Numerical tests
for scalar and distributed parameters are performed to validate the theoretical results.
Key words
Axiomatic approach to variable kernel density estimation
Variable kernel density estimation allows the approximation of a probability density by the mean of differently stretched and rotated kernels centered at given sampling points yn∈Rd, n=1,…,N. Up to now, the choice of the corresponding bandwidth matrices hn has relied mainly on asymptotic arguments, like the minimization of the asymptotic mean integrated squared error (AMISE), which work well for large numbers of sampling points. However, in practice, one is often confronted with small to moderately sized sample sets far below the asymptotic regime, which highly restricts the usability of such methods. As an alternative to this asymptotic reasoning we suggest an axiomatic approach which guarantees invariance of the density estimate under linear transformations of the original density (and the sampling points) as well as under splitting of the density into several `well-separated' parts. In order to still ensure proper asymptotic behavior of the estimate, we \emph{postulate} the typical dependence hn∝N−1/(d+4). Further, we derive a new bandwidths selection rule which satisfies these axioms and performs considerably better than conventional ones in an artificially intricate two-dimensional example as well as in a real life example
Convergence to Equilibrium in Energy-Reaction–Diffusion Systems Using Vector-Valued Functional Inequalities
Abstract We discuss how the recently developed energy dissipation methods for
reaction diffusion systems can be generalized to the non-isothermal case. For this, we
use concave entropies in terms of the densities of the species and the internal energy,
where the importance is that the equilibrium densities may depend on the internal
energy. Using the log-Sobolev estimate and variants for lower-order entropies as well
as estimates for the entropy production of the nonlinear reactions, we give two methods
to estimate the relative entropy by the total entropy production, namely a somewhat
restrictive convexity method, which provides explicit decay rates, and a very general,
but weaker compactness method
Toward a computationally-tractable maximum entropy principle for non-stationary financial time series
Statistical analysis of financial time series of equity returns can be hindered by various unobserved factors, resulting in a nonstationarity of the overall problem. Parametric methods approach the problem by restricting it to a certain (stationary) distribution class through various assumptions, which often result in a model misspecification when the underlying process does not belong to this predefined parametric class. Nonparametric methods are more general but can lead to ill-posed problems and computationally expensive numerical schemes. This paper presents a nonparametric methodology addressing these issues in a computationally tractable way by combining such key concepts as the maximum entropy principle for nonparametric density estimation and a Lasso regularization technique for numerical identification of redundant parameters and persistent latent regimes. In the context of volatility modeling, the presented approach identifies an a priori unknown nonparametric persistent regime transition process switching between distinct local nonparametric i.i.d. volatility distribution regimes with maximum entropy. Using historical return data for an equity index, we demonstrate that despite viewing the data as serially independent conditional on the latent regimes (serial dependence is introduced only via the stochastic latent regime), our methodology leads to identification of robust models that are superior to considered conditional heteroskedasticity models when compared to the model log-likelihood, Akaike, and Bayesian information criteria for the analyzed data.
Read More: https://epubs.siam.org/doi/abs/10.1137/17M1142600?journalCode=sjfmb