Freie Universität Berlin
Repository: Freie Universität Berlin (FU), Math Department (fu_mi_publications)Not a member yet
2251 research outputs found
Sort by
Dynamic Evolution of a Transient Supersonic Trailing Jet Induced by a Strong Incident Shock Wave
The dynamic evolution of a highly underexpanded transient supersonic jet at the exit of a pulse detonation engine is investigated via high-resolution time-resolved schlieren and numerical simulations. Experimental evidence is provided for the presence of a second triple shock configuration along with a shocklet between the reflected shock and the slipstream, which has no analog in a steady-state underexpanded jet. A pseudo-steady model is developed, which allows for the determination of the postshock flow condition for a transient propagating oblique shock. This model is applied to the numerical simulations to reveal the mechanism leading to the formation of the second triple point. Accordingly, the formation of the triple point is initiated by the transient motion of the reflected shock, which is induced by the convection of the vortex ring. While the vortex ring embedded shock move essentially as a translating strong oblique shock, the reflected shock is rotating towards its steady-state position. This results in a pressure discontinuity that must be resolved by the formation of a shocklet
Concentration and chemical form of dietary zinc shape the porcine colon microbiome, its functional capacity and antibiotic resistance gene repertoire
Despite a well-documented effect of high dietary zinc oxide on the pig intestinal microbiota composition less is it yet known about changes in microbial functional properties or the effect of organic zinc sources. Forty weaning piglets in four groups were fed diets supplemented with 40 or 110 ppm zinc as zinc oxide, 110 ppm as Zn-Lysinate, or 2500 ppm as zinc oxide. Host zinc homeostasis, intestinal zinc fractions, and ileal nutrient digestibility were determined as main nutritional and physiological factors putatively driving colon microbial ecology. Metagenomic sequencing of colon microbiota revealed only clear differences at genus level for the group receiving 2500 ppm zinc oxide. However, a clear group differentiation according to dietary zinc concentration and source was observed at species level. Functional analysis revealed significant differences in genes related to stress response, mineral, and carbohydrate metabolism. Taxonomic and functional gene differences were accompanied with clear effects in microbial metabolite concentration. Finally, a selection of certain antibiotic resistance genes by dietary zinc was observed. This study sheds further light onto the consequences of concentration and chemical form of dietary zinc on microbial ecology measures and the resistome in the porcine colon
Objective priors in the empirical Bayes framework
When dealing with Bayesian inference the choice of the prior often remains a debatable question. Empirical Bayes methods offer a data-driven solution to this problem by estimating the prior itself from an ensemble of data. In the nonparametric case, the maximum likelihood estimate is known to overfit the data, an issue that is commonly tackled by regularization. However, the majority of regularizations are ad hoc choices which lack invariance under reparametrization of the model and result in inconsistent estimates for equivalent models. We introduce a nonparametric, transformation-invariant estimator for the prior distribution. Being defined in terms of the missing information similar to the reference prior, it can be seen as an extension of the latter to the data-driven setting. This implies a natural interpretation as a trade-off between choosing the least informative prior and incorporating the information provided by the data, a symbiosis between the objective and empirical Bayes methodologies
Extending Transition Path Theory: Periodically Driven and Finite-Time Dynamics
Abstract
Given two distinct subsets A, B in the state space of some dynamical system, transition
path theory (TPT) was successfully used to describe the statistical behavior of
transitions from A to B in the ergodic limit of the stationary system.We derive generalizations
of TPT that remove the requirements of stationarity and of the ergodic limit
and provide this powerful tool for the analysis of other dynamical scenarios: periodically
forced dynamics and time-dependent finite-time systems. This is partially
motivated by studying applications such as climate, ocean, and social dynamics. On
simple model examples, we show how the new tools are able to deliver quantitative
understanding about the statistical behavior of such systems.We also point out explicit
cases where the more general dynamical regimes show different behaviors to their stationary
counterparts, linking these tools directly to bifurcations in non-deterministic
systems
The Impact of Halogenated Phenylalanine Derivatives on NFGAIL Amyloid Formation
The hexapeptide hIAPP22–27 (NFGAIL) is known as a crucial
amyloid core sequence of the human islet amyloid polypeptide
(hIAPP) whose aggregates can be used to better understand
the wild-type hIAPP’s toxicity to β-cell death. In amyloid
research, the role of hydrophobic and aromatic-aromatic
interactions as potential driving forces during the aggregation
process is controversially discussed not only in case of NFGAIL,
but also for amyloidogenic peptides in general. We have used
halogenation of the aromatic residue as a strategy to modulate
hydrophobic and aromatic-aromatic interactions and prepared
a library of NFGAIL variants containing fluorinated and
iodinated phenylalanine analogues. We used thioflavin T
staining, transmission electron microscopy (TEM) and smallangle
X-ray scattering (SAXS) to study the impact of side-chain
halogenation on NFGAIL amyloid formation kinetics. Our data
revealed a synergy between aggregation behavior and hydrophobicity
of the phenylalanine residue. This study introduces
systematic fluorination as a toolbox to further investigate the
nature of the amyloid self-assembly process
"Particle-Continuum Coupling and its Scaling Regimes: Theory and Applications"
This work is motivated by the goal of designing simulation software for technical devices that, at their functional core, rely on atomistic‐scale processes embedded in a larger‐scale fluid environment. The core of the problem is the conceptual and technical approach for coupling particle and continuum representations of a fluid. The state of the art for key aspects including physical modeling, mathematical formalization, computational implementation, and applications, is discussed and organized in a consistent picture across the relevant physical regimes
Multiscale Shear Forcing of Turbulence in the Nocturnal Boundary Layer: A Statistical Analysis
The lower nocturnal boundary layer is governed by intermittent turbulence which is thought to be triggered by sporadic activity of so-called sub-mesoscale motions in a complex way. We analyze intermittent turbulence based on an assumed relation between the vertical gradients of the sub-mean scales and turbulence kinetic energy. We analyze high-resolution nocturnal eddy-correlation data from 30-m tower collected during the Fluxes over Snow Surfaces II field program. The non-turbulent velocity signal is decomposed using a discrete wavelet transform into three ranges of scales interpreted as the mean, jet and sub-mesoscales. The vertical gradients of the sub-mean scales are estimated using finite differences. The turbulence kinetic energy is modelled as a discrete-time autoregressive process with exogenous variables, where the latter ones are the vertical gradients of the sub-mean scales. The parameters of the discrete model evolve in time depending on the locally-dominant turbulence-production scales. The three regimes with averaged model parameters are estimated using a subspace-clustering algorithm which illustrates a weak bimodal distribution in the energy phase space of turbulence and sub-mesoscale motions for the very stable boundary layer. One mode indicates turbulence modulated by sub-mesoscale motions. Furthermore, intermittent turbulence appears if the sub-mesoscale intensity exceeds 10% of the mean kinetic energy in strong stratification
ganon: precise metagenomics classification against large and up-to-date sets of reference sequences
Motivation:
The exponential growth of assembled genome sequences greatly benefits metagenomics studies. However, currently available methods struggle to manage the increasing amount of sequences and their frequent updates. Indexing the current RefSeq can take days and hundreds of GB of memory on large servers. Few methods address these issues thus far, and even though many can theoretically handle large amounts of references, time/memory requirements are prohibitive in practice. As a result, many studies that require sequence classification use often outdated and almost never truly up-to-date indices.
Results:
Motivated by those limitations, we created ganon, a k-mer-based read classification tool that uses Interleaved Bloom Filters in conjunction with a taxonomic clustering and a k-mer counting/filtering scheme. Ganon provides an efficient method for indexing references, keeping them updated. It requires <55 min to index the complete RefSeq of bacteria, archaea, fungi and viruses. The tool can further keep these indices up-to-date in a fraction of the time necessary to create them. Ganon makes it possible to query against very large reference sets and therefore it classifies significantly more reads and identifies more species than similar methods. When classifying a high-complexity CAMI challenge dataset against complete genomes from RefSeq, ganon shows strongly increased precision with equal or better sensitivity compared with state-of-the-art tools. With the same dataset against the complete RefSeq, ganon improved the F1-score by 65% at the genus level. It supports taxonomy- and assembly-level classification, multiple indices and hierarchical classification.
Availability and implementation:
The software is open-source and available at: https://gitlab.com/rki_bioinformatics/ganon
GenMap: ultra-fast computation of genome mappability
Motivation:
Computing the uniqueness of k-mers for each position of a genome while allowing for up to e mismatches is computationally challenging. However, it is crucial for many biological applications such as the design of guide RNA for CRISPR experiments. More formally, the uniqueness or (k, e)-mappability can be described for every position as the reciprocal value of how often this k-mer occurs approximately in the genome, i.e. with up to e mismatches.
Results:
We present a fast method GenMap to compute the (k, e)-mappability. We extend the mappability algorithm, such that it can also be computed across multiple genomes where a k-mer occurrence is only counted once per genome. This allows for the computation of marker sequences or finding candidates for probe design by identifying approximate k-mers that are unique to a genome or that are present in all genomes. GenMap supports different formats such as binary output, wig and bed files as well as csv files to export the location of all approximate k-mers for each genomic position.
Availability and implementation:
GenMap can be installed via bioconda. Binaries and C++ source code are available on https://github.com/cpockrandt/genmap
On a Scalable Entropic Breaching of the Overtting Barrier for Small Data Problems in Machine Learning.
Overfitting and treatment of small data are among the most challenging problems in machine learning (ML), when a relatively small data statistics size T is not enough to provide a robust ML fit for a relatively large data feature dimension D. Deploying a massively parallel ML analysis of generic classification problems for different D and T, we demonstrate the existence of statistically significant linear overfitting barriers for common ML methods. The results reveal that for a robust classification of bioinformatics-motivated generic problems with the long short-term memory deep learning classifier (LSTM), one needs in the best case a statistics T that is at least 13.8 times larger than the feature dimension D. We show that this overfitting barrier can be breached at a 10−12 fraction of the computational cost by means of the entropy-optimal scalable probabilistic approximations algorithm (eSPA), performing a joint solution of the entropy-optimal Bayesian network inference and feature space segmentation problems. Application of eSPA to experimental single cell RNA sequencing data exhibits a 30-fold classification performance boost when compared to standard bioinformatics tools and a 7-fold boost when compared to the deep learning LSTM classifier