1,721,053 research outputs found
Impact of phylogeny on inference from protein sequences: from models to natural data
Proteins are macromolecules considered as the building blocks of cells because they are at the heart of almost every cellular task, ranging from chemical to mechanical processes. For instance, proteins play a role as enzymes for digestion, transporters of oxygen in the blood, antibodies against viruses, propellers for bacteria, etc. Understanding proteins allows a systems-level understanding of the cell, thus providing insights on a large number of phenomena in biology or medicine. A protein can be described as a polymer, i.e.\ a linear chain of monomers, which are amino acids. In addition, the polymer folds into a specific three-dimensional structure that is crucial to the protein's function. Inversely, a given function can induce a conformational change in its structure. Importantly, the function of a protein is often mediated through interactions with other proteins or molecules. Proteins with similar function and structure are present in many different species. This means that there is an evolutionary relatedness, termed phylogeny, between proteins that is reflected in their composition.
In this thesis, we aim at studying the impact of phylogeny on inference methods from sequences of homologous protein families. We are interested in methods predicting structural contacts in proteins, but also partners of interactions, and functional groups of amino acids, termed sectors. We use methods that rely on correlations in multiple sequence alignments of proteins, specifically Potts models which are global statistical models, and local methods such as covariance or mutual information. A challenge is that it is difficult to disentangle phylogenetic from functional or structural correlations in natural sequences. To overcome this, we start by generating synthetic sequences using a minimal model where the amount of phylogeny and selection for a function can be tuned. These sequences allow to assess the performance of prediction methods while tuning phylogenetic or functional correlations. We then show that we recover the findings from synthetic sequences in natural or realistic sequences.
For the inference of structural contacts, we find that phylogenetic correlations are deleterious for local and global methods, but global methods are more robust to them. In contrast, for the inference of partners of interaction, we find that performance is improved by combining phylogenetic and functional correlations. For the inference of sectors, we find that performance is decreased when including phylogenetic correlations on top of functional ones. Finally, we observe that out-of-equilibrium noise associated to variations of selection pressure can impact positively the inference of structural contacts.
These findings show the interplay of phylogeny and functional or structural constraints in protein sequences and in inference methods from protein sequences. They illustrate the complexity and the rich structure of biological data, and show that it can be dissected and understood.UPBITBO
Impact of spatial structure and finite size on the evolution and ecology of asexual microbial populations
Evolution is the process by which organisms are modified over time. The fundamental mechanisms underlying evolution are mutations, natural selection and genetic drift. While natural selection favors fitter individuals, genetic drift, which arises from finite size effects, yields random changes in allele frequencies. A key factor impacting the amplitude of these fluctuations, and more broadly evolutionary dynamics, is the spatial structure of populations. When a population is subdivided into demes of finite size connected by migrations, the way it explores its fitness landscape differs from the one of well-mixed populations, in a way that depends on the specific characteristics of the spatial structure. Indeed, the fixation probability of the lineage of a mutant is structure- dependent. In addition, spatial structure affects ecological interactions between individuals, which ultimately impacts evolutionary dynamics. Typically, the more subdivided a population, the sparser the physical interactions between individuals, which tends to reduce competition between them.
In this thesis, we aim to assess the influence of spatial structure and finite size effects on the evolution and ecology of asexual microbial populations. We first examine the impact of population size on the exploration of rugged fitness landscapes by well-mixed populations, using a biased random walk model. Then, we turn our attention to the evolution of structured populations in rugged fitness landscapes and investigate the impact of spatial structure on the way bacterial populations explore their fitness landscapes. Finally, we develop a lattice-gas agent-based model to describe a system of Vibrio cholerae bacteria in liquid conditions. These bacteria interact with each other through type IV pili (appendages used by cells to bind to each other) and the type VI secretion system (a kin-discriminatory molecular weapon). In this system, the spatial organization is key to understanding the extent of non-kin depletion caused by cells carrying a functioning type VI secretion system. Therefore, spatial structure can strongly impact the ecology and evolution of V. cholerae.
Regarding the evolution of well-mixed populations, we find that adaptation is often more efficient at finite population size. Specifically, there is often a finite population size optimizing the search for high fitness peaks. This result highlights the significant role of finite size effects in evolution. In the case of spatially structured populations, we find that adaptation is often more efficient when there is a migration asymmetry associated with the presence of suppression of selection within the structure. We further find that the more suppression of selection, the smaller the effective population size for steady-state properties which matter for long-term evolution. Finally, the experimental observations about the interplay between type IV pili and the type VI secretion system, which yields a predator-prey dynamics in the system considered, are well reproduced by simulations from our model. Our results suggest that the spatial organization of the population, which is shaped by type-IV-pili-mediated interactions, is key to understanding the extent of the killing of prey by predators. Typically, if predators do not bind to prey, killing is largely prevented.UPBITBO
Revealing and exploiting coevolution through protein language models
Protein sequences carry rich evolutionary information that reflects both structural and functional constraints, as well as shared ancestry. In this thesis, I explore how protein language models (pLMs) can uncover and leverage these signals, particularly coevolutionary information, through the use of multiple sequence alignments (MSAs). I also investigate how these models can be made homology-aware in alignment-independent ways, thus expanding their applicability in protein modeling. First, I show that transformer-based pLMs trained on MSAs capture detailed phylogenetic relationships. Specifically, column attention patterns in MSA Transformer correlate strongly with sequence similarity, revealing a natural separation of phylogenetic and coevolutionary signals within the model's architecture. Next, also building on MSA Transformer, I develop an iterative generation method that produces realistic and diverse protein sequences. These synthetic sequences not only match natural sequences in evaluation metrics but also outperform those generated by traditional models, particularly when working with shallow MSAs. Next, I address the problem of pairing interacting protein sequences, which is crucial for predicting the structures of protein complexes. I introduce DiffPALM, a method that pairs interacting paralogs by minimizing the masked language modeling loss of paired MSAs, in a differentiable way. This unsupervised approach improves pairing accuracy and enhances the structure prediction of some protein complexes when used as input to AlphaFold-Multimer. Alignment methods are often imperfect. To overcome the limitations of MSAs, I introduce ProtMamba, a lightweight, alignment-free model capable of processing long concatenations of homologous sequences. ProtMamba matches or surpasses the performance of larger transformer models on tasks such as sequence generation and fitness prediction, while featuring greater computational efficiency. Finally, I present RAG-ESM, a retrieval-augmented framework that adds homology awareness to pretrained single-sequence pLMs by conditioning on retrieved homologs through cross-attention. RAG-ESM achieves improved prediction and generation performance with minimal computational overhead and reveals emergent sequence alignment capabilities, making it a strong candidate for scalable protein design applications. Together, these contributions demonstrate how protein language models can reveal, disentangle, and harness coevolutionary signals in different ways, and offer new paths forward for computational protein science.UPBITBO
Quantifying bacterial responses to antibiotics at the single-cell level
The emergence of pathogen resistance to antimicrobials is putting modern medicine at risk. One of the main challenges in the treatment of bacterial infections is that current antibiotics often fail to eradicate the whole bacterial population, driven by the constant misuse and overuse of these compounds. Even genetically identical cells can take on highly heterogeneous physiological states resulting in bacterial subpopulations being less susceptible to the treatment. In particular, some cells are in physiological states that allow them to survive antibiotic treatment without any resistance mutations. For instance, slow-growing cells have been reported to be more tolerant to some antibiotics, and this increased tolerance can facilitate the subsequent fixation of resistance mutations. Most methods for the discovery and study of antimicrobial compounds are based on liquid cultures which focus on the antibiotic response at the population level, being blind to single-cell dynamics. Consequently, the determinants of sensitivity to antibiotics are only poorly understood at the single-cell level due to the lack of quantitative data.
In recent years, powerful methods have been developed to quantitatively measure behaviour and responses in single bacterial cells. By combining microfluidics with time-lapse microscopy, it is possible to track growth, gene expression, division, and death within lineages of single cells. An especially attractive microfluidic design is the Mother Machine, a device where bacteria grow within narrow growth channels that are perpendicularly connected to a main flow channel, which supplies nutrients and washes away cells growing out of the growth channels. In this work, we investigate the response of bacterial cells to antibiotics using an integrated microfluidic and computational setup: the dual-input Mother Machine (DIMM). The DIMM allows arbitrary time-varying mixtures of two input media, such that cells can be exposed to a controlled set of varying external conditions. Using the companion image analysis software Mother Machine Analyser (MoMA) we can segment and track cell lineages from phase-contrast images with high throughput and accuracy.
In light of this work, we developed new multiplexed microfluidic designs using the new PyMicrofluidics tool, which facilitates the drawing and handling of complex circuits. These new designs, enable the study of multiple conditions in parallel. Furthermore, we introduced filtering structures in the Mother Machine channels. The media carrying nutrients or antibiotics flows from the main channel through the channels where the cells are trapped, without letting the cells escape. This grants the ability to load cells inside the Mother Machine faster, as well as reduce the nutrient-gradient effect caused by the delivery of media only through diffusion in classical dead-end channels.
We use this integrated setup to quantify how the antibiotic response of individual bacteria depends on their physiological state at the time the treatment commences. The time resolution achieved with such technology allows to track how the bacteria evolve during treatment and after the treatment to assess survival. For this, we focus on treating Escherichia coli (E. coli) with a variety of clinically relevant antibiotics at different concentrations. The methods we present in this work will allow the identification of antimicrobial compounds that specifically target these resistant subpopulations, which could in the future complement existing treatment strategies and have a potential impact on antimicrobial drug discovery and treatment design
Statistique et dynamique des membranes complexes
This thesis deals with biological membranes, studied from the point of view of theoretical physics. We focus on some generic effects of the presence of one or two membrane inclusions, e.g., proteins, or of a local chemical change of the environment of the membrane. First, we study the Casimir-like interaction between two membrane inclusions, which arises from the constraints imposed by the inclusions on the thermal fluctuations of the shape of the membrane. We calculate the fluctuations of the Casimir-like force between two point-like inclusions. We clarify the definition of the force exerted by a correlated fluid on an inclusion, in a microstate of the fluid. This definition plays a key art in studies of the Casimir-like force beyond its thermal equilibrium value. We also study the Casimir-like interaction between rod-shaped membrane inclusions. Then, we investigate membrane elasticity at the nanoscale, which is involved in local membrane thickness deformations in the vicinity of proteins. We put forward the importance of an energetic term that is neglected in existing models. Finally, we present a theoretical description of the dynamics of a membrane submitted to a local chemical perturbation of its environment, starting from first principles. We compare our theoretical predictions to new experimental results regarding the dynamical deformation of a biomimetic membrane submitted to a local pH increase.Cette thèse porte sur les membranes biologiques, étudiées du point de vue de la physique théorique. Nous nous intéressons aux effets génériques de la présence d'une ou deux inclusions membranaires, par exemple des protéines, et à ceux d'une modification chimique locale de l'environnement de la membrane. Tout d'abord, nous étudions une interaction entre deux inclusions membranaires, qui est analogue à la force de Casimir : elle provient des contraintes que les inclusions imposent aux fluctuations thermiques de la forme de la membrane. Nous calculons les fluctuations de cette force entre deux inclusions ponctuelles. Nous clarifions la définition de la force exercée par un fluide corrélé sur une inclusion, dans un micro-état du fluide. Cette définition joue un rôle clé dans l'étude de la force de Casimir au-delà de sa valeur moyenne à l'équilibre. Nous étudions également les interactions de Casimir entre des inclusions membranaires de forme allongée. Ensuite, nous nous intéressons à l'élasticité membranaire à l'échelle nanométrique, qui est mise en jeu dans les déformations locales de l'épaisseur de la membrane à proximité de certaines protéines. Nous soulignons l'importance d'un terme énergétique qui a été négligé jusqu'à présent. Enfin, nous présentons une description théorique, développée à partir de principes fondamentaux, de la dynamique d'une membrane soumise à une perturbation chimique locale de son environnement. Nous comparons nos prévisions théoriques à de nouveaux résultats expérimentaux portant sur la déformation dynamique d'une membrane biomimétique soumise à une augmentation locale de pH
Inferring interaction partners from protein sequences using mutual information
Functional protein-protein interactions are crucial in most cellular processes. They enable multi-protein complexes to assemble and to remain stable, and they allow signal transduction in various pathways. Functional interactions between proteins result in coevolution between the interacting partners, and thus in correlations between their sequences. Pairwise maximum-entropy based models have enabled successful inference of pairs of amino-acid residues that are in contact in the three-dimensional structure of multi-protein complexes, starting from the correlations in the sequence data of known interaction partners. Recently, algorithms inspired by these methods have been developed to identify which proteins are functional interaction partners among the paralogous proteins of two families, starting from sequence data alone. Here, we demonstrate that a slightly higher performance for partner identification can be reached by an approximate maximization of the mutual information between the sequence alignments of the two protein families. Our mutual information-based method also provides signatures of the existence of interactions between protein families. These results stand in contrast with structure prediction of proteins and of multi-protein complexes from sequence data, where pairwise maximum-entropy based global statistical models substantially improve performance compared to mutual information. Our findings entail that the statistical dependences allowing interaction partner prediction from sequence data are not restricted to the residue pairs that are in direct contact at the interface between the partner proteins.UPBITBO
eLife Assessment: A differentiable Gillespie algorithm for simulating chemical kinetics, parameter estimation, and designing synthetic biological circuits
UPBITBO
eLife Assessment: Exploring the repository of de novo-designed bifunctional antimicrobial peptides through deep learning
UPBITBO
Optimization and historical contingency in protein sequences
Protein sequences are shaped by functional optimization on the one hand and by evolutionary history, i.e. phylogeny, on the other hand. A multiple sequence alignment of homologous proteins contains sequences which evolved from the same ancestral sequence and have similar structure and function. In such an alignment, correlations in amino acid usage at different sites can arise from structural and functional constraints due to coevolution, but also from historical contingency. Correlations arising from phylogeny often confound coevolution signal from functional or structural optimization, impairing the inference of structural contacts from sequences. However, inferred Potts models are more robust than local statistics to these effects, which may explain their success. Dedicated corrections can further increase this robustness. Moreover, phylogenetic correlations can in fact provide useful information for some inference tasks, especially to infer interaction partners from sequences among the paralogs of two protein families. In this case, signal from phylogeny and signal from constraints combine constructively, and explicitly exploiting both further improves inference performance. Protein language models have recently been applied to sequence data, greatly advancing structure, function and mutational effect prediction. Language models trained on multiple sequence alignments capture coevolution and structural contacts, but also phylogenetic relationships. They are able to disentangle signal from structural constraints and from phylogeny more efficiently than Potts models, and they have promising generative properties. Furthermore, they allow predicting interacting partners from protein sequences, outperforming traditional coevolution methods on difficult datasets.UPBITBO
- …
