1,721,091 research outputs found

    Dataset for Shear fronts in shear-thickening suspensions

    No full text
    This dataset contains the data which are used for generating Fig.1 to Fig.4 in the reearch paper below. Data is given in separate Excel files, with seperate worksheets for the subfigures. The dataset supports the publication: Endao Han, Matthieu Wyart, Ivo Peters and Heinrich Jaeger (2018) &#39;Shear fronts in shear-thickening suspensions&#39; in Physical Review Fluids</span

    Shear fronts in shear-thickening suspensions

    No full text
    We study the fronts that appear when a shear-thickening suspension is submitted to a sudden driving force at a boundary. Using a quasi-one-dimensional experimental geometry, we extract the front shape and the propagation speed from the suspension flow field and map out their dependence on applied shear. We find that the relation between stress and velocity is quadratic, as is generally true for inertial effects in liquids, but with a pre-factor that can be much larger than the material density. We show that these experimental findings can be explained by an extension of a phenomenological model originally developed to describe steady-state shear-thickening. This is achieved by introducing a sole additional parameter: the characteristic strain scale that controls the crossover from start-up response to steady-state behavior. The theoretical framework we obtain points out a linkage between transient and steady-state properties of shear-thickening materials

    How data structures affect generalization in Kernel Methods and Deep Learning

    No full text
    Artificial Intelligence has revolutionized numerous fields, driving advancements in healthcare, finance, and autonomous systems. At the core of this revolution lies Deep Learning---algorithms that learn representations from data through computational layers of artificial neurons. For deep learning to succeed, it must extract meaningful information from data rather than merely memorizing it. Otherwise, it would require exponentially large training datasets to learn a task---an impractical scenario called curse of dimensionality. Deep networks must overcome this by extracting relevant features from the structure of the training data, enabling them to generalize to new test data. What types of data structures do deep networks learn? How do they use these structures to overcome the curse of dimensionality? To investigate these questions, we employ toy models that abstract real-world data features, allowing for theoretical analysis. These simplified models capture the phenomena observed on real data, offering insights into their behavior and generalization. In a simple classification task where data are clustered by labels, existing theories predict the performance of deep networks, in a regime where they do not learn features. However, these theories are justified for high-dimensional data, and their validity in lower dimensions is uncertain. In low dimensions, we identify a crossover between the success and failure of these theories, also observed in real-world data. Nevertheless, on this task, deep networks exhibit poor generalization, constrained by the curse of dimensionality. To explain their performance on real data, we consider additional structures that networks can leverage in the feature-learning regime. Our hypothesis is that these structures introduce invariances, enabling deep networks to simplify tasks. For instance, in image classification, recognizing a dog does not require to match an idealized representation; a distorted image is often sufficient. Recent findings suggest that the best-performing networks are those less sensitive to small deformations. We hypothesize this is because images are composed of local and sparse features like edges and patterns, whose exact relative positions are irrelevant for classification. This irrelevance introduces invariance to small deformations. To test this, we construct toy models of local and sparse features, and quantify how invariance to small deformations develops with training. To get good performance, we hypothesize that deep networks combine these features to infer the meaning of images in a hierarchical way. For example, recognizing a dog involves assembling edges into limbs and facial components, which combine to form a complete structure. Such hierarchical structures allow for synonymic variations---for instance, whether the dog's eyes are open or closed---without altering the task. To explore whether deep networks learn invariances from hierarchical structures and relate them to invariance to small deformations, we introduce synthetic datasets exhibiting both invariances. Our findings show that deep networks learn these invariances layer by layer, using exactly the same training set size required to learn the task. Our theoretical investigation quantifies this number, and demonstrates how deep networks are able to generalize well by constructing internal representations increasingly insensitive to the task-specific invariances.PCS

    Loss landscape and symmetries in Neural Networks

    No full text
    Neural networks (NNs) have been very successful in a variety of tasks ranging from machine translation to image classification. Despite their success, the reasons for their performance are still not well-understood. This thesis explores two main themes: loss landscapes and symmetries present in data. Machine learning consists of training models on data by optimizing the model parameters. This optimization is done by minimizing a loss function. NNs, a family of machine learning models, are created by composing functions, called layers. Informally, they can be visualized as a set of interconnected neurons. Ten years ago, NNs became the most popular models of machine learning. With their success come many open questions. For example, neural networks and glassy systems both have many degrees of freedom and highly non-convex objective or energy functions, respectively. However, glassy systems get stuck in local minima near where they are initialized, whereas neural networks avoid this even when they 100s of times more parameters than the number of data use to train them? (i) What drives this difference in behavior? (ii) How is it then that NNs do not become too specialized to the training data (overfitting)? In the first part of this thesis, we show that in classification tasks, NNs undergo a jamming transition dependent on the number of parameters, NN. This answers (i): With a sufficiently high NN above a critical number NN^*, local minima are avoided. Then, we establish a "double-descent" behavior in the test error of classification tasks: It decreases twice as a function of NN, before NN^* but also after, until infinity, where it converges to its minimum. We answer (ii) by explaining the origins of this double-descent. Finally, we introduce a phase diagram that describes the landscape of the loss function and unifies the two limits in which a neural network can converge when sending NN to infinity. In the second part of this thesis, we explore the issue of the curse of dimensionality (CD): Sampling a dd-dimensional space requires an exponential number of points PP. However, NNs perform well even for Pexp(d)P \ll \exp(d). Symmetries in the data play a role in this conundrum. For example, to process images we use convolutional NNs (CNNs) which have the property of being locally connected and equivariant with respect to translations, i.e., a translation in the input leads to a corresponding translation in the output. Although empirical experience suggests that locality and equivariance contribute to the success of CNNs, it is difficult to understand how. Indeed, equivariance reduces the dimensionality of the data only slightly. Stability toward diffeomorphisms however might be the key to CD. We studied how NNs are affected by images distorted by diffeomorphisms. Our results suggest that locality and equivariance allow, during learning, to develop stability towards diffeomorphisms \textit{relative} to other generic transformations. Following this intuition, we have created new architectures by extending CNNs properties to 3D rotations. Our work contributes to the current understanding of the behavior of neural networks empirically observed by machine learning practitioners. Moreover, the architectures developed for 3D rotation problems are currently being applied to a wide range of domains.PCS

    Mechanics and co-evolution of allosteric materials and proteins

    No full text
    The regulation of several processes inside and outside the cell depends on the action of a particular class of enzymes, called allosteric. In allosteric macromolecules, binding a ligand at one site affects the binding activity at a distal functional site, providing a reliable tool to regulate the corresponding function. The physical mechanisms underpinning allostery and its long-range communication are not yet fully understood, despite a great number of advances were made possible by significant works spanning 60 years, in between biology, bioinformatics and physics. In physics terms, proteins can be viewed as amorphous materials that however underwent billions of years of evolution to be functional as observed today. The framework introduced in this dissertation allows to explore how the structural organisation of an allosteric system is constrained by the function that it has evolved to perform. It allows a classification of allosteric architectures and suggests a physical explanation behind the emergence of such long-range allosteric coupling. Furthermore, it is also apt to build in silico a large amount of allosteric architectures that share the same evolutionary history. The constraints imprinted by evolution on sequences that share a common ancestor motivate the exploration of inference methods that try to predict the fitness of a protein solely from the knowledge of sequences. The strategy used to build this framework is to resort to a coarse-grained model of a protein, on the line of elastic network models resulted successful in the description of the large-scale dynamics of proteins. Ideas on how to pursue these research directions further are discussed throughout the chapters. Firstly, we introduce an in-silico model for the evolution of allosteric behaviour in discrete lattices of harmonic springs. The in-silico evolution is performed for two different allosteric tasks: one optimising for the transmission of strain between the allosteric and active site, while the other maximising the cooperative binding energy between the two. To optimise the transmission of strain, the network develops a lever that amplifies the response at the active site. In such a way, our model proposes a novel allosteric architecture, potentially in use in proteins as well. Cooperative architectures show, among others, hinge and shear motions and rationalise the observation of a low energy mode that describes conformational changes in proteins. Indeed, to achieve proper function, the mode is predicted to get softer as the size of the system increases. This prediction is tested by collecting a database of 34 high resolution structures of allosteric proteins and is proven valid even when elastic nonlinearities are introduced. Secondly, the sequences generated with the in-silico model serve to benchmark existing methods that infer co-evolutionary couplings between amino acids, proven to be successful in predicting local structural constraints, but with unclear performance in the presence of global allosteric constraints. These models do predict local features reflecting structure, but fail in the prediction of long-range functional dependencies and are not able to generate synthetic sequences that function as native ones. Thus, the exploration of new directions is needed.PCS

    Local excitations in amorphous solids

    No full text
    Amorphous solids are structurally disordered. They are very common and include glasses, colloids, and granular materials, but are far less understood than crystalline solids. Key aspects of these materials are controlled by the presence of excitations in which a group of particles rearranges. This motion can be triggered by (a) quantum fluctuations associated with two-level systems (TLS), which dominate the low temperature properties of conventional glasses and have practical importance on superconducting qubits; by (b) thermal fluctuations associated with activations, which are related to the famous and challenging ``glass transition'' problem; or by (c) exerting an external stress or strain associated with shear transformations, which control the plasticity. Hence, it is important to understand how temperature and system preparation determines the density and geometry of these excitations. The possible unification of these excitations into a common description is also a fundamental problem. These local excitations are thought to have a close relationship with ``Quasi-localised modes (QLMs)'' which are present in the low-frequency vibrational spectrum in amorphous solids. Understanding the properties of QLMs and clarifying the relation between QLMs and these local excitations are important to the study of the latter. In this thesis: (1) we provide a theory for the QLMs, D_L(omega) ~ omega^alpha, that establishes the link between QLMs and shear transformations for systems under quasi-static loading. It predicts two regimes depending on the density of shear transformations P(x)~ x^theta (with x the additional stress needed to trigger a shear transformation). If theta>1/4, alpha=4 and a finite fraction of quasi-localised modes form shear transformations, whose amplitudes vanish at low frequencies. If theta<1/4, alpha=3+ 4 theta and all QLMs form shear transformations with a finite amplitude at vanishing frequencies. We confirm our predictions numerically. (2) We present a protocol to generate extremely stable computer glasses at minimal computational cost. It consists of an instantaneous quench in an augmented potential energy landscape, with particle radii as additional degrees of freedom. (3) We propose a unification of theories predicting a gap in the spectrum of QLMs of the Hessian (Stiffness Matrix) that grows upon cooling, with others predict a pseudo-gap D_L(omega)} ~ omega^alpha. Specifically, we generate glassy configurations of controlled gap magnitude omega_c at temperature T=0, using so-called `breathing' particles, and study how such gapped states respond to thermal fluctuations. We propose an interpretation of mean-field theories of the glass transition, in which the modes beyond the gap act as an excitation reservoir, from which a pseudo-gap distribution is populated with its magnitude rapidly decreasing at lower T. (4) Preliminary results on the local excitations are presented for glasses (realistically prepared glasses) obtained by an instantaneous regular quench from the equilibrated configurations at low temperature T_p obtained by the SWAP Monte Carlo method. (5) The relationship between the density of TLS and the density of QLMs is built up.PCS

    Breaking the Curse of Dimensionality in Deep Neural Networks by Learning Invariant Representations

    No full text
    Artificial intelligence, particularly the subfield of machine learning, has seen a paradigm shift towards data-driven models that learn from and adapt to data. This has resulted in unprecedented advancements in various domains such as natural language processing and computer vision, largely attributed to deep learning, a special class of machine learning models. Deep learning arguably surpasses traditional approaches by learning the relevant features from raw data through a series of computational layers. This thesis explores the theoretical foundations of deep learning by studying the relationship between the architecture of these models and the inherent structures found within the data they process. In particular, we ask: What drives the efficacy of deep learning algorithms and allows them to beat the so-called curse of dimensionalityâ i.e. the difficulty of generally learning functions in high dimensions due to the exponentially increasing need for data points with increased dimensionality? Is it their ability to learn relevant representations of the data by exploiting their structure? How do different architectures exploit different data structures? In order to address these questions, we push forward the idea that the structure of the data can be effectively characterized by its invariancesâ i.e. aspects that are irrelevant for the task at hand. Our methodology takes an empirical approach to deep learning, combining experimental studies with physics-inspired toy models. These simplified models allow us to investigate and interpret the complex behaviors we observe in deep learning systems, offering insights into their inner workings, with the far-reaching goal of bridging the gap between theory and practice. Specifically, we compute tight generalization error rates of shallow fully connected networks demonstrating that they are capable of performing well by learning linear invariances, i.e. becoming insensitive to irrelevant linear directions in input space. However, we show that these network architectures can perform poorly in learning non-linear invariances such as rotation invariance or the invariance with respect to smooth deformations of the input. This result illustrates that, if a chosen architecture is not suitable for a task, it might overfit, making a kernel method, for which representations are not learned, potentially a better choice. Modern architectures like convolutional neural networks, however, are particularly well-fitted to learn the non-linear invariances that are present in real data. In image classification, for example, the exact position of an object or feature might not be crucial for recognizing it. This property gives rise to an invariance with respect to small deformations. Our findings show that the neural networks that are more invariant to deformations tend to have higher performance, underlying the importance of exploiting such invariance. Another key property that gives structure to real data is the fact that high-level features are a hierarchical composition of lower-level featuresâ a dog is made of a head and limbs, the head is made of eyes, nose, and mouth, which are then made of simple textures and edges. These features can be realized in multiple synonymous ways, giving rise to an invariance. To investigate the synonymic invariance that arises from the hierarchical structure of data, we introduce a toy data model that allows us to examine how features are extracted and combined to form incrPCS

    The Physics of Data and Tasks: Theories of Locality and Compositionality in Deep Learning

    No full text
    Deep neural networks have achieved remarkable success, yet our understanding of how they learn remains limited. These models can learn high-dimensional tasks, which is generally statistically intractable due to the curse of dimensionality. This apparent paradox suggests that learnable data must have an underlying latent structure. What is the nature of this structure? How do neural networks encode and exploit it, and how does it quantitatively impact performance - for instance, how does generalization improve with the number of training examples? This thesis addresses these questions by studying the roles of locality and compositionality in data, tasks, and deep learning representations. We begin by analyzing convolutional neural networks in the limit of infinite width, where the learning dynamics simplifies and becomes analytically tractable. Using tools from statistical physics and learning theory, we characterize their generalization abilities and show that they can overcome the curse of dimensionality if the target function is local by adapting to its spatial scale. We then turn to more complex structures in which features are composed hierarchically, with elements at larger scales built from sub-features at smaller ones. We model such data using simple probabilistic context-free grammars - tree-like graphical models used to describe data such as language and images. Within this framework, we study how diffusion-based generative models compose new data by assembling features learned from examples. This theory of composition predicts a phase transition in the generative process, which we confirm empirically in both image and language modalities, providing support for the compositional structure of natural data. We further demonstrate that the sample complexity for learning these grammars scales polynomially with data dimension, providing a mechanism by which diffusion models avoid the curse of dimensionality by learning to hierarchically compose new data. These results offer a theoretical grounding for how generative models learn to generalize, and ultimately, become creative. Finally, we shift our analysis from the structure of data in the input space to the structure of tasks in the model's parameter space. Here, we investigate a novel form of compositionality, where tasks and skills themselves can be composed. In particular, we empirically demonstrate that distinct directions in the weight space of large pre-trained models are associated with localized, semantic task-specific areas in function space, and how this modular structure enables task arithmetic and model editing at scale.PCS

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
    corecore