1,721,253 research outputs found
From Twentieth Annual Computational Neuroscience Meeting: CNS*2011 Stockholm, Sweden. 23-28 July 2011
Published by BioMed Central
Schultze-Kraft, Matthias ; Diesmann, Markus ; Grün, Sonja ; Helias, Moritz : Correlation transmission of spiking neurons is boosted by synchronous input : From Twentieth Annual Computational Neuroscience Meeting: CNS*2011 Stockholm, Sweden. 23-28 July 2011. - In: BMC Neuroscience. - ISSN 1471-2202 (online). - 12 (2011), suppl. 1, P144. - doi:10.1186/1471-2202-12-S1-P144
A statistical perspective on learning of time series in neural networks
The notion of neural networks encompasses two major domains: In neuroscience, their source of inspiration is the brain, where billions of neurons combine external and internal information to continuously solve tasks, and in machine learning, they are used to process vast amounts of data to solve complex tasks. The goal of this thesis is to provide a statistical perspective on learning in neural networks, while considering their biological inspiration and machine learning application. A statistical perspective implies stochasticity, which in our case comes from two sources: Firstly, biological neural activity is intrinsically noisy, and highly complex background activity influences the processing of stimuli. Secondly, natural stimuli themselves inherit stochastic features. Here, we target both scenarios using tools from statistical physics. While we take up the viewpoints of both biological and machine learning, we focus on processing of stimuli as present in the brain; to wit, everything is time-dependent. In the brain, signals reverberate due to the recurrent neural connections, creating natural interactions between inputs stemming from different time points. Because of the non-linear nature of those interactions, however, understanding their dynamical state forms an intricate challenge. For weakly non-linear interactions, we develop a method to unfold the recurrent dynamics into an effective feed-forward system. We thereby obtain an analytically tractable approximation of the network state distribution using perturbation theory. We utilize the solution of the network dynamics to find the optimal input and readout projections for classification in a random recurrent reservoir, improving the network performance. The optimal classifier in this framework changes, however, when independent background activity is present. For linear interactions, we derive the empirical risk minimizer for the input and output mapping with noisy dynamics. We find that the optimal solution employs a trade-off between stability and performance and we compare it to the noise-free case. But how does the non-linearity of the interactions shape the statistical processing of stimuli? We employ a single-layer feedforward model to answer this question, and connect the statistical features of the input and output layer. Creating a classifier with trainable gain function, we find a direct relation between the non-linearity and representation and processing of higher-order statistics. To conclude, we move from learning individual statistical features to the data distribution itself. Using an invertible type of feedforward neural network, we learn the non-linear manifold from samples and extract the most informative modes from the data. In this way, we obtain a fully adaptable mechanism to uncover structure, dimensionality, and meaningful latent features at once in an unsupervised fashion
A Neuron Model Independent Path Integral Explored via Binary Assemblies
We present the basic exploration of a novel path integral formulation for models of biological neuronal networks that allows to keep the neuron model unspecified until quantities are explicitly calculated. This is done at the example of a binary neuron network containing an assembly, which is suited to discuss the limits of standard mean field theory and the feasibility of a path integral approach, while also being of some neuroscientific interest. Advanced theoretical approaches to the description of neuronal network activity are still in their infancy, but much needed, due to its nonlinearities, statistical nature, and nonequilibrium dynamics. There is thus now renewed interest in the transfer of mathematical tools developed in statistical physics to the theory of neural networks. We model a Hebbian cell assembly as a group of O(100) excitatory binary neurons with increased coupling embedded in a larger balanced random network. Using standard mean-field theory and simulation, the system properties and parameter dependencies are analysed, especially emergence of the high activity state, spontaneous transitions, and pairwise correlations. We then introduce the path integral formulation, applying it to a single population of binary neurons. We show the relation of the tree level approximation to mean field theory, calculate propagators and a 1-loop diagram, and generalize to multiple populations. The formulation specifically uses generic properties of neuronal networks, which allows the formal description of the systems properties before an effective neuron model is fixed. It is analytically feasible for rate, binary and, possibly, spiking neurons. The implications of the results for the assembly model and the relation of our formulation to other path integral approaches are discussed. We stress that a functional form unlocks tools to treat critical phenomena, large fluctuations, disorder and may lead eventually to effective coarse grained theories by using renormalization group methods
Field theoretic approaches to computation in neuronal networks
This thesis is centered around the application of statistical field theory to the question of how computation is performed by neuronal networks. The spiking activity in dense networks of neurons, such as in the brain, tends to be strongly chaotic. How, then, can these circuits reliably process information? Past investigations have studied mainly weakly chaotic firing-rate models. Here we demonstrate a universal mechanism explaining how even strongly chaotic activity can support powerful computations. The calculations use a novel unified theoretical framework that allows to compare models of neural networks across scales of model complexity. Here, the framework is applied to two common model classes: binary neurons, which switch between a pulse-emitting and non-emitting state, and rate neurons, which describe just the number of pulses per second. These implement two different assumptions about the substrate of computation, commonly referred to as spike-coding vs. rate-coding. We calculate the transition to chaos in random binary networks and show that each chaotic binary network corresponds to an equivalent rate network with the same activity statistics, but with nonchaotic dynamics. Therefore results on the well-studied edge-of-chaos in firing-rate models cannot be directly transferred to spiking-type networks. Next, considering strongly chaotic regimes, we show that the activity transiently promotes the separability of different input stimuli. This effect arises because state trajectories for different inputs diverge from one another in a stereotypical, distance-dependent manner. Binary networks and pulse-coupled networks offer a particularly fast separation that can be exploited for fast, event-based computation, which, however, requires control of the initial conditions. These results provide predictions for experimental recordings in brain circuits and invite research on the use of chaotic dynamics in artificial neural networks. We further generalize the theoretical framework, which can serve as a bridge between many types of existing neural-network models and provides a systematic method to derive self-consistent, time-dependent Gaussian approximations and perturbation corrections for such systems. Furthermore, a parallel line of work is presented using the same type of techniques to study how the data representation is transformed in the process of computation by recurrent reservoir networks and trained artificial feed-forward networks. Because deep networks can exploit interactions between all scales in the data, these networks are difficult to understand based on their microscopic structure. We find that for close to Gaussian data classes, the computation can be captured by a Gaussian theory for the high-dimensional activity in each layer. Nonetheless, it remains a fundamental challenge to extend such a theory to strongly non-Gaussian distributions, and a graphical intuition to describe transformations of high-dimensional structured probability distributions is largely lacking. Therefore, inspired by our field-theoretic work we develop a graphical explanation for the transformations learned in classification tasks. We demonstrate how the transformations of the data manifold can be linked to folding operations which have a low-dimensional intuition that stays valid in the high-dimensional case, thereby opening an exiting link between the mathematics of folding algorithms and neuronal networks
Mean-field theory for complex neuronal networks
Understanding the working principles of the brain constitutes the major challenge in computational neuroscience. Reaching this goal is particularly difficult since the brain displays complex dynamics on various spatial and temporal scales, emerging from the interaction between a tremendous number of individual components, the neurons. Therefore a crucial step is to understand how collective dynamics in the brain arises from single neurons dynamics and the connectivity between the neurons. To better understand the collective dynamics in network of neurons, theoretical neuroscience has developed analytical tools, which reduce the complexity of the dynamics by averaging over spatial or temporal scales. In particular, this so-called mean-field theory relates single neuron properties and connectivity to network activity, which allows for a systematic understanding of the emerging activity. In this thesis we advance mean-field theory both on the level of single neuron properties and on the level of network structure. One building block of mean-field theory is the transfer function characterizing how a modulation of the neuron’s input is transmitted to a modulation of its output. A Neuron in the brain receives signals from a large number of presynaptic neurons, effectively causing a noisy input. Realistic activation kinetics of synaptic transmission amounts to a low-pass filtering of this input. Formally, the mathematical analysis of the transfer function therefore leads to a difficult colored-noise problem. We develop a general method to reduce this colored-noise to a white-noise system by capturing the color of the noise by effective boundary conditions. We apply this formalism to one of the most used neuron models, revealing a novel analytical expression for its transfer function. We further show how these results contribute to the understanding of fluctuations in a network of model neurons. On the level of structural connectivity we investigate the impact of connections between different parts of the brain on the stability of the network activity. To this end, we extend mean-field theory to devise a method allowing us to shape the phase space of a large-scale model by systematically refining its experimentally obtained connectivity map. Fundamental constraints on the activity, i.e., prohibiting quiescence and requiring global stability, prove sufficient to obtain realistic activity. Thus, the method contributes to the data integration process by constraining the empirical connectivity map to a realization that is compatible with physiological experiments. Moreover, dynamical mean-field theory relates connectivity to network generated fluctuations in the activity. We extend dynamical mean-field theory to stochastic systems where the noise mimics the spike emission process of neurons. This allows us to derive the self-consistent statistics of the activity as well as the point where the neural network displays a transition to chaos. We find that noise suppresses chaotic fluctuations by a dynamic mechanism effectively reducing the sensitivity of the network to small perturbations. The quantitative validation of the various mean-field approaches requires simulations of two model classes, i.e., rate-based and spike-based model neurons. We develop a unified simulation framework supporting both models, which facilitates this validation and increases reliability
Mechanics of deep neural networks beyond the Gaussian limit
Current developments in the field of artificial intelligence and the neural network technology supersede our theoretical understanding of these networks. In the limit of infinite width, networks at initialization are well described by the neural network Gaussian process (NNGP): the distribution of outputs is a zero-mean Gaussian characterized by its covariance or kernel across data samples. Going to the lazy learning regime, where network parameters change only slightly from their initial values, the neural tangent kernel characterizes networks trained with gradient descent. Despite the success of these Gaussian limits for deep neural networks, they do not capture important properties such as network trainability or feature learning. In this work, we go beyond Gaussian limits of deep neural networks by obtaining higher-order corrections from field-theoretic descriptions of neural networks. From a statistical point of view, two complimentary averages have to be considered: the distribution over data samples and the distribution over network parameters. We investigate both cases, gaining insights into the working mechanisms of deep neural networks. In the former case, we study how data statistics are transformed across network layers to solve classification tasks. We find that, while the hidden layers are well described by a non-linear mapping of the Gaussian statistics, the input layer extracts information from higher-order cumulants of the data. The developed theoretical framework allows us to investigate the relevance of different cumulant orders for classification: On MNIST, Gaussian statistics account for most of the classification performance, and higher-order cumulants are required to fine-tune the networks for the last few percentages. In contrast, more complex data sets such as CIFAR-10 require the inclusion of higher-order cumulants for reasonable performance values, giving an explanation for why fully-connected networks perform subpar compared to convolutional networks. In the latter case, we investigate two different aspects: First, we derive the network kernels for the Bayesian network posterior of fully-connected networks and observe a non-linear adaptation of the kernels to the target, which is not present in the NNGP. These feature corrections result from fluctuation corrections to the NNGP in finite-size networks, which allow the networks to adapt to the data. While fluctuations become larger near criticality, we uncover a trade-off between criticality and feature learning scales in networks as a driving mechanism for feature learning. Second, we study network trainability of residual networks by deriving the network prior at initialization. From this, we obtain the response function as a leading-order correction to the NNGP, which describes the signal propagation in networks. We find that scaling the residual branch by a hyperparameter improves signal propagation since it avoids saturation of the non-linearity and thus information loss. Finally, we observe a strong dependence of the optimal scaling of the residual branch on the network depth but only a weak dependence on other network hyperparameters, giving an explanation for the universal success of depth-dependent scaling of the residual branch. Overall, we derive statistical field theories for deep neural networks that allow us to obtain systematic corrections to the Gaussian limits. In this way, we take a step towards a better mechanistic understanding of information processing and data representations in neural networks
Path integral methods for correlated activity in neuronal networks
Nervous systems of highly developed organisms consists of very many cells. The human brain, to name a very complex example, is composed of nearly 100 billion neurons, that are connected via up to thousand trillion synapses. It is an essential aim of theoretical neuroscience to discover the functioning of this complicated system on the basis of the interaction of its individual parts. Many methods for achieving this goal are borrowed from many particle physics - classical (non-quantum-mechanical) statistical physics, to be precise. An important common property of biological neuronal networks and the usual subjects of statistical physics is that both can be described by models of stochastic (“noisy”) processes. The calculation of measurable quantities from these models is often difficult, for which reason approximate solutions are sought for. Here, the role of mean-field theory deserves to be emphasized, in which fluctuations are treated in a strongly simplified form. Even though in most cases many effects are neglected by this approach, it often yields quantitatively correct results in neuroscience. In this work, we use statistical field theory to derive this and related approximations for different systems, apply them to concrete problems and examine methods to improve them. To capture the interaction amongst different neurons, the description of the activity of a single neuron is often reduced to the question if it is active or not (binary model neuron). By means of its mean-field theory, we describe by which mechanisms the correlations between pairs of neurons change, when a network is driven by a stimulus varying in time. For inferences about the connections in an examined network from experimentally detected activity, the binary representation of neuronal activity is frequently used, as well. This method relies on the Ising model whose mean-field theory we derive using Feynman diagrams. We extend the formalism needed for this purpose to include expansions around non-Gaussian theories like the Ising model without coupling. Furthermore, we examine the statistics of the neuronal activity in an disordered network and its susceptibility to perturbation in mean-field theory. A generalized framework enables us to compare these results with the statistics and dynamics of networks consisting of rate model neurons. In the latter model, each nerve cell is solely characterized by the rate, which indicates how frequently it gets active. We use it in a different context to compare different path integral formalisms representing neuronal activity described by stochastic differential equations. Here we show how mean-field theory can be systematically corrected by the so called loop expansion and how the emerging correction terms can be interpreted in case mean-field theory should prove insufficient for a certain set of parameters. Another possibility to improve mean-field approximations is given be the functional Renormalization Group, whose application to simple models of biological networks we demonstrate for an example
Decomposition of Deep Neural Networks into Correlation Functions
Recent years have shown a great success of deep neural networks. One active field of research investigates the functioning mechanisms of such networks with respect to the network expressivity as well as information processing within the network. In this thesis, we describe the input-output mapping implemented by deep neural networks in terms of correlation functions. To trace the transformation of correlation functions within neural networks, we make use of methods from statistical physics. Using a quadratic approximation for non-linear activation functions, we obtain recursive relations in a perturbative manner by means of Feynman diagrams. Our results yield a characterization of the network as a non-linear mapping of mean and covariance, which can be extended by including corrections from higher order correlations. Furthermore, re-expressing the training objective in terms of data correlations allows us to study their role for solutions to a given task. First, we investigate an adaptation of the XOR problem, in which case the solutions implemented by neural networks can largely be described in terms of mean and covariance of each class. Furthermore, we study the MNIST database as an example of a non-synthetic dataset. For MNIST, solutions based on empirical estimates for mean and covariance of each class already capture a large amount of the variability within the dataset, but still exhibit a non-negligible performance gap in comparison to solutions based on the actual dataset. Lastly, we introduce an example task where higher order correlations exclusively encode class membership, which allows us to explore their role for solutions found by neural networks. Finally, our framework also allows us to make predictions regarding the correlation functions that are inferable from data, yielding insights into the network expressivity. This work thereby creates a link between statistical physics and machine learning, aiming towards explainable AI
Interactions on structured networks
Structured systems appear ubiquitously in nature. Indubitably, the structure of a system determines its characteristic behavior. However, predicting the behavior of a system given its structure, or vice versa, is not straightforward. We here demonstrate that the mapping from structure to behavior can be tackled using a systematic fluctuation expansion, and develop a new method to infer structure given observations of the system. Often, structure can be represented as a network of nodes, where the nodes represent the agents, the elementary degrees of freedom of the system, and the connections define their interactions. One common feature of structured systems are hubs: nodes with significantly more connections than average, which are expected to be key to the observed overall system behavior. To understand the influence of hubs, we investigate to which extent the hubs of a scale-free network can drive a system of binary agents into an ordered or disordered state. We find that a typical mean-field approach to these systems introduces a nonphysical process: the signal sent by a node to its neighbors may travel back and influence the same node, leading to a self-feedback loop. The phenomenon is most prominent in the presence of hubs; their accumulated self-feedback grows with the number of connections. We show that a second-order fluctuation correction eliminates this spurious self-feedback. These insights are then translated to a model of disease spreading: We investigate the SIR model, where each agent can be in one of three states (susceptible, infected, or recovered), and transitions between these states follow a stochastic process. A typical approach in literature to predict average infection curves is to assume that all agents are statistically independent, introducing self-feedback artificially into the system, which yields inflated infection curves. We use a dynamical Plefka expansion to calculate a fluctuation correction, which eliminates the self-feedback effect, leading to more accurate predictions on the spread of disease. We then approach the reverse direction: inferring pairwise and higher-order interactions from data, these interactions constitute the structure of the underlying system. In principle, inference problems require an optimization over the space of all possible interactions, whose number increases exponentially with the system size. Nevertheless, machine learning models can infer structures efficiently from data. Typically, however, the inferred structure is hidden in the parameters of the trained mdoel. We here show how to extract the learned structure, formulated in terms of interactions up to the fourth order. This process uncovers how the model hierarchically constructs interactions via nonlinear transformations of pairwise relations. This yields a fully understandable AI-powered tool for inference. Thus, we close the loop, demonstrating how collective behavior can emerge from structure and vice versa
Phase transitions in classical systems : anisotropic models, computational methods, and universality predictions
This thesis is mostly concerned with phase transitions in classical systems, with a focus on the anisotropic Ising model in two dimensions and on parallelogram lattices. After first discussing the history of phase transitions and specifically how anisotropies were treated in the renormalization group approach, an introduction to multi-parameter universality is given. This is followed by a derivation, using anti-commuting Grassmann variables, of the exact solution of the fully anisotropic 2d Ising model on a finite parallelogram lattice for all temperatures and couplings; from there, the scaling function near the critical point in the ferromagnetic regime is recovered. Additionally, some predictions made by multi-parameter universality regarding non-universal prefactors, modular invariance and behavior at criticality are confirmed. Finally, the strip limit of the model is discussed and connections to previous results of more restricted cases are made. In the next chapter, the investigation of anisotropic systems in 2d is continued, now by discussing the q-state Potts model and attempting to measure its angle dependent correlation lengths, a characterizing quantity according to multi-parameter universality, via an tensor network approach. More specifically, the Corner Transfer Matrix Renormalization Group (CTMRG) algorithm is used to numerically extract the quantities of interest. A range of checks and comparisons to the few exactly known results are made to ensure a continued high accuracy of the simulation method. Finally, the discrete to continuous crossover behavior in a modified 3d clock model, a relative of the Potts model, is investigated. This model exhibits a first order phase transition between an ordered and disordered phase and, based on prior work, predictions can be made for how much the phases contribute at the transition point when the clock has either three different states or, on the other extreme, infinitely many. This behavior is simulated at and between these extremes using the Wang-Landau Monte Carlo algorithm, which is very well suited for systems that exhibit complicated energy distributions, as present near and at first order phase transitions. A wide range of system sizes are simulated and care is taken to carefully determine the bulk transition temperature on which the accuracy of the final results depends very crucially
- …
