1,720,960 research outputs found
The role of the information bottleneck in representation learning
A grand challenge in representation learning is thedevelopment of computational algorithms that learn the differentexplanatory factors of variation behind high-dimensional data.Encoder models are usually determined to optimize performanceon training data when the real objective is to generalize well toother (unseen) data. Although numerical evidence suggests thatnoise injection at the level of representations might improve thegeneralization ability of the resulting encoders, an informationtheoretic justification of this principle remains elusive. In thiswork, we derive an upper bound to the so-called generalizationgap corresponding to the cross-entropy loss and show that whenthis bound times a suitable multiplier and the empirical riskare minimized jointly, the problem is equivalent to optimizingthe Information Bottleneck objective with respect to the empirical data-distribution. We specialize our general conclusionsto analyze the dropout regularization method in deep neuralnetworks, explaining how this regularizer helps to decrease thegeneralization gap.Fil: Vera, Matías Alejandro. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; ArgentinaFil: Piantanida, Pablo. Université Paris Sud; FranciaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaIEEE International Symposium on Information TheoryColoradoEstados UnidosInstitute of Electrical and Electronics Engineer
The two-way cooperative Information Bottleneck
The two-way Information Bottleneck problem, where two nodes exchange information iteratively about two arbitrarily dependent memoryless sources, is considered. Based on the observations and the information exchange, each node is required to extract "relevant information", measured in terms of the normalized mutual information, from two arbitrarily dependent hidden sources. The optimal trade-off between rates of relevance and complexity, and the number of exchange rounds, is obtained through a single-letter characterization. We further extend the results to the Gaussian case. Applications of our setup arise in the development of collaborative clustering algorithms.Fil: Vera, Matías Alejandro. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Piantanida, Pablo. Université Paris Saclay; CanadáIEEE International Symposium on Information TheoryHong KongChinaInstitute of Electrical and Electronics Engineer
Information flow in Deep Restricted Boltzmann Machines: An analysis of mutual information between inputs and outputs
Empirical evidence suggests the existence of an entangled relationship between the information flow from inputs features to hidden representations of a deep neural network and its ability to generalize from training samples to unobserved data. For instance, regularization techniques often used to control statistical generalization, are expected to impact this information flow. In this work, we study MI (mutual information) between inputs and representation outputs, and its relationship with various regularization methods commonly used in Restricted Boltzmann Machines (RBM) and their generalizations: Deep Belief Networks and Deep Boltzmann Machines. Our theoretical findings show the existence of fundamental connections between the hyperparameters associated with the regularization and the MI, including relevant practical ingredients such as: network dimension, matrix norms and dropout probability, which are well-known to influence the generalization ability of the network. These results are experimentally corroborated on various visual datasets.Fil: Vera, Matías Alejandro. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; ArgentinaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; ArgentinaFil: Piantanida, Pablo. Centre National de la Recherche Scientifique; Franci
Collaborative Information Bottleneck
This paper investigates a multi-terminal source coding problem under a logarithmic loss fidelity which does not necessarily lead to an additive distortion measure. The problem is motivated by an extension of the information bottleneck method to a multi-source scenario where several encoders have to build cooperatively rate-limited descriptions of their sources in order to maximize information with respect to other unobserved (hidden) sources. More precisely, we study fundamental informationtheoretic limits of the so-called: 1) two-way collaborative information bottleneck (TW-CIB) and 2) the collaborative distributed information bottleneck (CDIB) problems. The TW-CIB problem consists of two distant encoders that separately observe marginal (dependent) components X1 and X2 and can cooperate through multiple exchanges of limited information with the aim of extracting information about hidden variables (Y1, Y2), which can be arbitrarily dependent on (X1, X2). On the other hand, in CDIB, there are two cooperating encoders which separately observe X1 and X2 and a third node which can listen to the exchanges between the two encoders in order to obtain information about a hidden variable Y. The relevance (figureof-merit) is measured in terms of a normalized (per-sample) multi-letter mutual information metric (log-loss fidelity), and an interesting tradeoff arises by constraining the complexity of descriptions, measured in terms of the rates needed for the exchanges between the encoders and decoders involved. Inner and outer bounds to the complexity-relevance region of these problems are derived from which optimality is characterized for several cases of interest. Our resulting theoretical complexityrelevance regions are finally evaluated for binary symmetric and Gaussian statistical models, showing theoretical tradeoffs between the complexity-constrained descriptions and their relevance with respect to the hidden variablesFil: Vera, Matías Alejandro. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Piantanida, Pablo. Université Paris Sud; Francia. Centre National de la Recherche Scientifique; Franci
Diffusion assisted image reconstruction in optoacoustic tomography
In this paper we consider the problem of acoustic inversion in the context of the optoacoustic tomography image reconstruction problem. By leveraging the ability of the recently proposed diffusion models for image generative tasks among others, we devise an image reconstruction architecture based on a conditional diffusion process. The scheme makes use of an initial image reconstruction, which is preprocessed by an autoencoder to generate an adequate representation. This representation is used as conditional information in a generative diffusion process. Although the computational requirements for training and implementing the architecture are not low, several design choices discussed in the work were made to keep them manageable. Numerical results show that the conditional information allows to properly bias the parameters of the diffusion model to improve the quality of the initial reconstructed image, eliminating artifacts or even reconstructing finer details of the ground-truth image that are not recoverable by the initial image reconstruction method. We also tested the proposal under experimental conditions and the obtained results were in line with those corresponding to the numerical simulations. Improvements in image quality up to 17% in terms of peak signal-to-noise ratio were observed.Fil: González, Martín Germán. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Física; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas; ArgentinaFil: Vera, Matías Alejandro. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Dreszman, Alan. Universidad de Buenos Aires. Facultad de Ingeniería; ArgentinaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; Argentin
Information and Regularization in Restricted Boltzmann Machines
Recent works suggests an interesting interplay between the information flow between inputs features and hidden representations of a learning and the ability of the algorithm to generalize from trained samples to unobserved data. For instance, some of regularization techniques used to control generalization are expected to impact the corresponding information metrics. In this work, we study mutual information in Restricted Boltzmann Machines (RBM) and its relationship with the different regularization techniques. Our results show some evidence on interesting connections between the mutual information (inputs and its representations) with relevant parameters such as: network dimension, matrix norms and dropout probability, which are known to influence the generalization ability of the network. Results are empirically corroborated with a numerical study.Fil: Vera, Matías Alejandro. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; ArgentinaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Piantanida, Pablo. Universite Paris-Saclay;International Conference on Acoustics, Speech and Signal ProcessingTorontoCanadáInstitute of Electrical and Electronics Engineer
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
- …
