1,721,256 research outputs found
Information Bottleneck and Representation Learning
A grand challenge in representation learning is the development of computational algorithms that learn the different explanatory factors of variation behind high-dimensional data. Representation models (usually referred to as encoders) are often determined for optimizing performance on training data when the real objective is to generalize well to other (unseen) data. The first part of this chapter is devoted to provide an overview of and introduction to fundamental concepts in statistical learning theory and the Information Bottleneck principle. It serves as a mathematical basis for the technical results given in the second part, in which an upper bound to the generalization gap corresponding to the cross-entropy risk is given. When this penalty term times a suitable multiplier and the cross entropy empirical risk are minimized jointly, the problem is equivalent to optimizing the Information Bottleneck objective with respect to the empirical data distribution. This result provides an interesting connection between mutual information and generalization, and helps to explain why noise injection during the training phase can improve the generalization ability of encoder models and enforce invariances in the resulting representations.Fil: Piantanida, Pablo. No especifíca;Fil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; Argentin
The role of the information bottleneck in representation learning
A grand challenge in representation learning is thedevelopment of computational algorithms that learn the differentexplanatory factors of variation behind high-dimensional data.Encoder models are usually determined to optimize performanceon training data when the real objective is to generalize well toother (unseen) data. Although numerical evidence suggests thatnoise injection at the level of representations might improve thegeneralization ability of the resulting encoders, an informationtheoretic justification of this principle remains elusive. In thiswork, we derive an upper bound to the so-called generalizationgap corresponding to the cross-entropy loss and show that whenthis bound times a suitable multiplier and the empirical riskare minimized jointly, the problem is equivalent to optimizingthe Information Bottleneck objective with respect to the empirical data-distribution. We specialize our general conclusionsto analyze the dropout regularization method in deep neuralnetworks, explaining how this regularizer helps to decrease thegeneralization gap.Fil: Vera, Matías Alejandro. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; Argentina. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; ArgentinaFil: Piantanida, Pablo. Université Paris Sud; FranciaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaIEEE International Symposium on Information TheoryColoradoEstados UnidosInstitute of Electrical and Electronics Engineer
The two-way cooperative Information Bottleneck
The two-way Information Bottleneck problem, where two nodes exchange information iteratively about two arbitrarily dependent memoryless sources, is considered. Based on the observations and the information exchange, each node is required to extract "relevant information", measured in terms of the normalized mutual information, from two arbitrarily dependent hidden sources. The optimal trade-off between rates of relevance and complexity, and the number of exchange rounds, is obtained through a single-letter characterization. We further extend the results to the Gaussian case. Applications of our setup arise in the development of collaborative clustering algorithms.Fil: Vera, Matías Alejandro. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Piantanida, Pablo. Université Paris Saclay; CanadáIEEE International Symposium on Information TheoryHong KongChinaInstitute of Electrical and Electronics Engineer
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Information flow in Deep Restricted Boltzmann Machines: An analysis of mutual information between inputs and outputs
Empirical evidence suggests the existence of an entangled relationship between the information flow from inputs features to hidden representations of a deep neural network and its ability to generalize from training samples to unobserved data. For instance, regularization techniques often used to control statistical generalization, are expected to impact this information flow. In this work, we study MI (mutual information) between inputs and representation outputs, and its relationship with various regularization methods commonly used in Restricted Boltzmann Machines (RBM) and their generalizations: Deep Belief Networks and Deep Boltzmann Machines. Our theoretical findings show the existence of fundamental connections between the hyperparameters associated with the regularization and the MI, including relevant practical ingredients such as: network dimension, matrix norms and dropout probability, which are well-known to influence the generalization ability of the network. These results are experimentally corroborated on various visual datasets.Fil: Vera, Matías Alejandro. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; ArgentinaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; ArgentinaFil: Piantanida, Pablo. Centre National de la Recherche Scientifique; Franci
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Collaborative Information Bottleneck
This paper investigates a multi-terminal source coding problem under a logarithmic loss fidelity which does not necessarily lead to an additive distortion measure. The problem is motivated by an extension of the information bottleneck method to a multi-source scenario where several encoders have to build cooperatively rate-limited descriptions of their sources in order to maximize information with respect to other unobserved (hidden) sources. More precisely, we study fundamental informationtheoretic limits of the so-called: 1) two-way collaborative information bottleneck (TW-CIB) and 2) the collaborative distributed information bottleneck (CDIB) problems. The TW-CIB problem consists of two distant encoders that separately observe marginal (dependent) components X1 and X2 and can cooperate through multiple exchanges of limited information with the aim of extracting information about hidden variables (Y1, Y2), which can be arbitrarily dependent on (X1, X2). On the other hand, in CDIB, there are two cooperating encoders which separately observe X1 and X2 and a third node which can listen to the exchanges between the two encoders in order to obtain information about a hidden variable Y. The relevance (figureof-merit) is measured in terms of a normalized (per-sample) multi-letter mutual information metric (log-loss fidelity), and an interesting tradeoff arises by constraining the complexity of descriptions, measured in terms of the rates needed for the exchanges between the encoders and decoders involved. Inner and outer bounds to the complexity-relevance region of these problems are derived from which optimality is characterized for several cases of interest. Our resulting theoretical complexityrelevance regions are finally evaluated for binary symmetric and Gaussian statistical models, showing theoretical tradeoffs between the complexity-constrained descriptions and their relevance with respect to the hidden variablesFil: Vera, Matías Alejandro. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Rey Vega, Leonardo Javier. Consejo Nacional de Investigaciones Científicas y Técnicas. Oficina de Coordinación Administrativa Parque Centenario. Centro de Simulación Computacional para Aplicaciones Tecnológicas; Argentina. Universidad de Buenos Aires. Facultad de Ingeniería. Departamento de Electronica; ArgentinaFil: Piantanida, Pablo. Université Paris Sud; Francia. Centre National de la Recherche Scientifique; Franci
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
