1,721,047 research outputs found
Relaxation towards negative temperatures in bosonic systems: Generalized Gibbs ensembles and beyond integrability
Motivated by the recent experimental observation of negative absolute temperature states in systems of ultracold atomic gases in optical lattices [Braun et al., Science 339, 52 (2013)], we investigate theoretically the formation of these states. More specifically, we consider the relaxation after a sudden inversion of the external parabolic confining potential in the one-dimensional inhomogeneous Bose-Hubbard model. First, we focus on the integrable hard-core boson limit, which allows us to treat large systems and arbitrarily long times, providing convincing numerical evidence for relaxation to a generalized Gibbs ensemble at negative temperature T < 0, a notion we define in this context. Second, going beyond one dimension, we demonstrate that the emergence of negative temperature states can be understood in a dual way in terms of positive temperatures, which relies on a dynamic symmetry of the Hubbard model. We complement the study by exact diagonalization simulations at finite values of the on-site interaction
Unifying Feature-Based Explanations with Functional ANOVA and Cooperative Game Theory
Feature-based explanations, using perturbations or gradients, are a prevalent tool to understand decisions of black box machine learning models. Yet, differences between these methods still remain mostly unknown, which limits their applicability for practitioners. In this work, we introduce a unified framework for local and global feature-based explanations using two well-established concepts: functional ANOVA (fANOVA) from statistics, and the notion of value and interaction from cooperative game theory. We introduce three fANOVA decompositions that determine the influence of feature distributions, and use game-theoretic measures, such as the Shapley value and interactions, to specify the influence of higher-order interactions. Our framework combines these two dimensions to uncover similarities and differences between a wide range of explanation techniques for features and groups of features. We then empirically showcase the usefulness of our framework on synthetic and real-world datasets
Recommended from our members
Lossy Compression with Machine Learning: Techniques and Fundamental Limits
The exponential growth in global data creation and transmission necessitates increasingly powerful data compression algorithms. Neural compression, which leverages neural networks and deep learning techniques, presents a promising new frontier by learning compression algorithms end-to-end from data. This approach has shown great potential across various data types, outperforming traditional codecs that have been hand-designed over decades. Despite its promise, neural compression still faces many open problems and roadblocks. This thesis focuses on neural lossy compression, and tackles the problems of inference sub-optimality, high computation complexity, and establishing fundamental bounds for performance evaluation. Conceptually our contributions are two-fold, concerning both the practical performance and theoretical performance limit of neural lossy compression:First, drawing on techniques from deep learning and probabilistic inference, we develop methods for improving the practical performance of neural lossy compression, pushing its boundaries in rate-distortion performance and computation efficiency. We identify approximation gaps in the existing nonlinear transform coding [Ballé et al., 2021] approach, and propose inference optimization algorithms that close these gaps using ideas from variational optimization, the Gumbel-softmax trick, and bits-back coding. The proposed Stochastic Gumbel Annealing (SGA), allows for plug-and-play optimization of the discrete latent representation at inference time, and is shown to significantly improve the rate-distortion performance of the base approach. The idea of improved inference (or encoding) also lends itself to a novel asymmetric autoencoder architecture, which shifts much of the decoding computation cost to a higher encoding cost up-front. We design a lightweight and shallow decoder inspired by traditional transform coding algorithms like JPEG, and pair it with a powerful encoding procedure (such as SGA), demonstrating rate-distortion performance competitive with the Mean-Scale Hyperprior architecture [Minnen et al., 2018] at a fraction of the decoding computation cost.
Second, building on the interplay between lossy compression, statistical estimation, and entropic optimal transport, we present computational methods for estimating the rate-distortion (R-D) function, a fundamental quantity in the theory of lossy compression which delineates the performance ceiling of lossy compression. We develop novel variational bounds which sandwich the R-D function and are amenable to deep learning tools, as well as a neural-network-free, particle-based upper bound of the R-D function using optimal transport (OT) ideas. These algorithms allow for the estimation of the R-D function on a much larger class of problems than previously possible with the classic Blahut-Arimoto [Blahut, 1972, Arimoto, 1972] algorithm, such as on continuous real-world data sources. We empirically obtain tight sandwich bounds on the R-D functions of real-world sources from particle physics, speech, and GAN-generated images. We also obtain an R-D upper bound for high-resolution natural images, and shed light on the optimality of recent lossy image compression algorithms. On the theoretical side, we characterize the convergence and sample complexity of our particle-based method, and offer new insight into the solution of the R-D problem in the continuous setting based on connections to entropic optimal transport and maximum-likelihood deconvolution.
While lossless compression and machine learning have long been understood as two sides of the same coin — any probability model of discrete data yields a source code for lossless compression and vice versa [Kraft, 1949, McMillan, 1956], and compression-related metrics such as cross-entropy and bits-per-dimension routinely appear in machine learning [Bishop, 2006, Theis et al., 2016] — less known is the connection between lossy compression and machine learning. This thesis explores and further develops the parallel between lossy compression and machine learning, where a rate-distortion loss is equivalently viewed as a variational inference/learning objective; this objective can be naturally motivated by the statistical problem of maximum-likelihood deconvolution, and is further related to the cost of entropic-OT projection. It is our goal and hope for these connections to inspire and guide the development of new learning-based compression algorithms, as well as facilitate the exchange of ideas between the fields of compression, information theory, and machine learning
Recommended from our members
Generative Compression: Bridging Creation and Representation Learning in Image and Video Processing
The rapid evolution of generative models has revolutionized image and video processing, paving the way for innovative approaches in creation and optimization. This thesis explores the intersection of generative modeling and neural compression, emphasizing their roles in addressing the growing demands of high-quality media processing. Building upon three key works, we synthesize insights to present a unified perspective.
First, the foundational study on neural video compression through generative modeling demonstrates the utility of variational autoencoders and autoregressive flows for efficient rate-distortion tradeoffs. By introducing structured priors and temporal dependencies, this work establishes the viability of generative models in surpassing traditional codecs, marking a significant leap in neural video coding. Second, leveraging the strengths of diffusion probabilistic models for video generation, we address the challenge of forecasting high-dimensional video dynamics. This approach integrates deterministic and stochastic components, refining frame predictions with diffusion-based residual modeling to enhance perceptual quality and probabilistic accuracy. Third, the innovative application of conditional diffusion models in lossy image compression underscores a paradigm shift from deterministic decoders to diffusion-based architectures. By encoding semantic content as latent variables and reconstructing texture via diffusion, this framework achieves remarkable perceptual gains while maintaining competitive distortion metrics.Collectively, these contributions illustrate the convergence of generative techniques in optimizing rate-distortion performance, enabling high-fidelity reconstructions, and fostering perceptually appealing outputs. This work situates itself within the broader research landscape by addressing the dual objectives of efficiency and realism in media compression and generation. Applications span diverse domains, including video conferencing, autonomous systems, and content delivery networks, underscoring the transformative potential of generative modeling in a data-driven era. Through the synthesis of creation and optimization, this thesis lays the groundwork for novel multimedia processing systems that are both generative and compression-ready
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Recommended from our members
Deep Anomaly Detection and Distribution Shifts
Anomaly detection is important in various applications, from cyber-security, transportation, industry, and finance to healthcare. The anomaly detection problem is to identify anomalies originating from a different data-generating process from normal data. The rare occurrence of anomalies and their unknown causes makes it hard to collect and model them. Thus, anomaly detection methods utilize normal data to build anomaly detectors. In this dissertation, we apply deep anomaly detection methods--methods that apply deep learning techniques--to solve anomaly detection problems. We contribute multiple generic frameworks for various anomaly detection setups. First, we challenge the common clean training data assumption (free of anomalies) and stress that practical training data is often contaminated with unnoticed anomalies. We propose a novel unsupervised training strategy for training an anomaly detector in the presence of unlabeled anomalies that is compatible with a broad class of models. Second, selecting informative data points for expert feedback can significantly improve anomaly detection performance. The critical challenges are selecting the most informative samples for expert review and effectively incorporating their feedback to bolster anomaly detection capabilities. To address these challenges, we propose a new data labeling strategy and a new learning framework for active and semi-supervised anomaly detection. Third, real-world applications may face distribution shifts. We consider the online learning problem where the shifts occur at unknown positions and with unknown intensities. We derive a new Bayesian online inference approach to automatically infer these distribution shifts and adapt the model to the detected changes. This approach applies to both supervised and unsupervised learning settings. We also consider the problem of adapting an anomaly detector to drift in the normal data distribution, especially when no training data is available for the “new normal.” This setting is called zero-shot anomaly detection. We propose a simple yet effective method that combines batch normalization and meta-training for zero-shot anomaly detection. The learning frameworks introduced in this dissertation are model-agnostic and apply to various data types. Extensive experiments demonstrate the efficacy of our proposed approaches
Recommended from our members
On the Efficient Marginalization of Probabilistic Sequence Models
Real-world data often exhibits sequential dependence, across diverse domains such as human behavior, medicine, finance, and climate modeling. Probabilistic methods capture the inherent uncertainty associated with prediction in these contexts, with autoregressive models being especially prominent. This dissertation focuses on using autoregressive models to answer complex probabilistic queries that go beyond single-step prediction, such as the timing of future events or the likelihood of a specific event occurring before another. In particular, we develop a broad class of novel and efficient approximation techniques for marginalization in sequential models that are model-agnostic. These techniques rely solely on access to and sampling from next-step conditional distributions of a pre-trained autoregressive model, including both traditional parametric models as well as more recent neural autoregressive models. Specific approaches are presented for discrete sequential models, for marked temporal point processes, and for stochastic jump processes, each tailored to a well-defined class of informative, long-range probabilistic queries
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
- …
