1,720,956 research outputs found

    Advances in Bayesian cross-validation and tensor modelling

    Get PDF
    The first chapter presents a work I undertook with my supervisor Giacomo Zanella. We propose a novel estimator for the log pointwise predictive density(lppd) of a Bayesian model. The naive method to calculate this quantity would require running n markov chain Monte Carlo (MCMC) chains, resulting in an unfeasible computational cost. A classical approach to overcome such a problem is to leverage importance sampling, which although solves the computational hurdle results in a estimator that can be potentially very unstable. We propose to generate the samples from a particular mixture of the leave-one-out posteriors within the typical importance sampling framework. We provide theoretical proofs of the stability of our estimator for a broad family of models also in the challenging asymptotic regime of infinite dimensionality of the data, as well as both synthetic and real data applications displaying the superior performance of our estimator compared to competitors in the literature. The second chapter presents a project I undertook under the supervision of Giacomo Zanella and Omiros Papaspiliopoulos. In this work, we focus on studying how the local convergence rate of alternating least squares(ALS) algorithms behaves as we increase the size of the problem in the context of matrix completion. We first show that, under full design, classical ALS is non-scalable. Studying the geometric properties of the optimization landscape, we propose a modification to the classical ALS algorithm, which we term ALS with Gram Calibration, and we show that such an algorithm is scalable under full design. We then provide empirical evidence that such a behavior is maintained in various sparsity scenarios. The third and fourth chapters present the works I conducted during the visiting student period at Duke University under the supervision of David Dunson and Peter Hoff, respectively. Both works concentrate on Bayesian formulations of the Candecomp/Parafac (CP) decomposition. In the third chapter, which focuses on modelling dynamically evolving binary networks, we leverage the CP decomposition as a building block to propose a novel non-parametric tensor decomposition. We prove that such a decomposition is flexible enough to represent any underlying tensor, and we also show that our prior has full support. We then provide empirical evidence of our model capabilities, both for a synthetic design and a real dataset from ecology. In the fourth chapter, we develop a Bayesian hierarchical CP model with multiplicative error and apply it to a dataset of excitation-emission matrices (EEMs) from different sources of the Neuse River. Compared to classical optimization-only procedures, our proposed model allows our proposed model allows the borrowing of information across sources and the incorporation of available knowledge through the prior. Moreover, the multiplicative error term explicitly models the positivity of the data. We show some very promising initial results, with future research looking to extend those and leverage the generative capabilities of our model in other tasks. The fifth chapter presents a I carried out with Christoph Feinauer, Barthelemy Meynard-Piganeau and Carlo Lucibello. It focuses on protein inverse folding, in which the task is to generate a sequence of amino acids that will fold into a desired three-dimensional functioning protein. Such a problem is highly complex, also because such a mapping is notoriously many-to-one, meaning that many sequences fold into the same three-dimensional structure in nature. Typical deep-learning approaches, though, focus solely on mapping the native sequence to the structure, failing to model this diversity. We hence propose a novel deep learning architecture, which we term InvMSAFold , that explicitly models this diversity. We show the benefits of this modelling choice with various experiments, demonstrating how our work could be helpful in many protein-engineering pipelines

    Robust Leave-One-Out Cross-Validation for High-Dimensional Bayesian Models

    No full text
    Leave-one-out cross-validation (LOO-CV) is a popular method for estimating out-of-sample predictive accuracy. However, computing LOO-CV criteria can be computationally expensive due to the need to fit the model multiple times. In the Bayesian context, importance sampling provides a possible solution but classical approaches can easily produce estimators whose asymptotic variance is infinite, making them potentially unreliable. Here we propose and analyze a novel mixture estimator to compute Bayesian LOO-CV criteria. Our method retains the simplicity and computational convenience of classical approaches, while guaranteeing finite asymptotic variance of the resulting estimators. Both theoretical and numerical results are provided to illustrate the improved robustness and efficiency. The computational benefits are particularly significant in high-dimensional problems, allowing to perform Bayesian LOO-CV for a broader range of models, and datasets with highly influential observations. The proposed methodology is easily implementable in standard probabilistic programming software and has a computational cost roughly equivalent to fitting the original model once. Supplementary materials for this article are available online

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis

    Dispelling the Myths Behind First-author Citation Counts

    Get PDF
    We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more sophisticated methods

    Author Index

    No full text
    Nao informado

    koamabayili/VECTRON-author-checklist: VECTRON author checklist

    No full text
    We have done our best to complete the author checklist relating to the use of animals in the hut study. Note that the objective for the hut study was to evaluate the IRS treatment applications for residual efficacy against Anopheles mosquitoes, including the local An. coluzzii mosquito population. Cows were only used to attract mosquitoes into the huts and no tests were carried out directly on the cows. The author checklist is intended for use with studies where experiments are carried out on animals, which is why we have had such difficulty in completing this for the hut study, as many of the questions do not relate to how the cows were used
    corecore