1,721,129 research outputs found

    Estimation of multi-way tables subject to coherence constraints

    Get PDF
    Nowadays, traditional population censuses based on total enumeration of the population are being accompanied by sample surveys. Sampling within censuses allows to reduce costs and workload of authorities involved in censuses operations, along with the statistical burden for the people involved in the enumeration. In this paper, we deal with estimation of multi-way contingency tables involving variables measured both via census and sampling. In this framework, two main issues need to be addressed: first of all, sample size for estimating some of the entries of the contingency tables may be too small, delivering estimates prone to huge sampling variability. On the other hand, since estimates of the joint distribution need to be coherent with the marginal distribution of the variable collected via a census, estimation methods need to be coherent with the constraint imposed by marginal distribution of variables measured via census. The problem is tackled via a model-based approach that allows to comply with all coherence constraints following a fairly simple procedure. The merit of the proposed methodology is illustrated by means of a simulation study

    A bivariate model for improving the estimation of relative risk

    No full text
    Disease mapping studies have been widely performed at univariate level, that is considering only one disease in the estimated models. Nonetheless, simultaneous modeling of different diseases can be a valuable tool both from the epidemiological and from the statistical point of view. In this paper we propose a model for bivariate disease mapping that generalises the univariate CAR distribution. The proposed model is proven to be an effective alternative to existing bivariate models, mainly because it overcome some restrictive hypotheses underlying models previously proposed in this context. Model performances are checked via a simulation study and via application to some real case studies

    Studying Russian individual consumption in small areas of the Emilia Romagna Region

    No full text
    Policy makers ask for spatially detailed statistical information in the assessment of the tourism impact, often asking for ad-hoc surveys for gathering additional information. To design marketing strategies and to foster the development of tourism destinations, these information should be provided with respect to the characteristics generally recognized as main determinants of individuals’ behaviour spending: socio-economic, demographic, psychographic and behavioural variables. In this chapter, we adopt data gathered from the National Frontier Survey and made available by Banca d’Italia to deliver estimates of average individual expenditure by tourist nationality, adjusting for trip-related variables, in small areas of the Emilia Romagna Region, with particular focus on Russian tourists. A Bayesian approach is adopted in order to manage estimation based on small sample sizes by borrowing strength in space and time. The approach appears promising in small area estimation for producing the spatially detailed information sought by policy makers

    Spatio-temporal regression on compositional covariates: modeling vegetation in a gypsum outcrop

    Get PDF
    Investigating the relationship between vegetation cover and substrate typologies is important for habitat conservation. To study these relationships, common practice in modern ecological surveys is to collect information regarding vegetation cover and substrate typology over fine regular lattices, as derived from digital ground photos. Information on substrate typologies is often available as compositional measures, e.g., the area proportion occupied by a certain substrate. Two primary issues are of interest for ecologists: first, how much substrate typologies differ in terms of relative suitability for vegetation cover and, second, whether suitability varies over time. This paper develops a procedure for managing compositional covariates within a Bayesian hierarchical framework to effectively address the aforementioned issues. A spatio-temporal model is adopted to estimate the temporal pattern characterizing substrate relative suitability for vegetation cover and, at the same time, to account for spatio-temporal correlation. Relative suitability is modeled by time-varying regression coefficients, and spatial, temporal and spatio-temporal random effects are modeled using Gaussian Markov Random Field models

    A note on auxiliary mixture sampling for Bayesian Poisson models

    Get PDF
    Bayesian hierarchical Poisson models are an essential tool for analyzing count data. However, designing efficient algorithms to sample from the posterior distribution of the target parameters remains a challenging task. Auxiliary mixture sampling algorithms have been proposed to this aim. They involve two steps of data augmentation: the first leverages the theory of Poisson processes, and the second approximates the residual distribution of the resulting model through a mixture of Gaussian distributions. In this way, an approximate Gibbs sampler can be implemented. This strategy is particularly beneficial for latent Gaussian models, as it allows one to exploit the sparsity of the precision matrix associated with the random effects and to efficiently incorporate linear constraints. In this paper, we focus on the accuracy of the approximation step, highlighting scenarios where the mixture fails to represent accurately the true underlying distribution, leading to a lack of convergence in the algorithm. We outline key features to monitor, in order to assess if the approximation performs as intended. Building on this, we propose a robust version of the auxiliary mixture sampling algorithm. Our approach includes mechanisms for detecting approximation failures and introduces an enhanced approximation of the right tail of the auxiliary variable distribution, supplemented by a Metropolis-Hastings correction step when needed. Finally, we evaluate the proposed algorithm together with the original mixture sampling algorithms on both simulated and real datasets

    The Mellin Transform to Manage Quadratic Forms in Normal Random Variables

    Get PDF
    The problem of computing the distribution of quadratic forms in normal variables has a long tradition in the statistical literature. Well-established numerical algorithms that deal with this task rely on the inversion of Fourier transforms or series representations. In this article, the Mellin transform is proposed as a tool to compute both the density and the cumulative distribution functions of a positive definite quadratic form: an outline of the numerical algorithm is presented, providing details on the error analysis. The algorithm’s characteristics allow us to propose an efficient way to compute the random variables’ quantiles. From the theoretical point of view, the analytic properties of the Mellin transform are exploited to provide a novel representation of the distribution of the ratio of independent quadratic forms as a mixture of beta random variables of the second kind. Moreover, algorithms are proposed for computations related to ratios of both independent and dependent quadratic forms. The methods are tested and compared to popular numerical algorithms in terms of computational times and accuracy. The R package QF implementing all the proposed algorithms is also made available. Supplementary materials for this article are available online

    Regression on compositional covariates: assessing substrate suitability for vegetation

    No full text
    Investigating the relationship between vegetation cover and substrate typologies is important in habitat conservation and management. We focus on a modern ecological survey, where information regarding vegetation cover are derived from digital ground photos taken at different times. The aim is to estimate the effect of different substrate typologies on vegetation cover (substrate suitability). As it is often the case in ground cover imaging, information on substrate typologies are available as compositional data, e.g., the area proportion occupied by a certain substrate. We develop a novel procedure for managing compositional covariates within a Bayesian hierarchical framework and illustrate it with data from a gypsum outcrop located in the Emilia Romagna region, Italy

    Non-parametric regression on compositional covariates using Bayesian P-splines

    Get PDF
    Methods to perform regression on compositional covariates have recently been proposed using isometric log-ratios (ilr) representation of compositional parts. This approach consists of first applying standard regression on ilr coordinates and second, transforming the estimated ilr coefficients into their contrast log-ratio counterparts. This gives easy-to-interpret parameters indicating the relative effect of each compositional part. In this work we present an extension of this framework, where compositional covariate effects are allowed to be smooth in the ilr domain. This is achieved by fitting a smooth function over the multidimensional ilr space, using Bayesian Psplines. Smoothness is achieved by assuming random walk priors on spline coefficients in a hierarchical Bayesian framework. The proposed methodology is applied to spatial data from an ecological survey on a gypsum outcrop located in the Emilia Romagna Region, Italy

    Design and Structure Dependent Priors for Scale Parameters in Latent Gaussian Models

    Get PDF
    Bayesian inference in latent Gaussian models necessitates the specification of prior distributions for scale parameters, which govern the behavior of model components. This task is particularly delicate and many contributions in the literature are devoted to the topic. We show that the scale parameter plays a crucial role in determining the prior variability of the model components, which is influenced by factors such as correlation structure, design matrices, and potential linear constraints. This intricate relationship adds complexity, making it difficult to interpret and compare priors across diverse applications. To tackle this challenge, we propose a novel approach for prior specification based on the theory of distribution of quadratic forms. Our strategy involves the use of design and structure-dependent (DSD) priors, which ensure a consistent interpretation across diverse applications. By introducing a single parameter that governs the prior variability of the linear predictor, we simplify the process of prior specification, making it more manageable and interpretable. We derive analytical expressions for DSD priors on scale parameters and establish conditions that guarantee their existence. To demonstrate the efficacy of our proposed prior elicitation strategy, we conduct a simulation study, examining the sampling properties of the estimators. Additionally, we explore several real data applications to investigate prior sensitivity and the allocation of explained variance among model components

    On the specification of prior distributions for variance components in disease mapping models

    Get PDF
    In this paper, we consider the problem of specifying priors for the variance components in the Bayesian analysis of the Besag-York-Mollié model, a model that is popular among epidemiologists for disease mapping. The model encompasses two sets of random effects: one spatially structured to model spatial autocorrelation and the other spatially unstructured to describe residual heterogeneity. In this model, prior specification for variance components is an important problem because these priors maintain their influence on the posterior distributions of relative risks when mapping rare diseases. We propose using generalised inverse Gaussian priors, a broad class of distributions that includes many distributions commonly used as priors in this context, such as inverse gammas. We discuss the prior parameter choice with the aim of balancing the prior weight of the two sets of random effects on total variation and controlling the amount of shrinkage. The suggested prior specification strategy is compared to popular alternatives using a simulation exercise and applications to real data sets
    corecore