1,720,987 research outputs found

    Markov switching multiple-equation tensor regressions

    No full text
    A new flexible tensor model for multiple-equation regressions that accounts for latent regime changes is proposed. The model allows for dynamic coefficients and multi-dimensional covariates that vary across equations. The coefficients are driven by a common hidden Markov process that addresses structural breaks to enhance the model flexibility and preserve parsimony. A new soft PARAFAC hierarchical prior is introduced to achieve dimensionality reduction while preserving the structural information of the covariate tensor. The proposed prior includes a new multi-way shrinking effect to address over-parametrization issues while preserving interpretability and model tractability. Theoretical results are derived to help with the choice of the hyperparameters. An efficient Markov chain Monte Carlo (MCMC) algorithm based on random scan Gibbs and back-fitting strategy is designed with priority placed on computational scalability of the posterior sampling. The validity of the MCMC algorithm is demonstrated theoretically, and its computational efficiency is studied using numerical experiments in different parameter settings. The effectiveness of the model framework is illustrated using two original real data analyses. The proposed model exhibits superior performance compared to the current benchmark, Lasso regression

    Markov Switching Tensor Regressions

    No full text
    A new flexible tensor-on-tenor regression model that accounts for latent regime changes is proposed. The coefficients are driven by a common hidden Markov process that addresses structural breaks to enhance the model's flexibility and preserve parsimony. A new soft PARAFAC hierarchical prior is introduced to achieve dimensionality reduction while preserving the structural information of the covariate tensor. The proposed prior includes a new multi-way shrinking effect to address over-parametrization issues while preserving interpretability and model tractability. An efficient MCMC algorithm is introduced based on a random scan Gibbs and back-fitting strategy. The model framework’s effectiveness is illustrated using financial and commodity market volatility data. The proposed model exhibits superior performance compared to the current benchmark, Lasso regression

    Embarrassingly Parallel Sequential Markov-chain Monte Carlo for Large Sets of Time Series

    No full text
    Bayesian computation crucially relies on Markov chain Monte Carlo (MCMC) algorithms. In the case of massive data sets, running the Metropolis-Hastings sampler to draw from the posterior distribution becomes prohibitive due to the large number of likelihood terms that need to be calculated at each iteration. In order to perform Bayesian inference for a large set of time series, we consider an algorithm that combines “divide and conquer” ideas previously used to design MCMC algorithms for big data with a sequential MCMC strategy. The performance of the method is illustrated using a large set of financial data

    Conditional Copula Inference and Efficient Approximate MCMC

    No full text
    This thesis consists of two main parts. The first part focuses on parametric conditional copula models that allow the copula parameters to vary with a set of covariates according to an unknown calibration function. Flexible Bayesian inference for the calibration function of a bivariate conditional copula is introduced. The prior distribution over the set of smooth calibration functions is built using a sparse Gaussian Process prior for the Single Index Model. The estimation of parameters from the marginal distributions and the calibration function is done jointly via Markov Chain Monte Carlo sampling from the full posterior distribution. A new Conditional Cross Validated Pseudo-Marginal criterion is used to perform copula selection and is modified using a permutation-based procedure to assess data support for the simplifying assumption. The first part concludes with methods for establishing data support for the simplifying assumption in a bivariate conditional copula model. After splitting the observed data into training and test sets, the method proposed will use a flexible Bayesian model fit to the training data to define tests based on randomization and standard asymptotic theory. I discuss theoretical justification for the method and implementations in alternative models of interest: Gaussian, Logistic and Quantile regressions. The performance is studied via simulated data. The second part of the thesis focuses on approximate Bayesian methods. Approximate Bayesian Computation (ABC) and Bayesian Synthetic Likelihood (BSL) are popular simulation based methods for sampling from the posterior distribution when the likelihood is not tractable but simulations for each parameter are easily available. However these methods can be computationally inefficient since a large number of pseudo-data simulations is required. I propose to use perturbed MCMC versions of ABC and BSL algorithms and attempt to significantly accelerate these samplers. The main idea of the proposed strategy is to utilize past samples with k-Nearest-Neighbor approach for likelihood approximation. This general method works for ABC and BSL and greatly reduces computational cost and number of required simulations for these samplers. Performance and computational advantage are examined via series of simulation examples. The second part concludes with theoretical justifications and convergence properties of the proposed strategies.Ph.D

    Bayesian Inference for Bivariate Conditional Copula Models with Continuous or Mixed Outcomes

    Get PDF
    The main goal of this thesis is to develop Bayesian model for studying the influence of covariate on dependence between random variables. Conditional copula models are flexible tools for modelling complex dependence structures. We construct Bayesian inference for the conditional copula model adapted to regression settings in which the bivariate outcome is continuous or mixed (binary and continuous) and the copula parameter varies with covariate values. The functional relationship between the copula parameter and the covariate is modelled using cubic splines. We also extend our work to additive models which would allow us to handle more than one covariate while keeping the computational burden within reasonable limits. We perform the proposed joint Bayesian inference via adaptive Markov chain Monte Carlo sampling. The deviance information criterion and cross-validated marginal log-likelihood criterion are employed for three model selection problems: 1) choosing the copula family that best fits the data, 2) selecting the calibration function, i.e., checking if parametric form for copula parameter is suitable and 3) determining the number of independent variables in the additive model. The performance of the estimation and model selection techniques are investigated via simulations and demonstrated on two data sets: 1) Matched Multiple Birth and 2) Burn Injury. In which of interest is the influence of gestational age and maternal age on twin birth weights in the former data, whereas in the later data we are interested in investigating how patient’s age affects the severity of burn injury and the probability of death.Ph

    Copulas: New Theory and Methods

    Get PDF
    In this thesis, we present new theory and methods related to copulas, which are mathematical objects that allow one to model many complex dependence structures and to understand the dependence structure underlying a random vector independently of its marginals. Our first contribution is to develop a new class of absolutely continuous copulas parameterized by a function space. We prove a number of desirable properties of this family, and demonstrate its utility with an application that defines patterns of neural connectivity. In addition, we show that the family includes a highly irregular copula, thereby illustrating that although intuitively one would expect them to be well-behaved, absolutely continuous copulas can be quite pathological. Our second contribution is to integrate copulas within hidden Markov models to develop highly flexible models that allow both the copulas and marginals that generate the observed multivariate data to vary across hidden states. We develop from first principles a new efficient estimation algorithm for this model, examine the asymptotic properties of the resulting estimator, and show -- via simulations and an application that predicts room occupancy -- that the model is informative both for the goal of state prediction and for the goal of understanding the data-generating process. Our final contribution turns to the Bayesian regime, where multivariate "surrogate" data provides information about a latent variable of interest, both through the marginals of the surrogate vector and the dependence within it. We set up a Bayesian hierarchical model and develop an efficient Gibbs sampler to sample from the posterior distribution, and we use simulations and an obstetrical application that monitors fetal health during labour and delivery to show that considering the copula is again informative for understanding the data-generating process.Ph.D

    Conditional Copula Inference and Efficient Approximate MCMC

    Get PDF
    This thesis consists of two main parts. The first part focuses on parametric conditional copula models that allow the copula parameters to vary with a set of covariates according to an unknown calibration function. Flexible Bayesian inference for the calibration function of a bivariate conditional copula is introduced. The prior distribution over the set of smooth calibration functions is built using a sparse Gaussian Process prior for the Single Index Model. The estimation of parameters from the marginal distributions and the calibration function is done jointly via Markov Chain Monte Carlo sampling from the full posterior distribution. A new Conditional Cross Validated Pseudo-Marginal criterion is used to perform copula selection and is modified using a permutation-based procedure to assess data support for the simplifying assumption. The first part concludes with methods for establishing data support for the simplifying assumption in a bivariate conditional copula model. After splitting the observed data into training and test sets, the method proposed will use a flexible Bayesian model fit to the training data to define tests based on randomization and standard asymptotic theory. I discuss theoretical justification for the method and implementations in alternative models of interest: Gaussian, Logistic and Quantile regressions. The performance is studied via simulated data. The second part of the thesis focuses on approximate Bayesian methods. Approximate Bayesian Computation (ABC) and Bayesian Synthetic Likelihood (BSL) are popular simulation based methods for sampling from the posterior distribution when the likelihood is not tractable but simulations for each parameter are easily available. However these methods can be computationally inefficient since a large number of pseudo-data simulations is required. I propose to use perturbed MCMC versions of ABC and BSL algorithms and attempt to significantly accelerate these samplers. The main idea of the proposed strategy is to utilize past samples with k-Nearest-Neighbor approach for likelihood approximation. This general method works for ABC and BSL and greatly reduces computational cost and number of required simulations for these samplers. Performance and computational advantage are examined via series of simulation examples. The second part concludes with theoretical justifications and convergence properties of the proposed strategies.Ph.D

    Bayesian Inference for Bivariate Conditional Copula Models with Continuous or Mixed Outcomes

    No full text
    The main goal of this thesis is to develop Bayesian model for studying the influence of covariate on dependence between random variables. Conditional copula models are flexible tools for modelling complex dependence structures. We construct Bayesian inference for the conditional copula model adapted to regression settings in which the bivariate outcome is continuous or mixed (binary and continuous) and the copula parameter varies with covariate values. The functional relationship between the copula parameter and the covariate is modelled using cubic splines. We also extend our work to additive models which would allow us to handle more than one covariate while keeping the computational burden within reasonable limits. We perform the proposed joint Bayesian inference via adaptive Markov chain Monte Carlo sampling. The deviance information criterion and cross-validated marginal log-likelihood criterion are employed for three model selection problems: 1) choosing the copula family that best fits the data, 2) selecting the calibration function, i.e., checking if parametric form for copula parameter is suitable and 3) determining the number of independent variables in the additive model. The performance of the estimation and model selection techniques are investigated via simulations and demonstrated on two data sets: 1) Matched Multiple Birth and 2) Burn Injury. In which of interest is the influence of gestational age and maternal age on twin birth weights in the former data, whereas in the later data we are interested in investigating how patient’s age affects the severity of burn injury and the probability of death.Ph

    Living on the Edge: An Unified Approach to Antithetic Sampling

    Get PDF
    We identify recurrent ingredients in the antithetic sampling literature leading to a unified sampling framework. We introduce a new class of antithetic schemes that includes the most used antithetic proposals. This perspective enables the derivation of new properties of the sampling schemes: i) optimality in the Kullback--Leibler sense; ii) closed-form multivariate Kendall's τ\tau and Spearman's ρ\rho; iii) ranking in concordance order and iv) a central limit theorem that characterizes stochastic behaviour of Monte Carlo estimators when the sample size tends to infinity. The proposed simulation framework inherits the simplicity of the standard antithetic sampling method, requiring the definition of a set of reference points in the sampling space and the generation of uniform numbers on the segments joining the points. We provide applications to Monte Carlo integration and Markov Chain Monte Carlo Bayesian estimation

    A Variance-centric Approach to Statistical Analyses of Large-scale Genetic Data

    Get PDF
    The ever more complex and larger datasets that statisticians can routinely access have prompted the modi cation of traditional statistical methodology. Notably, genomic data have different representations at the gene, exon, and sequence level, which require both customized quality controls to remove technological artefacts and tailored statistical methods to make real discoveries. This thesis focuses on understanding the sources of, and decomposing variance as well as its role in modelling high-dimensional genetic data. Eigenvalue decomposition, factor analysis, and singular value decomposition are variations of variance based transformation techniques that help us understand how data are structured according to variance contribution. A crucial aspect of these techniques is the number of transformed variables to retain, which is closely related to other statistical constructs such as a low-rank approximation, dimension reduction, effective degrees of freedom, and effective sample size. I propose to tackle the estimation of dimension of high-dimensional genomics data via a penalized profile likelihood criterion within the probabilistic principal component analysis framework. Leveraging the existing two-stage regression framework for detecting variance heterogeneity, I propose a simple-to-implement testing strategy for X-chromosome that is robust to confounding and can be used to indirectly infer the presence of genetic interactions. Finally, building on the singular value decomposition, I introduce a novel severity measure of multi-collinearity for high-dimensional data that can be used to construct high-dimensional variance estimators.Ph.D
    corecore