1,720,984 research outputs found
An adaptive empirical likelihood test for parametric time series regression models
Song Xi Chen and Jiti Ga
Two sample inference for high dimensional data and nonparametric variable selection for census data
In the first part of this thesis, we address the question of how new testing methods can be developed for two sample inference for high dimensional data. Particularly, chapter 2 focuses on testing the equality of two high dimensional covariance matrices, which can be directly applied to evaluating the difference in genetic correlation for
different populations subject to various biological conditions. As we will demonstrate in chapter 2 , the test we propose has no normality assumption and also allows the dimension to be much larger than the sample sizes. These two aspects surpass the capacity of the classical tests such as the likelihood ratio test. Testing the equality of high dimensional mean vectors is another important two-sample testing problem. Most tests for the equality of two mean vectors are not powerful against sparse alternative in the sense that the difference of two population mean vectors only spreads out over a small number of coordinates. In chapter 3, we propose two tests designed to obtain better power performance against sparse alternative by conducting both variance reduction and signal enhancement through thresholding and transformation, respectively.
The second part of this thesis is on variable selection for census data. Human populations are heterogeneous in that the probability of enumerating an individual depends on the characteristics of the individual. For the US Census, a group of variables is chosen to reflect much of the heterogeneity and the relevance of these variables to the enumeration function needs to be investigated. In chapter 4, we introduce a nonparametric variable selection method based on the optimal bandwidths obtained by minimizing the cross- validation function. The relevance of each variable to the enumeration function is reflected by the asymptotic convergence of associated optimal bandwidth. Also to formally test the significance of each variable, a bootstrap procedure is introduced.</p
Statistical inference for high-dimensional data
High-dimensional data, where the number of variables p is large compared to the sample size n, are widely available from microarray studies, finance and many other sources. This dissertation focuses on the effects of high dimensionality on some aspects of statistical inference. A two-sample test for means of high-dimensional data
proposed in this dissertation allows p to be much larger than n. We will show that in the simulation study the proposed test statistic performed consistently better than the other existing methods. Two distributions sharing the same mean may differ in many other aspects. We therefore consider a two-sample test for high-dimensional distributions. The proposed test statistic is based on empirical distribution functions
and is a natural extension to our two-sample test statistic for means. Empirical likelihood has many important applications in nonparametric or semiparametric statistical inference. In this dissertation, we further study the effects of data dimension on the asymptotic normality of the empirical likelihood ratio for high-dimensional data under a general multivariate model.</p
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Topics in matrix completion and genomic prediction
This dissertation consists of three projects focused on low-rank modeling to deal with matrix completion problems and genomic prediction by adjusting spatial effects. One big challenge in matrix completion is that the real data arising are high-dimensional, low-rank and have many missing entries. In the first project (Chapter 2), we propose a column-space-decomposition model with the utilization of some additional covariate information. This helps us both in improving the prediction of ratings and understanding how the covariates affect the missingness and ratings. The proposed estimation method is shown to provide efficient estimators and achieve computational efficiency. In the second project (Chapter 3), we are motivated by a general low-rank missing mechanism rather than the specific missing-at-random mechanism assumed in the first project. We consider an additive model with mean effect to estimate the linear predictors which are further used to estimate the probabilities of observations. To get the prediction of ratings under non-uniform missingness, we adopt a weighted objective function and apply constraints to the estimator of probabilities to avoid issues with extreme values. Both the asymptotic convergence rates and numerical efficiencies of the proposed estimators of probabilities and ratings are studied.
In the third project (Chapter 4), we address challenges that arise when phenotypes measured on plants grown in fields are spatially correlated. We focus on a Gaussian random field (GRF) model with an additive covariance matrix structure that incorporates the genotype effects, spatial effects and subpopulation effects to predict phenotypes from a huge number of marker genotypes, accounting for the spatial dependence among measurements. Two datasets are studied by using the GRF model to show the benefits of spatial effects adjustments. Further, we apply the proposed GRF method to help choose the best plants in terms of a specific phenotype.</p
New aspects of statistical methods for missing data problems, with applications in bioinformatics and genetics
As missing data problems become more commonplace in biological research and other areas, a method with relaxed assumptions while flexible enough to accommodate a wide range of situations is highly desired. We propose a nonparametric imputation method for data with missing values. The inference on the parameter defined by general estimating equations is performed using an empirical likelihood method. It is shown that the nonparametric imputation method together with empirical likelihood can reduce bias and improve efficiency of the estimate relative to inference using only complete cases of the dataset. The confidence regions obtained by empirical likelihood demonstrate good coverage properties. Since our method is valid under very weak assumptions while also possessing the flexibility inherent to estimating equations and empirical likelihood, it can be applied to a wide range of problems. An example is given using mouse eye weight and gene expression data;Missing data methods are also highly valuable from an experimental design point of view. We proposed a selective transcriptional profiling approach in improving the efficiency and affordability of genetical genomics research. The high cost of microarrays tends to limit the adoption of the standard genetical genomics approach. Our method is derived in a missing data framework, in which only a subset of objects are subjected to microarray experiments. It is shown that this approach can significantly reduce experimental cost while still achieving satisfactory power. To address the need for a nonparametric method, we developed empirical likelihood based inference for multi-sample comparison problems using data with surrogate variables. By applying this result to selective transcriptional profiling, we show that the idea of using relatively inexpensive trait data on extra individuals to improve the power of test for association between a QTL and gene transcriptional abundance also applies to the empirical likelihood based method.</p
Topics in statistical inference for massive data and high-dimensional data
This dissertation consists of three research papers that deal with three different problems in statistics concerning high-volume datasets. The first paper studies the distributed statistical inference for massive data. With the increasing size of the data, computational complexity and feasibility should be taken into consideration for statistical analyses. We investigate the statistical efficiency of the distributed version of a general class of statistics. Distributed bootstrap algorithms are proposed to approximate the distribution of the distributed statistics. These approaches relief the computational burdens of conventional methods while preserving adequate statistical efficiency. The second paper deals with testing the identity and sphericity hypotheses problem regarding high-dimensional covariance matrices, with a focus on improving the power of existing methods. By taking advantage of the sparsity in the underlying covariance matrices, the power improvement is accomplished by utilizing the banding estimator for the covariance matrices, which leads to a significant reduction in the variance of the test statistics. The last paper considers variable selection for high-dimensional data. Distance-based variable importance measures are proposed to rank and select variables with dependence structures being taken into consideration. The importance measures are inspired by the multi-response permutation procedure (MRPP) and the energy distance. A backward selection algorithm is developed to discover important variables and to improve the power of the original MRPP for high-dimensional data.</p
- …
