1,721,114 research outputs found
A journey in single steps: robust one-step M-estimation
We present a unified treatment of different types of one-step M-estimation in regression models which incorporates the Newton–Raphson, method of scoring and iteratively reweighted least squares forms of one-step estimator. We use higher order expansions to distinguish between the different forms of estimator and the effects of different initial estimators. We show that the Newton–Raphson form has better properties than the method of scoring form which, in turn, has better properties than the iteratively reweighted least squares form. We also show that the best choice of initial estimator is a smooth, robust estimator which converges at the rate n?1/2. These results have important consequences for the common data-analytic strategy of using a least squares analysis on "clean" data obtained by deleting observations with extreme residuals from an initial least squares fit. It is shown that the resulting estimator is an iteratively reweighted least squares one-step estimator with least squares as the initial estimator, giving it the worst performance of the one-step estimators we consider: inferences resulting from this strategy are neither valid nor robust
Resistant techniques for nonparametric regression, generalized linear and additive models
The motivation for our research is issued from the importance of modeling in statistical work. In fact, the description of the relation between two or more qualitative or quantitative variables is one of the major concern of a statistician, but it is also of interest to applied researchers in economics, social and medical sciences, biology and so on. In this thesis, we develop parametric and nonparametric statistical tools for modern regression from the point of view of robust statistics. The work can essentially be summarized in three parts. In the first one, we address the problem of the automatic selection of the smoothing parameter in nonparametric regression and propose two resistant criteria which can, for example, be applied to M-type smoothing splines. The proposals are based on robust predictive error criteria. The second part of the thesis is devoted to robust estimation and testing in the parametric framework of generalized linear models. A class of M-estimators and a family of robust test statistics are developed. The estimators proposed are a generalization of the classical quasi-likelihood estimators and the test statistics are based on a robust deviance function, which makes them a generalization of the quasi-deviance tests. The statistical properties of both estimators and tests are derived. Finally, the last part of this work merges the techniques developed in parametric and nonparametric regression into the framework of generalized additive models. In particular, the techniques for univariate resistant smoothing are used as building blocks in the generalized additive model, and we transfer the family of test statistics developed for generalized linear models to the setting of generalized additive models. The possibility of implementation of these techniques has been our concern through all these three steps
Accurate finite sample inference for generalized linear models and models on overidentifying moment conditions
Classical inference in statistic and econometric models is typically carried out by means of asymptotic approximations to the sampling distribution of estimators and test statistics. These approximations often do not provide accurate p-values and confidences intervals, especially when the sample size is small. Moreover, even if the sample size is large, the accuracy can be poor due to model misspecification (nonrobustness). Several alternative techniques have been proposed in the statistic and econometric literature to improve the accuracy of clasical inference. In general, these alternatives address either the accuracy of the first-order approximations or the nonrobustness issue. However, the development of general procedures which are both robust and second order accurate is still an open question. In this thesis, we propose an alternative statistical test wich has both robustness and small sample properties for two large and important classes of models: Generalized Linear Models (GLM) and models on overidentifying moments conditions
Robust methods for personal income distribution models
In the present thesis, robust statistical techniques are applied and developed for the economic problem of the analysis of personal income distributions and inequality measures. We follow the approach based on influence functions in order to develop robust estimators for the parametric models describing personal income distributions when the data are censored and when they are grouped. We also build a robust procedure for a test of choice between two models and analyse the robustness properties of goodness-of-fit tests. The link between economic and robustness properties is studied through the analysis of inequality measures. We begin our discussion by presenting the economic framework from which the statistical developments are made, namely the study of the personal income distribution and inequality measures. We then discuss the robust concepts that serve as basis for the following steps and compute optimal bounded-influence estimators for different personal income distribution models when the data are continuous and complete. In a third step, we study the case of censored data and propose a generalization of the EM algorithm with robust estimators. For grouped data, Hampel's theorem is extended in order to build optimally bounded-influence estimators for grouped data. We then focus on tests for model choice and develop a robust generalized Cox-type statistic. We also analyse the robustness properties of a wide class of goodness-of-fit statistics by computing their level influence functions. Finally, we study the robustness properties of inequality measures and relate our findings with some economic properties these measures should fulfil. Our motivation for the development of these new robust procedures comes from our interest in the field of income distribution and inequality measurement. However, it should be stressed that the new estimators and tests procedures we propose do not only apply in this particular field, but they can be used in or extended to any parametric problem in which density estimation, incomplete information, grouped or discrete data, model choice, goodness-of-fit, concentration index, is one of the key words
Confidence sets for model selection
At first glance the goals of model selection might seem clear. Out of a set of possible models, we want to select the ”best” or a subset of ”best” models. This notion of ”best” however is not well defined, since it obviously depends on the initial goals of the selection. In order to study the uncertainty in model selection, we introduce a new definition of a model, where the models are no longer defined through zero and non-zero components but through irrelevant and relevant component. Then inspired by confidence intervals for estimated parameters, we propose a method to build confidence sets for model selection in a parametric setting, i.e. create sets of models within which the true model is included with a certain confidence. This allows us to perform inference on model selection. We discuss the computational challenges with such a method, how to find p-values (for the model), consistency in model selection and through a data set and a simulation study show the implications of this new method
Robust penalized M-estimators for generalized linear and additive models
Generalized linear models (GLM) and generalized additive models (GAM) are popular statistical methods for modelling continuous and discrete data both parametrically and nonparametrically. In this general framework, we consider the problem of variable selection by studying a wide class of penalized M-estimators that are particularly well suited for high dimensional scenarios where the number of covariates is very large relative to the sample size . We focus on resistance issues in the presence of deviations from the stochastic assumptions of the postulated models and highlight the weaknesses of widely used estimators. We advocate the need for robust estimators and propose several penalized quasilikelihood estimators that achieve both good statistical properties at the assumed model and stability in a neighborhood of it. Specifically, we provide careful asymptotic analyses of our robust estimators for GLM and GAM when the number of parameters increases with the sample size. We start by revisiting the asymptotics of M-estimators for GLM with a diverging number of parameters. We establish asymptotic normality of these estimators and reexamine distributional results for likelihood ratio type and Wald type tests based on them. We then consider penalized M-estimators for high dimensional set ups where . In the GLM setting we show that our estimators are consistent, asymptotically normally distributed and variable selection consistent under regularity conditions. Furthermore they have a bounded bias in a neighborhood of the model. In the GAM setting we establish an -norm consistency result for the nonparametric components which achieves the optimal rates of convergence. In addition, the proposed penalized estimator is able to select the correct model consistently. We propose new algorithms for the implementation of our penalized M-estimators and illustrate the finite sample performance of our methods, at the model and under contamination, in simulation studies. An important contribution of this thesis is to formally study the local robustness properties of general nondifferentiable penalized M-estimators. In particular, we propose a framework that allows us to define rigorously the influence function as the limiting influence function of a sequence of approximating functionals. We show that this influence function can be used to characterize the robustness properties of a wide range of sparse estimators and that it can be viewed as a derivative in the sense of distribution theory. At the end of this thesis, we discuss some extensions of our work and give an overview of the future challenges of robust statistics in high dimensions
Validity and accuracy of posterior distributions in Bayesian statistics
In this thesis I investigate the validity and the accuracy properties of the posterior quantiles in Bayesian statistics when replacing the parametric likelihood with the Cressie-Read empirical likelihoods based on a set of unbiased M-estimating equations. At first order I study the validity of the empirical posterior distribution derived from the pseudo-likelihood constructed with profiled weights and estimated at a minimum distance from the empirical distribution in the Cressie-Read family of divergences, indexed by γ. The bias in coverage of the resulting empirical posterior quantile is inversely proportional to the asymptotic efficiency of the estimator corresponding to the set of M-estimating functions. By comparing different members of the Cressie-Read family of empirical likelihoods for models in the exponential family, I establish a hierarchy in the accuracy of the quantile function of the empirical posterior distribution depending on the index parameter γ
Three essays in econometrics: robust model and moment selection in GMM and an application of semi-parametric Taylor rules
The first two chapters of this thesis develop a new methodology in the Generalized Method of Moments. Typically, researchers assume that the data come from an unknown ideal distribution. In the first two chapters, we relax this assumption by assuming shrinking neighborhoods of this ideal distribution. We show the conditions that are needed for GMM estimators to have a stable behaviour in these neighborhoods and propose a robust GMM estimator based on these conditions. Finally, we show how to perform Moment and Model Selection based on the robust GMM estimators. The third chapter is an extension of the framework introduced by J. Taylor in 1993. We propose a more flexible setting that can capture asymmetric preferences of the Central Bank between the macroeconomic fundamentals as well as a possibly nonlinear economy
Small sample asymptotics: a review with applications to robust statistics
The aim of this paper is to review concepts, theory, and applications of small sample asymptotic techniques. The striking characteristic of these techniques is that they give uniformly very accurate approximation in the tails of the distribution of a statistics based on n observations even for very small sample sizes. These ideas will be applied to various classes of estimators and tests, including L-estimators, rank procedures, maximum likelihood estimators for general models, and multivariate M-estimators. Other applications to robust statistics, density estimation, and to an efficiency's criterion discussed by Rao will also be discussed
- …
