1,721,051 research outputs found

    Performance of robust regression methods in real-time polymerase chain reaction calibration

    No full text
    Introduction Controlled calibration enables determining the unknown concentration of a particular substance in a given solution. The assay method of concern here is the real-time polymerase chain reaction (PCR). Ordinary Least Squares (OLS) method is routinely used to estimate the standard curve and the concentration of an unknown sample. It was found that many laboratories participating to a project of External Quality Control concerning quantitative real-time PCR presented outliers, most frequently in the lowest concentrations. Objectives This work aims at investigating and comparing the performance of OLS, Least Absolute Deviation (LAD) method [1] and biweight MM-estimator [2,3] in real-time PCR calibration via a Monte Carlo simulation. Methods The statistical model corresponding to the standard curve is a simple linear regression. Calibration enables predicting the unknown x0 (nucleic acid content of the unknown sample), for a given y0 (value of the fractional cycle where a threshold amount of amplified nucleic acid is produced in the sample), through inverse regression: x . The Monte Carlo simulation mimicked the design adopted in the Italian project of External Quality Control. Outliers were introduced by contamination. The 90% confidence intervals for x0 were obtained by resorting to the approximate variance of x computed by the Delta method and their coverage probability was investigated. Results When contamination is absent, the coverage of OLS estimator interval is at the nominal level and its width is the smallest, MM-estimator performance is similar to OLS, whereas LAD interval has an acceptable coverage at the expense of width. On the contrary the contamination of all concentration levels makes the OLS performance worse: even though the width tends to increase when contamination increases, the coverage tends to decrease and to appear unacceptable; as regards the robust methods we can observe a tradeoff between width and coverage: MM estimator seems resistant to contamination, as the intervals widths remain short, but this is associated with reduction in coverages, whereas LAD intervals widths are constantly larger so that their coverages are acceptable at the nominal level. With reference to contamination of the lowest concentrations only, when x0=2.5 or x0=3.5 all the methods tend to underestimate the nominal coverage, however the performance of LAD method compares favourably to that of the other two. On the contrary when x0=5.5 or x0=6.5 we observe an improper increase of the confidence intervals width for all methods studied and consequently highly overestimating coverages, particularly for MM estimator and LAD. Conclusions In performing real-time PCR calibration, it seems reasonable to postulate a heterogeneous distribution of measurements errors; hence robust regression methods should be implemented to estimate x0 and to compute its confidence interval. However to select the most appropriate robust procedure a thorough investigation of the features of the process generating measurements in each laboratory is needed

    Presenza di outlier nei saggi biologici con PCR real-time: L1-norm (LAD) or L2-norm (OLS)?

    No full text
    Real-time PCR is a laboratory technique which is used to amplify and simultaneously quantify the concentration of biological markers based on sequences of nucleic acid (DNA or RNA). We investigated data from an external quality control of real-time PCR and we observed the presence of outliers. The aim of this work is to compare the performance of L2-norm (OLS) regression method with L1-norm (LAD) robust regression method in estimating the nucleic acid concentration in the presence of outliers. For this purpose a simulation study was planned and two different procedures (OLS and LAD) were applied to the simulated data. Both non contaminated and contaminated (with outliers) situations were considered. In the absence of outliers OLS procedure enables obtaining a precise and accurate estimate of the nucleic acid concentration, whereas the presence of outliers leads to a biased estimate, which can be improved by using LAD robust method

    Performance of robust regression methods in real-time polymerase chain reaction calibration

    No full text
    The ordinary least squares (OLS) method is routinely used to estimate the unknown concentration of nucleic acids in a given solution by means of calibration. However, when outliers are present it could appear sensible to resort to robust regression methods. We analyzed data from an External Quality Control program concerning quantitative real-time PCR and we found that 24 laboratories out of 40 presented outliers, which occurred most frequently at the lowest concentrations. In this article we investigated and compared the performance of the OLS method, the least absolute deviation (LAD) method, and the biweight MM-estimator in real-time PCR calibration via a Monte Carlo simulation. Outliers were introduced by replacement contamination. When contamination was absent the coverages of OLS and MM-estimator intervals were acceptable and their widths small, whereas LAD intervals had acceptable coverages at the expense of higher widths. In the presence of contamination we observed a trade-off between width and coverage: the OLS performance got worse, the MM-estimator intervals widths remained short (but this was associated with a reduction in coverages), while LAD intervals widths were constantly larger with acceptable coverages at the nominal level

    A reaction to a challenging example in multiple regression analysis

    No full text
    In a very stimulating paper, Preece gives an artificial dataset useful to illustrate the hazard of multiple regression and challenges the reader to spot the simple inbuilt features of these data. The present note aims at finding how Preece generated the whole set of data. First of all OLS regression model is fitted to the data; after checking for model assumptions some doubts arise on the validity of OLS regression; thus robust regression estimators are considered as a proper alternative. The latter give discordant coefficient estimates, but after a deep analysis, they agree in highlighting the presence of two subsets within the dataset: 9 cases being generated by one model, and the remaining 8 cases being generated by a second model. This particular pattern of the data is recognized by the mixture model as well

    A performance counterexample of Billor-Chatterjee-Hadi procedure and an improvement proposal for robust regression

    No full text
    The BCH procedure introduced by Billor et al. for fitting linear models was found to be inefficient for y-outliers in the presence of a high perturbation level. We propose to modify the first step of the BCH procedure, so that the robust distances are computed on the matrix Z=(y, X) of the basic subset. The performance of the present note procedure (PNP), as compared to the BCH procedure and the OLS method, was studied by processing several datasets used in the literature for robust regression and by performing a Monte Carlo experiment. PNP performs better particularly with datasets having high perturbation

    SURVIVAL ANALYSIS AND REGRESSION MODELS IN THE PRESENCE OF COMPETING AND SEMI-COMPETING RISKS

    Get PDF
    Evaluation of a therapeutic strategy is complex when the course of a disease is characterized by the occurrence of different kinds of events. Competing risks arise when the occurrence of specific events prevents the observation of other events. Different survival or incidence functions can be defined in the presence of competing risks and a relevant issue is an adequate knowledge of the methodological background in order to apply a suitable statistical analysis for the study aims. This work aims at presenting different estimates of survival or incidence probabilities used in this framework. From clinical application, it emerges that crude cumulative incidence is widely diffuses both to estimate incidence probabilities and to evaluate covariate effects. On the contrary net survival functions, although of clinical interest, are not diffused because of more difficult model structure and lack of software availability. If the independence assumption between different events is tenable, Kaplan-Meier method can be used to estimate net survival. Otherwise multivariate distribution of times based on Copulas can be adopted. In the case of different causes of death, relative survival can be interpreted as net survival only under specific assumptions on the mortality pattern. A particular case on competing risks arises when only fatal events can prevent the observation of the non fatal ones, but not vice versa (semi-competing risks). The estimate of an interpretable measure of association between times to non fatal and fatal event is often of biological interest, in order to understand the disease progression. In the statistical literature some approaches have been proposed to estimate the association between two independently doubly censored failure times, but more specific approach have to be applied in the presence of semi-competing risks. After estimating the association parameter, the survival function of a non terminal event can be estimated after fixing a copula structure by means of the semi parametric methods proposed by Fine, Jiang and Chappell or the copula graphic estimator. Furthermore when the interest is to evaluate the effect of different therapeutic strategies or covariates on the occurrence of a non terminal event in a semi-competing risks setting, specific regression model have to be adopted. I propose here to adopt the methodology based on pseudo-observations, having the advantage to be implemented by generalized linear models approaches. Simulation studies are performed to compare the performances of methods to estimate the association between events, of methods based on Copulas models to estimate net survival and of regression method for net survival in the presence of semi-competing risks. A case series of breast cancer patients is used to illustrate different methods of estimating net survival functions on the causes of deaths and on the severe non fatal events in the presence of competing and semi-competing risks framework

    Esimating relapse free survival as a net probability: regression models and graphical representation

    No full text
    In several experimental or observational clinical studies, the evaluation of the effect of a therapy and the impact of prognostic factors is based on relapse-free survival and the suited regression models. Relapse free survival is a net survival and needs to be interpreted as the survival probability that would be observed if all patients experienced relapse sooner or later. Death without evidence of relapse prevents the subsequent observation of relapse, acting in a competing risks framework. Relapse free survival is often estimated by standard regression models after censoring times to death. The association between relapse and death is thus ignored. However to better estimate relapse free survival a bivariate distribution of times to events needs to be considered, for example by means of copula models, with ad hoc estimating procedures. We concentrate here on the copula graphic estimator, for which a pertinent regression model has been developed (Lo and Wilke). The advantage of this approach is based on the relationship between net survival, overall survival and cause specific hazard. Regression models can be fitted for the latter quantity by standard statistical methods and the estimates can be used to compute net survival through a copula structure. Parametric models are preferred. To avoid the constraint of parametric distribution, we propose piecewise regression models. A consistent estimate of the association parameter for the copula model can be obtained by considering the semi-competing risks framework, because death can be observed after relapse. The drawback of the copula graphic regression model is that no direct parametric estimation of the regression coefficient for the covariates is available. To obtain an overall view of the association between covariate levels and net relapse free survival we propose a multivariate visualisation approach through Multiple Correspondence Analysis. This approach has been applied to two case series of patients with breast cancer and extremity soft tissue sarcoma respectively, in order to compare the results obtained by piecewise exponential model on cause specific hazard and net relapse free survival computed through copula graphic estimator

    Detecting outliers and/or leverage points: a robust two-stage procedure with bootstrap cut-off points

    No full text
    This paper presents a robust two-stage procedure for identification of outlying observations in regression analysis. The exploratory stage identifies leverage points and vertical outliers through a robust distance estimator based on Minimum Covariance Determinant (MCD). After deletion of these points, the confirmatory stage carries out an Ordinary Least Squares (OLS) analysis on the remaining subset of data and investigates the effect of adding back in the previously deleted observations. Cut-off points pertinent to different diagnostics are generated by bootstrapping and the cases are definitely labelled as good-leverage, bad-leverage, vertical outliers and typical cases. The procedure is applied to four examples.</p
    corecore