1,721,051 research outputs found
Performance of robust regression methods in real-time polymerase chain reaction calibration
Introduction Controlled calibration enables determining the unknown concentration of a
particular substance in a given solution. The assay method of concern here is the real-time
polymerase chain reaction (PCR). Ordinary Least Squares (OLS) method is routinely used
to estimate the standard curve and the concentration of an unknown sample. It was found
that many laboratories participating to a project of External Quality Control concerning
quantitative real-time PCR presented outliers, most frequently in the lowest concentrations.
Objectives This work aims at investigating and comparing the performance of OLS, Least
Absolute Deviation (LAD) method [1] and biweight MM-estimator [2,3] in real-time PCR
calibration via a Monte Carlo simulation.
Methods The statistical model corresponding to the standard curve is a simple linear
regression. Calibration enables predicting the unknown x0 (nucleic acid content of the
unknown sample), for a given y0 (value of the fractional cycle where a threshold amount of
amplified nucleic acid is produced in the sample), through inverse regression: x
.
The Monte Carlo simulation mimicked the design adopted in the Italian project of External
Quality Control. Outliers were introduced by contamination. The 90% confidence intervals
for x0 were obtained by resorting to the approximate variance of x computed by the Delta
method and their coverage probability was investigated.
Results When contamination is absent, the coverage of OLS estimator interval is at the
nominal level and its width is the smallest, MM-estimator performance is similar to OLS,
whereas LAD interval has an acceptable coverage at the expense of width. On the contrary
the contamination of all concentration levels makes the OLS performance worse: even
though the width tends to increase when contamination increases, the coverage tends to
decrease and to appear unacceptable; as regards the robust methods we can observe a tradeoff
between width and coverage: MM estimator seems resistant to contamination, as the
intervals widths remain short, but this is associated with reduction in coverages, whereas
LAD intervals widths are constantly larger so that their coverages are acceptable at the
nominal level. With reference to contamination of the lowest concentrations only, when
x0=2.5 or x0=3.5 all the methods tend to underestimate the nominal coverage, however the
performance of LAD method compares favourably to that of the other two. On the contrary
when x0=5.5 or x0=6.5 we observe an improper increase of the confidence intervals width
for all methods studied and consequently highly overestimating coverages, particularly for
MM estimator and LAD.
Conclusions In performing real-time PCR calibration, it seems reasonable to postulate a
heterogeneous distribution of measurements errors; hence robust regression methods should
be implemented to estimate x0 and to compute its confidence interval. However to select the
most appropriate robust procedure a thorough investigation of the features of the process
generating measurements in each laboratory is needed
Presenza di outlier nei saggi biologici con PCR real-time: L1-norm (LAD) or L2-norm (OLS)?
Real-time PCR is a laboratory technique which is used to amplify and simultaneously quantify
the concentration of biological markers based on sequences of nucleic acid (DNA or RNA).
We investigated data from an external quality control of real-time PCR and we observed the
presence of outliers.
The aim of this work is to compare the performance of L2-norm (OLS) regression method with
L1-norm (LAD) robust regression method in estimating the nucleic acid concentration in the
presence of outliers.
For this purpose a simulation study was planned and two different procedures (OLS and LAD)
were applied to the simulated data. Both non contaminated and contaminated (with outliers)
situations were considered.
In the absence of outliers OLS procedure enables obtaining a precise and accurate estimate of
the nucleic acid concentration, whereas the presence of outliers leads to a biased estimate, which
can be improved by using LAD robust method
Pinpointing outliers and influential cases in regression analysis : a robust method at work
Performance of robust regression methods in real-time polymerase chain reaction calibration
The ordinary least squares (OLS) method is routinely used to estimate the unknown concentration of nucleic acids in a given solution by means of calibration. However, when outliers are present it could appear sensible to resort to robust regression methods. We analyzed data from an External Quality Control program concerning quantitative real-time PCR and we found that 24 laboratories out of 40 presented outliers, which occurred most frequently at the lowest concentrations. In this article we investigated and compared the performance of the OLS method, the least absolute deviation (LAD) method, and the biweight MM-estimator in real-time PCR calibration via a Monte Carlo simulation. Outliers were introduced by replacement contamination. When contamination was absent the coverages of OLS and MM-estimator intervals were acceptable and their widths small, whereas LAD intervals had acceptable coverages at the expense of higher widths. In the presence of contamination we observed a trade-off between width and coverage: the OLS performance got worse, the MM-estimator intervals widths remained short (but this was associated with a reduction in coverages), while LAD intervals widths were constantly larger with acceptable coverages at the nominal level
A reaction to a challenging example in multiple regression analysis
In a very stimulating paper, Preece gives an artificial dataset useful to illustrate the hazard of multiple regression and challenges the reader to spot the simple inbuilt features of these data. The present note aims at finding how Preece generated the whole set of data. First of all OLS regression model is fitted to the data; after checking for model assumptions some doubts arise on the validity of OLS regression; thus robust regression estimators are considered as a proper alternative. The latter give discordant coefficient estimates, but after a deep analysis, they agree in highlighting the presence of two subsets within the dataset: 9 cases being generated by one model, and the remaining 8 cases being generated by a second model. This particular pattern of the data is recognized by the mixture model as well
A performance counterexample of Billor-Chatterjee-Hadi procedure and an improvement proposal for robust regression
The BCH procedure introduced by Billor et al. for fitting linear models was found to be inefficient for y-outliers in the presence of a high perturbation level. We propose to modify the first step of the BCH procedure, so that the robust distances are computed on the matrix Z=(y, X) of the basic subset. The performance of the present note procedure (PNP), as compared to the BCH procedure and the OLS method, was studied by processing several datasets used in the literature for robust regression and by performing a Monte Carlo experiment. PNP performs better particularly with datasets having high perturbation
SURVIVAL ANALYSIS AND REGRESSION MODELS IN THE PRESENCE OF COMPETING AND SEMI-COMPETING RISKS
Evaluation of a therapeutic strategy is complex when the course of a disease is characterized by the occurrence of different kinds of events. Competing risks arise when the occurrence of specific events prevents the observation of other events. Different survival or incidence functions can be defined in the presence of competing risks and a relevant issue is an adequate knowledge of the methodological background in order to apply a suitable statistical analysis for the study aims. This work aims at presenting different estimates of survival or incidence probabilities used in this framework. From clinical application, it emerges that crude cumulative incidence is widely diffuses both to estimate incidence probabilities and to evaluate covariate effects. On the contrary net survival functions, although of clinical interest, are not diffused because of more difficult model structure and lack of software availability.
If the independence assumption between different events is tenable, Kaplan-Meier method can be used to estimate net survival. Otherwise multivariate distribution of times based on Copulas can be adopted. In the case of different causes of death, relative survival can be interpreted as net survival only under specific assumptions on the mortality pattern.
A particular case on competing risks arises when only fatal events can prevent the observation of the non fatal ones, but not vice versa (semi-competing risks). The estimate of an interpretable measure of association between times to non fatal and fatal event is often of biological interest, in order to understand the disease progression. In the statistical literature some approaches have been proposed to estimate the association between two independently doubly censored failure times, but more specific approach have to be applied in the presence of semi-competing risks. After estimating the association parameter, the survival function of a non terminal event can be estimated after fixing a copula structure by means of the semi parametric methods proposed by Fine, Jiang and Chappell or the copula graphic estimator.
Furthermore when the interest is to evaluate the effect of different therapeutic strategies or covariates on the occurrence of a non terminal event in a semi-competing risks setting, specific regression model have to be adopted. I propose here to adopt the methodology based on pseudo-observations, having the advantage to be implemented by generalized linear models approaches.
Simulation studies are performed to compare the performances of methods to estimate the association between events, of methods based on Copulas models to estimate net survival and of regression method for net survival in the presence of semi-competing risks.
A case series of breast cancer patients is used to illustrate different methods of estimating net survival functions on the causes of deaths and on the severe non fatal events in the presence of competing and semi-competing risks framework
Esimating relapse free survival as a net probability: regression models and graphical representation
In several experimental or observational clinical studies, the evaluation of the effect of a therapy and the impact of prognostic factors is based on relapse-free survival and the suited regression models. Relapse free survival is a net survival and needs to be interpreted as the survival probability that would be observed if all patients experienced relapse sooner or later.
Death without evidence of relapse prevents the subsequent observation of relapse, acting in a competing risks framework.
Relapse free survival is often estimated by standard regression models after censoring times to death. The association between relapse and death is thus ignored. However to better estimate relapse free survival a bivariate distribution of times to events needs to be considered, for example by means of copula models, with ad hoc estimating procedures. We concentrate here on the copula graphic estimator, for which a pertinent regression model has been developed (Lo and Wilke). The advantage of this approach is based on the relationship between net survival, overall survival and cause specific hazard. Regression models can be fitted for the latter quantity by standard statistical methods and the estimates can be used to compute net survival through a copula structure. Parametric models are preferred. To avoid the constraint of parametric distribution, we propose piecewise regression models. A consistent estimate of the association parameter for the copula model can be obtained by considering the semi-competing risks framework, because death can be observed after relapse. The drawback of the copula graphic regression model is that no direct parametric estimation of the regression coefficient for the covariates is available. To obtain an overall view of the association between covariate levels and net relapse free survival we propose a multivariate visualisation approach through Multiple Correspondence Analysis.
This approach has been applied to two case series of patients with breast cancer and extremity soft tissue sarcoma respectively, in order to compare the results obtained by piecewise exponential model on cause specific hazard and net relapse free survival computed through copula graphic estimator
Detecting outliers and/or leverage points: a robust two-stage procedure with bootstrap cut-off points
This paper presents a robust two-stage procedure for identification of outlying observations in regression analysis. The exploratory stage identifies leverage points and vertical outliers through a robust distance estimator based on Minimum Covariance Determinant (MCD). After deletion of these points, the confirmatory stage carries out an Ordinary Least Squares (OLS) analysis on the remaining subset of data and investigates the effect of adding back in the previously deleted observations. Cut-off points pertinent to different diagnostics are generated by bootstrapping and the cases are definitely labelled as good-leverage, bad-leverage, vertical outliers and typical cases. The procedure is applied to four examples.</p
- …
