Journal of Statistical Software
Not a member yet
1629 research outputs found
Sort by
clinicalsignificance: Clinical Significance Analyses of Intervention Studies in R
The analysis of clinical significance is helpful to decide if an intervention leads to practically relevant or meaningful changes for individual patients which is clearly different from the analysis of statistical significance. However, the framework is used rarely and inconsistently. We introduce the R package clinicalsignificance to harness the use of clinical significance analysis of intervention trials in clinical research. This package provides all relevant methods to calculate and present analyses of clinical significance in a consistent form and easy to use implementation. Despite its shortcomings, clinical significance analyses are a valuable tool to gain more insight into intended and potential unintended intervention effects and they may improve the interpretation and comparability of intervention trial results. Lastly, analyses of clinical significance may guide researchers and policy makers in determining which interventions are clinically effective
BEKKs: An R Package for Estimation of Conditional Volatility of Multivariate Time Series
We describe the R package BEKKs, which implements the estimation and diagnostic analysis of a prominent family of multivariate generalized autoregressive conditionally heteroskedastic (MGARCH) processes, the so-called BEKK models. Unlike existing software packages, we make use of analytical derivatives implemented in efficient C++ code for nonlinear log-likelihood optimization. This allows fast parameter estimation even in higher model dimensions N > 3. The baseline BEKK model is complemented with an asymmetric parameterization that allows for a flexible modeling of conditional (co)variances. Furthermore, we provide the user with the simplified scalar and diagonal BEKK models to deal with high dimensionality of heteroskedastic time series. The package is designed in an object-oriented way featuring a comprehensive toolbox of methods to investigate and interpret, for instance, volatility impulse response functions, risk estimation and forecasting (VaR) and a backtesting algorithm to compare the forecasting performance of alternative BEKK models. For illustrative purposes, we analyze a bivariate ETF return series (S&P, US treasury bonds) and a four-dimensional system comprising, in addition, a gold ETF and changes of a log oil price by means of the suggested package. We find that the BEKKs package is more than 100 times faster for time series systems of dimension N > 3 than other existing packages
GET: Global Envelopes in R
This work describes the R package GET that implements global envelopes for a general set of d-dimensional vectors T in various applications. A 100(1 - α)% global envelope is a band bounded by two vectors such that the probability that T falls outside this envelope in any of the d points is equal to α. The term 'global' means that this probability is controlled simultaneously for all the d elements of the vectors. The global envelopes can be employed for central regions of functional or multivariate data, for graphical Monte Carlo and permutation tests where the test statistic is multivariate or functional, and for global confidence and prediction bands. Intrinsic graphical interpretation property is introduced for global envelopes. The global envelopes included in the GET package have this property, which particularly helps to interpret test results, by providing a graphical interpretation that shows the reasons of rejection of the tested hypothesis. Examples of different uses of global envelopes and their implementation in the GET package are presented, including global envelopes for single and several one- or two-dimensional functions, Monte Carlo goodness-of-fit tests for simple and composite hypotheses, comparison of distributions, functional analysis of variance, functional linear model, and confidence bands in polynomial regression
How to Interpret Statistical Models Using marginaleffects for R and Python
The parameters of a statistical model can sometimes be difficult to interpret substantively, especially when that model includes nonlinear components, interactions, or transformations. Analysts who fit such complex models often seek to transform raw parameter estimates into quantities that are easier for domain experts and stakeholders to understand. This article presents a simple conceptual framework to describe a vast array of such quantities of interest, which are reported under imprecise and inconsistent terminology across disciplines: predictions, marginal predictions, marginal means, marginal effects, conditional effects, slopes, contrasts, risk ratios, etc. We introduce marginaleffects, a package for R and Python which offers a simple and powerful interface to compute all of those quantities, and to conduct (non-)linear hypothesis and equivalence tests on them. marginaleffects is lightweight; extensible; it works well in combination with other R and Python packages; and it supports over 100 classes of models, including linear, generalized linear, generalized additive, mixed effects, Bayesian, and several machine learning models
makemyprior: Intuitive Construction of Joint Priors for Variance Parameters in R
Priors allow us to robustify inference and to incorporate expert knowledge in Bayesian hierarchical models. This is particularly important when there are random effects that are hard to identify based on observed data. The challenge lies in understanding and controlling the joint influence of the priors for the variance parameters, and makemyprior is an R package that guides the formulation of joint prior distributions for variances. A joint prior distribution is constructed based on a hierarchical decomposition of the total variance in the model along a tree, and takes the entire model structure into account. Users input their prior beliefs or express ignorance at each level of the tree. Prior beliefs can be general ideas about reasonable ranges of variance values and need not be detailed expert knowledge. The constructed priors lead to robust inference and guarantee proper posteriors. A graphical user interface facilitates construction and assessment of different choices of priors through visualization of the tree and joint prior. The package aims to expand the toolbox of applied researchers and make priors an active component in their Bayesian workflow
bayesnec: An R Package for Concentration-Response Modeling and Estimation of Toxicity Metrics
The bayesnec package has been developed for R to fit concentration (dose)-response curves (CR) to toxicity data for the purpose of deriving no-effect-concentration (NEC), no-significant-effect-concentration (NSEC), and effect-concentration (of specified percentage "x", ECx) thresholds from non-linear models fitted using Bayesian Hamiltonian Monte Carlo (HMC) via R packages brms and rstan or cmdstanr. In bayesnec it is possible to fit a single model, custom model-set, specific model-set or all of the available models. When multiple models are specified, the bnec() function returns a model weighted average estimate of predicted posterior values. A range of support functions and methods is also included to work with the returned single, or multi-model objects that allow extraction of raw, or model averaged predicted, NEC, NSEC and ECx values and to interrogate the fitted model or model-set. By combining Bayesian methods with model averaging, bayesnec provides a single estimate of toxicity and associated uncertainty that can be directly integrated into risk assessment frameworks
CRTFASTGEEPWR: A SAS Macro for Power of Generalized Estimating Equations Analysis of Multi-Period Cluster Randomized Trials with Application to Stepped Wedge Designs
Multi-period cluster randomized trials (CRTs) are increasingly used for the evaluation of interventions delivered at the group level. While generalized estimating equations (GEE) are commonly used to provide population-averaged inference in CRTs, there is a gap of general methods and statistical software tools for power calculation based on multi-parameter, within-cluster correlation structures suitable for multi-period CRTs that can accommodate both complete and incomplete designs. A computationally fast, nonsimulation procedure for determining statistical power is described for the GEE analysis of complete and incomplete multi-period cluster randomized trials. The procedure is implemented via a SAS macro, CRTFASTGEEPWR, which is applicable to binary, count and continuous responses and several correlation structures in multi-period CRTs. The SAS macro is illustrated in the power calculation of two complete and two incomplete stepped wedge cluster randomized trial scenarios under different specifications of marginal mean model and within-cluster correlation structure. The proposed GEE power method is quite general as demonstrated in the SAS macro with numerous input options. The power procedure and macro can also be used in the planning of parallel and crossover CRTs in addition to cross-sectional and closed cohort stepped wedge trials
anomaly: Detection of Anomalous Structure in Time Series Data
One of the contemporary challenges in anomaly detection is the ability to detect, and differentiate between, both point and collective anomalies within a data sequence or time series. The anomaly package has been developed to provide users with a choice of anomaly detection methods and, in particular, provides an implementation of the recently proposed collective and point anomaly family of anomaly detection algorithms. This article describes the methods implemented whilst also highlighting their application to simulated data as well as real data examples contained in the package
Modeling Big, Heterogeneous, Non-Gaussian Spatial and Spatio-Temporal Data Using FRK
Non-Gaussian spatial and spatio-temporal data are becoming increasingly prevalent, and their analysis is needed in a variety of disciplines. FRK is an R package for spatial and spatio-temporal modeling and prediction with very large data sets that, to date, has only supported linear process models and Gaussian data models. In this paper, we describe a major upgrade to FRK that allows for non-Gaussian data to be analyzed in a generalized linear mixed model framework. These vastly more general spatial and spatio-temporal models are fitted using the Laplace approximation via the software TMB. The existing functionality of FRK is retained with this advance into non-Gaussian models; in particular, it allows for automatic basis-function construction, it can handle both point-referenced and areal data simultaneously, and it can predict process values at any spatial support from these data. This new version of FRK also allows for the use of a large number of basis functions when modeling the spatial process, and thus it is often able to achieve more accurate predictions than previous versions of the package in a Gaussian setting. We demonstrate innovative features in this new version of FRK, highlight its ease of use, and compare it to alternative packages using both simulated and real data sets
melt: Multiple Empirical Likelihood Tests in R
Empirical likelihood enables a nonparametric, likelihood-driven style of inference without relying on assumptions frequently made in parametric models. Empirical likelihood-based tests are asymptotically pivotal and thus avoid explicit studentization. This paper presents the R package melt that provides a unified framework for data analysis with empirical likelihood methods. A collection of functions are available to perform multiple empirical likelihood tests for linear and generalized linear models in R. The package melt offers an easy-to-use interface and flexibility in specifying hypotheses and calibration methods, extending the framework to simultaneous inferences. Hypothesis testing uses a projected gradient algorithm to solve constrained empirical likelihood optimization problems. The core computational routines are implemented in C++, with OpenMP for parallel computation