1,721,029 research outputs found
Recommended from our members
A Statistical Approach to Detecting Patterns in Behavioral Event Sequences
The identification of recurring patterns within a sequence of events is an important task in behavior research. In this thesis, we develop a probabilistic framework for identifying such patterns from behavioral data. This framework allows us to distinguish between events that belong to a pattern and events that occur as part of background or unstructured behavior. Stochastic processes are introduced to describe the incidence of both background events and events that belong to recurring patterns. The behavioral events are modeled together using a competing risks framework combining the stochastic processes. We develop an inference procedure to detect the sequences present in observed data. The motivation for this work comes from a large scale longitudinal study to assess the impact of fragmented and unpredictable maternal behavior on emotional and cognitive development of children. We describe our results on both simulated data and the maternal data.We also consider extensions to our model. We develop a model to study population level behavior, allowing us to compare separate populations as well as pool information across individuals. To perform this analysis, we describe how to extend both our model and inference procedure to a hierarchical setting. We describe results for both simulated data and the maternal data.Finally, we explore the distributional assumptions inherent in our model. We consider a variety of parametric forms for our model, as well as a nonparametric approach that is both flexible and computationally efficient. In addition to improved model fit, these approaches allow us to better describe background behaviors, such as behaviors that occur in bursts or with strict regularity. We also explore the effect of distributional assumptions on both simulated and real data
Recommended from our members
Methods for Optimal Covariate Balance in Observational Studies for Causal Inference
The most basic approach to causal inference measures the response of a system or population to different exposures, or treatments, and compares one or more summaries of the responses. Thinking formally about causal inference requires that we consider what the potential outcomes would have been under a set of alternative exposures. Though it is only possible in practice to observe a single outcome for each unit, the principles of good experimental design, including random treatment assignment, allow a valid comparison of average responses to each exposure because effects from extraneous factors are minimized. In observational settings, where exposures or treatments arise "naturally", i.e., without experimental manipulation, a common strategy for estimating causal effects is to find units that are similar based upon a set of covariates, but receiving different exposures, and then compare their outcomes. This strategy is challenging if there are many covariates. Balancing scores, a low-dimensional summary of the relevant covariate space, can facilitate causal inference for observational data in settings with many covariates. Propensity scores which measure the probability of receiving a particular exposure or treatment are one example of a balancing score. To estimate treatment effects, balancing scores are used to to group individuals from different exposure groups to compare their response levels, or functions of balancing scores are used to re-weight the sample. This thesis explores novel methods for obtaining covariate balance in observational studies through the use of balancing scores and weighting methodology. The dissertation begins by providing an overview of the potential outcome framework for causal inference in observational studies and required background knowledge for the methods developed. The first methodological contribution is the optimally balanced Gaussian process propensity score approach that applies a binary regression framework using Gaussian processes for estimating the propensity score. The hyperparameters of the process are selected to minimize a metric of covariate imbalance. The next methodological contribution is the development of targeted balancing weights for both binary and multi-treatment settings, where a covariate imbalance metric is created with respect to a covariate density of interest (this could be the distribution within the full population under study or within a specific subpopulation of interest) and unit weights are selected that minimize this metric, without an explicit assumption on the functional form of the weights. Each method is evaluated against competing methods from the causal inference literature through series of simulations and against a benchmark causal inference data set. The dissertation concludes with suggestions for future work. A contribution to measurement in observational studies is included as an appendix
Recommended from our members
Non-Parametric Tests for Treatment Effect Heterogeneity in Randomized Experiments and Observational Studies
Comprehensively assessing the effect of a treatment usually includes two objectives, estimating the average treatment effect across the whole target population and evaluating variability of the treatment effect across different subpopulations in an effort to provide more precise treatment recommendations. A common way to identify treatment effect heterogeneity is to split the sample into several strata based on one or more baseline covariates which may be relevant to the effect of treatment, and then compare the localized or stratum-specific treatment effects across those strata. Parametric approaches have been proposed to compare average treatment effects across several strata. One approach is testing interactions between treatment indicator and group indicators in linear regressions (Allison, 1977), and another approach is the likelihood ratio test proposed by Gail and Simon(1985). Both of them require parametric assumptions of outcome distributions. When the parametric assumptions fail, the test may be invalid or the power may be negatively impacted. Thus there is a need for non-parametric tests that can better adapt to various outcome distributions.
Randomized experiments are considered to be the gold standard for assessing treatment effects, as all baseline covariates are expected to be well balanced in treatment groups after randomization. However, randomized experiments are not always feasible due to various obstacles, e.g., ethical concerns and high expense. Therefore researchers turn to observational studies. Not only can they avoid the obstacles faced by randomized experiments, but because there are often fewer exclusionary characteristics, the results of observational studies may generalize better to the target population. A main challenge of observational studies is controlling confounding variables. There is a considerable literature on causal inference in observational studies that has been developed targeting this challenge. Many of the proposed procedures balance the observed variables and then rely on the unconfoundedness assumption, i.e., all confounding variables are observed. This assumption is not only strong but also unverifiable. The violation of this assumption can invalidate causal conclusions. A variety of approaches to assessing the sensitivity of causal conclusions to violations of the unconfoundedness assumption have been proposed. By assessing the extent of the assumption violation required to change the conclusion and evaluating the possibility of such a violation based on domain knowledge, it is possible to provide more reliable conclusions.
In this dissertation, we describe three contributions we have made to the goal of comprehensively evaluating the effectiveness of treatments. The first two contributions are non-parametric U-statistic-based tests examining the variability of treatment effects across different subpopulations. The first procedure can be appropriately applied in cases, like randomized experiments, where all baseline covariates are well balanced within each stratum; the second procedure adjusts unbalanced confounding variables using propensity scores. Compared to their parametric counterparts, likelihood ratio tests, our non-parametric tests are more powerful when the distributions of study outcomes depart substantially from the distributions assumed by likelihood ratio tests. The third contribution is a sensitivity analysis that addresses the concern of possible violation of the unconfoundedness assumption for the adjusted Mann-Whitney test, a non-parametric test that evaluates the existence of treatment effects in observational studies
Recommended from our members
A Time-Varying Low-Dimensional Representation for Spatio-Temporal Data
We were motivated by the two major limitations of the current research approaches on the North Atlantic Oscillation (NAO) based on empirical orthogonal functions (EOF) analysis: (i) long-term stationary assumptions; (ii) lack of measures of uncertainty, and proposed and developed a time-varying low-dimensional representation for spatio-temporal data in this thesis. The low-dimensional representation is based on a structured spatial covariance matrix using a certain number of structured basis functions with certain parametric forms. Initially, we developed the Parametric Basis Function (PBF) spatial covariance model in a stationary scenario and provided the statistical inference in both maximum likelihood and Bayesian analysis frameworks. We further extended the model by introducing time-varying parameters to develop the time-varying parametric basis function (TV-PBF) model in the state space model framework. The Bayesian approach with MCMC techniques was used to make inference for the TV-PBF model. The model is able to provide smoothly changing patterns of the 1st EOFs NAO over time which can serve as an alternative representation for the spatio-temporal NAO data
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
