1,721,007 research outputs found
The Generalized Linear Mixed Model for Finite Normal Mixtures with Application to Tendon Fibrilogenesis Data
We propose the generalized linear mixed model for finite normal mixtures (GLMFM), as well as the estimation procedures for the GLMFM model, which are widely applicable to the hierarchical dataset with small number of individual units and multi-modal distributions at the lowest level of clustering. The modeling task is two-fold: (a). to model the lowest level cluster as a finite mixtures of the normal distribution; and (b). to model the properly transformed mixture proportions, means and standard deviations of the lowest-level cluster as a linear hierarchical structure. We propose the robust generalized weighted likelihood estimators and the new cubic-inverse weight for the estimation of the finite mixture model (Zhan et al., 2011). We propose two robust methods for estimating the GLMFM model, which accommodate the contaminations on all clustering levels, the standard-two-stage approach (Chervoneva et al., 2011, co-authored) and a robust joint estimation. Our research was motivated by the data obtained from the tendon fibril experiment reported in Zhang et al. (2006). Our statistical methodology is quite general and has potential application in a variety of relatively complex statistical modeling situations.Statistic
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
The use of temporally aggregated data on detecting a structural change of a time series process
A time series process can be influenced by an interruptive event which starts at a certain time point and so a structural break in either mean or variance may occur before and after the event time. However, the traditional statistical tests of two independent samples, such as the t-test for a mean difference and the F-test for a variance difference, cannot be directly used for detecting the structural breaks because it is almost certainly impossible that two random samples exist in a time series. As alternative methods, the likelihood ratio (LR) test for a mean change and the cumulative sum (CUSUM) of squares test for a variance change have been widely employed in literature. Another point of interest is temporal aggregation in a time series. Most published time series data are temporally aggregated from the original observations of a small time unit to the cumulative records of a large time unit. However, it is known that temporal aggregation has substantial effects on process properties because it transforms a high frequency nonaggregate process into a low frequency aggregate process. In this research, we investigate the effects of temporal aggregation on the LR test and the CUSUM test, through the ARIMA model transformation. First, we derive the proper transformation of ARIMA model orders and parameters when a time series is temporally aggregated. For the LR test for a mean change, its test statistic is associated with model parameters and errors. The parameters and errors in the statistic should be changed when an AR(p) process transforms upon the mth order temporal aggregation to an ARMA(P,Q) process. Using the property, we propose a modified LR test when a time series is aggregated. Through Monte Carlo simulations and empirical examples, we show that the aggregation leads the null distribution of the modified LR test statistic being shifted to the left. Hence, the test power increases as the order of aggregation increases. For the CUSUM test for a variance change, we show that two aggregation terms will appear in the test statistic and have negative effects on test results when an ARIMA(p,d,q) process transforms upon the mth order temporal aggregation to an ARIMA(P,d,Q) process. Then, we propose a modified CUSUM test to control the terms which are interpreted as the aggregation effects. Through Monte Carlo simulations and empirical examples, the modified CUSUM test shows better performance and higher test powers to detect a variance change in an aggregated time series than the original CUSUM test.Statistic
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
A Nonparametric Test for Deviation from Randomness
There are many existing tests used to determine if a series consists of a random sample. Often these tests have restrictive distributional assumptions, size distortions, or low power for key useful alternative situations. The interest of this dissertation lies in developing an alternative nonparametric test to determine whether a series consists of a random sample. The proposed test detects deviations from randomness, without a priori distributional assumption, when observations are not independent and identically distributed (i.i.d.), which is suitable for our motivating stock market index data. Departures from i.i.d. are tested by subdividing data into subintervals and then using a conditional probability measure within intervals as a binomial test. This nonparametric test is designed to detect deviations of neighboring observations from randomness when the data set consists of time series observations. Simulation results confirm correct test size for varied distributions and good power for detecting alternative cases. This test is compared to a number of other popular methods and shown to be a competitive alternative. Although the proposed test may be applicable to multiple areas, this dissertation is mostly interested in applications to stock market and regression data. The proposed test is effectively illustrated with the common three stock market index data sets using a newly created transformation, and shown to perform exceptionally well.Statistic
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
Statistics of Tumor Micro-environment
Introduction: Immune cells play a prominent role in keeping tumors suppressed, but how the distribution of these immune cells within a tumor’s microenvironment remains poorly understood. The long-term goal of this project is to study how statistical spatial distributions of different immune cells is associated with clinical outcome. The first objective is developing an algorithm for identifying different types of immune cells.
Methods: The data motivating this project includes spatial localization information (x-y coordinates) and expression levels of immune cell CD markers quantified by immunofluorescence immunohistochemistry (IF-IHC) in ~1,500 cases of invasive breast cancer. Using expression levels of CD markers in cancer cells (viewed as background noise), we compute upper nonparametric tolerance limits for CD expression in cancer cells. The stroma cells with CD expression above this tolerance limit are considered to be immune cells of the corresponding CD marker type.
Results: We have developed a Python program allowing us to quickly process a dataset of x-y coordinates of various cells that took up IHC stain, and creates a dataset of coordinates that are true immune cells. We have additionally analyzed multiple parameters for the development of a tolerance interval and concluded that a combination of 95%-confidence 99%-content allows for a minuscule chance of including the stroma cells that are not immune while maintaining enough data for analysis. Exploratory analysis of spatial point patterns of identified immune cell populations and their association with progression-free survival is in progress.
Discussion: We have developed an algorithm for the identification of different types of immune cells associated with type-specific CD markers quantified using IF-IHC. These tools enable further studies of spatial arrangement of immune cells in the tumor tissue and relating them to clinical outcome
- …
