1,721,137 research outputs found
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Multiple testing with test statistics following heavy-tailed distributions
In multiple testing problems where the components come from a mixture model of noise and true effect, we seek to first test for the existence of the non-zero components, and then identify the true alternatives under a fixed significance level . Two parameters, namely the fraction of the non-null components and the size of the effects , characterise the two-point mixture model under the global alternative. When the number of hypotheses goes to infinity, we are interested in an asymptotic framework where the fraction of the non-null components is vanishing, and the true effects need to be sizable to be detected. Donoho and Jin give an explicit form of the asymptotic detectable boundary based on the Gaussian mixture model under the classic calibration of the parameters of the mixture model. We prove the analogous results for the Cauchy mixture distribution as an example heavy-tailed case. This requires a different formulation of the parameters, which reflects the added difficulties.
We also propose a multiple testing procedure based on a filtering approach that can discover the true alternatives.
Benjamini and Hochberg (BH) compare the observed -values to a linear threshold curve and reject the null hypotheses from the minimum up to the last up-crossing, and prove the false discovery rate (FDR) is controlled.
However, there is an intrinsic difference in heavy-tailed settings. Were we to use the BH procedure we would get a highly variable positive false discovery rate (pFDR). In our study we analyse the distribution of the -values and devise a new multiple testing procedure to combine the usual case and the heavy-tailed case based on the empirical properties of the -values. The filtering approach is designed to eliminate most -values that are more likely to be uniform, while preserving most of the true alternatives. Based on the filtered -values, we estimate the mode and define the rejection region such that the most informative -values are included. The length is chosen by controlling the data-dependent estimation of FDR at a desired level.STA
Presentation and study of robustness for several methods to classify individuals based on their gene expressions
Several studies have shown that it is possible to detect cancer tissues based on gene expressions using methods of machine learning. The main problem with classifying gene expression data is to obtain accurate rules that are easy to interpret and provide indications for follow up studies. Indeed high accuracy is hard to achieve due to the small number of observations and the large amount of genes in the human genome. Some methods of machine learning are based on an important quantity of genes, which lead to decision rules that are usually difficult to interpret. These methods were tested on different samples and their results were compared. Most of them provided good results with a high accuracy. Among these methods for gene classification one distanced itself from the others by producing transparents results which were readily interpretable and were very useful for follow up studies. It highlighted pair of genes that were the most efficient to classify individuals with respect to their gene expressions. This is the so called Top Scoring Pair (TSP) classifier. This method achieves prediction rates that are as high as those of the other methods. In contrast to other classifiers which use considerably more genes and more complicated procedures, the TSP has an easy and quick implementation and involves very few genes, namely only two. This provides very easy rules that are accurate and transparent. Finally, the TSP is paramter-free, which avoids overfitting and inflation of the estimation of the prediction rate.SB-SM
- …
