1,721,134 research outputs found
A Hidden Markov Model Approach to Testing Multiple Hypotheses on a Tree-Transformed Gene Ontology Graph
Gene category testing problems involve testing hundreds of null hypotheses that correspond to nodes in a directed acyclic graph. The logical relationships among the nodes in the graph imply that only some configurations of true and false null hypotheses are possible and that a test for a given node should depend on data from neighboring nodes. We developed a method based on a hidden Markov model that takes the whole graph into account and provides coherent decisions in this structured multiple hypothesis testing problem. The method is illustrated by testing Gene Ontology terms for evidence of differential expression.This is a manuscript of an article published as Liang, Kun, and Dan Nettleton. "A hidden Markov model approach to testing multiple hypotheses on a tree-transformed gene ontology graph." Journal of the American Statistical Association 105, no. 492 (2010): 1444-1454. doi: 10.1198/jasa.2010.tm10195. Posted with permission.</p
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
rmRNAseq: differential expression analysis for repeated-measures RNA-seq data
Motivation: With the reduction in price of next generation sequencing technologies, gene expression profiling using RNA-seq has increased the scope of sequencing experiments to include more complex designs, such as designs involving repeated measures. In such designs, RNA samples are extracted from each experimental unit at multiple time points. The read counts that result from RNA sequencing of the samples extracted from the same experimental unit tend to be temporally correlated. Although there are many methods for RNA-seq differential expression analysis, existing methods do not properly account for within-unit correlations that arise in repeated-measures designs.
Results: We address this shortcoming by using normalized log-transformed counts and associated precision weights in a general linear model pipeline with continuous autoregressive structure to account for the correlation among observations within each experimental unit. We then utilize parametric bootstrap to conduct differential expression inference. Simulation studies show the advantages of our method over alternatives that do not account for the correlation among observations within experimental units.
Availability:We provide anRpackage rmRNAseq implementing our proposed method (function TC_CAR1) at https://cran.r-project.org/web/packages/rmRNAseq/index.html. Reproducible R codes for data analysis and simulation are available at https://github.com/ntyet/rmRNAseq/ tree/master/simulation.This is a manuscript of an article published as Nguyen, Yet, and Dan Nettleton. "rmRNAseq: Differential Expression Analysis for Repeated-measures RNA-seq Data." Bioinformatics (2020). doi: 10.1093/bioinformatics/btaa525. Posted with permission.</p
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
New statistical methods in bioinformatics: for the analysis of quantitative trait loci (QTL), microarrays, and eQTLs
This thesis focuses on new statistical methods in the area of bioinformatics which uses computers and statistics to solve biological problems. The first study discusses a method for detecting a quantitative trait locus (QTL) when the trait of interest has a zero-inflated Poisson (ZIP) distribution. Though existing methods based on normality may be reasonably applied to some ZIP distributions, the characteristics of other ZIP distributions make such an application inappropriate. We compare our method to an existing non-parametric approach, and we illustrate our method using QTL data collected on two ecotypes of the Arabidopsis thaliana plant where the trait of interest is shoot count;The second study discusses a method to detect differentially expressed genes in an unreplicated multiple-treatment microarray timecourse experiment. In a two-sample setting, differential expression is well defined as non-equal means, but in the present setting, there are numerous expression patterns that may qualify as differential expression, and that may be of interest to the researcher. This method provides the researcher with a list of significant genes, an associated false discovery rate for that list, and a 'best model' choice for every gene. The model choice component is relevant because the alternative hypothesis of differential expression does not dictate one specific alternative expression pattern. In fact, in this type of experiment, there are many possible expression patterns of interest to the researcher. Using simulations, we provide information on the specificity and sensitivity of detection under a variety of true expression patterns using receiver operating characteristic curves. The method is illustrated using an Arabidopsis thaliana microarray experiment with five time points and three treatment groups;The third study discusses a new type of analysis, called eQTL analysis. This analysis brings together the methods of microarray and QTL analyses in order to detect locations on the genome that control gene expression. These controlling loci are called expression QTL, or eQTL. Locating eQTL can help researchers uncover complex networks in biological systems. The method is illustrated using an Arabidopsis thaliana eQTL experiment with 22,787 genes and 288 markers.</p
Model estimation, identification and inference for next-generation functional data and spatial data
This dissertation is composed of three research projects focused on model estimation, identification, and inference for next-generation functional data and spatial data.
The first project deals with data that are collected on a count or binary response with spatial covariate information. In this project, we introduce a new class of generalized geoadditive models (GGAMs) for spatial data distributed over complex domains. Through a link function, the proposed GGAM assumes that the mean of the discrete response variable depends on additive univariate functions of explanatory variables and a bivariate function to adjust for the spatial effect. We propose a two-stage approach for estimating and making inferences of the components in the GGAM. In the first stage, the univariate components and the geographical component in the model are approximated via univariate polynomial splines and bivariate penalized splines over triangulation, respectively. In the second stage, local polynomial smoothing is applied to the cleaned univariate data to average out the variation of the first-stage estimators. We investigate the consistency of the proposed estimators and the asymptotic normality of the univariate components. We also establish the simultaneous confidence band for each of the univariate components. The performance of the proposed method is evaluated by two simulation studies and the crash counts data in the Tampa-St. Petersburg urbanized area in Florida.
In the second project, motivated by recent work of analyzing data in the biomedical imaging studies, we consider a class of image-on-scalar regression models for imaging responses and scalar predictors. We propose to use flexible multivariate splines over triangulations to handle the irregular domain of the objects of interest on the images and other characteristics of images. The proposed estimators of the coefficient functions are proved to be root- consistent and asymptotically normal under some regularity conditions. We also provide a consistent and computationally efficient estimator of the covariance function. Asymptotic pointwise confidence intervals (PCIs) and data-driven simultaneous confidence corridors (SCCs) for the coefficient functions are constructed. A highly efficient and scalable estimation algorithm is developed. Monte Carlo simulation studies are conducted to examine the finite-sample performance of the proposed method. The proposed method is applied to the spatially normalized Positron Emission Tomography (PET) data of Alzheimer's Disease Neuroimaging Initiative (ADNI).
In the third project, we propose a heterogeneous functional linear model to simultaneously estimate multiple coefficient functions and identify groups, such that coefficient functions are identical within groups and distinct across groups. By borrowing information from relevant subgroups, our method enhances estimation efficiency while preserving heterogeneity. We use an adaptive fused lasso penalty to shrink subgroup coefficients to shared common values within each group. We also establish the theoretical properties of our adaptive fused lasso estimators. To enhance the computation efficiency and incorporate neighborhood information, we propose to use a graph-constrained adaptive lasso. A highly efficient and scalable estimation algorithm is developed. Monte Carlo simulation studies are conducted to examine the finite-sample performance of the proposed method. The proposed method is applied to a dataset of hybrid maize grain yields from the Genomes to Fields consortium.</p
- …
