1,721,001 research outputs found
A Falsifiability Characterization of Double Robustness Through Logical Operators
We address the characterization of problems in which a consistent estimator exists in a union of two models, also termed as a doubly robust estimator. Such estimators are important in missing information, including causal inference problems. Existing characterizations, based on the semiparametric theory of projections, have seen sufficient progress, but can still leave one’s understanding less than satisfied as to when and especially why such estimation works. We explore here a different, explanatory characterization – an exegesis based on logical operators. We show that double robustness exists if and only if we can produce consistent estimators for each contributing model based on an “AND” estimator, i. e., an estimator whose consistency generally needs both models to be correct. We show how this characterization explains double robustness through falsifiability
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Statistical methods and applications in medicine and public health
This thesis can be divided into two parts. The first (Chapter 1), concerns the development of a novel method for identifying highly-benefited patients in randomized clinical trials. The second (Chapters 2 and 3) aims to deepen our understanding of the spatial transmission of influenza in the United States. Here we review the relevant background for each topic, and discuss similarities in the statistical framework applied throughout.
Chapter 1 of this thesis concerns the statistical problem of identifying the covariate profiles of patients associated with a substantial, clinically-meaningful response to a particular treatment. “Personalized medicine”, the idea of therapeutic options tailored to specific individuals based on patient-level information (and potentially genetic data), is now a mainstream idea in clinical medicine. At present, randomized controlled trials remain the gold standard evidence-base by which clinicians make treatment decisions. Clinical trials, however, are generally performed with highly heterogeneous samples of patients, in part due to the logistics of patient recruitment, and in part because, often, a deep understanding of the biological mechanisms driving diseases processes remain elusive. Treatment effects, a comparison of responses between two treatment groups, may vary in magnitude among the heterogeneous population, and in extreme cases, may even be in the opposite direction for different group of patients. Methods capable of identifying such groups of patients reliably are of major clinical and statistical importance. A number of methods for identifying subgroups of patients with large estimated treatment effects (which then correspond to individualized treatment regimes) have been proposed. However, a central problem with existing approaches is that, in general, they ignore the size of the identified subgroup. The subgroup size is of crucial importance for clinicians wishing to apply such individualized treatment regimes, and for researchers seeking to validate the identified covariate profiles in a prospective fashion. In Chapter 1, we propose a novel method of characterizing highly benefited patients. In short, our method attempts to identify the largest group of patients whose non-parametrically estimated treatment effect is greater than an a priori defined clinical threshold. In order to efficiently search across numerous possible subgroups, we link the model-based (parametric) estimation of treatment effects directly with the goal of maximizing subgroup size. We apply this method to the Citalopram for Agitation in Alzheimer's Disease trial, and demonstrate that our new method can identify larger subgroups of patients than traditional methods, while maintaining the non-parametrically estimated treatment effects in the identified subgroups.
Chapters 2 and 3 concern the spatial transmission of influenza in the US. Influenza is an RNA virus that infects the epithelial cells lining the nose, throat and lungs of humans and other mammals. When infected, the human immune system is capable of mounting a strong response to eradicate infection. Unfortunately, at present, we lack a complete understanding of immunity to influenza. One important aspect of the human immune response to influenza infection is the production of antibodies capable of recognizing particular antigenic epitopes of surface influenza glycoproteins hemagglutinin (HA) and neuraminidase (NA), thereby acting to prevent infection. In response to this antibody production (and perhaps cell-mediated immunity), viruses with slightly altered antigenic epitopes experience high selection pressure, and thus the population of influenza viruses capable of evading immune response proliferate. The rapid evolution of influenza viruses in this fashion, termed antigenic drift, is responsible for the yearly circulation of novel influenza viruses in the Northern Hemisphere, and for the need to update the influenza vaccine on a yearly basis. Of note, T-cell recognition of influenza-infected cells remains poorly understood, but likely contributes substantially to long-term immunity against certain strains of influenza. Pandemic influenza occurs due to antigenic shift, when new HA and/or NA proteins occur in influenza viruses capable of infecting humans. These new influenza viruses typically emerge from an animal population and have protein structures that are sufficiently different from existing influenza strains infecting humans that most people do not have substantial baseline immunity to these novel viruses.
Understanding patterns in the spatial spread of novel influenza viruses is of great public health importance, and is tightly connect to the spatial and temporal dynamics of immunity against particular influenza strains. Vaccine development and roll-out in pandemic scenarios relies on a deep understanding of the spatial transmission of influenza and is a substantial area of research and public policy development. In Chapter 2, through a collaborative agreement with a private-sector data warehouse company, IMS Health, we analyze patterns of influenza spatial transmission across ~300 cities in the US. This is the most spatially-resolved study of influenza transmission to date. We develop and fit a stochastic model to the data to quantify patterns of spread at the city-level and factors associated with the risk of infection, including human mobility indices such as work commutes and domestic air traffic. We find that although the speed and pathways of spread varied across seasons, seven of eight epidemics studied were seeded in the Southern US. Each epidemic was associated with 1-5 early long-range transmission events, half of which sparked onward transmission. Parameter estimates from the stochastic model of influenza transmission indicate a sharp decay in transmission with the distance between infectious and susceptible cities, consistent with spread dominated by work commutes rather than air traffic. Two seasons associated with antigenic novelty had particularly localized mode of spread, suggesting that novel strains may spread more diffusely than previously anticipated. Finally, sensitivity analyses demonstrate that analyses of data at coarser geographic scales (e.g. state-level data) lead to biased inferences about the spatial dynamics of influenza spread.
In Chapter 3, to address the overwhelming lack of highly spatially-resolved data on influenza transmission, we develop a straightforward framework for reconstructing the spatial spread of an epidemic using multiple data sources. We focus on the 2003/2004 A/H3N2 epidemic, which was marked by a substantial antigenic drift event, using data from outpatient influenza-like-illness visits and influenza-coded hospitalizations. We were able to reconstruct the spatial spread across approximately 600 cities, confirming the localized nature of spread observed during this season across two separate data collecting systems. In this chapter, we also consider different network-based connectivity metrics between counties, in an attempt to explain the county-level transmission of influenza. We find that a network-distance metric based on work commutes does not outperform a model based solely on geographic distance.
While disparate in their objectives, the two problems under consideration in this work are related with respect to their statistical demands in the following ways: first, both problems involve the appropriate formulation of a specific objective function to be optimized that directly captures the essence of the scientific problem at hand; second, optimization of these objective functions is not always straightforward and requires the appropriate translation of estimated parameters into a scientifically meaningful and coherent hypothesis; and finally, both problems involve the analysis of relatively limited amounts of data and a balance must be struck between careful, rigorous analysis and the ability to best utilize the data to answer the scientific questions of interest
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
koamabayili/VECTRON-author-checklist: VECTRON author checklist
We have done our best to complete the author checklist relating to the use of animals in the hut study. Note that the objective for the hut study was to evaluate the IRS treatment applications for residual efficacy against Anopheles mosquitoes, including the local An. coluzzii mosquito population. Cows were only used to attract mosquitoes into the huts and no tests were carried out directly on the cows. The author checklist is intended for use with studies where experiments are carried out on animals, which is why we have had such difficulty in completing this for the hut study, as many of the questions do not relate to how the cows were used
Author-wise bibliometric analysis based on entropy.
Author-wise bibliometric analysis based on entropy.</p
- …
