1,721,032 research outputs found

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Projection Based Models for High Dimensional Data

    Get PDF
    In recent years, many machine learning applications have arisen which deal with the problem of finding patterns in high dimensional data. Principal component analysis (PCA) has become ubiquitous in this setting. PCA performs dimensionality reduction by estimating latent factors which minimise the reconstruction error between the original data and its low-dimensional projection. We initially consider a situation where influential observations exist within the dataset which have a large, adverse affect on the estimated PCA model. We propose a measure of “predictive influence” to detect these points based on the contribution of each point to the leave-one-out reconstruction error of the model using an analytic PRedicted REsidual Sum of Squares (PRESS) statistic. We then develop a robust alternative to PCA to deal with the presence of influential observations and outliers which minimizes the predictive reconstruction error. In some applications there may be unobserved clusters in the data, for which fitting PCA models to subsets of the data would provide a better fit. This is known as the subspace clustering problem. We develop a novel algorithm for subspace clustering which iteratively fits PCA models to subsets of the data and assigns observations to clusters based on their predictive influence on the reconstruction error. We study the convergence of the algorithm and compare its performance to a number of subspace clustering methods on simulated data and in real applications from computer vision involving clustering object trajectories in video sequences and images of faces. We extend our predictive clustering framework to a setting where two high-dimensional views of data have been obtained. Often, only either clustering or predictive modelling is performed between the views. Instead, we aim to recover clusters which are maximally predictive between the views. In this setting two block partial least squares (TB-PLS) is a useful model. TB-PLS performs dimensionality reduction in both views by estimating latent factors that are highly predictive. We fit TB-PLS models to subsets of data and assign points to clusters based on their predictive influence under each model which is evaluated using a PRESS statistic. We compare our method to state of the art algorithms in real applications in webpage and document clustering and find that our approach to predictive clustering yields superior results. Finally, we propose a method for dynamically tracking multivariate data streams based on PLS. Our method learns a linear regression function from multivariate input and output streaming data in an incremental fashion while also performing dimensionality reduction and variable selection. Moreover, the recursive regression model is able to adapt to sudden changes in the data generating mechanism and also identifies the number of latent factors. We apply our method to the enhanced index tracking problem in computational finance

    Advanced bayesian modelling for the analysis of outbreaks and shifting epidemiological dynamics

    No full text
    The emergence, spread, and establishment of an infectious disease within a population brings about a plethora of challenges for public health organisations, whose aim is to reduce disease burden while having access to limited information. In this thesis, we develop statistical models and analyses to support public health response, addressing uncertainties that are inherent to epidemics. The work is divided into two parts, focusing on the last century’s biggest pandemics. In the first part, we focus on the emergence of novel pathogens and variants of concern, with applications to SARS-CoV-2. Firstly, we develop an adjustment to early reproduction number estimates, when generations of infections have not been reported. Our adjustment is shown to reduce early biases in simulation studies. Secondly, we develop a multi-strain Bayesian model to describe fluctuations in hospital fatality rates in Brazil following the emergence of the Gamma variant. By synthesising data from separate sources, we estimate the proportion of patients with either variant in hospitals across Brazil, and quantify the impact of healthcare pressures, variant, and location effects. In the second part of the thesis, we describe changes in transmission dynamics and burden of HIV, using data from the Rakai Community Cohort Study. In the first project, we develop a phylogenetic pipeline to estimate HIV time since infection from viral sequences, and develop statistical models to refine estimates by incorporating testing histories and known transmission network. By dating transmissions, we are able to describe changes in transmission patterns, highlighting shifts in the age-profile of the sources. Finally, we provide detailed descriptions of shifts in the age and gender compositions of the burden of HIV and viraemia in Uganda. We obtain estimates at the age level by developing non-parametric models sharing information across age groups. We conclude by proposing novel metrics to inform prevention strategies.Open Acces

    Bayesian point processes models with applications in the COVID-19 pandemic

    Get PDF
    A point process is a set of points randomly located in a space, such as time or abstract spaces. Point process models have found numerous applications in epidemiology, ecology, geophysics, social networks and many other areas. The Poisson process is the most widely known point process. Poisson intensity estimation is a vital task in various applications including medical imaging, astrophysics and network traffic analysis. A Bayesian Additive Regression Trees (BART) scheme for estimating the intensity of inhomogeneous Poisson processes is introduced. The new approach enables full posterior inference of the intensity in a non-parametric regression setting. The performance of the novel scheme is demonstrated through simulation studies on synthetic and real datasets up to five dimensions, and the new scheme is compared with alternative approaches. A drawback of the proposed algorithm is its axis-alignment nature. We discuss this problem and suggest alternative approaches to remedy the drawback. The novel coronavirus disease (COVID-19) has been declared a Global Health Emergency of International Concern with over 557 million cases and 6.36 million deaths as of 3 August 2022 according to the World Health Organization. Understanding the spread of COVID-19 has been the subject of numerous studies, highlighting the significance of reliable epidemic models. We introduce a novel epidemic model using a latent Hawkes process with temporal covariates for modelling the infections. Unlike other Hawkes models, we model the reported cases via a probability distribution driven by the underlying Hawkes process. Modelling the infections via a Hawkes process allows us to estimate by whom an infected individual was infected. We propose a Kernel Density Particle Filter (KDPF) for inference of both latent cases and reproduction number and for predicting new cases in the near future. The computational effort is proportional to the number of infections making it possible to use particle filter-type algorithms, such as the KDPF. We demonstrate the performance of the proposed algorithm on synthetic data sets and COVID-19 reported cases in various local authorities in the UK, and benchmark our model to alternative approaches. We extend the unstructured homogeneously mixing epidemic model considering a finite population stratified by age bands. We model the actual unobserved infections using a latent marked Hawkes process and the reported aggregated infections as random quantities driven by the underlying Hawkes process. We apply a Kernel Density Particle Filter (KDPF) to infer the marked counting process, the instantaneous reproduction number for each age group and forecast the epidemic’s future trajectory in the near future. We demonstrate the performance of the proposed inference algorithm on synthetic data sets and COVID-19 reported cases in various local authorities in the UK. Taking into account the individual heterogeneity in age provides a real-time measurement of interventions and behavioural changes.Open Acces

    Spatio-temporal models of west Pacific tropical cyclones

    Get PDF
    Tropical cyclones are devastating destructive forces of nature that can cause loss of life and catastrophic damage. Understanding the factors that affect tropical cyclone genesis and track and having accurate models to simulate them is of great interest to the financial and meteorological industries, as well as governmental agencies. This thesis develops the area of statistical-dynamic models for tropical cyclone genesis and track. Firstly, we describe the potential predictors of genesis and track independently, marginally against genesis and track respectively and jointly for genesis. This provides insight into the ranges of the predictors that are required for genesis. There are a variety of approaches to modelling genesis and track in the literature, however, due to a lack of standard quantitative measures, comparatively assessing the performance of these models is difficult. In this thesis, we develop a range of parametric and non-parametric models as well as spatial and spatio-temporal models for both genesis and track. For one of the approaches, we use a Bayesian conditional logistic regression model for genesis which has not been previously used before and performs well. The Bayesian approach results in a significant reduction on the bias between the observed and simulated mean annual frequency out-of-sample. As an example, there is a 94% reduction in the bias between a Bayesian model and its frequentist counterpart. A Bayesian approach to modelling the track of a tropical cyclone has not been introduced in the literature and we develop a Bayesian model for the track that outperforms other models particularly out-of-sample. Using the main characteristics of tropical cyclone genesis and track, we develop a set of metrics in which we compare our models to each other and to observation. Such specific characteristics are the mean annual frequency, the spatial distribution and the average track. This approach provides transparency and insight into the various modelling approaches which are fragmented in the literature. It also allows models to be quantitatively assessed to expose their weaknesses and strengths respectively. For example, we observe a reduction in the bias of landfall in South East China by more than 76% using the Bayesian track model developed in this thesis compared a resampling model in the literature. We apply our models to a likely sea surface temperature scenario that is expected to occur under climate change. We assess the potential impact on the change of genesis and landfall. We also propose an alternative two-layer modelling approach to overcome a key shortcoming in current genesis models, that is obtaining the year-to-year fluctuations in tropical cyclone frequency. New potential predictors of the frequency of events in the upcoming tropical cyclone season are defined, with correlations in the region of 0.6 and a model with out-of-sample correlation of 0.67 with the annual frequency. We describe the larger-scale phenomenon including the El Nino Modoki pattern that influences the frequency of genesis.Open Acces

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Latent factor representations of dynamic networks with applications in cyber-security

    Get PDF
    Dynamic networks frequently arise in nature, representing, for example, neuronal connectivity in the brain, internet connections between computers, and human interactions within social networks. Motivated by specific characteristics of computer networks, this thesis gives novel contributions to three important branches of statistical analysis of dynamic graphs: modelling of connectivity events, graph clustering, and link prediction. Nodes and edges within the network will be assumed to have latent, unobserved characteristics, estimated using statistical latent factor models. The first part of this work proposes edge-based models for individual connection events. A model for dynamic network evolution based on the Pitman-Yor process is proposed, which is used for network-wide anomaly detection. Furthermore, a model for classification of periodic arrivals in event time data is proposed, to separate human and automated activity within individual edges. Correct classification of these two types of activity is fundamental for effective intrusion detection. In its second part, this work proposes novel methodologies for graph clustering. Finding nodes with similar connectivity patterns is important for reliable network monitoring. A model for simultaneous estimation of the latent dimension and number of communities in spectral clustering is proposed, under a generalised random dot product graph interpretation of the stochastic blockmodel. The model is then extended to allow for heterogeneous within-community degree distributions, proposing a novel spectral clustering algorithm under the degree corrected stochastic blockmodel. Finally, in the third part of this thesis, link prediction methods are discussed. The ability to correctly associate anomaly scores with the connections in a network is crucial for the cyber-defence of an organisation. In this work, it is demonstrated that random dot product graphs are a powerful and scalable tool for link prediction in large graph. Furthermore, an extension of the popular Poisson matrix factorisation model is proposed, which correctly models binary matrices and allows the effects of nodal covariates to be estimated. Scalable inferential techniques are also discussed. Overall, this work gives contributions towards a unified model for cyber-security applications, in a Bayesian framework, with the ultimate aim to dynamically predict the connectivity of IP addresses, or users and hosts, in a computer network.Open Acces

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis

    Monitoring in survival analysis and rare event simulation

    Get PDF
    Monte Carlo methods are a fundamental tool in many areas of statistics. In this thesis, we will examine these methods, especially for rare event simulation. We are mainly interested in the computation of multivariate normal probabilities and in constructing hitting thresholds in survival analysis models. Firstly, we develop an algorithm for computing high dimensional normal probabilities. These kinds of probabilities are a fundamental tool in many statistical applications. The new algorithm exploits the diagonalisation of the covariance matrix and uses various variance reduction techniques. Its performance is evaluated via a simulation study. The new method is designed for computing small exceedance probabilities. Secondly, we introduce a new omnibus cumulative sum chart for monitoring in survival analysis models. By omnibus we mean that it is able to detect any change. This chart exploits the absolute differences between the Kaplan-Meier estimator and the in-control distribution over specific time intervals. A simulation study is presented that evaluates the performance of our proposed chart and compares it to existing methods. Thirdly, we apply the method of adaptive multilevel splitting for the estimation of hitting probabilities and hitting thresholds for the survival analysis cumulative sum charts. Simulation results are presented evaluating the benefits of adaptive multilevel splitting. Finally, we extend the idea of adaptive multilevel splitting by estimating not just a hitting probability, but the whole distribution function up to a certain point. A theoretical result is proved that is used to construct confidence bands for the distribution function conditioned on lying in a closed interval
    corecore