1,721,032 research outputs found
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Projection Based Models for High Dimensional Data
In recent years, many machine learning applications have arisen which deal with the
problem of finding patterns in high dimensional data. Principal component analysis
(PCA) has become ubiquitous in this setting. PCA performs dimensionality reduction
by estimating latent factors which minimise the reconstruction error between
the original data and its low-dimensional projection. We initially consider a situation
where influential observations exist within the dataset which have a large,
adverse affect on the estimated PCA model. We propose a measure of “predictive
influence” to detect these points based on the contribution of each point to the
leave-one-out reconstruction error of the model using an analytic PRedicted REsidual
Sum of Squares (PRESS) statistic. We then develop a robust alternative to PCA
to deal with the presence of influential observations and outliers which minimizes
the predictive reconstruction error.
In some applications there may be unobserved clusters in the data, for which
fitting PCA models to subsets of the data would provide a better fit. This is known
as the subspace clustering problem. We develop a novel algorithm for subspace
clustering which iteratively fits PCA models to subsets of the data and assigns observations
to clusters based on their predictive influence on the reconstruction error.
We study the convergence of the algorithm and compare its performance to a number
of subspace clustering methods on simulated data and in real applications from
computer vision involving clustering object trajectories in video sequences and images
of faces.
We extend our predictive clustering framework to a setting where two high-dimensional
views of data have been obtained. Often, only either clustering or predictive modelling is performed between the views. Instead, we aim to recover
clusters which are maximally predictive between the views. In this setting two block
partial least squares (TB-PLS) is a useful model. TB-PLS performs dimensionality
reduction in both views by estimating latent factors that are highly predictive. We
fit TB-PLS models to subsets of data and assign points to clusters based on their
predictive influence under each model which is evaluated using a PRESS statistic.
We compare our method to state of the art algorithms in real applications in webpage
and document clustering and find that our approach to predictive clustering
yields superior results.
Finally, we propose a method for dynamically tracking multivariate data streams
based on PLS. Our method learns a linear regression function from multivariate
input and output streaming data in an incremental fashion while also performing
dimensionality reduction and variable selection. Moreover, the recursive regression
model is able to adapt to sudden changes in the data generating mechanism and also
identifies the number of latent factors. We apply our method to the enhanced index
tracking problem in computational finance
Advanced bayesian modelling for the analysis of outbreaks and shifting epidemiological dynamics
The emergence, spread, and establishment of an infectious disease within a population brings about a plethora of challenges for public health organisations, whose aim is to reduce disease burden while having access to limited information. In this thesis, we develop statistical models and analyses to support public health response, addressing uncertainties that are inherent to epidemics. The work is divided into two parts, focusing on the last century’s biggest pandemics.
In the first part, we focus on the emergence of novel pathogens and variants of concern, with applications to SARS-CoV-2. Firstly, we develop an adjustment to early reproduction number estimates, when generations of infections have not been reported. Our adjustment is shown to reduce early biases in simulation studies. Secondly, we develop a multi-strain Bayesian model to describe fluctuations in hospital fatality rates in Brazil following the emergence of the Gamma variant. By synthesising data from separate sources, we estimate the proportion of patients with either variant in hospitals across Brazil, and quantify the impact of healthcare pressures, variant, and location effects.
In the second part of the thesis, we describe changes in transmission dynamics and burden of HIV, using data from the Rakai Community Cohort Study. In the first project, we develop a phylogenetic pipeline to estimate HIV time since infection from viral sequences, and develop statistical models to refine estimates by incorporating testing histories and known transmission network. By dating transmissions, we are able to describe changes in transmission patterns, highlighting shifts in the age-profile of the sources. Finally, we provide detailed descriptions of shifts in the age and gender compositions of the burden of HIV and viraemia in Uganda. We obtain estimates at the age level by developing non-parametric models sharing information across age groups.
We conclude by proposing novel metrics to inform prevention strategies.Open Acces
Bayesian point processes models with applications in the COVID-19 pandemic
A point process is a set of points randomly located in a space, such as time or abstract
spaces. Point process models have found numerous applications in epidemiology,
ecology, geophysics, social networks and many other areas.
The Poisson process is the most widely known point process. Poisson intensity
estimation is a vital task in various applications including medical imaging,
astrophysics and network traffic analysis. A Bayesian Additive Regression Trees
(BART) scheme for estimating the intensity of inhomogeneous Poisson processes
is introduced. The new approach enables full posterior inference of the intensity
in a non-parametric regression setting. The performance of the novel scheme is
demonstrated through simulation studies on synthetic and real datasets up to five
dimensions, and the new scheme is compared with alternative approaches. A drawback
of the proposed algorithm is its axis-alignment nature. We discuss this problem
and suggest alternative approaches to remedy the drawback.
The novel coronavirus disease (COVID-19) has been declared a Global Health
Emergency of International Concern with over 557 million cases and 6.36 million
deaths as of 3 August 2022 according to the World Health Organization. Understanding
the spread of COVID-19 has been the subject of numerous studies, highlighting
the significance of reliable epidemic models. We introduce a novel epidemic
model using a latent Hawkes process with temporal covariates for modelling the infections.
Unlike other Hawkes models, we model the reported cases via a probability
distribution driven by the underlying Hawkes process. Modelling the infections via
a Hawkes process allows us to estimate by whom an infected individual was infected.
We propose a Kernel Density Particle Filter (KDPF) for inference of both latent
cases and reproduction number and for predicting new cases in the near future. The
computational effort is proportional to the number of infections making it possible
to use particle filter-type algorithms, such as the KDPF. We demonstrate the performance
of the proposed algorithm on synthetic data sets and COVID-19 reported
cases in various local authorities in the UK, and benchmark our model to alternative
approaches.
We extend the unstructured homogeneously mixing epidemic model considering
a finite population stratified by age bands. We model the actual unobserved infections
using a latent marked Hawkes process and the reported aggregated infections
as random quantities driven by the underlying Hawkes process. We apply a Kernel
Density Particle Filter (KDPF) to infer the marked counting process, the instantaneous
reproduction number for each age group and forecast the epidemic’s future
trajectory in the near future. We demonstrate the performance of the proposed
inference algorithm on synthetic data sets and COVID-19 reported cases in various
local authorities in the UK. Taking into account the individual heterogeneity in age
provides a real-time measurement of interventions and behavioural changes.Open Acces
Spatio-temporal models of west Pacific tropical cyclones
Tropical cyclones are devastating destructive forces of nature that can cause loss of life and
catastrophic damage. Understanding the factors that affect tropical cyclone genesis and track
and having accurate models to simulate them is of great interest to the financial and meteorological industries, as well as governmental agencies.
This thesis develops the area of statistical-dynamic models for tropical cyclone genesis and track.
Firstly, we describe the potential predictors of genesis and track independently, marginally
against genesis and track respectively and jointly for genesis. This provides insight into the
ranges of the predictors that are required for genesis.
There are a variety of approaches to modelling genesis and track in the literature, however,
due to a lack of standard quantitative measures, comparatively assessing the performance of
these models is difficult. In this thesis, we develop a range of parametric and non-parametric
models as well as spatial and spatio-temporal models for both genesis and track. For one of the
approaches, we use a Bayesian conditional logistic regression model for genesis which has not
been previously used before and performs well. The Bayesian approach results in a significant
reduction on the bias between the observed and simulated mean annual frequency out-of-sample.
As an example, there is a 94% reduction in the bias between a Bayesian model and its frequentist
counterpart.
A Bayesian approach to modelling the track of a tropical cyclone has not been introduced in
the literature and we develop a Bayesian model for the track that outperforms other models
particularly out-of-sample. Using the main characteristics of tropical cyclone genesis and track,
we develop a set of metrics in which we compare our models to each other and to observation.
Such specific characteristics are the mean annual frequency, the spatial distribution and the
average track. This approach provides transparency and insight into the various modelling
approaches which are fragmented in the literature. It also allows models to be quantitatively
assessed to expose their weaknesses and strengths respectively. For example, we observe a
reduction in the bias of landfall in South East China by more than 76% using the Bayesian
track model developed in this thesis compared a resampling model in the literature.
We apply our models to a likely sea surface temperature scenario that is expected to occur under
climate change. We assess the potential impact on the change of genesis and landfall. We also
propose an alternative two-layer modelling approach to overcome a key shortcoming in current
genesis models, that is obtaining the year-to-year fluctuations in tropical cyclone frequency.
New potential predictors of the frequency of events in the upcoming tropical cyclone season
are defined, with correlations in the region of 0.6 and a model with out-of-sample correlation
of 0.67 with the annual frequency. We describe the larger-scale phenomenon including the El
Nino Modoki pattern that influences the frequency of genesis.Open Acces
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Latent factor representations of dynamic networks with applications in cyber-security
Dynamic networks frequently arise in nature, representing, for example, neuronal connectivity in the brain, internet connections between computers, and human interactions within social networks. Motivated by specific characteristics of computer networks, this thesis gives novel contributions to three important branches of statistical analysis of dynamic graphs: modelling of connectivity events, graph clustering, and link prediction. Nodes and edges within the network will be assumed to have latent, unobserved characteristics, estimated using statistical latent factor models.
The first part of this work proposes edge-based models for individual connection events. A model for dynamic network evolution based on the Pitman-Yor process is proposed, which is used for network-wide anomaly detection. Furthermore, a model for classification of periodic arrivals in event time data is proposed, to separate human and automated activity within individual edges. Correct classification of these two types of activity is fundamental for effective intrusion detection.
In its second part, this work proposes novel methodologies for graph clustering. Finding nodes with similar connectivity patterns is important for reliable network monitoring. A model for simultaneous estimation of the latent dimension and number of communities in spectral clustering is proposed, under a generalised random dot product graph interpretation of the stochastic blockmodel. The model is then extended to allow for heterogeneous within-community degree distributions, proposing a novel spectral clustering algorithm under the degree corrected stochastic blockmodel.
Finally, in the third part of this thesis, link prediction methods are discussed. The ability to correctly associate anomaly scores with the connections in a network is crucial for the cyber-defence of an organisation. In this work, it is demonstrated that random dot product graphs are a powerful and scalable tool for link prediction in large graph. Furthermore, an extension of the popular Poisson matrix factorisation model is proposed, which correctly models binary matrices and allows the effects of nodal covariates to be estimated. Scalable inferential techniques are also discussed.
Overall, this work gives contributions towards a unified model for cyber-security applications, in a Bayesian framework, with the ultimate aim to dynamically predict the connectivity of IP addresses, or users and hosts, in a computer network.Open Acces
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Monitoring in survival analysis and rare event simulation
Monte Carlo methods are a fundamental tool in many areas of statistics. In this thesis,
we will examine these methods, especially for rare event simulation. We are mainly
interested in the computation of multivariate normal probabilities and in constructing
hitting thresholds in survival analysis models.
Firstly, we develop an algorithm for computing high dimensional normal probabilities.
These kinds of probabilities are a fundamental tool in many statistical
applications. The new algorithm exploits the diagonalisation of the covariance matrix
and uses various variance reduction techniques. Its performance is evaluated via
a simulation study. The new method is designed for computing small exceedance
probabilities.
Secondly, we introduce a new omnibus cumulative sum chart for monitoring in
survival analysis models. By omnibus we mean that it is able to detect any change.
This chart exploits the absolute differences between the Kaplan-Meier estimator and
the in-control distribution over specific time intervals. A simulation study is presented
that evaluates the performance of our proposed chart and compares it to existing
methods.
Thirdly, we apply the method of adaptive multilevel splitting for the estimation of
hitting probabilities and hitting thresholds for the survival analysis cumulative sum
charts. Simulation results are presented evaluating the benefits of adaptive multilevel
splitting.
Finally, we extend the idea of adaptive multilevel splitting by estimating not
just a hitting probability, but the whole distribution function up to a certain point.
A theoretical result is proved that is used to construct confidence bands for the distribution function conditioned on lying in a closed interval
- …
