1,720,961 research outputs found
Generative machine learning for multivariate angular simulation
With the recent development of new geometric and angular-radial frameworks for multivariate
extremes, reliably simulating from angular variables in moderate-to-high dimensions is of increasing
importance. Empirical approaches have the benefit of simplicity, and work reasonably well in low
dimensions, but as the number of variables increases, they can lack the required flexibility and scalability. Classical parametric models for angular variables, such as the von Mises–Fisher distribution
(vMF), provide an alternative. Exploiting finite mixtures of vMF distributions increases their flexibility, but there are cases where, without letting the number of mixture components grow considerably,
a mixture model with a fixed number of components is not sufficient to capture the intricate features that can arise in data. Owing to their flexibility, generative deep learning methods are able to
capture complex data structures; they therefore have the potential to be useful in the simulation of
multivariate angular variables. In this paper, we introduce a range of deep learning approaches for
this task, including generative adversarial networks, normalizing flows and flow matching. We assess
their performance via a range of metrics, and make comparisons to the more classical approach of
using a finite mixture of vMF distributions. The methods are also applied to a metocean data set,
with diagnostics indicating strong performance, demonstrating the applicability of such techniques to
real-world, complex data structures
New estimation methods for extremal bivariate return curves
In the multivariate setting, estimates of extremal risk measures are important in many contexts, such as environmental planning and structural engineering. In this paper, we propose new estimation methods for extremal bivariate return curves, a risk measure that is the natural bivariate extension to a return level. Unlike several existing techniques, our estimates are based on bivariate extreme value models that can capture both key forms of extremal dependence. We devise tools for validating return curve estimates, as well as representing their uncertainty, and compare a selection of curve estimation techniques through simulation studies. We apply the methodology to two met-ocean data sets, with diagnostics indicating generally good performance
Novel methodology for the estimation of extremal bivariate return curves
The aim of this thesis is to develop novel methodology for estimating an extreme risk measure, known as a return curve, for bivariate random vectors. In doing this, we also aim to develop novel techniques for estimating the extremal dependence structure of bivariate random vectors, and then to compare these techniques to existing methodology. In many practical applications, understanding joint extreme risks from pairs of random variables is crucial for ensuring robust risk analyses and informed decision making. Return curves provide a means of both quantifying and visualising such risks. However, estimation of these curves has not been well studied, particularly in the case when data exhibits asymptotic independence. Furthermore, under the influence of climate change, the joint extremal behaviour for pairs of environmental variables is likely to change; techniques are required to ensure such trends are captured when estimating return curves. We first propose a range of novel estimation methods for return curves; unlike several existing techniques, our estimates are based on bivariate extreme value models that can capture both key forms of extremal dependence. We devise tools for validating return curve estimates, as well as representing their uncertainty, and compare a selection of curve estimation techniques through simulation studies. Curve estimates are obtained for two metocean data sets, with diagnostics indicating generally good performance. In the context of extremes, few methods have been proposed for modelling trends in extremal dependence, even though capturing this feature is important for quantifying joint tail behaviour. Motivated by observed dependence trends in data from the UK Climate Projections, we propose a novel semi-parametric modelling framework for non-stationary, bivariate extremal dependence structures. This framework allows us to capture a wide variety of dependence trends for datasets exhibiting asymptotic independence. We compare our model to an existing technique through a simulation study, obtaining competitive results over a range of extremal dependence structures. When applied to a climate projection dataset, our model is able to capture observed dependence trends and, in combination with models for marginal non-stationarity, can be used to produce estimates of return curves in future climates. Whilst asymptotic independence is frequently observed in practice, the majority of approaches for bivariate extremes are based on the framework of regular variation. In practice, this is problematic since this framework is is unable to accurately extrapolate into the joint tail for data sets exhibiting this class of extremal dependence. Motivated by this shortcoming, we introduce a range of novel estimators for the so-called `angular dependence function', a quantity which summarises the dependence structure for asymptotically independent variables. We compare the proposed estimators to existing techniques through a systematic simulation study, obtaining competitive results in many cases. The proposed methodology is also applied to river flow data from the north of England, UK, and used to obtain return curve estimates for different pairs of gauge sites
Improving estimation for asymptotically independent bivariate extremes via global estimators for the angular dependence function
Modelling the extremal dependence of bivariate variables is important in a wide variety of practical applications, including environmental planning, catastrophe modelling and hydrology. The majority of these approaches are based on the framework of bivariate regular variation, and a wide range of literature is available for estimating the dependence structure in this setting. However, such procedures are only applicable to variables exhibiting asymptotic dependence, even though asymptotic independence is often observed in practice. In this paper, we consider the so-called ‘angular dependence function’; this quantity summarises the extremal dependence structure for asymptotically independent variables. Until recently, only pointwise estimators of the angular dependence function have been available. We introduce a range of global estimators and compare them to another recently introduced technique for global estimation through a systematic simulation study, and a case study on river flow data from the north of England, UK
A marginal modelling approach for predicting wildfire extremes across the contiguous United States
This paper details a methodology proposed for the EVA 2021 conference data challenge. The aim of this challenge was to predict the number and size of wildfires over the contiguous US between 1993 and 2015, with more importance placed on extreme events. In the data set provided, over 14\% of both wildfire count and burnt area observations are missing; the objective of the data challenge was to estimate a range of marginal probabilities from the distribution functions of these missing observations. To enable this prediction, we make the assumption that the marginal distribution of a missing observation can be informed using non-missing data from neighbouring locations. In our method, we select spatial neighbourhoods for each missing observation and fit marginal models to non-missing observations in these regions. For the wildfire counts, we assume the compiled data sets follow a zero-inflated negative binomial distribution, while for burnt area values, we model the bulk and tail of each compiled data set using non-parametric and parametric techniques, respectively. Cross validation is used to select tuning parameters, and the resulting predictions are shown to significantly outperform the benchmark method proposed in the challenge outline. We conclude with a discussion of our modelling framework, and evaluate ways in which it could be extended.20 pages, 4 figure
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
