30793 research outputs found
Sort by
New Bayesian regression models for massive data and extreme longitudinal data
This thesis was submitted for the award of Doctor of Philosophy and was awarded by Brunel University LondonThe phenomena of heavy-tailedness and asymmetry are ubiquitous in a variety
of practical applications. The intriguing property of heavy-tailedness implies
that the underlying distribution is capable of producing anomalous observations
which deviate too far from the main body of observations. Down-weighing
such extreme observations for an asymmetric distribution can sacrifice inherited
information and introduce considerable bias on parameter estimation. Over the
past decades, two main approaches have emerged to tackle the distributional deviation
caused by heavy-tailedness. The first procedure adopts mixture models
to accommodate the heterogeneity in the distribution of the data. The second
technique considers appropriate distributions to take care of the majority as
well as the heavy tail of the data.
This thesis aims to make some novel contributions to the following three issues
related to massive data and extreme longitudinal data exhibiting heavy-tailed
characteristics. First, the multitude of existing literature coping with continuous
distributions with heavy-tailedness contrasts sharply with the scarcity of
integer-valued distributions. This especially applies to integer-valued time series
modelling. Second, heavy tails can considerably shadow the nature of the
dependence between the response and the covariates of interest, calling the normality
assumption and conventional linear models into question. The quantile
regression (QR) approach, which is robust to outlier contamination associated
with heavy-tailed errors, serves as a remedy for these hurdles. Bayesian quantile
regression (BQR) has received increasing attention from both theoretical
and empirical viewpoints with wide applications and variants, but little attention
has been paid to BQR for big data analysis. Third, the phenomena of
heavy-tailedness often arise with the semi-continuous data, which are commonly
characterized by a mixture of zero values and continuously distributed positive
values. This conceptual framework leads to the formulation of a two-part model.
The literature on two-part models, especially in Bayesian paradigms, for investigating
quantiles of semi-continuous longitudinal data with bounded support
such as the standard unit interval p0, 1q, is relatively limited.
This thesis encapsulates three themes to address the above-mentioned challenges:
Bayesian integer-valued time series modelling with heavy-tailedness
characteristics, Bayesian quantile regression for big data analysis and Bayesian quantile parametric mixed regression for semi-continuous longitudinal data with
bounded support. The main contributions are elaborated as below:
• Chapter 2 gives rise to the Bayesian inference for log-linear Beta–negative
binomial integer-valued generalized autoregressive conditional heteroscedastic
(BNB-INGARCH) models and conducts parameter estimations within
adaptive Markov chain Monte Carlo frameworks. The conditions for the
posterior distribution of the full model parameter to be proper given some
general priors have been presented.
• Chapter 3 contributes to a new approach of Bayesian quantile regression
for big data. This chapter introduces the structure link between Bayesian
scale mixtures of normals linear regression and BQR via normal-inversegamma
(NIG) distribution type of likelihood function, prior distribution
and posterior distribution. The big data based algorithms for BQR and
Bayesian LASSO quantile regression are provided and the proposed algorithms
are demonstrated via simulations and a real-world data analysis.
• Chapter 4 introduces a two-part latent class Kumaraswamy quantile mixed
regression with Bayesian inference for bounded longitudinal data that exhibit
a large spike at zeros. Correlated random effects with class-specific
covariance structures are formulated for the binary and the bounded positive
components to account for both zero inflation and unobserved heterogeneity.
The developed method portrays the trajectory of distinct latent
class evolutions in the underlying outcome process, which provides
valuable insights into the latent cluster structure at various quantiles encompassing
the tails and caters to the exploration of skewed longitudinal
data with bounded support.Partially from OptiRisk System
Optimising integration of renewable energy sources to achieve net zero
This research paper develops a novel data-driven method to optimise the connection and integration of renewable energy resources (RES) in the United Kingdom (UK), aiming to address the ambitious target to achieve net-zero greenhouse gas emissions in the UK by 2050, according to 2019 legislation. This commitment has led to a major transition from using fossil fuels to high penetration of RES. A future transition to clean energy is achievable through a combination of RES, their optimal location, and the reform of existing market arrangements. Firstly, the paper aims to establish locations with optimal weather conditions for deploying photovoltaic (PV) and wind power generation using the k-means clustering technique. The datasets used for this research were collected from available weather data stations located across the UK. The power generation of PV installations heavily depends on solar irradiance, whereas onshore wind generators require stable and strong wind conditions. However, the choice of location is crucial for the connection of such RES. The strategic placement of these energy sources can play a major role in relieving congestion and enhancing grid efficiency. The research as presented in this paper suggests locations with favourable weather conditions for the integration of RES. Secondly, the paper uses a 36-zone reduced model of the Great Britain transmission system to conduct direct current power flow studies across the network to determine areas with network constraints. The findings of this study have implications for policymakers and investors as they provide the necessary signals for the optimal integration of RES to achieve net-zero.This work is supported by the following organisations and funding bodies: the British Academy, the Council for At-Risk Academics, the Petroleum Technology Development Fund
A Novel Methodology to Investigate the Impact of Electricity Market Reforms on Future Electricity Prices
Electricity market reforms are essential to the United Kingdom (UK) government's plan to achieve net zero by 2035. However, the impact of these reforms on future electricity prices remains challenging due to the complex and intermittent nature of renewable energy sources. This research paper presents a novel methodology to investigate the impact of current electricity market reforms in the UK on future electricity prices. The methodology integrates advanced simulation techniques with a detailed analysis of market mechanisms, providing a comprehensive framework to assess the impacts of regulatory changes. Locational pricing is a proposed alternative to single wholesale market pricing in the recent Review of Electricity Markets Arrangements (REMA) consultation. To conduct the research, the study uses a 36-bus reduced version Great Britain (GB) transmission system using PowerFactory to calculate the prices in each zone based on the Future Energy Scenarios (FES). The findings of this study reveal that adopting locational pricing could result in significant geographical disparities in electricity prices across different zones.10.13039/501100009614-Petroleum Technology Development Fund (PTDF)
Lockdowns and vaccinations: Could COVID-19 interventions reduce long-term COVID-19 consequences in Ghana?
Abstract of a presentation made at the Workshop on the Economics of Pandemic Preparedness, 19 June 2024 - 20 June 2024. Stockholm, Sweden. The presentation formed part of Theme 2: Economic Impact of Public Health Interventions.Shirley Crankson from Brunel University used an agent-based model to simulate long-term economic and epidemiological effects of lockdowns and vaccination during the COVID-19 pandemic in Ghana. Shirley Crankson’s work provided a perspective on pandemic responses in resource-constrained settings. The study found that a combination of whole-population vaccination and periodic lockdowns could reduce COVID-19 related health outcomes by over 90%, and targeted vaccination of high-risk groups could further reduce mortality by 13%. These findings provide critical insights for policymakers in Ghana, suggesting that maintaining a balance between broad and tailored interventions can offer substantial health and economic gains when resources are limited
Data-Driven Delta Machine Learning Models for Improved Extrapolation
This work demonstrates a method of testing a machine learning model’s extrapolation accuracy, a capability that is significant to efficiently aid with discovery of improved alloys and presents the application of a pure data-driven method to make steps to reduce this extrapolation error. By using linear models to capture general trends in the data and then the subsequent application of more complex machine learning methods, extrapolation capabilities can be reduced. Being purely data-driven, this type of model can be coupled with other Delta-Machine Learning techniques such as those that utilize physics-domain knowledge, and coupled with active learning methods, with better extrapolation capabilities reducing the number of iterations needed to outperform existing alloys.This work was completed as part of an industrial CASE studentship funded by both the Engineering and Physical Sciences Research Council (EPSRC) and industrial partner Constellium
Semgrep∗: Improving the Limited Performance of Static Application Security Testing (SAST) Tools
Vulnerabilities in code should be detected and patched quickly to reduce the time in which they can be exploited. There are many automated approaches to assist developers in detecting vulnerabilities, most notably Static Application Security Testing (SAST) tools. However, no single tool detects all vulnerabilities and so relying on any one tool may leave vulnerabilities dormant in code. In this study, we use a manually curated dataset to evaluate four SAST tools on production code with known vulnerabilities. Our results show that the vulnerability detection rates of individual tools range from 11.2% to 26.5%, but combining these four tools can detect 38.8% of vulnerabilities. We investigate why SAST tools are unable to detect 61.2% of vulnerabilities and identify missing vulnerable code patterns from tool rule sets. Based on our findings, we create new rules for Semgrep, a popular configurable SAST tool. Our newly configured Semgrep tool detects 44.7% of vulnerabilities, more than using a combination of tools, and a 181% improvement in Semgrep’s detection rate
Information disclosure vs. information learning via Google search
We decompose the Google Trends Search Volume Index into naïve and sophisticated searches and examine their impacts on mortgage default, respectively. Using U.S. data from 2006 to 2018, we find that the sophisticated search activity has a positive and robust relationship with the change in the percentage of mortgages in 90+ days of delinquency. However, foreclosure starts are positively related to naïve search activity in the short term, but negatively related to sophisticated search activity in the long term. Borrowers are more likely learn from sophisticated online searches than from naïve online searches, and they can use the information to avoid foreclosure starts and keep their houses. The relationship between Google search
activity and mortgage default outcomes are significantly stronger in states that experienced substantial house price drops in the recent year. Our findings are robust to a battery of alternative settings
Influence of Binding Energies on Required Process Conditions in Aerosol Deposition
With the high interest in aerosol deposition in order to form high-quality coatings by solid-state impact, there is an increasing demand for developing general guidelines to estimate needed particle velocities and thus process parameter sets for successful deposition of ceramic materials. By using modeling approaches, rather different material properties in first instance can be expressed in terms of binding energies. Needed velocities for possible bonding can then derived by impact simulations and compared to experimental results from the literature. In order to study the role of binding energy on the impact behavior of ceramic particles in aerosol deposition, a molecular dynamics study is presented. Single-particle impacts are simulated for a range of binding energies, particle sizes and impact velocities. The results show that increasing the binding energy from 0.22 to 0.96 eV results in up to three times higher characteristic velocities corresponding to the threshold of bonding or grain size-dependent fragmentation of the particles. However, regardless of the binding energy, exceeding the characteristic velocities results in a similar deformation and fragmentation pattern. This allows for a general representation of the impact behavior as a function of normalized impact velocity for different ceramic materials. Apart from dealing with prerequisites for bonding of different materials by aerosol deposition, this study could also be generally relevant to the high-velocity deformation behavior of ceramics with different grain sizes.Open Access funding enabled and organized by Projekt DEAL
Spongy-looking microfabrics in the earliest named stromatolite represent deep burial alteration and incipient metamorphism
Data availability:
Data is provided within the manuscript and the supplementary information file
available online at: https://www.nature.com/articles/s41598-024-83359-7#Sec14 .The earliest named stromatolite Cryptozoon Hall, 1884 (Late Cambrian, ca. 490 Ma, eastern New York State), was recently re-interpreted as an interlayered microbial mat and non-spiculate (keratosan) sponge deposit. This “classic stromatolite” is prominent in a fundamental debate concerning the significance or even existence of non-spiculate sponges in carbonate rocks from the Neoproterozoic (Tonian) onwards. Cryptozoon has three types of microbially-induced carbonate layers: clotted-pelletoidal micrite with microbial filaments, clotted-pelletoidal micrite with vesicular structure, and dense microcrystalline laminae. A fourth, stratiform to patchy fabric comprises suspect sponges. Using contextual fabric analysis, elemental mapping, cathodoluminescence, fluid inclusions, electron backscatter diffraction, U–Pb dating, and burial history, the sponge interpretation is denied. Neither a distinct sponge body outline nor a canal system is identifiable. Instead, the suspect fabric is secondary in origin, and best explained as a product of Carboniferous (Mississippian) deep burial alteration associated with basement reactivation. Key petrographic observations include heterogenous recrystallization via aggrading Ostwald ripening with interfingering reaction fronts typical for partially miscible fluids, a granoblastic calcite texture (incipient metamorphism), and subsequent hypidioblastic white mica (arguably Carboniferous/Permian, Alleghenian orogeny). Topotype Cryptozoon is a stromatolite altered to sub-greenschist metacarbonate. The published Tonian to Phanerozoic record of interpreted non-spiculate sponges requires reassessment
Religious women receive more allomaternal support from non-partner kin in two low-fertility countries
Data availability:
The data associated with this research are available on the project's OSF page at: https://osf.io/rg235/ .Supplementary data are available online at: https://www.sciencedirect.com/science/article/pii/S1090513824000369?via=ihub#s0085 .In low fertility settings, religious people tend to have larger families than non-religious people. One way religious individuals may achieve larger relative family sizes is through support from their families. In this paper, we investigate the relationships between religiosity, kin contact, allomaternal investment from relatives, and fertility in two high income low fertility settings: the United Kingdom and the United States. Data for this pre-registered research come from an online survey of 609 women living in the US and 919 women living in the UK, recruited through Prolific, who answered questions about their religious practices, childbirth histories, social networks, and allomaternal networks. We find that, compared with less religious peers, more religious women: 1) have more geographically diffuse kin networks (particularly in the UK) but have social networks that are equally kin-dense; 2) receive more allomaternal support from kin beyond their partner, particularly help with household tasks, though the countries differ in the exhibited relationship between religiosity and partner support; and 3) have higher fertility in both countries. We do not find strong evidence for a mediating role of allomaternal support on the relationship between religiosity and fertility. Our study highlights important variation in the relationship between religion and fertility across two high income low fertility countries and raises new questions about the role that religion plays in allomaternal support networks in these settings.Funding for this study was provided by the John Templeton Foundation [grant number 61426], and the Templeton Religion Trust [grant number TRT-2022-30378]