1,720,989 research outputs found
An optimized intermolecular force field for hydrogen bonded organic molecular crystals using atomic multipole electrostatics
We present a re-parameterization of the a popular intermolecular force field for describing intermolecular interactions in the organic solid state. Specifically, we optimize the performance of the exp-6 force field when used in conjunction with atomic multipole electrostatics. We also parameterize force fields that are optimized for use with multipoles derived from polarized molecular electron densities, to account for induction effects in molecular crystals. Parameterization is performed against a set of 186 experimentally determined, low temperature crystal structures and 53 measured sublimation enthalpies of hydrogen bonding organic molecules. The resulting force fields are tested on a validation set of 129 crystal structures and show improved reproduction of the structures and lattice energies of a range of organic molecular crystals compared to the original force field with atomic partial charge electrostatics. Unit cell dimensions of the validation set are typically reproduced to within 3% with the re-parameterized force fields. Lattice energies, which were all included during parameterisation, are systematically underestimated when compared to measured sublimation enthalpies, with mean absolute errors of between 7.4 and 9.0%
Accelerating computational discovery of porous solids through improved navigation of energy structure function maps
While energy-structure-function (ESF) maps are a powerful new tool for in silico materials design, the cost of acquiring an ESF map for many properties is too high for routine integration into high-throughput virtual screening workflows. Here, we propose the next evolution of the ESF map. This uses parallel Bayesian optimization to selectively acquire energy and property data, generating the same levels of insight at a fraction of the computational cost. We use this approach to obtain a two orders of magnitude speedup on an ESF study that focused on the discovery of molecular crystals for methane capture, saving more than 500,000 central processing unit hours from the original protocol. By accelerating the acquisition of insight from ESF maps, we pave the way for the use of these maps in automated ultrahigh-throughput screening pipelines by greatly reducing the opportunity risk associated with the choice of system to calculate.</p
Data-driven approach of discovering organic photocatalysts and developing molecular force field by machine learning
Machine learning techniques are becoming more prevalent in chemistry research as they
offer an effective approach for handling large, complex chemical datasets generated from
high-throughput experiments and molecular simulations. To gain a comprehensive under-
standing of datasets, it is crucial to employ efficient methods for data representation and
analysis. This PhD project utilized classical machine learning algorithms to effectively
visualize high-dimensional chemical data, ascertain connections between chemical struc-
ture and properties, facilitate the discovery of novel organic catalysts, and developing a
machine learning potential to describe intermolecular interactions
Scalable Sequential Monte Carlo Samplers for Numerical Bayesian Inference
Sequential Monte Carlo (SMC) samplers offer a promising alternative to Markov Chain Monte Carlo (MCMC) methods for inferring a target density associated with a static dataset. Empirical evidence suggests that SMC samplers may require a shorter burn-in period than MCMC. Moreover, SMC samplers can be parallelised, provide estimates of the normalising constant, and support a variety of proposal distributions. However, current state-of-the-art SMC configurations have notable limitations: they suffer from the curse of dimensionality, are computationally inefficient, operate sequentially, and have long run times due to expensive tempering schemes.
Recent studies have shown that SMC samplers can effectively utilise clusters of CPUs. Yet, these samplers fail to exploit hardware accelerators and are often restricted to specialist high-performance computing facilities. These challenges highlight the need for SMC samplers that can handle high-dimensional problems, avoid costly tempering schemes, leverage hardware accelerators, and operate on commodity computing resources.
This thesis begins by outlining the Bayesian framework and MCMC methods before focusing on the importance sampling framework and SMC methods. A novel SMC sampler is introduced, designed to tackle high-dimensional problems and exhibit strong parallel scaling behaviour. Particle recycling schemes are then explored to enhance the computational efficiency of the sampler. A mechanism for selecting iterations to recycle and forming unbiased recycled estimates of high-order statistics is proposed. The proposed SMC sampler is parallelised on shared and distributed memory architectures, and various parallel computing frameworks are evaluated.
Subsequently, a framework for distributing the SMC sampler on an opportunistic computing environment, comprising a heterogeneous collection of commodity computing resources, is presented. This framework offers a cost-effective and efficient way for practitioners to benefit from SMC samplers. A novel SMC sampler is then introduced, providing an unbiased, lower-error alternative to multiple short parallel MCMC chains. The thesis concludes with a summary of contributions and recommendations for future work
Controlling the crystallization of porous organic cages: molecular analogs of isoreticular frameworks using shape-specific directing solvents
Small structural changes in organic molecules can have a large influence on solid-state crystal packing, and this often thwarts attempts to produce isostructural series of crystalline solids. For metal–organic frameworks and covalent organic frameworks, this has been addressed by using strong, directional intermolecular bonding to create families of isoreticular solids. Here, we show that an organic directing solvent, 1,4-dioxane, has a dominant effect on the lattice energy for a series of organic cage molecules. Inclusion of dioxane directs the crystal packing for these cages away from their lowest-energy polymorphs to form isostructural, 3-dimensional diamondoid pore channels. This is a unique function of the size, chemical function, and geometry of 1,4-dioxane, and hence, a noncovalent auxiliary interaction assumes the role of directional coordination bonding or covalent bonding in extended crystalline frameworks. For a new cage, CC13, a dual, interpenetrating pore structure is formed that doubles the gas uptake and the surface area in the resulting dioxane-directed crystals
Privacy-Preserving Gaussian Process Regression – A Modular Approach to the Application of Homomorphic Encryption
Much of machine learning relies on the use of large amounts of data to train models to make predictions. When this data comes from multiple sources, for example when evaluation of data against a machine learning model is offered as a service, there can be privacy issues and legal concerns over the sharing of data. Fully homomorphic encryption (FHE) allows data to be computed on whilst encrypted, which can provide a solution to the problem of data privacy. However, FHE is both slow and restrictive, so existing algorithms must be manipulated to make them work efficiently under the FHE paradigm. Some commonly used machine learning algorithms, such as Gaussian process regression, are poorly suited to FHE and cannot be manipulated to work both efficiently and accurately. In this paper, we show that a modular approach, which applies FHE to only the sensitive steps of a workflow that need protection, allows one party to make predictions on their data using a Gaussian process regression model built from another party's data, without either party gaining access to the other's data, in a way which is both accurate and efficient. This construction is, to our knowledge, the first example of an effectively encrypted Gaussian process
- …
