DSpace@RPI (Rensselaer Polytechnic Institute)
Not a member yet
6809 research outputs found
Sort by
System analysis with signal temporal logic: monitoring, motion planning, and data mining
May 2023School of EngineeringThis dissertation is about the analysis of systems using temporal logics, focusing on three perspectives: temporal logic monitoring, motion planning with temporal logic specifications, and temporal logic inference from time series data. Temporal logic monitoring checks if all the given simulation traces satisfy a particular temporal logic formula describing a desired or undesired system attribute, such as correctness or safety criteria. This dissertation is about temporal logic monitoring for swarm robotic systems in particular. Motion planning with temporal logic specifications involves designing control inputs for the robotic system, given its dynamic model and a temporal logic formula, in order to satisfy the desired temporal logic specifications. Temporal logic inference from time series data addresses the challenge of learning a temporal logic formula from a given dataset, allowing for the description of data characteristics in an interpretable format. To begin with, we first introduce the Signal Temporal Logic (STL) and Swarm Signal Temporal Logic (SwarmSTL). STL describes behaviors of individual agents, whereas SwarmSTL describes behaviors of robot swarms on the swarm level. SwarmSTL is defined on the generalized moments (GMs) that represent swarm features. For SwarmSTL, we propose a centralized, sampling-based monitoring algorithm as well as a distributed monitoring algorithm for temporal logic monitoring of robot swarms. We use the attribute of GMs being the mean of a polynomial function to design a Generalized Moment Consensus Algorithm (GMCA), allowing each agent to estimate the GMs and track their satisfaction with respect to SwarmSTL formulas. Next, we present the work of distributed motion planning for robot swarms under SwarmSTL specifications. The motion planning problem is formulated as a mixed-integer quadratic programming (MIQP) problem with the SwarmSTL formulas encoded as linear constraints. A distributed branch and bound algorithm (DBB) executed by the agents is proposed to solve the MIQP problem, and the agents can achieve consensus on the optimal trajectory of the generalized moments, from which the agents can compute the trajectories of their own through inverse projection. Finally, we present the work of temporal logic inference from time series data. The inference work can be divided into two categories from the perspective of the methods to find the optimal parameters of the formulas. The first category of temporal logic inference involves SwarmSTL inference from swarm execution traces. The SwarmSTL inference problem is posed as optimizing parameters for given SwarmSTL formula structures such that the inferred SwarmSTL formula can best describe the swarm data. We propose two algorithms for SwarmSTL inference. One is SwarmSTL inference with entire swarm data, and another is SwarmSTL inference via sampling. To overcome the traditional STL inference algorithm's limitations of nonsmooth robustness degrees and low computation efficiency, we propose the second category of temporal logic inference algorithms. We introduce weights into STL and propose a weighted STL (wSTL). An end-to-end differentiable neuro-symbolic model called Signal Temporal lOgic Neural nEtwork (STONE) is developed to learn wSTL formulas, which improves the computation efficiency of learning wSTL formulas. STONE's design is based on the robustness degree, which is a real-valued number reflecting the degree of satisfaction of wSTL formulas over time series data. To better accommodate the design of STONE to the convention of neural networks and its real-world application, novel quantitative semantics based on the notion of truth degree are proposed for wSTL, which generates the development of a signal temporal logic neural network for electroencephalogram (EEG) signal analysis (EEG-STONE) and seizure detection. Although STONE and EEG-STONE can tackle the computation efficiency limitation of traditional signal temporal logic inference algorithms, they can only achieve binary time series classification tasks. To expand their capability to multi-class time series classification, we develop a decision tree-based time series classifier based on EEG-STONE to classify multi-class time series data, where every node in the tree classifier is an EEG-STONE. In addition to the above neuro-symbolic models, we also explore the effectiveness of similar neuro-symbolic models for modeling multivariate point process. A novel framework for modeling temporal point processes called clock logic neural networks (CLNN) which learn weighted clock logic (wCL) formulas as interpretable temporal logic rules by which some events promote or inhibit the occurrence of other events. Unlike conventional approaches of searching for generative rules through expensive combinatorial optimization, we design smooth activation functions for components of wCL formulas that enable a continuous relaxation of the discrete search space and efficient learning of wCL formulas using gradient-based methods.Ph
Pushing the boundaries of nonlinearities using structured light
May 2023School of ScienceSpatial light modulators give us the ability to manipulate the fundamental constituents of light (amplitude and phase). This ability has ushered in a new way to approach light matter interaction (to control light matter interactions) by modulating the wavefront incident on a material. Light can carry two types of angular momentum, spin angular momentum and orbital angular momentum. The former is linked to the rotation of the polarization vector whereas the latter stems from a helical phase structure. Some structured light beams are specially designed to have this rotating phase structured resulting in beams which carry and have the ability to transfer orbital angular momentum. In this work we use structured light fields shaped by spatial light modulators in conjunction with nonlinear phenomena to create new light beams. The nonlinear processes of concern are second harmonic generation and spontaneous parametric down-conversion . Both have been extensively studied using Gaussian beam profiles, but not structured light fields. The research presented here focuses on developing new, interesting beams and more efficient, brighter sources of entangled photons. To create structured light fields such as a Laguerre-Gaussian or a Bessel-Gaussian, physical optical elements such as a spatial phase plate or an axicon had to be used to produce the respective beam. As these physical elements can only produce a single mode, this limits what can be studied. However, by encoding holographic phase masks onto spatial light modulators one can use a single tool to create structured light fields. We use the versatility afforded to us by the SLM to study the second harmonic signal pumped by three different types of structured light fields: the Bessel-Gaussian, the Laguerre-Gaussian, and the Hermite-Gaussian. With the Hermite-Gaussian beam, we present the difference in mode-conservation of the SH depending on the focusing of the pump. For Laguerre-Gaussian beams, we examine the conservation of total angular momentum in the second harmonic process. Then, as Hermite-Gaussian and Laguerre-Gaussian modes can be inter-converted using astigmatic transform, we study the effect of such a transformation of the pump on the resultant second harmonic signal. Lastly, we present the novel propagation behavior of a frequency doubled Bessel-Gaussian which can be likened to the scattering of solitons.
Spontaneous parametric down-conversion is used to generate entangled photons. Commonly, the surface area of the crystal is larger than the incident laser on it. Therefore, we devised a method to sub-divide the beam before incident on the crystal such that each sub-divided beam drives its own down-conversion process. Instead of a single source of entangled photons per crystal, we have multiple sources from the same crystal. Thereby, when multiple sources or a more complex entanglement product is required, our method is more efficient.
In order to generate a brighter source of entangled photons or single photons, we combine wavefront shaping with spontaneous parametric down-conversion. Through a simple feedback loop formed by the spatial light modulator and a camera capturing the down-conversion. We optimize the wavefront incident on the nonlinear medium to increase the intensity of the down-conversion at a desired location (and its conjugate).Ph
Adapting Emotion Detection to Analyze Influence Campaigns on Social Media
Social media is an extremely potent tool for influencing public opinion, particularly during important events such as elections, pandemics, and national conflicts. Emotions are a crucial aspect of this influence, but detecting them accurately in the political domain is a significant challenge due to the lack of suitable emotion labels and training datasets. In this paper, we present a generalized approach to emotion detection that can be adapted to the political domain with minimal performance sacrifice. Our approach is designed to be easily integrated into existing models without the need for additional training or fine-tuning. We demonstrate the zero-shot and few-shot performance of our model on the 2017 French presidential elections and propose efficient emotion groupings that would aid in effectively analyzing influence campaigns and agendas on social media
Development of a high-order cut-cell method for cfd-based design optimization
August 2023School of EngineeringComputational fluid dynamics (CFD)-based aerodynamic design optimization requires automation, particularly when dealing with complex geometries that require boundary conforming meshes. To address this issue and help automate the design optimization process, non-boundary conforming meshes can be used, leading to cut-cell methods. However, cut-cell methods introduce their own challenges, one of which is a potentially ill-conditioned discretization due to some cut-cells being orders of magnitude smaller than regular shaped cells. Various techniques have been developed to tackle this issue, most of which require special treatment for dealing with small cut-cells. This thesis introduces a cut-cell method based on the discontinuous Galerkin difference discretization (cut-DGD) that does not require any special treatment of cut-cells to eliminate ill-conditioning. The first part of this thesis presents the cut-DGD discretization method, which uses an underlying discontinuous Galerkin (DG) discretization. The role of DGD basis functions and the stencil construction in maintaining a well-conditioned cut-DGD discretization is discussed, and numerical experiments are conducted to show that the condition numbers of the cut-DGD mass and stiffness matrices remain bounded as the size of the cut cell approaches zero. Furthermore, the accuracy of the cut-DGD discretization in solving the one-dimensional steady-state linear advection equation and the two-dimensional Poisson partial differential equation (PDE) is demonstrated in the presence of small cut-cells. The second part of the thesis focuses on performing numerical integration of arbitrarily shaped cut-cells. A level-set formulation is developed to approximate the geometries of general shapes, such as airfoils, which are of interest in aerodynamic applications. The impact of increasing the number of boundary points and applying high-order corrections to improve the accuracy of the level-set approximation is discussed. The level-set definition of a given geometry is necessary to apply the quadrature algorithms used in this work. Two different quadrature algorithms are explored: one that requires user-defined bounds on the level-set function, which can be challenging for geometries with discontinuities (e.g., trailing edges of airfoils), and another that only requires polynomial-type level-set functions without such bounds. Fluid-flow accuracy studies are conducted to demonstrate the high-order accuracy of cut-DGD, using both exact and approximate level-set functions, by solving two-dimensional Euler equations. The final part of the thesis presents contributions towards gradient-based design optimization. Specifically, the sensitivity analysis of cut-DGD is derived, and the semi-automatic differentiation of a cut-cell quadrature algorithm is performed. These derivatives are then verified against a finite-difference approximation. A discrete adjoint approach is used to perform sensitivity analysis and to compute adjoint variables necessary for solving functional sensitivity derivatives, which are further required for a gradient-based optimization algorithm. The residual sensitivity derivative is verified against finite-difference approximations for a model flow problem.Ph
Computational modelling of ti-6al-4v microstructure evolution during thermomechanical processing
December 2022School of EngineeringThe microstructure that results from thermomechanical processing of metals and alloys is directly responsible for the material’s mechanical, electrical, and thermal properties. Hence, the goal of this work is to model and simulate Ti-6Al-4V microstructure under thermo-mechanical loading in order to understand the mechanisms that guide the evolution of themicrostructure during processing. It is a dual-phase alloy that is characterized by a vanadium stabilized body-centered cubic (BCC) β phase and an aluminum stabilized hexagonal close-packed (HCP) α phase at room temperature. Ti-6Al-4V is typically deformed at elevated temperatures, followed by a prescribed heat treatment schedule designed to generate
a suitable microstructure. On heating above 600℃, the α phase starts transforming into β phase. The phase transformation during this process is governed by the Burgers orientation relationship, where the α → β (heating) transformation can result in six possible β orientation variants and the β → α (cooling) transformation can result in twelve possible α variants due to crystal symmetry. Experimental observations show that variant selection, i.e. an apparent preference for certain variants over others, during heating (α → β) is negligible. The final microstructure is largely dominated by the β → α transformation where the initial α grains show features of the final texture, meaning that the mechanisms leading to variant selection occur mainly during the initial stages of cooling. Hence, it is vital to understand this process in order to understand why certain microstructures form under a given processing condition. As the material is cooling down, the β → α phase transformation can lead to significant deformation, well beyond the elastic limit of the surrounding β grain. A vast majority of current models, however, use elastic analysis to compute the effects such deformation may have on the local energy and α growth. The first part of this work looks at the impact of plastic relaxation during transformation-induced deformation on the subsequent growth of α. A model is developed to simulate the deformation caused by phase transformation using a finite element crystal plasticity method. Transformation of a single lath of α is simulated and the phase transformation at each time step is introduced as a deformation gradient. Elastic and plastic deformations that can accommodate such a deformation are computed. The resulting strain energy, which contributes to the local free energy of the system and thus affects the nucleation and growth of α, is compared for simulations with and without plasticity to determine the impact of plastic relaxation on the driving energy in phase transformation. We show that when the surrounding β crystals are allowed to undergo plastic relaxation, the resulting strain energy is significantly lower when compared to the results from elastic analysis. This shows that existing models need to consider elastic-plastic deformation to accurately compute the driving energy of transformation. It was also found
that increasing the growth rate of α increases strain energy density in and around the lath – suggesting that deformation due to transformation may also be an important consideration to understand microstructure morphology as a result of processing conditions. While it is important to study the cooling cycle, it is also necessary to look at the heating cycle and the warm working regime for a better understanding of processes that end at warm deformation and do not heat the material above β-transus. In order to study the microstructure evolution during the warm deformation conditions, following industriallyrelevant processing conditions, both the effects of annealing and deformation need to be studied simultaneously. The crystal plasticity model is calibrated for both phases at 800℃, and the deformation of an equiaxed α + β microstructure is simulated at 800℃ under 15% compressive load at a strain rate of 10−3 s−1. The Monte-Carlo Potts model, used to model
grain-growth during annealing, is calibrated for various temperatures using literature data. For temperatures above β-transus, the model is calibrated using grain growth data at 1088℃. For temperatures below β-transus, the model is calibrated using grain growth data at three different temperatures – 650℃, 775℃, and 815℃, in order to capture the growth behavior in the warm working temperature range. To integrate these models, the polycrystal used in deformation simulations is used as the starting microstructure for the Monte Carlo model and the strain energy as a result of the deformation is mapped on to the Monte Carlo grid for the grain growth simulations. Grain growth simulations are conducted with and without strain energy density to compare the impact of prior deformation on grain growth. In presence of strain from deformation simulation, the grain growth gets more localized to the regions where there is significant gradient in energy across grain boundaries whereas without the strain energy, the grains grow irrespective of the phase to minimize the grain boundary energy. The soft β grains having relaxed first, have significantly lower strain energy than the surrounding α resulting the grain growth getting concentrated at the α/β interfaces. This work provides a process of integrating the stored elastic energy from deformation in Monte Carlo framework to study the effect of prior deformation on grain growth, with further work on multi-phase grain and phase evolution this integrated approach can be used to predict
microstructure evolution during warm working accurately.Ph
A computational chemist’s guide to the development of designer material systems through the utilization of past, present, and developmental computational techniques
December 2023School of ScienceComputational chemistry broadly investigates chemical problems using computer aided simulations/models. The exploration of material properties through computational means requires traversing length scales, from atomic to macroscopic, and large amounts computing time and energy. The implementation of computational techniques; such as machine learning, molecular dynamics simulations, and density functional theory calculations, helps to bridge this gap. Two major computational technique were broached; Quantum Theory of Atoms in Molecules (QTAIM or AIM) and Quantitative Structure-Activity (or Property) Relationship modeling (QSAR or QSPR), in the development of metal-based TAE (Transferable Atom Equivalent) atom types, the utilization of a descriptor-based representation DNA-Pixels, and the evaluation of dielectric properties of nanocomposite systems. TAE descriptors encode the distributions of electron density of the atomic subsections producing a number of properties; Kinetic energy densities, local average ionization potentials, electrostatic potentials, and other atomic charge density derived properties. The development of metal-based TAE atom types allows for the utilization of these electron density derived properties for structural activity and molecular property modeling. The TAE framework provides an adaptive infrastructure for the creation of molecular electron density from TAE atom type fragments with low computational cost. The objective of this work was the development and incorporation of a Platinum TAE atom type library into the existing TAE library of atom types. Another application of the TAE methodology has be used in the generation of a descriptor-based representation of DNA formed from the investigations of the short-range effects of flanking base pairs. DNA-Pixel (DIXEL) represent electron density features such as Electrostatic Potential (EP) and Politzer’s local average Ionization Potential (PIP) on the accessible surfaces of the major or minor groove. DIXELs provide the user with the ability to create a raster of alignment-free spatially resolved interaction points on the major and minor grooves of an DNA sequence, allowing for bioinformatic application. DIXELs were developed using DNA triplets, three basepair combinations of Cytosine, Guanine, Thymine, and Adenine. Methylated Cytosine and Adenine were also included in these combinations. A DNA triplet contains the central basepair of interest flanked by its two adjacent basepairs, producing a total of 216 DNA triplet combinations. Evaluating the electrostatic potential surface of two Adenine triplets, AAA (non methylated) and AA*A (*methylated Adenine), the methylated AA*A triplet contains a region of decreased electrostatic potential on the surface. This decrease is directly opposite of the point of methylation, within the minor groove of the DNA system. This phenomenon demonstrates the effects points of methylation has on DNA electron density. An advantage of DIXEL is its ability to be utilized for a comparative approach evaluating the complementary of electron density surface features of DNA sequences and small molecules. This ligand based similarity approach creates an \emph{ad hoc} method for DNA ligand docking which could be apply to other techniques such as machine learning. These DIXEL descriptors were used to probe sequence-specific interactions of preferential binding of amino acid functionalized amikabeads to methylated over unmethylated regions of DNA. With results indicating that the methylation of Adenine is the primary driver of selected methylated DNA ligand lead of sequence-specific DNA interactions. Next, the evaluation of dielectric properties of high energy-density nanocomposite systems was done using a combination of machine learning, density functional theory calculations, and a large collaborative effort. Predicting the dielectric properties of high energy-dense nanocomposite materials requires a useful model of the electron trapping and mobility. This work investigated the electron trapping phenomenon at interfacial regions of nano-composites, nanoparticle surface, and polymeric bulk through the evaluation of the electronic structure at the interface using several quantum mechanical methods; density functional theory calculations for single molecules (surface functionalization groups) zero-point energy calculations, local density of state calculations for -Quartz Silica electron trapping, and local-density approximation studies for kinetic study of electron/hole mobility in polyethylene systems. Experimental work performed in Dr. Linda Schadler's group concluded the grafting of functional groups to nanoparticles increased the breakdown strength of resulting polymer nanocomposites. To investigate the underlying phenomenon quantum computation of electron affinity (EA) and ionization energy (IE) of isolated single molecules (surface functional groups) was done to explore the possible correlation between these two properties and dielectric breakdown of the resulting functionalized nanoparticle in a composite system. Even though the physical nature of electron traps and the mechanism of nanocomposite breakdown are not clearly understood; intuitively the relationship between the charged states of polymer chains and functionalized nanoparticles contributes to the restriction of free electrons and stabilization of the system. Our focus was on investigating the effect of the incorporation of different functional groups on the surface of a nanoparticle with the belief that systems with smaller band gaps (EA+IE value) would produce higher stability points for electron trapping, depressing likelihood for dielectric breakdown to occur within the system. Assuming the energy difference between the lower energy state (valence band) produced by the electron deficient state and the higher energy state (conduction band) produced by the enriched state. An -quartz silica model system was used to represent functional nanofiller particles and local density of state calculations were performed. These local density of state calculations were used to examine the changes to the band structures that occur when traversing from a fully coordinated silica system to an under-coordinated / functionalized system. At these interfacial regions, the attached functional chains create both deep and shallow electron traps that can affect electron mobility. To better simulate the experimental behavior of amorphous systems at interfacial regions, an amorphous model analog was developed and evaluated. A comparison of the defect states between amorphous and crystalline systems and their effects on the band structure, when functionalized, allow new insights into the role that surface modification/functionalization plays in dielectric breakdown. These model silica analogs enabled the investigation of electron trapping phenomenon at interfacial regions of nanocomposites, nanoparticle surfaces, and polymeric bulk through the evaluation of the electronic structure at the interface. This electron trapping and mobility study allows for the prediction of dielectric properties of new high energy-density nanocomposite materials. Local-density approximation studies for kinetic study of electron/hole mobility in polyethylene systems was done to investigate "voltage stabilizers", molecular additives that have been shown to provide deep traps for hopping carriers. These deep traps slows down the transport as carriers drop into the deep potential well formed by these species. The detrapping rate can be calculated based on the semi-classical thermally assisted tunneling model. To replicate Sato’s methodology for determination of hole mobility, a coreshell structure C8H18 polyethylene oligomers was constructed. Using the core-shell model, density functional theory computations were performed for polyethylene to construct Electron/Hole mobility model.Ph
The mechanochemical output of the kinesin-2 motor kif3ac is tuned by heterodimerization
May 2019School of ScienceThe kinesin-2 family of molecular motors is unusual among kinesins in that it contains both hetero- and homodimeric motors. It is of importance to understand the impact and purpose of heterodimerization in the kinesin-2 family. Mammals express two kinesin-2 heterodimers, KIF3AB and KIF3AC, resulting from 3 genes, kif3a, kif3b, and kif3c, as well as homodimeric KIF17 from the kif17 gene. KIF3AC is of particular interest because KIFA and KIF3C have a 30-fold difference in velocity when expressed as engineered homodimers. The work presented here seeks to define the potential transport capabilities of KIF3AC. This was accomplished by optical trapping techniques and iSCAT (interferometric scattering) microscopy in order to define the behavior of these distinct motor domains during a processive run. The results show that KIF3AC is strikingly more sensitive to external load than KIF3AB, and that KIF3AC steps are characterized by a single observed rate constant and not two independent rate constants as one may expect for fast KIF3A and slow KIF3C. Together, these results suggest that it is unlikely that KIF3AC can transport cargo as one or a few motors and that the KIF3AC heterodimer shows an emergent mechanochemistry in which KIF3A is slowed and KIF3C is accelerated. I argue that this change in kinetics is mediated by interhead tension.Ph
Neural models for causal information extraction using domain adaptation
August 2023School of EngineeringThe task of identifying causality related events or actions from text is an important step in building knowledge graph of events and consequences from the vast amount of unlabeled documents available to us in the digital age. As there has been very little work on this problem, we perform a comparative study on different neural models on 4 different data sets with causal labels. We train sequence tagging and span based models to extract causally related events from their textual description. Our experiments affirm the fact that large pre-trained language models, i.e. BERT, can be fine-tuned on labeled data sets to outperform traditional deep learning models like LSTM. Our results show that span based models are better at classifying spans of words as cause or effect compared to sequence tagging models using the same pre-trained weights from BERT. The length of the spans labeled as causes and effects in a data set also has a significant impact on the advantage of using a span based model. Towards the goal of developing a general purpose model for extraction of causal knowledge from text, we focus on the unsupervised domain adaptation (UDA) scenario where we adapt a model trained on a source domain to a new domain without any label. Several studies on the UDA task for text classification have shown the effectiveness of the adversarial domain adaptation method. We investigate the effect of integrating linguistic information in the adversarial domain adaptation framework for the causal information extraction task. We show the advantage of leveraging the word dependecy relationship in adapting word based neural models like LSTM to new domains. We also find that guiding the adversarial domain classifier to adapt the specific task classifier output is more effective than requiring the encoder model outputs to have similar distributions in two domains.M
Three essays in financial intermediation: a regulatory perspective
August 2023School of ManagementThis dissertation consists of three essays on the role of added regulation on financial advisory firms and bank holding companies. Compliance with stringent added regulations can create ethical dilemmas in financial institutions, causing them to employ unsavory activities such as financial misconduct or corporate tax planning to increase their corporate bottom line. In the first essay, I study how a change in regulatory oversight of mid-size advisory firms from the SEC to the local state regulator can result in increased fraud, emboldened by stronger social connections with local regulators. In the second essay, I study how added transparency requirements create an environment of regulatory uncertainty, which can further lead to increased tax planning in banks. The added cash from tax planning can help banks comply with the minimum capital requirements of bank stress tests. While tax planning itself is not an illegal activity, investors and regulators alike do not always approve of it. While the first two chapters of this dissertation focus on the unintended consequences of regulation, the third chapter highlights an improvement in bank reporting quality following the global financial crisis. The first essay of my dissertation investigates how strong social ties affect agent-regulator relationships and financial misconduct in U.S. registered investment advisories. Using difference-in-differences to exploit the quasi-experimental properties of the Dodd-Frank Act, I find that the change in regulatory purview increases the fraudulent malpractice in mid-sized advisory firms compared to large advisories. Financial misconduct increases at advisories more socially connected to regulators, even after controlling for geographical distance. The results show a disrupting agent-regulator relationship effect when the officer belongs to the largest homogeneous group, white males, and insignificance among female and under-represented minority groups. Consequently, the change in regulators and misconduct affect advisory firm performance and services. In the second essay, using U.S. bank stress tests and regression discontinuity, I find that stress tests have unintended consequences of intensifying tax planning and increasing tax avoidance. Stress-test banks increase tax avoidance by accelerating charge-offs, net interest, and non-interest expenses. However, this increase in tax planning is not optimally maximized, leading to lower effective tax planning compared to non-stress-test banks. Banks with a substantial increase in tax avoidance under the Dodd-Frank Act tend to increase their risk, investing in high-risk-weight assets and lending in riskier loan categories. These findings are consistent with tax minimization conditions under added regulatory attention and policy uncertainty. In the third essay, I study the effects of transparency disclosures on the reporting quality, i.e., restatements and accounting and advisory-related expenses of U.S. banks. I exploit the quasi-experimental set-up provided by the Dodd-Frank stress tests, which mandate policy thresholds according to bank asset size. Using a difference-in-difference design, I find that stress-test banks encounter substantially increased quality in reporting oversight in the form of decreased restatements issued, and face increased advisory and consulting fees to comply with improved reporting standards. The findings, consistent with the literature on increased transparency, indicate improved reporting standards and show a decrease in restatements issued, increased consulting and advisory expenses, while overall legal expenses decline in the immediate years following the regulation.Ph
Associative learning for text and graph data
May 2023School of ScienceThis dissertation focuses on the application of associative learning into text data and graphstructured data. Associative learning is the process of learning to associate two stimuli.
If the connection between two events are repeatedly strengthened, or some similar events
happen over and over again, our brain memorizes those patterns which are frequently shown
together. When we encounter a similar pattern, or one of the event pairs we have in our
memory, then we can retrieve the relevant associated pattern. This kind of phenomenon
is extensively studied in biology, and we want to apply this idea to develop biologically
plausible neural networks.
We identified three main challenges of associative learning. The first challenge is that
most of the early associative learning methods use local learning rules to update the weights.
Though the local leaning rule is efficient and biologically plausible, it may not be able
to extract features that are good enough for representation learning. On the other hand,
backpropagation can guide the network’s weights towards the loss we select, which can be
more helpful for the given task. How to train the weights effectively in the associative
memory network is still an open question. The second challenge is that some recent work
has shown the connection between Hopfield networks and attention module that is widely
used in modern deep neural network architectures. How to make the improvement of existing
attention module from the perspective of associative memory is worth exploring. And also
how to integrate the associative memory into modern deep neural network architectures is
an interesting task. The third challenge is to apply associative learning to different types of
data (e.g., text, graph, tables and so on). The network has to learn to extract the feature,
and learn to distinguish and associate between different features. The network also has to
let the memories remember diverse patterns to increase the model capacity given the limited
resources.
In this dissertation, we propose solutions to address the challenges mentioned above.
Specifically, to address the first challenge, we explore different ways to train the associative
memory networks. In our text application, we propose a way to learn the memory of the
model using a local update rule, and in our graph application, we focus on training the whole
network using backpropagation. To address the second challenge, we propose two different
ways to apply Modern Hopfield Networks in graph applications (for structure graph and
featured graph, respectively). To address the third challenge, we propose novel approaches
for associative learning in both text data and graph data. Our associative learning based
model is competitive with existing deep learning architectures, and can also help us interpret
the network from the angle of information retrieval and completion.Ph