12023 research outputs found
Sort by
Core Reinforcement Learning Computations Underlying Distinct Behavioral Strategies and their Implications in Psychiatry
Reinforcement learning (RL) models have shown great capabilities in characterizing learning and decision-making in the real world. The dual systems of the model-free (MF) and model-based (MB) algorithms have been proposed to describe the computational mechanism underlying a reflexive habitual control and a cognitive goal-directed control, respectively. Given the dual systems under control, it is worth asking how the choice of which system to use is made with the changing environmental statistics of rewards and states. In Chapter 2, three types of prediction error signals from the dual systems are found to guide the arbitration process in a reliability-based RL framework. Moreover, an exploratory analysis was conducted to test for alternative arbitration theories that utilize the cost-benefit analysis on the goal-directed (or MB) system. Understanding learning and decision-making would not be complete without knowing how our neural machinery implements these RL computations when a given system is engaged. The robustness and replicability of neural encoding of learning and decision signals from the MF and MB systems are essential to set a reassuring path for future neurocomputational work on dual systems. In Chapter 3, we address recent concerns over the existence of the MF system and its neural computations in a widely-used Markov decision task (two-step task). By applying a model-based functional magnetic resonance imaging (fMRI) approach to a large number of participants, we found both MF and MB learning signals in the human striatum and that neural patterns of decision utility across different RL-strategy groups further add to the evidence of ubiquitous MF computations in Markov decisions. It turns out the framework of dual systems could not only account for normal learning behaviors but also inform us of what actually goes wrong in mental disorders. In Chapter 4, we show that, via the reliability-based arbitration framework, the MF behavioral bias observed in participants with high obsessive-compulsive tendency could be attributed to an enhanced encoding of MB reward prediction error in the anterior cingulate cortex, a region previously implicated in the error-monitoring process. Chapter 1 introduces basic concepts and example algorithms in RL; we also review relevant theoretical and neuroscientific works to build the knowledge base for subsequent chapters. Chapter 5 discusses the significance of empirical findings in this thesis, the values of adopting some of the methodologies herein, potential future research directions on the dual systems, and implications in computational psychiatry
A Multispecies Perspective on the Evolution of Form Vision
In the mammalian visual system, photons captured by the retina are transformed into meaningful internal percepts of surroundings through a hierarchy of interconnected visual areas. Understanding the representation of visual information at each node of the hierarchy has been a central quest of visual systems neuroscience over the past 50 years. The primate visual system, with its over two dozen distinct areas broadly organized into a dorsal stream for visuo-motor transformations and a ventral stream for object recognition, has served as the gold standard for studying the organization of the visual system. Recent advances in artificial neural networks modeled on the primate visual system for object recognition have prompted the question, is hierarchical representation necessary, and if so, can we observe it across all highly visual mammalian species? Hierarchical organization appears to be a key architectural principle of both artificial and biological networks, enabling stepwise construction of a structured and compact representation from raw sensory input. Here we present a series of efforts to determine the cortical organization and connectivity of the tree shrew visual system and directly compare to that of the primate. This cross-species study sheds light on the evolution and mechanisms of vision in a close relative of primates. Using high-density Neuropixels recordings, we demonstrate that the tree shrew ventral visual pathway exhibits primate-like hierarchical processing, with progressively larger receptive fields, increasing response latencies, and enhanced selectivity for complex stimuli along the visual pathway. Area V2 in the tree shrew performs key functions similar to those of the primate inferotemporal (IT) cortex. Specifically, V2 contains strongly face-selective cells, supports a complete representation of high-level object space, and achieves the most accurate object identity decoding and reconstruction among all tree shrew visual areas. Yet we also found significant differences from the canonical template for hierarchical organization observed in the primate, including maintenance of relatively small, focal receptive fields throughout the hierarchy, and better decoding of latent variables in late deep neural network (DNN) layers by area V2 compared to other areas.
The hierarchical organization of the visual system describes the arrangement of areas but does not reveal how information flows between them. Understanding the type of processing carried out at each node raised the next question of whether information that is transmitted across nodes is differentiated between feedforward and feedback connections. To explore this, we combined electrical microstimulation and extracellular recordings to identify the directionality of projections which is applicable in various species. We used this technique to first study the connections between the first two nodes of the tree shrew cortical hierarchy, V1 and V2. We found that V2 feedback neurons carry a full visual representation on par to other V2 cells. These feedback neurons were distinct with regards to their spatial features, including distinct locations and sizes of their receptive fields. We also found that both feedforward and feedback V2 neurons were modulated by perceptual conflict arising when distinct textures were presented to each eye, suggesting they could refine V1 processing to perceptual inconsistencies.
These studies provide insights into how the tree shrew visual system generates object representations through a hierarchy of interconnected nodes, employing strategies adapted to its cortical constraints. In addition, by combining electrical microstimulation with electrophysiology we set the foundation for cross-species studies to determine the role of feedforward and feedback processing along the visual hierarchy. Together, this work reveals conserved principles of visual processing across species while showcasing unique adaptations in the tree shrew, offering insights into the evolutionary origins and functional organization of the primate visual system.</p
Ultrafast Computing with Nonlinear Photonics
Computers have revolutionized almost every facet of modern society, and as we approach the physical limits of digital electronics, it becomes imperative to investigate alternative computing hardware paradigms to enable the next generation of faster and more energy-efficient computers. This thesis embarks on building the foundation for a new kind of computer, based on ultrafast nonlinear photonics, aiming to overcome some of the limitations plaguing current computers. In particular, we primarily focus on the clock rate, which has stagnated at ∼5 GHz for conventional microprocessors over the past two decades.
We begin by identifying single nonlinear devices in lithium niobate nanophotonics that can act as essential building blocks for computers, showing a variety of nonlinear functions with operational speeds > 13 THz for artificial intelligence computing workloads. Then, we progress to small-scale photonic computing circuits combining both strong nonlinearity and memory feedback in a physical reservoir computer for temporal information processing with ∼10 GHz clock rates. Additionally, we explore unconventional computer architectures such as Cellular Automata, which reveals key system-level considerations that maximize the benefits of ultrafast nonlinear photonics in large-scale computers. This culminates in the demonstration of truly end-to-end and all-optical computing with > 100 GHz clock rates, which represents over an order-of-magnitude advancement compared to existing electronic computers. Finally, we prove mathematically how coupled nonlinear optical resonators are Turing-complete computers.
Overall, this work builds on the recent advances in nonlinear photonics and highlights a path for a new class of ultrafast photonic computers that can surpass the clock rate and latency limits of electronic computers, hence enabling nascent applications requiring real-time control or information processing at picosecond timescales.</p
Understanding and Improving Efficiency in Training of Deep Neural Networks
As deep neural networks (DNNs) continue to drive progress in fields like computer vision and natural language processing, their increasing complexity presents significant challenges for training efficiency, particularly in large language models (LLMs). These challenges include memory limitations, energy consumption, and bandwidth constraints during training.
In this thesis, I address these challenges by analyzing the training dynamics of DNNs and proposing hardware-efficient learning algorithms to enhance training efficiency. First, I focus on mitigating memory limitations in LLM training. Training large models like LLMs requires substantial memory for parameters, gradients, and optimizer states, often exceeding standard hardware capacity. To tackle this, I propose GaLore, a memory-efficient training algorithm that reduces the memory footprint of LLM training by up to 65.5% while preserving performance. Additionally, I introduce InRank, an incremental low-rank learning algorithm that further reduces memory usage by gradually increasing matrix rank.
Next, I address the issue of high energy consumption during training. Training large models like LLMs demands considerable energy, contributing to environmental impact. To mitigate this, I propose LNS-Madam, a low-precision training algorithm leveraging the logarithmic number system (LNS) to lower energy consumption without compromising accuracy. LNS-Madam achieves up to 90% energy savings compared to a full-precision baseline model.
Finally, I focus on bandwidth limitations in distributed training. Training LLMs often requires distributing computations across multiple devices to accelerate training. However, network bandwidth constraints can cause communication bottlenecks that slow down training. To resolve this, I introduce signSGD with Majority Vote, a communication-efficient training algorithm that reduces the overhead associated with distributed training.</p
Predictive Modeling of Architected Solids across Scales
Architected solids, comprising discrete or continuous materials and structures, are purposefully designed to achieve specific functional objectives, such as tailored mechanical properties or enhanced performance. The integration of architectural features and material science has revolutionized design and functionality across multiple length scales. However, experimental exploration of architected solids is often constrained by physical, financial, or technological limitations. To address these challenges, this study leverages computational models as powerful tools for validating and probing the behaviors of architected solids through three distinct case studies spanning different length scales.
The first case study focuses on capturing the seismic performance of multiblock concrete structures at CERN for radiation shielding. The Level Set Discrete Element Method (LS-DEM), combined with Monte Carlo sampling of material properties, is employed to benchmark the displacement profiles of four concrete configurations against experimental data. In the second case study, a bonded LS-DEM model is utilized to investigate the bending response of a woven topological interlocking material (TIM). After validation against experimental results, the model is employed to explore how friction and contact area influence the bending resistance of the TIM system. The third case study introduces a 3D translational tensegrity structure modeled using the Finite Element Method (FEM). This model captures the deformation responses of single cells, monolayers, and multicellular spheroids under various loading conditions. Additionally, a data-driven (DD) framework with multiscale analysis is implemented, offering accurate results with enhanced computational efficiency. Through these three case studies, this research illustrates the evolution of computational models from tools for validating known behaviors to frameworks for exploring new phenomena.</p
Towards Hybrid Physics-Machine Learning Parameterizations: Employing Data Assimilation for Online Learning of Turbulence and Convection Closures in a Unified Scheme
Despite advances in climate modeling, the spread in equilibrium climate sensitivity estimates has remained largely unchanged over generations of modeling, mainly due to uncertainties in cloud feedback mechanisms arising from subgrid-scale turbulence, convection, clouds, and the resulting cloud-radiation interactions. Misrepresentations of these processes affect both long-term climate projections and the simulation of short-term atmospheric phenomena, such as the diurnal cycle of precipitation. These limitations are most pronounced in regimes like stratocumulus clouds and their transition to cumulus over ocean basins---areas where climate models have the largest cloud biases in the historical record. This thesis aims to constrain the critical subgrid-scale physics of turbulence and convection by developing and calibrating a hybrid physics–machine learning parameterization using the Eddy-diffusivity Mass-flux (EDMF) framework. By integrating machine learning components into the EDMF and employing data assimilation techniques for online learning, we attempt to directly target some of the processes responsible for uncertainties in cloud feedbacks.
In this thesis, we employ ensemble Kalman inversion within a single-column setup to simultaneously perform online calibration of parameters in empirical closures and embedded neural networks, targeting large-eddy simulations as ground truth. The online learning framework ensures stability and physical consistency, as machine learning components are trained within the context of the full model dynamics. By directly targeting poorly constrained processes like lateral entrainment/detrainment and turbulent mixing lengths, we improve the representation of subgrid-scale fluxes and resulting cloud properties across various atmospheric regimes. We uncover limitations of traditional semi-empirical closures, providing insights for future model development. The calibrated hybrid parameterization outperforms existing schemes, particularly in regions where climate models have historically underperformed, and maintains accuracy in out-of-sample forcings from a warmer climate. This work demonstrates that integrating machine learning with physics-based parameterizations through data assimilation offers a systematic and robust approach for reducing biases in climate models and understanding the physics of elusive subgrid-scale closures
Dynamical Control of Many-Body Interactions in Driven Quantum Matter
Strongly driven Floquet systems have emerged as promising platforms for exotic non-equilibrium physics, but their instability to heating motivates practical questions about how Floquet engineering can be useful. Although drive-induced heating is often attributed to interactions, this thesis adopts a different perspective, identifying regimes where dissipative many-body dynamics can stabilize Floquet physics and define remarkable new drive-tunable properties. This principle enables highly tunable many-body steady states with minimal heating, leading to a novel regime where drive control over single-particle Floquet states can extend to many-body interactions. Our theoretical and experimental results in Parts II and III center around two themes. The first theme focuses on discovering controllable and stable many-body Floquet states. The second explores further into what the future holds--envisioning the prospects for unconventional Floquet physics with nontraditional driving fields and three-dimensional materials.
Part II of this thesis leverages kinematic constraints on low-dimensional many-body scattering as new principles for tuning and stabilizing Floquet phases. First, we predict that a circularly polarized laser can drive slow electrons of moiré systems into a subsonic regime where they decouple from the intrinsic 2D acoustic phonons of the system. This "slow-electron regime" enables optical control over the steady-state occupation of topological Floquet states and the resulting anomalous Hall conductivity. Second, we present experimental transport signatures of steady Floquet physics in graphene irradiated by a continuous-wave laser. Our experiment, performed at 3-4 K lattice temperatures with lasers off-resonant to optical phonons, creates electron-phonon scattering bottlenecks that stabilize persistent low-temperature phases with light-induced longitudinal transport characteristics. The long-lived many-body phase represents the first experimental signatures of steady Floquet physics in a metallic solid.
Part III presents emerging opportunities for many-body Floquet engineering beyond traditional optically-driven, low-dimensional materials. We first explore beyond-optical driving fields, revealing the emergence of quantized charge transport in 1D systems driven by coherent phonons. Incoherent phonons relax electrons into a topological spatiotemporal Floquet state with quantized group velocity set by the coherent phonon, realizing topological charge pumping in a highly non-adiabatic setting. Finally, we address the topological effects of time-periodic drives beyond low-dimensional systems, revealing that THz-frequency, circularly polarized light can induce topological chiral plasmons in Weyl semimetals with band anisotropy, broken time-reversal symmetry, and broken inversion symmetry.
The theoretical and experimental work in this thesis represent key progress towards realizing persistent Floquet physics for diverse applications in quantum device engineering.</p
Computational Methods for Nucleic Acid Structure Prediction and G Protein-Coupled Receptor Mechanism Investigation
Molecular dynamics (MD) simulation is a powerful tool to characterize molecular structure. In this thesis, we use MD to solve three problems in biological chemistry: the prediction of secondary nucleic acid structure, the prediction of the structure of a siRNA-based tool, and the understanding of the activation mechanism of the sweet taste receptor.
First, we use MD to computationally parameterize nucleic acid secondary structure models. Current models are parameterized using experimental data that was limited and time consuming to generate. In this work, we present a workflow to select an ensemble of base pairing reactions using machine learning, perform MD simulations to obtain their free energy and enthalpy profiles, and perform nonlinear regression on the energies and enthalpies to yield thermodynamic nearest neighbor parameters. This computational framework parameterizes secondary structure models with comparable accuracy to experimental data, and can be used to expand the current models to a range of materials and experimental conditions beyond the specific experimental conditions the parameters were originally generated in.
Next, we predict the structure of a RNA therapeutic tool, conditionally activated small interfering RNAs (Cond-siRNAs). We evaluate how two structural modifications to the two double helix topology affect the ability of the construct to perform its function in the RNAi pathway, finding that a short sensor overhang and a carbon chain linker between the two duplexes reduce undesired interactions in the structure. We also characterize the structure for a new topology with a single linker between the duplexes, and show how the position of terminal modifiers causes structural distortions and can disrupt its function. These insights will guide the design of Cond-siRNAs and other RNA tools with similar non-canonical modifications.
Lastly, we investigate the mechanism of activation of the TAS1R2/TAS1R3 sweet taste receptor, coupled to the G protein gustducin. We use metadynamics simulations to estimate the reduction of the free energy of opening the G alpha subunit in the presence of a positive allosteric modulator and steviol glycosides of varying sweetness. We also uncover insights into how the modulator induces the activation of the G alpha, leading to the partial release of GDP; and how the steviol glycosides RebM, RebD and IsoRebM of high sweetness induce an interdomain twist in the Venus Fly Trap domains. These results further our understanding of the activation mechanism of the class C sweet taste receptor and can be used for the development of new sweeteners.</p
Dust in Astrophysical Systems: Impacts on Dynamics, Plasma Physics, and Thermochemistry
The role of astrophysical dust is often simplified in models of galaxy and star formation, where it is typically treated as passively tracing the gas. However, dust actively influences the dynamics, thermodynamics, and observable properties of diverse environments. This thesis explores how explicitly modeling dust grain dynamics and their interactions with gas, radiation, and electromagnetic fields alters the behavior of three key astrophysical regimes: active galactic nuclei (AGN), star-forming giant molecular clouds (GMCs), and the early Universe.
Radiation pressure on dust is widely regarded as a driver of large-scale outflows in AGN, though the dynamics of this interaction remain poorly constrained. Using radiation–dust–magnetohydrodynamic (RDMHD) simulations, we show that radiation efficiently drives supersonic, dust-laden outflows, which are unstable to resonant drag instabilities (RDIs). These instabilities generate turbulence and restructure the dust into clumpy, anisotropic forms, accounting for the torus's patchiness and producing time-variable reprocessed emission in the infrared and optical.
In GMCs, dust dynamics play a crucial role in shaping both the chemistry and thermodynamics of the gas. Leveraging the STARFORGE framework with live dust dynamics and non-equilibrium thermochemistry, we show that stellar radiation redistributes dust, reducing dust accretion during the main mass growth phases and leading to substantial abundance variations among co-natal stars. Statistically, these variations align with observational data, providing an alternative mechanism for driving abundance fluctuations in co-natal stars, beyond interpretations focused solely on post-formation processes like planet accretion. We find that, for a fixed dust mass, grain size variations significantly affect the thermodynamics, influencing local opacity, radiative transport, thermal balance, and ionization structure, thereby suppressing SFE by up to an order of magnitude.
Finally, we introduce a novel mechanism for magnetogenesis in the early Universe via radiatively accelerated, charged dust grains. This "dust battery" takes advantage of the large stopping lengths of dust grains to produce significant charge separation over large distances, thereby driving electric fields and seeding magnetic fields. Unlike conventional mechanisms (e.g., Biermann battery, Weibel instability), which rely on short-range electron-ion separation, this process operates efficiently and generates magnetic fields several orders of magnitude stronger than those produced by traditional mechanisms. We derive the underlying theory, develop a Magnetohydrodynamic-Particle-In-Cell (MHD-PIC) model, and a sub-grid model suitable for implementation in cosmological contexts.
These insights underscore the importance of incorporating dust dynamics into astrophysical models to enhance our understanding of the formation and evolution of galaxies, stars, and the interstellar medium.</p
On Multiple SLE Systems and their Deterministic Limits
In this thesis, we study multiple radial SLE(k) systems -- a family of random multi-curve systems in a simply-connected domain Ω, with marked boundary points z₁....zₙ \in ∂Ω and a marked interior point q, where parameter k > 0 measures the randomness of the system. We also study the multiple radial SLE(0) systems as the deterministic limit of multiple radial SLE(k) systems.
As a consequence of domain Markov property and conformal invariance, we derive that a multiple radial SLE(k) system is characterized by a conformally covariant partition function satisfying the null vector equations--a second-order PDE system. On the other hand, using the Coulomb gas method inspired by conformal field theory, we construct four types of solutions to the null vector equations, which can be classified according to topological link patterns.
We construct the multiple radial SLE(0) systems from stationary relations by heuristically taking the classical limit of partition functions as k > 0. By constructing the field integrals of motion for the Loewner dynamics, we show that the traces of multiple radial SLE(0) systems are the horizontal trajectories of an equivalence class of quadratic differentials. These trajectories have limiting ends at the growth points and form a radial link pattern.
The stationary relations connect the classification of multiple radial SLE(0) systems to the enumeration of critical points of the master function of trigonometric Knizhnik-Zamolodchikov (KZ) equations.
For k > 0$, the partition functions of multiple radial SLE(k) systems correspond to eigenstates of the quantum Calogero-Sutherland (CS) Hamiltonian beyond the fermionic states. In the deterministic case of k=0, we show that the Loewner dynamics with a common parametrization of capacity form a special class of classical CS systems, restricted to a submanifold of phase space defined by the Lax matrix.</p