UARK (University of Arkansas )
Not a member yet
19829 research outputs found
Sort by
The Returns to Education over time and the Effect of COVID-19
This paper examines the effects of the COVID-19 pandemic on the returns to education in the United States. Using data from the Current Population Survey 2011-2022, the analysis reveals that, after a period of decline, returns to education increased significantly because of COVID, particularly for men and those with university education. The returns to university for men increased by 1 percentage points. The results underscore the importance of continued investment in education to mitigate the adverse effects of future crises
Causal Discovery in Time Series Data using Deep Learning Techniques
Causal structure learning from observational data has been an active field of research over the past decades. In the literature, different algorithms and models have been proposed, such as constrained-based methods and score-based methods including the emerging deep learning-based methods. However, most of the approaches apply to static and non-dynamic data only. In many applications, the data is temporal. For example, monitoring systems, weather surveillance systems, and stock data, to name but a few. Incorporating temporal information is an important extension of the causal discovery field. With the growth of observational data these days, the discovery of causal relationships from time series has become possible by observing their behavior over time. However, existing methods for time series causal discovery usually suffer from one or more of the following limitations: (1) assume acyclicity, i.e., assume that the self-causes do not exist or always exist; (2) apply only to discrete time series; (3) based on linear models only; and (4) assume stationarity, i.e., causal dependencies are repeated with the same time lag at all time points. Thus, this thesis addresses the problem of learning a summary causal graph from multivariate time series data that addresses all the above limitations. We present causal discovery algorithms based on both constraint-based and score-based approaches. The goal of this dissertation is to evaluate different techniques of time-series causal discovery for multivariate data, that determine a summary causal graph which is a directed graph representing the underlying causal relationships among the variables in the data, without specifying the exact time lags. In this dissertation, we attempt to leverage the power of deep neural networks and incorporate them into causal discovery strategies including the constraint-based approach, the score-based and noise-based approach, the hybrid approach, and an LLM-based approach. For the constraint-based approach, based on the theory of µ-separation, we develop an algorithm called the µ-PC that extends the well-known PC algorithm to the time domain. It uses a conditional independence testing technique with a Recurrent Neural Network (RNN) to be applicable to both discrete and continuous time series. For the score-based approach, we develop a model named Neural Time-invariant Causal Discovery (NTiCD), which is based on the principle of Granger causality. NTiCD is a continuous optimization-based technique that leverages the power of deep neural networks to compute the score values. To this end, we use an LSTM to obtain the hidden non-linear representations of temporal variables in the time series data. Then, these features are aggregated using graph convolutional networks and decoded using an MLP that outputs the forecast of the future data values in the time series. The model is optimized based on a score function subject to regularized loss. The final output is a summary causal graph that captures the time-invariant causal relations within and between time series. Next, we propose a hybrid approach named Neural-HATS (Neural Hybrid Approach for Time Series Causal Discovery), an innovative framework that combines conditional independence (CI) testing with continuous optimization-based learning algorithms to uncover causal structures in time series data. The approach features an attention-based encoder-decoder architecture integrated with Kernel Conditional Independence (KCI) testing, enabling direct CI tests between time series. These tests are then integrated into continuous optimization learning algorithms for enhanced causal discovery. This integration not only refines the causal inference process but also expands the capabilities of continuous optimization algorithms, significantly improving their performance without the need for extensive computational resources. We evaluate the performance of all our algorithms on several synthetic and real datasets. Following these, next, we implement a noise-based approach called Granger Causal discovery using Residual Independence (GCRI). GCRI framework uses an autoencoder-based approach to uncover Granger causal relationships in time series and the distribution of exogenous variables. Here, we assume that the exogenous variables are mutually independent and impose constraints in the loss function of the encoder to ensure this independence. The encoder models abductive reasoning to derive mutually independent exogenous variables, while the decoder applies deductive reasoning to predict inputs using a limited window of past exogenous variables and time series data. In this way, GCRI effectively performs Granger causal discovery on multivariate time series data using a noise-based approach, where we show the performance using several synthetic and real datasets. We also demonstrate GCRI\u27s effectiveness in identifying the root cause of anomalies, presenting it through a case study. By modeling the distribution of exogenous variables in multivariate time series data, GCRI is able to detect exogenous interventions as the root cause of anomalies. Using synthetic non-linear time series, our framework accurately localizes the source of the anomaly, demonstrating high precision in root cause analysis. This case study underscores the practical applicability of GCRI for anomaly detection in complex time series data. Notably, all our above-mentioned proposed algorithms do not assume data linearity, stationarity, or acyclicity. Finally, we propose a framework that leverages large-language models (LLMs) to enhance the performance of both score-based and constraint-based causal discovery methods. Our LLM-guided initialization approach integrates LLMs with data-driven algorithms, allowing for more precise causal discovery while incorporating valuable domain knowledge. By utilizing LLMs, we generate an initial causal structure from real-world observational data, ensuring that expert knowledge is embedded without violating the core principles of temporal causal discovery. The LLM processes the dataset and its description to create an initial causal graph, which is then refined using traditional temporal causal discovery methods, producing a more accurate, robust, and interpretable causal structure
Indigenous Materials Mixtures and AC Parameters
This research focuses on the design and development of indigenous materials for Additive Construction (AC) mixtures. AC uses 3D printing techniques to extrude concrete layer-by-layer from a digital model to create 3D printed objects. In construction, AC can improve efficiency, speed, and sustainability, particularly in remote or resource-constrained environments and significantly reduce logistical demands and environmental impacts associated with traditional material transport. Key research objectives include identifying suitable additives for achieving printable soil-based mixtures and establishing practical, field-appropriate methods for assessing fresh properties that indicate printability. The findings highlight factors such as water content, cement content, additive type, and quantity and type of fines (i.e., particles smaller than 0.075 mm) affect a mix’s printability. Fines tended to improve the printability of a mix; however, they did cause a reduction in the compressive strength. The flow table test in combination with the hand squeeze test were shown to be simple, field-ready methods that could assess the pumpability of a mix. The values targeted did depend on the type of mix and the equipment used. A small batch flow chart was also developed to guide decisions regarding printability, incorporating flowability ranges and compressive strength data for different custom mixes. Future research will focus on refining mix designs, exploring new material combinations, and evaluating scalability to enhance guidelines for testing methods and acceptance thresholds for printability
Predicting Different Psychological Profiles Amongst Substance Users.
Substance use disorders (SUDs) are some of the most prevalent issues facing society. As SUDs are highly individualized in how they manifest, recent work has attempted to classify distinct profiles of drug users to better understand what factors put someone at risk for shifting from casual substance use to developing a substance use disorder (SUD). Common risk factors for SUD development include stressor exposure, difficulties with emotion regulation, personality, and risky decision-making. As many of the behaviors that constitute SUDs are exhibited prior to potential diagnosis, developing “profiles” that combine these traits and experiences could eventually predict SUDs. The current study (n = 266) used supervised machine learning to examine psychological traits and stressful life experiences as predictors into random forests that aimed to classify individuals as users or nonusers of specific substances. Each substance’s use was accurately predicted at a rate above chance. Additionally, the same variables used in the random forests were then considered as predictors of the severity of use of alcohol, cannabis, and nicotine using a more traditional statistical approach (i.e., ANOVAs). High degrees of stress and mismatches between developmental and recent stress, certain personality traits, and emotion regulation strategies were associated factors with severity of substance use. These findings represent a potential avenue for developing individualized profiles of addiction that could help in identifying how substance use escalates in severity
Carbon Dioxide Dynamics in the Carbonate Critical Zone, Savoy Experimental Watershed, Arkansas, USA
The dynamics of carbon dioxide (CO2) within the Earth’s critical zone (CZ) are essential for driving spatial and temporal variations of carbonate dissolution and precipitation, which are fundamental to subsurface porosity and permeability development. However, there is still a significant knowledge gap regarding the patterns of CO2 transport and production within the carbonate critical zone (CCZ). In this study, we investigated these dynamics within a mantled karst terrain at the Savoy Experimental Watershed (SEW) in Northwest Arkansas. Previous monitoring of partial pressure of CO2 (PCO2) and dissolution rates at karst springs in SEW’s Basin 1 has provided valuable insights into the temporal variations in CO2 within a CCZ system. To improve our understanding of the mechanisms that govern CO2 production and transport in the SEW-CCZ system, we analyze the variability of CO2 in areas with high soil CO2 production and elevated concentrations along flow pathways, such as springs and groundwater. Our hypothesis suggests that in karst landscapes, CO2 is generated through biological processes in the vadose zone and is primarily transported by advective mechanisms, resulting in spatially variable distribution patterns. Alternatively, diffusive transport may produce a more gradual and widespread dispersal of CO2 in the karst CZ. To gain a comprehensive understanding of the spatial and temporal dynamics of CO2 in the SEW-CCZ, we instrumented five PCO2 monitoring sites in SEW’s Basin 1 from 2022 – 2024. These monitoring stations span a range of subsurface depths, including “deep” groundwater (10 meters), the “intermediate” soil zone (1.8 meters), and the “shallow” soil zone (60 cm). Additionally, we monitored an epikarst spring and the main karst spring outlet of the basin, known as Langle Spring. Utilizing high-resolution PCO2 monitoring, we compared PCO2 data with hydrometeorological events and hydroclimatic parameters to identify controlling variables and examine their impact on CO2 production and transport in the SEW-CCZ system. Additionally, this study employed simultaneous PCO2 and partial pressure of oxygen (PO2) measurements in the soil zone to further explore how coupled biotic (e.g., soil respiration) and abiotic (e.g., calcite dissolution and precipitation) processes control CO2 dynamics. The PCO2 at Langle Spring are closely linked to the PCO2 in the soil, indicating that soil respiration is the primary source of CO2 in the SEW-CCZ. This study illuminates the ways in which soil moisture and temperature impact root respiration, while also exploring how respiration affects daily, storm-related, and seasonal CO2 variations in CO2 concentrations within the SEW-CCZ
Interaction-Sensitive Tree-Based Statistical Models
This dissertation introduces a tree-based framework to improve the interpretability and modeling of interaction effects among variables, essential in fields like biostatistics, healthcare, science and engineering. Traditional regression methods often fail to clearly capture complex interactions, while tree-based approaches, despite their interpretability, face performance limitations and overfitting concerns. Our proposed interaction-sensitive tree-based method, designed for seamless integration, combines various statistical techniques tailored to different data types, leveraging ensemble learning methods to enhance accuracy and mitigate overfitting. We present methods for regression, survival analysis, and classification, validated with case studies and benchmarked against traditional models using metrics like BIC and R-squared. The results highlight our framework’s balance between interpretability and predictive performance, offering a robust solution for analyzing complex data interactions
A CMOS LDO Voltage Regulator for Low Power Applications
This thesis presents the design, simulation, layout, and testing of a low dropout linear voltage regulator (LDO) in a 180 nm CMOS process. The LDO is intended for use in low-power, battery-operated applications, and it has an adjustable output for a variety of different implementations. It can supply up to 100 mA of current, which is enough to power many small electronic circuits. It has a dropout voltage of 150 mV, and it consumes less than 200 μA of current during operation, which is on par with many commercial regulators. The line regulation is also comparable to commercially available LDOs, at 0.035%, making it stable over a range of different supply conditions. The simulated load regulation is 0.038%, which is also excellent, however, the test circuit degraded the measured load regulation, so the measured figure is higher than expected. This test circuit error is discussed, and some errors resulting from device mismatch in fabrication are also presented and explained
Annual Report [College of Education], 2023-2024
Beginning in 2004/2005- issued in online format onl