arXiv.org e-Print Archive

arXiv.org e-Print Archive
Not a member yet
    623509 research outputs found

    Exploring and Benchmarking the Planning Capabilities of Large Language Models

    No full text
    Classical and natural language planning tasks remain a difficult domain for modern large language models (LLMs). In this work, we lay the foundations for improving planning capabilities of LLMs. First, we construct a comprehensive benchmark suite encompassing both classical planning benchmarks and natural language scenarios. This suite includes algorithms to methodically generate instances of tasks with varying levels of difficulty, allowing for rigorous and systematic evaluation of LLM performance. Next, we investigate the use of many-shot in-context learning to enhance LLM planning, exploring the relationship between increased context length and improved planning performance. In addition, we demonstrate the positive impact of fine-tuning LLMs on optimal planning paths. We also probe the efficacy of chain-of-thought reasoning methods to improve LLM planning performance. Moreover, we probe the performance of the proposed methods in out-of-distribution scenarios, assessing the ability to generalize to novel and unseen planning challenges. Finally, we investigate model\u27s failure modes and reveal insights that hold true across different benchmarks

    An Agglomerative Clustering of Simulation Output Distributions Using Regularized Wasserstein Distance

    No full text
    Using statistical learning methods to analyze stochastic simulation outputs can significantly enhance decision-making by uncovering relationships between different simulated systems and between a system\u27s inputs and outputs. We focus on clustering multivariate empirical distributions of simulation outputs to identify patterns and trade-offs among performance measures. We present a novel agglomerative clustering algorithm that utilizes the regularized Wasserstein distance to cluster these multivariate empirical distributions. This framework has several important use cases, including anomaly detection, pre-optimization, and online monitoring. In numerical experiments involving a call-center model, we demonstrate how this methodology can identify staffing plans that yield similar performance outcomes and inform policies for intervening when queue lengths signal potentially worsening system performance

    Lower-mass-gap Black Holes in Dense Star Clusters

    No full text
    The existence of compact stellar remnants in the mass range 25M2-5\,M_{\odot} has long been debated. This so-called lower mass gap was initially suggested by the lack of low-mass X-ray binary observations with accretors about 25M2-5\,M_{\odot}, but it has recently been called into question following newer observations, including a lower-mass-gap candidate with a millisecond pulsar companion in the dense globular cluster NGC 1851. Here we model NGC 1851 with a grid of similar dense star clusters utilizing the state-of-the-art Monte Carlo NN-body code \texttt{CMC}, and we specifically study the formation of lower-mass-gap black holes. We demonstrate that both massive star evolution and dynamical interactions can contribute to forming lower-mass-gap black holes. In general, the collapse of massive remnants formed through mergers of neutron stars or massive white dwarfs produces the largest number of lower-mass-gap black holes among all formation channels. However, in more massive clusters, supernova core collapse can contribute comparable numbers. Our NGC 1851-like models can reproduce millisecond pulsar -- lower-mass-gap black hole binaries similar to the observed system. Additionally, the lower-mass-gap black holes can also become components of dynamically assembled binaries, and some will be in merging black hole - neutron star systems similar to the recently detected gravitational wave source GW230529. However, the corresponding merger rate is probably 1 Gpc3yr1\lesssim 1~{\rm Gpc^{-3}\,yr^{-1}}.18 pages, 7 figures, 3 tables. Published at Ap

    Reciprocal Learning

    No full text
    We demonstrate that a wide array of machine learning algorithms are specific instances of one single paradigm: reciprocal learning. These instances range from active learning over multi-armed bandits to self-training. We show that all these algorithms do not only learn parameters from data but also vice versa: They iteratively alter training data in a way that depends on the current model fit. We introduce reciprocal learning as a generalization of these algorithms using the language of decision theory. This allows us to study under what conditions they converge. The key is to guarantee that reciprocal learning contracts such that the Banach fixed-point theorem applies. In this way, we find that reciprocal learning algorithms converge at linear rates to an approximately optimal model under relatively mild assumptions on the loss function, if their predictions are probabilistic and the sample adaption is both non-greedy and either randomized or regularized. We interpret these findings and provide corollaries that relate them to specific active learning, self-training, and bandit algorithms.Accepted at NeurIPS 2024. v2: fixed typos, added future work. v3: changed def. 4 and proof of thm. 4, added illustration

    On Diameters of Cayley Graphs over Matrix Groups

    No full text
    We establish for the matrix group G=SLn(Fp)G=\mathrm{SL}_{n}\left(\mathbb{F}_{p}\right) that there exist absolute constants c(0,1)c\in\left(0,1\right) and C>0C>0 such that any symmetric generating set AA, with AG1c\left|A\right|\geq\left|G\right|^{1-c} has a covering number Cn2.\leq Cn^{2}. This result is sharp up to the value of the constant C>0C>0

    Recovering chemical bimodalities in observed edge-on stellar disks: insights from AURIGA simulations

    No full text
    We assessed the ability to recover chemical bimodalities in integral-field spectroscopy (IFS) observations of edge-on galaxies, using 24 Milky Way-mass galaxies from the AURIGA zoom-in cosmological simulations. We first analyzed the distribution of single stellar particles in the [Mg/Fe] - [Fe/H] plane. Then we produced mock IFS [Mg/Fe] and [Fe/H] maps of galaxies seen edge on, and considered integrated stellar-population properties (projected and spatially binned). We investigated how the distribution of stars in the [Mg/Fe] - [Fe/H] plane is affected by edge-on projection and spatial binning. Bimodality is preserved while distributions change their shapes. Naturally, broad distributions of individual star particles are narrowed into smaller [Mg/Fe] and [Fe/H] ranges for spatial bins. We observe continuous distributions, bimodal in most cases. The overlap in [Fe/H] is small, and different [Mg/Fe] components show up as peaks instead of sequences (even when the latter are present for individual particles). The larger the spatial bins, the narrower the [Mg/Fe] - [Fe/H] distribution. This narrowing helps amplify the density of different [Mg/Fe] peaks, often leading to a clearer bimodality in mock IFS observations than for original star particles. We have also assessed the correspondence of chemical bimodalities with the distinction between geometric thick and thin disks. Their individual particles have different distributions but mostly overlap in [Mg/Fe] and [Fe/H]. However, integrated properties of geometric thick and thin disks in mock maps do mostly segregate into different regions of the [Mg/Fe] - [Fe/H] plane. In bimodal distributions, they correspond to the two distinct peaks. Our results show that this approach can be used for bimodality studies in future IFS observations of edge-on external galaxies.12 pages, 6 figures, accepted for publication in Astronomy & Astrophysic

    Hypergeometric Moments and Hecke Trace Formulas

    No full text
    Moments for hypergeometric functions over finite fields were studied in the work of Ono, Pujahari, Saad, and Saikia for several 2F1_{2}F_{1} and 3F2_{3}F_{2} cases. We generalize their work to prove results for new cases where the hypergeometric data is defined over Q\mathbb{Q} and primitive. These new moments are established using Hecke trace formulas of hypergeometric origin recently established by Hoffman, Li, Long, and Tu. We also obtain several algebraic formulas in the finite field setting and present conjectures for additional 2F1_{2}F_{1} and 3F2_{3}F_{2} moments.The exposition has improved and minor errors were fixe

    From Tokens to Materials: Leveraging Language Models for Scientific Discovery

    No full text
    Exploring the predictive capabilities of language models in material science is an ongoing interest. This study investigates the application of language model embeddings to enhance material property prediction in materials science. By evaluating various contextual embedding methods and pre-trained models, including Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformers (GPT), we demonstrate that domain-specific models, particularly MatBERT significantly outperform general-purpose models in extracting implicit knowledge from compound names and material properties. Our findings reveal that information-dense embeddings from the third layer of MatBERT, combined with a context-averaging approach, offer the most effective method for capturing material-property relationships from the scientific literature. We also identify a crucial tokenizer effect, highlighting the importance of specialized text processing techniques that preserve complete compound names while maintaining consistent token counts. These insights underscore the value of domain-specific training and tokenization in materials science applications and offer a promising pathway for accelerating the discovery and development of new materials through AI-driven approaches

    Little Giants: Synthesizing High-Quality Embedding Data at Scale

    No full text
    Synthetic data generation has become an increasingly popular way of training models without the need for large, manually labeled datasets. For tasks like text embedding, synthetic data offers diverse and scalable training examples, significantly reducing the cost of human annotation. However, most current approaches rely heavily on proprietary models like GPT-4, which are expensive and inefficient for generating large-scale embedding data. In this paper, we introduce SPEED, a framework that aligns open-source small models (8B) to efficiently generate large-scale synthetic embedding data. Through supervised fine-tuning, preference optimization, and self-improvement, SPEED enables small open-source models to produce high-quality data. Remarkably, SPEED uses only less than 1/10 of the GPT API calls, outperforming the state-of-the-art embedding model E5_mistral when both are trained solely on their synthetic data. Using this efficient generator, we conduct a comprehensive study on how various factors within the alignment pipeline impact data quality and reveal the scaling law for synthetic embedding data

    Planning-Aware Diffusion Networks for Enhanced Motion Forecasting in Autonomous Driving

    No full text
    Autonomous driving technology has seen significant advancements, but existing models often fail to fully capture the complexity of multi-agent environments, where interactions between dynamic agents are critical. To address this, we propose the Planning-Integrated Forecasting Model (PIFM), a novel framework inspired by neural mechanisms governing decision-making and multi-agent coordination in the brain. PIFM leverages rich contextual information, integrating road structures, traffic rules, and the behavior of surrounding vehicles to improve both the accuracy and interpretability of predictions. By adopting a diffusion-based architecture, akin to neural diffusion processes involved in predicting and planning, PIFM is able to forecast future trajectories of all agents within a scenario. This architecture enhances model transparency, as it parallels the brain\u27s method of dynamically adjusting predictions based on external stimuli and other agents\u27behaviors. Extensive experiments validate PIFM\u27s capacity to provide interpretable, neuroscience-driven solutions for safer and more efficient autonomous driving systems, with an extremely low number of parameters.The experimental verification section has some issue

    375,182

    full texts

    623,509

    metadata records
    Updated in last 30 days.
    arXiv.org e-Print Archive is based in United States
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇