119092 research outputs found
Sort by
Attentional Graph Neural Network Is All You Need for Robust Massive Network Localization
In this paper, we design Graph Neural Networks (GNNs) with attention mechanisms to tackle an important yet challenging nonlinear regression problem: massive network localization. We first review our previous network localization method based on Graph Convolutional Network (GCN), which can exhibit state-of-the-art localization accuracy, even under severe Non-Line-of-Sight (NLOS) conditions, by carefully preselecting a constant threshold for determining adjacency. As an extension, we propose a specially designed Attentional GNN (AGNN) model to resolve the sensitive thresholding issue of the GCN-based method and enhance the underlying model capacity. The AGNN comprises an Adjacency Learning Module (ALM) and Multiple Graph Attention Layers (MGALs), employing distinct attention architectures to systematically address the demerits of the GCN-based method, rendering it more practical for real-world applications. Comprehensive analyses are conducted to explain the superior performance of these methods, including a theoretical analysis of the AGNN's dynamic attention property and computational complexity, along with a systematic discussion of their robust characteristic against NLOS measurements. Extensive experimental results demonstrate the effectiveness of the GCN-based and AGNN-based network localization methods. Notably, integrating attention mechanisms into the AGNN yields substantial improvements in localization accuracy, approaching the fundamental lower bound and showing approximately 37% to 53% reduction in localization error compared to the vanilla GCN-based method across various NLOS noise configurations. Both methods outperform all competing approaches by far in terms of localization accuracy, robustness, and computational time, especially for considerably large network sizes
Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
Large-scale Transformer language models (LMs) trained solely on next-token prediction with web-scale data can solve a wide range of tasks after seeing just a few examples. The mechanism behind this capability, known as in-context learning (ICL), remains both controversial and poorly understood. Some studies argue that it is merely the result of memorizing vast amounts of data, while others contend that it reflects a fundamental, symbolic algorithmic development in LMs. In this work, we introduce a suite of investigative tasks and a novel method to systematically investigate ICL by leveraging the full Pythia scaling suite, including interim checkpoints that capture progressively larger amount of training data. By carefully exploring ICL performance on downstream tasks and simultaneously conducting a mechanistic analysis of the residual stream's subspace, we demonstrate that ICL extends beyond mere "memorization" of the training corpus, yet does not amount to the implementation of an independent symbolic algorithm. Our results also clarify several aspects of ICL, including the influence of training dynamics, model capabilities, and elements of mechanistic interpretability. Overall, our work advances the understanding of ICL and its implications, offering model developers insights into potential improvements and providing AI security practitioners with a basis for more informed guideline
Association of the delayed changes in glutamate levels and functional connectivity with the immediate network effects of S-ketamine
Ketamine shows rapid antidepressant effects peaking 24 h after administration. The antidepressant effects may occur through changes in glutamatergic metabolite levels and resting-state functional connectivity (rsFC) within the default mode network (DMN). A multistage drug effect of ketamine has been suggested, inducing acute effects on dysfunctional network configuration and delayed effects on homeostatic synaptic plasticity. Whether the DMN-centered delayed antidepressant-related changes are associated with the immediate changes remains unknown. Thirty-five healthy male participants (25.1 ± 4.2 years) underwent 7 T magnetic resonance spectroscopy (MRS) and resting-state functional magnetic resonance imaging (rsfMRI) before, during, and 24 h after a single S-ketamine or placebo infusion. Changes in glutamatergic measures and rsFC in the DMN node pregenual anterior cingulate cortex (pgACC) were examined. A delayed rsFC decrease of the pgACC to inferior parietal lobe (family-wise error corrected p (pFWEc) = 0.018) and dorsolateral prefrontal cortex (PFC; pFWEc = 0.002) was detected that was preceded by an immediate rsFC increase of the pgACC to medial PFC (pFWEc < 0.001) and dorsomedial PFC (pFWEc = 0.005). Additionally, the immediate rsFC reconfigurations correlated with the delayed pgACC glutamate (Glu) level increase (p = 0.024) after 24 h at trend level (p = 0.067). Baseline measures of rsFC and MRS were furthermore associated with the magnitude of the respective delayed changes (p’s < 0.05). In contrast, the delayed changes were not associated with acute psychotomimetic side effects or plasma concentrations of ketamine and its metabolites. This multimodal study suggests an association between immediate S-ketamine-induced network effects and delayed brain changes at a time point relevant in its clinical context
Lossless multi-scale constitutive elastic relations with artificial intelligence
A seamless and lossless transition of the constitutive description of the elastic response of materials between atomic and continuum scales has been so far elusive. Here we show how this problem can be overcome by using artificial intelligence (AI). A convolutional neural network (CNN) model is trained, by taking the structure image of a nanoporous material as input and the corresponding elasticity tensor, calculated from molecular statics (MS), as output. Trained with the atomistic data, the CNN model captures the size- and pore-dependency of the material’s elastic properties which, on the physics side, derive from its intrinsic stiffness as well as from surface relaxation and non-local effects. To demonstrate the accuracy and the efficiency of the trained CNN model, a finite element method (FEM)-based result of an elastically deformed nanoporous beam equipped with the CNN as constitutive law is compared with that obtained by a full atomistic simulation. The trained CNN model predicts the elasticity tensor in the test dataset with a root-mean-square error of 2.4 GPa (3.0% of the bulk modulus) when compared to atomistic calculations. On the other hand, the CNN model is about 230 times faster than the MS calculation and does not require changing simulation methods between different scales. The efficiency of the CNN evaluation together with the preservation of important atomistic effects makes the trained model an effective atomistically informed constitutive model for macroscopic simulations of nanoporous materials, optimization of nanostructures, and the solution of inverse problems
Five points to check when comparing visual perception in humans and machines
With the rise of machines to human-level performance in complex recognition tasks, a growing amount of work is directed toward comparing information processing in humans and machines. These studies are an exciting chance to learn about one system by studying the other. Here, we propose ideas on how to design, conduct, and interpret experiments such that they adequately support the investigation of mechanisms when comparing human and machine perception. We demonstrate and apply these ideas through three case studies. The first case study shows how human bias can affect the interpretation of results and that several analytic tools can help to overcome this human reference point. In the second case study, we highlight the difference between necessary and sufficient mechanisms in visual reasoning tasks. Thereby, we show that contrary to previous suggestions, feedback mechanisms might not be necessary for the tasks in question. The third case study highlights the importance of aligning experimental conditions. We find that a previously observed difference in object recognition does not hold when adapting the experiment to make conditions more equitable between humans and machines. In presenting a checklist for comparative studies of visual reasoning in humans and machines, we hope to highlight how to overcome potential pitfalls in design and inference
Analysis of Schedule and Layout Tuning for Sparse Matrices With Compound Entries on GPUs
Large sparse matrices with compound entries, i.e. complex and quaternionic matrices as well as matrices with dense blocks, are a core component of many algorithms in geometry processing, physically based animation and other areas of computer graphics. We generalize several matrix layouts and apply joint schedule and layout autotuning to improve the performance of the sparse matrix‐vector product on massively parallel graphics processing units. Compared to schedule tuning without layout tuning, we achieve speedups of up to 5.5 ×. In comparison to cuSPARSE, we achieve speedups of up to 4.7 ×
Tailoring the Switching Dynamics in Yttrium Oxide‐Based RRAM Devices by Oxygen Engineering: From Digital to Multi‐Level Quantization toward Analog Switching
This work investigates the transition from digital to gradual or analog resistive switching in yttrium oxide‐based resistive random‐access memory devices. It is shown that this transition is determined by the amount of oxygen in the functional layer. A homogeneous reduction of the oxygen content not only reduces the electroforming voltage, allowing for forming‐free devices, but also decreases the voltage operation window of switching, thereby reducing intra‐device variability. The most important effect as the dielectric becomes substoichiometric by oxygen engineering is that more intermediate (quantized) conduction states are accessible. A key factor for this reproducibly controllable behavior is the reduced local heat dissipation in the filament region due to the increased thermal conductivity of the oxygen depleted layer. The improved accessibility of quantized resistance states results in a semi‐gradual switching both for the set and reset processes, as strongly desired for multi‐bit storage and for an accurate definition of the synaptic weights in neuromorphic systems. A theoretical model based on the physics of mesoscopic structures describing current transport through a nano‐constriction including asymmetric potential drops at the electrodes and non‐linear conductance quantization is provided. The results contribute to a deeper understanding on how to tailor materials properties for novel memristive functionalities
Einfluss von Werkstoffzustand und chemischer Zusammensetzung auf die Eigenschaften plasmanitrierter austenitischer Stähle
Plasmanitrieren bietet großes Potenzial, die Verschleißeigenschaften von austenitischen Stählen zu verbessern. Es werden neben den Werkstoffen 1.4307 und 1.4404 die Titan-stabilisierten Güten 1.4541 und 1.4571 betrachtet, um insbesondere den Einfluss der Titanstabilisierung auf das Nitrierergebnis und das Korrosionsverhalten zu untersuchen. Außerdem wurde untersucht, inwieweit der Herabsetzung der Korrosionseigenschaften durch kaltumformungs-induzierte Defektstrukturen durch die Titanstabilisierung begegnet werden kann. Im Vergleich zu den Werkstoffen 1.4307 und 1.4404 wird bei beiden titanstabilisierten austenitischen Stählen weniger Stickstoff in Bereichen mit Umformmartensit und niedrigen Nitriertemperaturen eingebaut, während an Gleitbändern eine erhöhte Eindiffusion von Stickstoff zu beobachten ist. Die Korrosionsbeständigkeit verbessert sich generell durch die hier verwendeten Plasmanitrierparameter. Generell bewirkt eine höhere Dicke der beim Plasmanitrieren erzeugten S-Phase eine bessere Korrosionsbeständigkeit sowie eine höhere Oberflächenhärte. Die Titanstabilisierung hemmt die Stickstoffdiffusion bei hohen Umformmartensitgehalten und niedrigen Nitriertemperaturen und fördert die Diffusion an Gleitlinien