Archive ouverte HAL-LAAS
Not a member yet
12189 research outputs found
Sort by
Score-Aware Policy-Gradient and Performance Guarantees using Local Lyapunov Stability
International audienceIn this paper, we introduce a policy-gradient method for model-based reinforcement learning (RL) that exploits a type of stationary distributions commonly obtained from Markov decision processes (MDPs) in stochastic networks, queueing systems, and statistical mechanics. Specifically, when the stationary distribution of the MDP belongs to an exponential family that is parametrized by policy parameters, we can improve existing policy gradient methods for average-reward RL. Our key identification is a family of gradient estimators, called score-aware gradient estimators (SAGEs), that enable policy gradient estimation without relying on value-function estimation in the aforementioned setting. We show that SAGE-based policy-gradient locally converges, and we obtain its regret. This includes cases when the state space of the MDP is countable and unstable policies can exist. Under appropriate assumptions such as starting sufficiently close to a maximizer and the existence of a local Lyapunov function, the policy under SAGE-based stochastic gradient ascent has an overwhelming probability of converging to the associated optimal policy. Furthermore, we conduct a numerical comparison between a SAGE-based policy-gradient method and an actor–critic method on several examples inspired from stochastic networks, queueing systems, and models derived from statistical physics. Our results demonstrate that a SAGE-based method finds close–to–optimal policies faster than an actor–critic method
Online Matching via Reinforcement Learning: An Expert Policy Orchestration Strategy
Online matching problems arise in many complex systems, from cloud services and online marketplaces to organ exchange networks, where timely, principled decisions are critical for maintaining high system performance. Traditional heuristics in these settings are simple and interpretable but typically tailored to specific operating regimes, which can lead to inefficiencies when conditions change. We propose a reinforcement learning (RL) approach that learns to orchestrate a set of such expert policies, leveraging their complementary strengths in a data-driven, adaptive manner. Building on the Adv2 framework (Jonckheere et al., 2024), our method combines expert decisions through advantage-based weight updates and extends naturally to settings where only estimated value functions are available. We establish both expectation and high-probability regret guarantees and derive a novel finite-time bias bound for temporal-difference learning, enabling reliable advantage estimation even under constant step size and non-stationary dynamics. To support scalability, we introduce a neural actor-critic architecture that generalizes across large state spaces while preserving interpretability. Simulations on stochastic matching models, including an organ exchange scenario, show that the orchestrated policy converges faster and yields higher system level efficiency than both individual experts and conventional RL baselines. Our results highlight how structured, adaptive learning can improve the modeling and management of complex resource allocation and decision-making processes
Detecting Anomalies Using Graph Neural Networks: A Review
International audienceAnomaly detection is the process of identifying unusual behaviors in systems. In this wide-ranging field, graph neural networks (GNNs) are highly effective compared to the other proposed approaches in the literature. This article summarizes the representative GNN-based methods for anomaly detection and proposes a novel taxonomy based on how these methods predict anomalies.La détection d’anomalies vise à identifier des comportements atypiques au sein des systèmes complexes. Parmi les différentes approches développées dans ce domaine, les réseaux de neurones graphiques (GNN) se distinguent par leur efficacité. Dans cet article, nous proposons une revue des méthodes fondées sur les GNN pour la détection d’anomalies, et introduisons une nouvelle taxonomie, construite autour des mécanismes de prédiction d’anomalies utilisés
Static vs. dynamic characterization of p-GaN HEMTs: Discrepancies in electrical characteristics and their dependence on bias history
International audienceQuasi-static electrical characteristics of p-GaN HEMTs fluctuate with bias history. This study evidences that dynamic operation is fortunately highly reproducible without pre-conditioning. The original experimental setup highlights that quasi-static data alone is insufficient for modeling dynamic behavior, while allowing precise detection of discrepancies, enabling improved transient modeling
A Human Colon-based Microphysiological System with Mechanical Microenvironment Control for 3D Imaging
International audienceIn vitro artificial colonic micro-devices that better replicate complex in vivo systems are essential tools for advancing our understanding of human gut (patho)physiology. Micro-physiological systems (MPS) offer controlled environments, enabling precise manipulation of tissue topography, rigidity and nutrient flow. We present the EnView system configured as a human colon-based MPS that faithfully replicates the 3D topography and matrix stiffness of the human colonic environment using an interpenetrating network of polyacrylamide and collagen I hydrogels, and enabling the culture of human colonic epithelium for weeks. The EnView system integrates a microfluidic chamber with active control of apical/luminal and basal/stromal compartments, allowing for in situ imaging and monitoring. Human colonic epithelial Caco-2 cells cultured up to 21 days in this system follow the crypt topography and formed a polarized epithelial monolayer. This innovative MPS recapitulates in vitro a human colon epithelium with its 3D matrix topology and stiffness control, integrates a microfluidic chamber allowing active control of the apical/luminal and basal/stromal compartments by accurate injection, as well as in situ imaging. Teaser Human gut-on-chip reproducing tissue mechanical properties with active microfluidic control for 3D tissue characterization
Microarchitectural signals analysis platform for the implementation of Hardware Security Counters
International audienceDetecting malicious software or hardware behavior during the operation of a computer system requires observables from one or more abstraction layers of the system. This abstraction, however, tends to limit the ability to detect behavioral deviations, especially for attack classes that exploit vulnerabilities very close to the target hardware. Conversely, too low a level of abstraction tends to significantly increase the complexity of the system model, and therefore poses a number of difficulties for the extraction and selection of relevant observables for a given class of attack.Hardware performance counters in particular have been used as an indirect means of observing micro-architecture behavior and detecting software attempting to exploit hardware vulnerabilities. In order to improve the various detection methods, we propose the construction of hardware metrics designed from the outset for security, by studying the correlation between signals from the micro-architecture and the various classes of attack in the literature, targeting both conventional IT and industrial OT systems. By extension, this work aims to detect attacks originating from hardware Trojans, the latter having the effect of changing the behavior of a given micro-architecture
On the optimal control of birhythmic oscillatory PWA systems: an application to the p53-Mdm2 network
Accepted for publication in CDC 2025 - 64th IEEE Conference on Decision and Control, Dec 2025, Rio de Janeiro, BrazilIn this work, we tackle the problem of inducing optimal transfers between the two oscillatory regimes of a birhythmic genetic network, represented through a piecewise affine dynamical system. For that, we resort to an adaptation of Pontryagin's Maximum Principle to the hybrid setting, with a cost function that combines the transfer time and an L¹-control cost. We focus on a two-dimensional PWA model of the p53-Mdm2 network, a well-known tumor suppressor module that represents a key example of birhythmicity naturally found in mammalian cells. The resulting optimal control can be expressed in feedback form, and is able to remove an oscillatory mode of the system, allowing selection between low or high frequency oscillations of the bimodal genetic network
Learning Efficiency Meets Symmetry Breaking
International audienceLearning-based planners leveraging Graph Neural Networks can learn search guidance applicable to large search spaces, yet their potential to address symmetries remains largely unexplored. In this paper, we introduce a graph representation of planning problems allying learning efficiency with the ability to detect symmetries, along with two pruning methods, action pruning and state pruning, designed to manage symmetries during search. The integration of these techniques into Fast Downward achieves a first-time success over LAMA on the latest IPC learning track dataset
Disentangling redundant and synergistic interactions in the alignment between auditory brains and machines
Artificial neural networks (ANNs) have become increasingly useful for modeling how the brain builds representations from the natural world, yet the nature of their representational alignment with dynamic brain activity remains underexplored. Here, we introduce an informationtheoretic framework to decompose representational geometries into redundant and synergistic components using partial information decomposition (PID). Combining magnetoencephalography (MEG) recordings from participants listening to natural sounds, and two soundprocessing ANNs with categorical (CatDNN) and continuous (SemDNN) semantic outputs, we analyze timevarying brain-model alignment for two optimized stimulus sets. For low-agreement stimulus sets, where mutual information between models is minimized, SemDNN reveals higher mutual information with brain activity. PID further shows greater redundancy and synergy for SemDNN, suggesting sustained temporal integration of intermediate semantic features that can potentially afford a more accurate readout of the auditory environment. These results highlight the value of representational decomposition for detailing shared and complementary components of the alignment between brains and ANNs
Sequence‐ and Docking‐Site‐Dependent Contributions to Multi‐Site Phosphorylation of an Intrinsically Disordered MAPK Substrate
International audienceAbstract Protein kinases often rely on docking site motifs to enhance substrate interactions and facilitate phosphorylation. For example, mitogen‐activated protein kinases (MAPKs) utilize D‐ and F‐motifs, which frequently act in concert to enable bipartite substrate binding. While these motifs are known to modulate phosphorylation efficiency, their quantitative impact on target site phosphorylation within long intrinsically disordered substrates remains largely unexplored. Using NMR spectroscopy, JNK1‐dependent phosphorylation of JIP1, a 450‐amino acid disordered substrate, is investigated, identifying eleven phosphosites with distinct phosphorylation efficiencies. By selectively disrupting JNK1 binding to the D‐ and F‐motifs of JIP1, the determinants of phosphorylation efficiency are uncovered. Specifically, it is found that the D‐motif selectively enhances phosphorylation in the C‐terminal direction in a length‐dependent manner, impressively increasing the phosphorylation efficiencies of sites located at sequence distances exceeding 120 amino acids, while the F‐motif primarily promotes phosphorylation of a site located immediately N‐terminal to the F‐motif. Additionally, docking‐site‐independent phosphorylation is observed, whose efficiency is dictated by the intrinsic sequence preference of JNK1, as inferred from motif scores derived from positional scanning peptide arrays. The work highlights how docking site motifs and sequence context synergistically regulate phosphorylation efficiency, emphasizing the critical role of substrate architecture in determining MAPK‐mediated signaling outcomes