Archive ouverte HAL-LAAS
Not a member yet
    12189 research outputs found

    Policy Optimization via Adv2: Adversarial Learning on Advantage Functions

    No full text
    International audienceWe revisit the reduction of learning in adversarial Markov decision processes [MDPs] to adversarial learning based on QQ--values; this reduction has been considered in a number of recent articles as one building block to perform policy optimization. Namely, we first consider and extend this reduction in an ideal setting where an oracle provides value functions: it may involve any adversarial learning strategy (not just exponential weights) and it may be based indifferently on QQ--values or on advantage functions. We then present two extensions: on the one hand, convergence of the last iterate for a vast class of adversarial learning strategies (again, not just exponential weights), satisfying a property called monotonicity of weights; on the other hand, stronger regret criteria for learning in MDPs, inherited from the stronger regret criteria of adversarial learning called strongly adaptive regret and tracking regret. Third, we demonstrate how adversarial learning, also referred to as aggregation of experts, relates to aggregation (orchestration) of expert policies: we obtain stronger forms of performance guarantees in this setting than existing ones, via yet another, simple reduction. Finally, we discuss the impact of the reduction of learning in adversarial MDPs to adversarial learning in the practical scenarios where transition kernels are unknown and value functions must be learned. In particular, we review the literature and note that many strategies for policy optimization feature a policy-improvement step based on exponential weights with estimated QQ--values. Our main message is that this step may be replaced by the application of any adversarial learning strategy on estimated QQ--values or on estimated advantage functions. We leave the empirical evaluation of these twists for future research

    New method for quantifying power during wheelchair sports propulsion in the field

    No full text
    The importance of accelerating from a standstill is crucial in dynamic wheelchair sports, as it is closely tied to the ability to generate and apply significant power and net horizontal propulsion force. Assessing and quantifying para-athletes’ physical capabilities could enhance training to performance transition. This study aimed to propose a field method for quantifying total wheelchair propulsion forces and output power, while exploring the usability of the 1080 Motion Sprint. Five para-athletes from the national French wheelchair racing team and seven wheelchair tennis players from the national French team participated. Unloaded and resisted sprints of 50 m and 20 m were performed. Mono-exponential velocity function was deduced using photocells, IMUs, and the 1080 Motion Sprint velocity-time raw data. Net horizontal propulsion force was estimated from Newton’s second law and considered the loads applied by the 1080 Motion Sprint, rolling resistance, and aerodynamic drag. While no significant difference was observed between conditions for theoretical maximal force and maximum power developed, variations were evident in estimated power output and mechanical variables from force-velocity relationships, contingent on the athlete’s classification and sport specialty. The developed protocol can be used by trainers to assess physical capacities during training sessions, guiding subsequent training

    Flow-shop and job-shop robust scheduling problems with budgeted uncertainty

    No full text
    International audienceIn this paper, we study different solution methods for two two-stage robust, multi-machine scheduling problems, namely permutation flow-shop and job-shop scheduling problems under uncertainty budget. Compact formulations of the problems are proposed and two decomposition approaches are presented: a Benders decomposition approach and a column and constraint generation approach. Computational experiments show that for small-sized instances, a compact formulation of the problem quickly yields optimal solutions. However, for larger instances, decomposition methods, particularly the column and constraint generation method with a master problem solved using constraint programming, provide better quality solutions. An acceleration method for the column and constraint generation algorithm is proposed. This method is generic and can be applied to any two-stage robust optimisation problem

    Sensitivity-Aware Model Predictive Control for Robots with Parametric Uncertainty

    No full text
    International audienceThis paper introduces a computationally efficient robust Model Predictive Control (MPC) scheme for controlling nonlinear systems affected by parametric uncertainties in their models. The approach leverages the recent notion of closedloop state sensitivity and the associated ellipsoidal tubes of perturbed trajectories for taking into account online time-varying restrictions on state and input constraints. This makes the MPC controller "aware" of potential additional requirements needed to cope with parametric uncertainty, thus significantly improving the tracking performance and success rates during navigation in constrained environments. One key contribution lies in the introduction of a computationally efficient robust MPC formulation with a comparable computational complexity to a standard MPC (i.e., an MPC not explicitly dealing with parametric uncertainty). An extensive simulation campaign is presented to demonstrate the effectiveness of the proposed approach in handling parametric uncertainties and enhancing task performance, safety, and overall robustness. Furthermore, we also provide an experimental validation that shows the feasibility of the approach in real-world conditions and corroborates the statistical findings of the simulation campaign. The versatility and efficiency of the proposed method make it therefore a valuable tool for real-time control of robots subject to non-negligible uncertainty in their models

    Shannon-and von neumann-entropy regularizations of linear and semidefinite programs: Entropy regularization of LP and SDP

    No full text
    International audienceWe consider the LP in standard form min {c T x : Ax = b; x ≥ 0} and inspired by ε-regularization in Optimal Transport, we introduce its ε-regularization "min {c T x + ε f (x) : Ax = b; x ≥ 0}" via the (convex) Boltzmann-Shannon entropy f (x) := i x i ln x i . We also provide a similar regularization for the semidefinite program "min {Tr(C • X) : A(X) = b; X 0}" but with now the so-called Von Neumann entropy, as in Quantum Optimal Transport. Importantly, both are not barriers of the LP and SDP cones respectively. We show that this problem admits an equivalent unconstrained convex problem max λ∈R m Gε(λ) for an explicit concave differentiable function Gε in dual variables λ ∈ R m . As ε goes to zero, its optimal value converges to the optimal value of the initial LP. While it resembles the log-barrier formulation of interior point algorithm for the initial LP, it has a distinguishing advantage. Namely for fixed λ, Gε(λ) is obtained as a minimization over the whole space x ∈ R d (and not over x ≥ 0) to still obtain a nonnegative solution x(λ) ≥ 0, whence an explicit form of Gε very useful for its unconstrained maximization over R m

    Experimental transients of detonation re-initiation behind a single-hole obstacle: the effects of the tube cross-section shape

    No full text
    International audienceWe describe the scenarios of detonation re-initiation downstream of a single-hole obstacle in a straight tube depending on its cross-sectional shape, namely square or round. The tubes have the same crosssectional area of 16 cm 2 , and the holes have the same shape as that of the tubes, but different open area ratios. The reactive mixtures are 2 H 2 + O 2 + 2 Ar and 2 H 2 + O 2 . We used wall soot recordings, highspeed shadowgraphy or chemiluminescence imaging to obtain parietal and frontal views of the diffraction phenomena. We identify six behaviors, one subcritical, four critical, and one supercritical, depending on the initial pressure. The critical and supercritical behaviors are more easily obtained in the square tube than in the round tube, other parameters being equal. These transient effects of the cross-sectional shape can be observed where no effect would be observed for conditions of steady propagation. Inductively, this leads to discuss what dimensionality refers to experimentally for the global and local transients of detonation dynamics

    ACPAC bilan 2024

    No full text
    National audienceBilan des activités 2024 de l'action concertée packaging du PEPR électronique, présenté lors des journées scientifiques annuelles de ce PEP

    PIF: A PI-based AQM that ensures performance isolation in multi tenant cloud data centers

    No full text
    International audienceManaging fairness and performance isolation in multi-tenant cloud data centers is challenging using traditional flow-based methods like pure TCP settings without AQMs, which do not account for tenants generating multiple-flow traffic. Motivated by this observation, we explore PID (Proportional-Integral-Derivative) based AQMs and address this issue by first identifying and validating the system model through theoretical Analysis and simulations. We then implement a standard PI controller-based AQM and evaluate its performance, finding it insufficient in handling aggressive multi-flow tenant workloads.To improve this, we introduce a Modified PI Controller (PIF)based AQM with adaptive mechanisms tailored for multi-tenant environments. Simulations show that the PIF controller significantly enhances fairness and isolation compared to default settings without AQM control, offering a robust solution for effective resource management in complex cloud data centers.</p

    Impact of environmental factors on photovoltaic system performance degradation

    No full text
    International audienceThe rapid expansion of photovoltaic (PV) systems underscores the need to understand environmental factors affecting their performance, degradation, and economic viability. This study comprehensively reviews 175 articles, classifying environmental factors such as atmospheric deposits (dust, sea salt, pollen), meteorological conditions (wind, temperature, humidity, rainfall, snowfall, hailstorms), shading, and solar irradiation variability. A novel multilevel classification of degradation modes is introduced, identifying failure mechanisms and their impacts. Key findings reveal performance losses of up to 60%-70% due to combined factors, while mitigation strategies, such as wind-induced cooling, can improve power output by 14.25%, and snow accumulation results in up to 12% annual energy losses. Performance metrics like Performance Loss Rate (PLR) and Degradation Rate (DR) are evaluated to quantify long-term impacts, with economic implications including potential revenue losses and maintenance costs. For instance, addressing dust accumulation in arid regions could save 20%-30% in annual cleaning costs while reducing energy inefficiencies. Recent advancements in AI-driven predictive maintenance are highlighted as pivotal for optimizing system performance and minimizing costs. This integrated analysis provides actionable insights for researchers, engineers, and policymakers, emphasizing the need for tailored strategies to enhance PV resilience and economic sustainability. By addressing the interaction of environmental factors and introducing standardized metrics, this study fills critical research gaps, offering a roadmap for improving PV system reliability, reducing operational costs, and supporting the transition to sustainable energy under diverse environmental conditions

    The Decisive Role of Gate Voltage Overshoots on the Ron of p-GaN HEMTs in Realistic Dynamic Operation

    No full text
    International audienceAn original setup is used to explore the impact of gate voltage spikes on p-GaN HEMTs in parasitic-free dynamic operation. The component is maintained on after the voltage spikes, unlike in state-of-the-art studies, which reveals previously unobserved consequences of the spikes. Dynamic-only trapping effects are observed, and the spikes increase the Ron by up to 50 times

    0

    full texts

    12,189

    metadata records
    Updated in last 30 days.
    Archive ouverte HAL-LAAS
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇