Portail HAL Ensta
Not a member yet
    11080 research outputs found

    Exploration des algorithmes d'apprentissage par renforcement pour la perception et le controle d'un véhicule autonome par vision

    No full text
    Reinforcement learning is an approach to solve a sequential decision making problem. In this formalism, an autonomous agent interacts with an environment and receives rewards based on the decisions it makes. The goal of the agent is to maximize the total amount of rewards it receives. In the reinforcement learning paradigm, the agent learns by trial and error the policy (sequence of actions) that yields the best rewards.In this thesis, we focus on its application to the perception and control of an autonomous vehicle. To stay close to human driving, only the onboard camera is used as input sensor. We focus in particular on end-to-end training, i.e. a direct mapping between information from the environment and the action chosen by the agent. However, training end-to-end reinforcement learning for autonomous driving poses some challenges: the large dimensions of the state and action spaces as well as the instability and weakness of the reinforcement learning signal to train deep neural networks.The approaches we implemented are based on the use of semantic information (image segmentation). In particular, this work explores the joint training of semantic information and navigation.We show that these methods are promising and allow to overcome some limitations. On the one hand, combining segmentation supervised learning with navigation reinforcement learning improves the performance of the agent and its ability to generalize to an unknown environment. On the other hand, it enables to train an agent that will be more robust to unexpected events and able to make decisions limiting the risks.Experiments are conducted in simulation, and numerous comparisons with state of the art methods are made.L'apprentissage par renforcement est une approche permettant de résoudre un problème de prise de décision séquentielle. Dans ce formalisme, un agent autonome interagit avec un environnement et reçoit des récompenses en fonction des décisions qu'il prend. L'objectif de l'agent est de maximiser le montant total des récompenses qu'il obtient. Dans le paradigme de l'apprentissage par renforcement, l'agent apprend par essais-erreurs la politique (séquence d'actions) qui donne les meilleures récompenses.Dans cette thèse, nous nous concentrons sur son application à la perception et au contrôle d'un véhicule autonome. Pour rester proche des conditions d'un conducteur humain, seule la caméra embarquée est utilisée comme capteur d'entrée. Nous nous focalisons en particulier sur l'apprentissage de bout-en-bout de la conduite, c'est-à-dire une correspondance directe entre les informations provenant de l'environnement et l'action choisie par l'agent. Ce type d'apprentissage pose cependant certains défis : les grandes dimensions des espaces d'états et d'actions ainsi que l'instabilité et la faiblesse du signal de l'apprentissage par renforcement pour entraîner des réseaux de neurones profonds.Les approches que nous avons mises en oeuvre pour faire face à ces défis reposent sur l'utilisation de l'information sémantique (segmentation d'images). En particulier, nous explorons l'apprentissage conjoint de l'information sémantique et de la navigation.Nous montrons que ces méthodes sont prometteuses et permettent de lever certains verrous. D'une part combiner l'apprentissage supervisé de la segmentation à l'apprentissage par renforcement de la navigation améliore les performances de l'agent, ainsi que sa capacité à généraliser à un environnement inconnu. D'autre part, cela permet d'entraîner un agent qui sera plus robuste aux évènements inattendus et capable de prendre des décisions en limitant les risques.Les expériences sont menées en simulation, et de nombreuses comparaisons avec les méthodes de l'état de l'art sont effectuées

    EZIOTracer: unifying kernel and user space I/O tracing for data-intensive applications

    No full text
    International audienceTracing is a popular method for evaluating, investigating, and modeling the performance of today's storage systems. Tracing has become crucial with the increase in complexity of modern storage applications/systems, that are manipulating an ever-increasing amount of data and are subject to extreme performance requirements. There exists many tracing tools focusing either on the user-level or the kernel-level, however we observe the lack of a unified tracer targeting both levels: this prevents a comprehensive understanding of modern applications' storage performance profiles. In this paper, we present EZIOTracer, a unified I/O tracer for both (Linux) kernel and user spaces, targeting data intensive applications. EZIOTracer is composed of a userland as well as a kernel space tracer, complemented with a trace analysis framework able to merge the output of the two tracers, and in particular to relate user-level events to kernel-level ones, and vice-versa. On the kernel side, EZIOTracer relies on eBPF to offer safe, low-overhead, low memory footprint, and flexible tracing capabilities. We demonstrate using FIO benchmark the ability of EZIOTracer to track down I/O performance issues by relating events recorded at both the kernel and user levels. We show that this can be achieved with a relatively low overhead that ranges from 2% to 26% depending on the I/O intensity

    DLI-MOCVD Crx Cy coating to prevent Zr-based cladding from inner oxidation and secondary hydriding upon LOCA conditions

    No full text
    International audienceZirconium-based claddings with an outer chromium coating resistant to corrosion are studied and devel- oped as an evolutionary Enhanced Accident Tolerant Fuel (E-ATF) concept for light water reactors. How- ever, in hypothetical LOss-of-Coolant-Accident (LOCA) conditions, following clad ballooning and burst, the outer coating does not allow to protect the inner surface of the cladding from High Temperature (HT) steam oxidation and associated secondary hydriding due to steam starvation occurring within the gap between the clad inner surface and the nuclear fuel pellets. To address this issue, DLI-MOCVD (Direct Liquid Injection of Metal-Organic precursors - Chemical Vapor Deposition) Cr x C y coatings have been developed and successfully deposited onto the inner surface of Zr- based cladding tube prototypes. Then, preliminary two-sided oxidation tests have shown that such inner coating is able to increase the resistance to oxidation at HT of the inner clad surface. The present study aimed at performing new steam oxidation tests at 1200 °C on Zircaloy-4 clad proto- types with a 5–20μm-thick Cr x C y inner coating, in conditions more representative of LOCA, after a first internal pressure-induced burst step. Additionally, complementary two-sided steam oxidation tests have been carried out up to 1 h at 1200 °C, on short inner and/or outer-coated clad segments. Finally, Post- Quench (PQ) Ring Compression Tests (RCTs), fractographic analysis and deep metallurgical investigations including neutron-tomography have been performed to get more insights into the PQ behavior of the inner-coated clad. Among other results, it is shown that the inner Cr x C y coating makes it possible to reduce significantly the oxidation and the associated secondary hydriding of the clad inner surface, after ballooning and burst. After at least 600 s under steam at 1200 °C, the reference uncoated clad fails upon final water quenching while the inner-coated prototype keeps its integrity. PQ RCTs showed a higher strength of the inner- coated material, related to lower oxygen and hydrogen uptakes of the substrate

    Theoretical Reassessment and Model Validation of Some Kinetic Parameters Relevant to Si/Cl/H Systems

    No full text
    International audienc

    Wind-Forced Submesoscale Symmetric Instability around Deep Convection in the Northwestern Mediterranean Sea

    No full text
    International audienceDuring the winter from 2009 to 2013, the mixed layer reached the seafloor at about 2500minthe northwestern Mediterranean Sea. Intense fronts around the deep convection area were repeatedlysampled by autonomous gliders. Subduction down to 200–300 m, sometimes deeper, below themixed layer was regularly observed testifying of important frontal vertical movements. PotentialVorticity dynamics was diagnosed using glider observations and a high resolution realistic modelat 1-km resolution. During down-front wind events in winter, remarkable layers of negative PVwere observed in the upper 100m on the dense side of fronts surrounding the deep convection areaand successfully reproduced by the numerical model. Under such conditions, symmetric instabilitycan grow and overturn water along isopycnals within typically 1–5 km cross-frontal slanted cells.Two important hotpspots for the destruction of PV along the topographically-steered NorthernCurrent undergoing frequent down-front winds have been identified in the western part of Gulf ofLion and Ligurian Sea. Fronts were there symmetrically unstable for up to 30 days per winter inthe model, whereas localized instability events were found in the open sea, mostly influenced bymesoscale variability. The associated vertical circulations also had an important signature on oxygenand fluorescence, highlighting their under important role for the ventilation of intermediate layers,phytoplankton growth and carbon export

    Weak Input to state estimates for 2D damped wave equations with localized and non-linear damping

    No full text
    International audienceIn this paper, we study input-to-state (ISS) issues for damped wave equations with Dirichlet boundary conditions on a bounded domain of dimension two. The damping term is assumed to be non-linear and localized to an open subset of the domain. In a first step, we handle the undisturbed case as an extension of a previous work, where stability results are given with a damping term active on the full domain. Then, we address the case with disturbances and provide input-to-state types of results

    A numerical investigation of the effects of ice accretion on the aerodynamic and structural behavior of offshore wind turbine blade

    No full text
    International audienceIn recent years, several wind turbines have been installed in cold climate sites and are menaced by the icing phenomenon. This article focuses on two parts: the study of the aerodynamic and structural performances of wind turbines subject to atmospheric icing. Firstly, the aerodynamic analysis of NACA 4412 airfoil was obtained using QBlade software for a clean and iced profile. Finite element method (FEM) was employed using ABAQUS software to simulate the structural behavior of a wind turbine blade with 100 mm ice thickness. A comparative study of two composite materials and two blade positions were considered in this section. Hashin criterion was chosen to identify the failure modes and determine the most sensitive areas of the structure. It has been found that the aerodynamic and structural performance of the turbine were degraded when ice accumulated on the leading edge of the blade and changed the shape of its profile

    Distributed Competitive Decision Making Using Multi-Armed Bandit Algorithms

    No full text
    International audienceThis paper tackles the problem of Opportunistic Spectrum Access (OSA) in the Cognitive Radio (CR). The main challenge of a Secondary User (SU) in OSA is to learn the availability of existing channels in order to select and access the one with the highest vacancy probability. To reach this goal, we propose a novel Multi-Armed Bandit (MAB) algorithm called ϵ-UCB in order to enhance the spectrum learning of a SU and decrease the regret, i.e. the loss of reward by the selection of worst channels. We corroborate with simulations that the regret of the proposed algorithm has a logarithmic behavior. The last statement means that within a finite number of time slots, the SU can estimate the vacancy probability of targeted channels in order to select the best one for transmitting. Hereinafter, we extend ϵ-UCB to consider multiple priority users, where a SU can selfishly estimate and access the channels according to his prior rank. The simulation results show the superiority of the proposed algorithms for a single or multi-user cases compared to the existing MAB algorithms

    Comparison of Reduction Methods for Finite Element Geometrically Nonlinear Beam Structures

    No full text
    International audienceThe aim of this contribution is to present numerical comparisons of model-order reductionmethods for geometrically nonlinear structures in the general framework of finite element (FE)procedures. Three different methods are compared: the implicit condensation and expansion (ICE),the quadratic manifold computed from modal derivatives (MD), and the direct normal form (DNF)procedure, the latter expressing the reduced dynamics in an invariant-based span of the phasespace. The methods are first presented in order to underline their common points and differences,highlighting in particular that ICE and MD use reduction subspaces that are not invariant. Asimple analytical example is then used in order to analyze how the different treatments of quadraticnonlinearities by the three methods can affect the predictions. Finally, three beam examples areused to emphasize the ability of the methods to handle curvature (on a curved beam), 1:1 internalresonance (on a clamped-clamped beam with two polarizations), and inertia nonlinearity (on acantilever beam)

    Transformation of Uncertain Linear Systems with Real Eigenvalues into Cooperative Form: The Case of Constant and Time-Varying Bounded Parameters

    No full text
    International audienceContinuous-time linear systems with uncertain parameters are widely used for modeling real-life processes. The uncertain parameters, contained in the system and input matrices, can be constant or time-varying. In the latter case, they may represent state dependencies of these matrices. Assuming bounded uncertainties, interval methods become applicable for a verified reachability analysis, for feasibility analysis of feedback controllers, or for the design of robust set-valued state estimators. The evaluation of these system models becomes computationally efficient after a transformation into a cooperative state-space representation, where the dynamics satisfy certain monotonicity properties with respect to the initial conditions. To obtain such representations, similarity transformations are required which are not trivial to find for sufficiently wide a-priori bounds of the uncertain parameters. This paper deals with the derivation and algorithmic comparison of two different transformation techniques for which their applicability to processes with constant and time-varying parameters has to be distinguished. An interval-based reachability analysis of the states of a simple electric step-down converter concludes this paper

    0

    full texts

    11,080

    metadata records
    Updated in last 30 days.
    Portail HAL Ensta
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇