Portail HAL Ensta
Not a member yet
    11080 research outputs found

    Génération automatique de curriculum pour apprenants artificiels

    No full text
    A long-standing goal of Machine Learning (ML) and AI at large is to design autonomous agents able to efficiently interact with our world. Towards this, taking inspirations from the interactive nature of human and animal learning, several lines of works focused on building decision making agents embodied in real or virtual environments. In less than a decade, Deep Reinforcement Learning (DRL) established itself as one of the most powerful set of techniques to train such autonomous agents. DRL is based on the maximization of expert-defined reward functions that guide an agent’s learning towards a predefined target task or task set. In parallel, the Developmental Robotics field has been working on modelling cognitive development theories and integrating them into real or simulated robots. A core concept developed in this literature is the notion of intrinsic motivation: developmental robots explore and interact with their environment according to self-selected objectives in an open-ended learning fashion. Recently, similar ideas of self-motivation and open-ended learning started to grow within the DRL community, while the Developmental Robotics community started to consider DRL methods into their developmental systems. We propose to refer to this convergence of works as Developmental Machine Learning. Developmental ML regroups works on building embodied autonomous agents equipped with intrinsic-motivation mechanisms shaping open-ended learning trajectories. The present research aims to contribute within this emerging field. More specifically, the present research focuses on proposing and assessing the performance of a core algorithmic block of such developmental machine learners: Automatic Curriculum Learning (ACL) methods. ACL algorithms shape the learning trajectories of agents by challenging them with tasks adapted to their capacities. In recent years, they have been used to improve sample efficiency and asymptotic performance, to organize exploration, to encourage generalization or to solve sparse reward problems, among others. Despite impressive success in traditional supervised learning scenarios (e.g. image classification), large-scale and real-world applications of embodied machine learners are yet to come. The present research aims to contribute towards the creation of such agents by studying how to autonomously and efficiently scaffold them up to proficiency.Un objectif de longue date du Machine Learning (ML) et de l'IA en général est de concevoir des agents autonomes capables d'interagir efficacement avec notre monde. Dans cette optique, s'inspirant de la nature interactive de l'apprentissage humain et animal, plusieurs axes de travaux se sont concentrés sur la construction d'agents décisionnels incarnés dans des environnements réels ou virtuels. En moins d'une décennie, le Deep Reinforcement Learning (DRL) s'est imposé comme l'un des ensembles de techniques les plus puissants pour former de tels agents autonomes. Le DRL est basé sur la maximisation de fonctions de récompense définies par des experts qui guident l'apprentissage d'un agent vers une tâche ou un ensemble de tâches cible prédéfinies. En parallèle, la Robotique Développementale a travaillé sur la modélisation des théories du développement cognitif et de leur intégration dans des robots réels ou simulés. Un concept central développé dans cette littérature est la notion de motivation intrinsèque : les robots développementaux explorent et interagissent avec leur environnement selon des objectifs auto-sélectionnés dans un mode d'apprentissage ouvert. Récemment, des idées similaires d'auto-motivation et d'apprentissage ouvert ont commencé à se développer au sein de la communauté DRL, tandis que la Robotique Développementale a commencé à considérer les méthodes DRL dans leurs systèmes de développement. Nous proposons de désigner cette convergence de travaux sous le nom de Developmental Machine Learning. Developmental ML regroupe des travaux sur la construction d'agents autonomes incarnés équipés de mécanismes de motivation intrinsèque façonnant des trajectoires d'apprentissage ouvertes. La présente recherche vise à contribuer dans ce domaine émergent. Plus précisément, la présente recherche se concentre sur la proposition et l'évaluation des performances d'un bloc algorithmique de base de tels apprenants: les méthodes d'apprentissage automatique de curriculum (ACL). Les algorithmes ACL façonnent les trajectoires d'apprentissage des agents en les challengeant avec des tâches adaptées à leurs compétences. Ces dernières années, ils ont été utilisés pour améliorer la vitesse d’apprentissage et les performances asymptotiques, pour organiser l'exploration, pour encourager la généralisation ou pour résoudre des problèmes de récompense clairsemée, entre autres. Malgré un succès impressionnant dans les scénarios d'apprentissage supervisé traditionnels (par exemple, la classification d'images), les applications à grande échelle et dans le monde réel des apprenants automatiques incarnés sont encore à venir. La présente recherche vise à contribuer à la création de tels agents en étudiant comment guider leurs apprentissage de manière autonome

    Modelling of cyclic visco-elasto-plastic one dimensional behaviour of polyamide-based woven strap

    No full text
    International audienceThis study focuses on modelling of mechanical behaviour of threadlike woven materials or shaped in uniaxial form such as wires, ropes, woven lines, cables, straps, slings etc. The proposed one-dimensional model is based on the superimposition of two stress contributions: a non-Newtonian visco-elastic stress and a time-independent stress. The time-independent stress stands for a particular irreversible behaviour, linked to the loading history. This model neglects the thickness of the time independent hysteresis loops during the unloading-reloading processes while preserving the irreversible character of elastoplastic type behaviour. The model's predictions are compared to a set of experimental results, carried out on polyamide 6-6 (PA66) straps. The model describes the shape of the stress-strain hysteresis loops very well and predicts perfectly the direction of the strain or stress evolution during the creep or relaxation periods, regardless of their position in the first load or in the load-unload branches

    Towards Automata-Based Abstraction of Goals in Hierarchical Reinforcement Learning

    No full text
    International audienceHierarchical Reinforcement Learning (HRL) offers potential benefits for solving long horizon tasks, generally unhandled by standard Reinforcement Learning (RL) techniques, by decomposing the problem and combining simple policies to achieve the goal. They are however still held back by the curse of dimensionality and the ambiguity of selected subtasks. We explore relevant approaches in HRL while highlighting the key challenges of goal representation, high-level planning and propose a research outline tackling them

    An Online Interval-Based Inertial Navigation System for Control Purposes of Autonomous Boats

    No full text
    International audienceInterval analysis is a numerical tool classically used for solving nonlinear equations in a guaranteed way. It has been shown that it can be used to build reliable nonlinear state estimators for dynamical systems. Numerous simulations inspired from real-life applications have shown the applicability of the approach. This paper proposes to implement an interval-based INS (Inertial Navigation System) in an actual robot to estimate its orientation and position. It shows that some types of outliers can be naturally handled by the fusion algorithm, while the resulting controller can be both fast and reliable. Experiments with an actual autonomous boat conclude this article

    Coherently controlled ionization of gases by three-color femtosecond laser pulses

    No full text
    International audiencePhotoionization of atoms and molecules by intense femtosecond laser pulses is a fundamental process of strong-field physics. Using a three-color femtosecond laser scheme with attosecond phase control precision, we demonstrate coherently controlled ionization of nitrogen molecules with a modulation level up to 20% by varying the phase shifts between the fundamental laser frequency at 800 nm and its second and third harmonics. Furthermore, the phase dependence of the ionization degree qualitatively changes with the laser intensity ratios between the three colors. The observations are interpreted as a manifestation of the competition between different parametric channels contributing to the ionization process. Such coherent control of ionization opens new ways to finely tune and optimize various phenomena accompanying laser-material interactions: high-order harmonic and attosecond generation, nanofabrication, remote ablation of samples, and even guidance of discharge and control of lightning by lasers

    Robust subspace tracking algorithms using fast adaptive Mahalanobis distance

    No full text
    International audienceWe consider the problem of robust subspace tracking (RST) in burst noise which appears in some applications in array signal processing and wireless communication. Our approach is based on robust adaptive covariance matrix, thus avoiding parameter fine-tuning in existing methods in the literature. However, this involves the estimation of the Mahalanobis distance, which has a quadratic computational complexity. We propose a novel algorithm to efficiently estimate the adaptive Mahalanobis distance, with linear complexity thanks to exploiting noise characteristics and the data structure. This approach can be used to robustify many existing RST algorithms. Particularly, in this paper, based on two efficient but non-robust algorithms, called YAST and LORAF, we propose their robust counterparts – RYAST and ROBUSTQR –, robust to burst noise while having the same complexity order as that of the original ones. We illustrate the effectiveness of the proposed RST algorithms by comparing them to the state-of-the-art in different scenarios

    A Differential game control problem with state constraints

    No full text
    International audienceWe study the Hamilton-Jacobi (HJ) approach for a two-person zero-sum differential game with state constraints and where controls of the two players are coupled within the dynamics, the state constraints and the cost functions. It is known for such problems that the value function may be discontinuous and its characterization by means of an HJ equation requires some controllability assumptions involving the dynamics and the set of state constraints. In this work, we characterize this value function through an auxiliary differential game free of state constraints. Furthermore, we establish a link between the optimal strategies of the constrained problem and those of the auxiliary problem and we present a general approach allowing to construct approximated optimal feedbacks to the constrained differential game for both players. Finally, an aircraft landing problem in the presence of wind disturbances is given as an illustrative numerical example

    A non-overlapping domain decomposition method with perfectly matched layer transmission conditions for the Helmholtz equation

    No full text
    International audienceIt is well-known that the convergence rate of non-overlapping domain decomposition methods (DDMs) applied to the parallel finite-element solution of large-scale time-harmonic wave problems strongly depends on the transmission condition enforced at the interfaces between the subdomains. Transmission operators based on perfectly matched layers (PMLs) have proved to be well-suited for configurations with layered domain partitions. They are shown to be a good compromise between basic impedance conditions, which lead to suboptimal convergence, and computational expensive conditions based on the exact Dirichlet-to-Neumann (DtN) map related to the complementary of the subdomain. Unfortunately, the extension of the PML-based DDM for more general partitions with cross-points (where more than two subdomains meet) is rather tricky and requires some care.In this work, we present a non-overlapping substructured DDM with PML transmission conditions for checkerboard (Cartesian) decompositions that takes cross-points into account. In such decompositions, each subdomain is surrounded by PMLs associated to edges and corners. The continuity of Dirichlet traces at the interfaces between a subdomain and PMLs is enforced with Lagrange multipliers. This coupling strategy offers the benefit of naturally computing Neumann traces, which allows to use the PMLs as discrete operators approximating the exact Dirichlet-to-Neumann maps. Two possible Lagrange multiplier finite element spaces are presented, and the behavior of the corresponding DDM is analyzed on several numerical examples

    Fokker-Planck equations with terminal condition and related McKean probabilistic representation

    No full text
    International audienceUsually Fokker-Planck type partial differential equations (PDEs) are well-posed if the initial condition is specified. In this paper, alternatively, we consider the inverse problem which consists in prescribing final data: in particular we give sufficient conditions for existence and uniqueness. In the second part of the paper we provide a probabilistic representation of those PDEs in the form a solution of a McKean type equation corresponding to the time-reversal dynamics of a diffusion process

    Stable GSTC formulation for Maxwell’s equations

    No full text
    International audienceWe revisit the classical zero-thickness Generalized Sheet Transition Conditions (GSTCs) which are a key tool for efficiently designing metafilms able to control the flow of light in a desired way. It is shown that it is more convenient to use an enlarged formulation of the GSTC in which the original metafilm is replaced by GSTCs that exclude the layer from the physical or computational domain. These new "layer" transition conditions have the same form as their "sheet" analogues hence they do not necessitate additional complications in their use; their advantage is that they provide a well-posed problem hence guaranty the stability of numerical schemes in the timedomain. These assessments are demonstrated for an all-dielectric structure; the effective susceptibility tensors are derived thanks to asymptotic analysis combined with homogenization technique and bounds for the susceptibilities entering the balance of energy are provided. While negative constant susceptibilities appear in the classical zero-thickness GSTCs, their values in the enlarged formulation are always positive which ensure the stability of the effective problem. Validation of the effective model is provided by means of comparison with direct numerics in two and three dimensions

    0

    full texts

    11,080

    metadata records
    Updated in last 30 days.
    Portail HAL Ensta
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇