1,721,104 research outputs found

    Khamassi, Mehdi

    No full text

    Nouvelles approches en Robotique Cognitive

    No full text
    New Approaches in Cognitive Robotics . This issue presents a set of contributions showing recent evolutions of researches within the so-called “ Cognitive Robotics” field. This term refers to any scientific research aiming at having a robot realize a task which is considered, when performed by humans, to involve cognitive functions. The latters can be (without being exhaustive) learning, social interaction, perception, motricity, spatial cognition, reasoning, language, or awareness. The goal of this issue is to show how the specific objectives and methods of this research field have historically and progressively stepped away from other work in Robotics and Artificial Intelligence, in order to get closer to other disciplines within the field of Cognitive Sciences. As a corollary, the issue aims at making explicit some identified bridges that could enable these various disciplines to cross-fertilize. In particular, one of the objectives of this issue is to illustrate how robotics experimentation can serve as platforms to test hypotheses from other disciplines of Cognitive Sciences, and thus contribute to the study of biological cognition.Ce volume présente un ensemble de contributions montrant les évolutions récentes des recherches du domaine dit de «Robotique Cognitive ». Cette dénomination vaut dès lors que l’on cherche à faire réaliser au robot des tâches qui semblent nécessiter chez l’homme la mise en oeuvre de fonctions cognitives telles que (sans être exhaustifs) l’apprentissage, l’interaction sociale, la perception, la motricité, la cognition spatiale, le raisonnement, le langage ou encore la conscience. Le but du volume est de montrer en quoi les objectifs et méthodes particuliers de ce domaine se sont historiquement distingués d’autres travaux en Robotique ou en Intelligence Artificielle, pour se rapprocher des autres disciplines des Sciences Cognitives. En corollaire, le volume vise à expliciter certains des ponts possibles qui peuvent permettre à ces disciplines de se féconder mutuellement. En particulier, un des objectifs de ce volume est d’illustrer en quoi l’expérimentation robotique peut servir de plateforme de test d’hypothèses d’autres disciplines des Sciences Cognitives, et ainsi contribuer à l’étude de la cognition biologique.Réguigne-Khamassi Mehdi, Doncieux Stéphane. Nouvelles approches en Robotique Cognitive. In: Intellectica. Revue de l'Association pour la Recherche Cognitive, n°65, 2016/1. Nouvelles approches en Robotique Cognitive. pp. 7-25

    Modélisation computationnelle de la variabilité et de la régulation de l'apprentissage par renforcement et des signaux dopaminergiques associés chez le rat

    No full text
    Dans cette thèse, je traiterai de deux sujets principaux concernant les comportements d’apprentissage chez le Rat : premièrement, ce qu’il convient d’appeler le méta-apprentissage ("meta-learning" en anglais), c’est-à-dire la capacité d’autorégulation des paramètres comportementaux qui déterminent la prise de décision ;deuxièmement, la variabilité inter-individuelle quant au choix de la stratégie à employer dans le cadre d’une expérience de conditionnement pavlovien. Pour chacune de ces deux thématiques, j’adopterai un point de vue computationnel ancré dansles techniques de l’apprentissage par renforcement afin de modéliser des données expérimentales, tout en dressant des parallèles avec les fonctions dopaminergiques censées être associées à ces processus.In this work, I will discuss two main topics concerning learnt behaviour in Rats: firstly meta-learning, i.e. the regulation of learning and decision-making parameters; secondly, inter-individual variability in the strategies used in a simple Pavlovian conditioning experiment. In both cases, I will adopt a computational standpoint using reinforcement learning algorithms to model experimental data while also attending to related dopamine functions in the rat brain. If environmental access to food and reproductive opportunities evolves at a relatively stable pace, the learning abilities of an organism should be sufficient to keep track of this evolution and enable appropriate behaviour in response to these changes, but, if the environment should change unexpectedly, an additional process of meta-learning might be required to cope with this change. In particular, controlling the learning rate or speed with which state, stimulus or action values are updated in response to discrete environmental feedback, and balancing exploitation of what seems to be the best option with exploration of potentially better ones, could constitute two powerful meta-learning strategies when faced with a volatile environment. I will start my investigation of meta-learning by analysing the results of a three-armed bandit task with pharmacological inhibition of dopamine, a neurotransmitter suspected of regulating the exploration-exploitation trade-off by Humphries et al. (2012). After this, I will assess how well different models with meta-learning mechanisms regulating either the learning rate or explorationexploitation trade-off can explain long-term changes in behaviour observed between the control sessions of the same three-armed bandit task. Finally, in a Pavlovian conditioning task in which the appearance of a lever predicts food delivery, it is well known that two kinds of behaviour can appear in a rat population (Flagel et al., 2007). On the one hand, so-called sign-trackers become strongly attracted to the lever which they will approach and nibble, while goal-trackers will prefer to immediately go to the site of reward delivery. In parallel, there are differences in the associated dopamine signals, sign-trackers presenting a classical reward prediction error pattern, i.e. a burst of phasic activity which shifts from the time of reward delivery in the early stages of the task to the apparition of the lever-CS in later stages, contrary to goal-trackers whose dopamine signals are mostly stable throughout the task. A model attempting to explain these behavioural and neurological results was previously proposed by Lesaint et al. (2014a), and I will apply this model to new experimental findings based on a task with different inter-trial interval durations. This will result in adjustments to the previous model and propositions for going forward

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Modélisation des réactivations hippocampiques en navigation spatiale avec la théorie de l'apprentissage par renforcement : une approche neuroscientifique et robotique

    No full text
    L'expérience obtenue par les interactions avec le monde qui nous entoure est le principal moyen par lequel les animaux et les humains apprennent. Comme cela a été étudié dans la navigation spatiale chez les rongeurs, le cerveau des mammifères peut réélaborer l'expérience passée de manière contextuelle, grâce aux réactivations des cellules de lieu dans l'hippocampe. Récemment, la théorie de l'apprentissage par renforcement (AR) s'est avérée très efficace dans la modélisation de la navigation dirigée vers un but et des différents types de réactivations de l'hippocampe. Dans cette perspective, notre première contribution scientifique concerne la conception et la validation d'un modèle d'exploration spatiale inspiré par des données comportementales de rongeurs. Nous avons identifié des caractéristiques comportementales communes chez les rongeurs dans un contexte d'exploration libre, et les avons modélisées sous la forme d'un modèle de prise de décision. En exploitant ces données comportementales, nous avons proposé un modèle décisionnel d'exploration général, où la prise de décision repose sur la sécurité perçue d'un lieu dans le labyrinthe et sur le coût et la persistance biomécanique de la dynamique exploratoire de l'animal. Enfin, nous avons validé l'adoption du modèle proposé sur deux nouveaux groupes de données comportementales de rongeurs et nous avons discuté les possibles améliorations futures du modèle pour mieux l'adapter à un plus large éventail de situations. L'exploration libre correspond à un cas particulier du comportement exploratoire, où aucune condition externe n'affecte émotionnellement l'animal. Des recherches antérieures montrent que les réactivations hippocampiques représentent plus fréquemment des lieux à haut contenu émotionnel, mais il existe peu d'études traitant directement de l'apparition de réactivations hippocampiques à la suite d'un évènement aversif ou appétitif et qui comparent les mécanismes impliqués. Pour répondre à ce manque, notre deuxième contribution concerne l'extension du modèle d'exploration libre à la modélisation de ces mécanismes. En s'inspirant des modèles d'AR existants sur les réactivations hippocampiques, nous étendons notre modèle d'exploration libre avec un composant qui décrit la valence apprise et rejouée du conditionnement positif ou négatif. En exploitant des nouveaux données expérimentales, nous avons reproduit qualitativement la corrélation entre la quantité estimée de réactivations hippocampiques pendant le sommeil et l'occupation différentielle de la zone de conditionnement aversif, après et avant le conditionnement. De plus, elles soulèvent de nouvelles prédictions intéressantes, qui concernent la plus grande importance des réactivations pendant le sommeil pour reproduire au mieux le comportement des souris après un conditionnement négatif, plutôt qu'après un conditionnement positif. L'apprentissage des comportements dirigés vers un but est également essentiel dans la conception d'agents artificiels et de robots adaptatifs. En dépit de l'efficacité des méthodes d'AR pour l'apprentissage en ligne, celles-ci ne sont en général pas assez réactives pour répondre aux contraintes de la robotique réelle. De plus, avant même les premières études neurophysiologiques sur les réactivations hippocampiques, la recherche en apprentissage automatique proposait de rejouer l'expérience passée dans l'AR pour améliorer la vitesse d'apprentissage et l'adaptabilité des algorithmes. La dernière contribution scientifique de cette thèse concerne une analyse prospective des avantages et des inconvénients en neurorobotique de différents algorithmes d'AR inspirés par les réactivations hippocampiques. Nous avons proposé un modèle qui combine différentes stratégies de réactivations d'AR et avons testé leurs interactions dans différentes tâches de navigation robotique dirigée vers un but, en partant de simulations purement théoriques jusqu'à arriver à des expériences robotiques complètes.The experience gained by interacting with the surrounding world is the main mean by which animals and humans learn. The mammalian brain can re-elaborate past experiences and contextually organize them, through neural circuitries which involve the hippocampus. Hippocampal reactivations of place cells seem to exploit experience to infer the outcome of new situations, as it has been studied in rodent spatial navigation experiments. Recently, the Reinforcement Learning (RL) theory has been proved to be very efficient in modeling goal-directed navigation and the contributions of different types of hippocampal replay. However, how it can account for the richness of exploratory behaviors is still a matter of debate. Thus, our first contribution has been designing and validating a data-driven exploration model for rodents. Our first research interest was identifying common behavioral characteristics in rodent free exploration and modeling them as a valued-based decision-making model. Starting from observations and data analyses performed on a new rodent dataset, we propose a parametrized general decision-making model, where decisions are based on the perceived safety of a location and the biomechanical cost and persistence of the animal's exploratory dynamic. Eventually, we validate the adoption of the same model on two new rodent datasets, freely exploring different mazes for different periods, and we discuss future improvements of the model to better adapt it to a wider range of situations. Free exploration represents a particular case of exploratory behavior when no external conditioning emotionally affects the animal. When animals experience positive or negative external stimuli, their exploratory behavior changes. Research studies have shown that hippocampal reactivations represent emotional-related locations more frequently, but there exist very few studies directly addressing this phenomenon following aversive and appetitive stimuli and comparing the mechanisms involved. Our second contribution concerns the extension of the free exploration model to describe these mechanisms. Starting from existing RL models of hippocampal replay, we extended our free exploration model with a component accounting for learned and replayed stimulus valence. Our results can qualitatively reproduce the correlation between the estimated amount of hippocampal replay during sleep and the differential occupancy of the shock zone in post- and pre-conditioning found in the experimental data by our collaborators. Moreover, they raise new interesting experimental predictions concerning the increasing relevance of sleep replay in proper learning the post-conditioning behavior of the animal in negative conditioning, compared to their relevance in the positive conditioning case. Learning goal-directed behaviors is also key in designing adaptive artificial agents and robots. Autonomous robots usually have limited knowledge of the stochastic nature of the real world surrounding them and one of the most powerful sets of online learning algorithms, RL, is often neither responsive enough nor time-efficient for the constraints imposed by real robotic applications. Even before the first neurophysiological studies on hippocampal reactivations, machine learning research proposed mechanisms for experience replay in RL algorithms, to enhance the speed of learning and the adaptability of the already existing algorithms. The last scientific contribution of this thesis concerns a prospective analysis of the possible benefits and disadvantages of different state-of-the-art hippocampal replay-inspired RL algorithms in neurorobotics. Since the impact of different types of replay in neurorobotics scenarios has only recently started to be investigated, we test a model combining different RL replay strategies and test their interaction in different goal-oriented robotic navigation tasks, going from a pure theoretical simulation to a complete robotic experiment

    From flexible to habitual behaviors : neuro-inspired meta-learning for autonomous robots

    No full text
    Dans cette thèse, nous proposons d'intégrer la notion d'habitude comportementale au sein d'une architecture de contrôle robotique, et d'étudier son interaction avec les mécanismes générant le comportement planifié. Les architectures de contrôle robotiques permettent à ce dernier d'être utilisé efficacement dans le monde réel et au robot de rester réactif aux changements dans son environnement, tout en étant capable de prendre des décisions pour accomplir des buts à long terme (Kortenkamp et Simmons, 2008). Or, ces architectures sont rarement dotées de capacités d'apprentissage leur permettant d'intégrer les expériences précédentes du robot. En neurosciences et en psychologie, l'étude des différents types d'apprentissage montre pour que ces derniers sont une capacité essentielle pour adapter le comportement des mammifères à des contextes changeants, mais également pour exploiter au mieux les contextes stables (Dickinson, 1985). Ces apprentissages sont modélisés par des algorithmes d'apprentissage par renforcement direct et indirect (Sutton et Barto, 1998), combinés pour exploiter leurs propriétés au mieux en fonction du contexte (Daw et al., 2005). Nous montrons que l'architecture proposée, qui s'inspire de ces modèles du comportement, améliore la robustesse de la performance lors d'un changement de contexte dans une tâche simulée. Si aucune des méthodes de combinaison évaluées ne se démarque des autres, elles permettent d'identifier les contraintes sur le processus de planification. Enfin, l'extension de l'étude de notre architecture à deux tâches (dont l'une sur robot réel) confirme que la combinaison permet l'amélioration de l'apprentissage du robot.In this work, we study how the notion of behavioral habit, inspired from the study of biology, can benefit to robots. Robot control architectures allow the robot to be able to plan to reach long term goals while staying reactive to events happening in the environment (Kortenkamp et Simmons, 2008). However, these architectures are rarely provided with learning capabilities that would allow them to acquire knowledge from experience. On the other hand, learning has been shown as an essential abiilty for behavioral adaptation in mammals. It permits flexible adaptation to new contexts but also efficient behavior in known contexts (Dickinson, 1985). The learning mechanisms are modeled as model-based (planning) and model-free (habitual) reinforcement learning algorithms (Sutton et Barto, 1998) which are combined into a global model of behavior (Daw et al., 2005). We proposed a robotic control architecture that take inspiration from this model of behavior and embed the two kinds of algorithms, and studied its performance in a robotic simulated task. None of the several methods for combining the algorithm we studied gave satisfying results, however, it allowed to identify some properties required for the planning process in a robotic task. We extended our study to two other tasks (one being on a real robot) and confirmed that combining the algorithms improves learning of the robot's behavior

    Computational modeling of the role of dopamine in the cortico-striatal loops in learning and action selection's regulation

    No full text
    Dans ce travail de thèse, nous avons modélisé le rôle de la dopamine dans l'apprentissage et dans les processus de sélection de l'action en lien avec les ganglions de la base. L'activité des neurones dopaminergiques présente de nombreuses similarités avec l'erreur de prédiction de la récompense utilisée par les algorithmes d'apprentissage par renforcement. Ainsi, ces neurones sont supposés guider le processus de sélection de l'action.Dans une première partie, nous avons analysé l'information encodée par les neurones dopaminergiques dans une tâche à choix multiples en la comparant à différentes informations utilisées par les modèles d'apprentissage par renforcement. Nos résultats suggèrent que l'information encodée par les neurones dopaminergiques enregistrer dans la tâche n'est que partiellement compatible avec une erreur de prédiction et semble en partie dissociée du comportement.Dans une deuxième partie, nous avons simulé l'effet de la dopamine sur un modèle des ganglions de la base prenant en compte des connections existant chez le primate, souvent négligées dans la littérature. La plupart des modèles actuels font en effet l'hypothèse d'une séparation stricte de deux chemins dans les ganglions de la base : le chemin direct lié à la récompense et le chemin indirect lié à la punition. Cependant des études anatomiques remettent en question cette dissociation, en particulier chez le primate. Nous proposons ainsi d'étudier comment différents niveaux de dopamine, dans le contexte de la maladie de Parkinson, affectent l'apprentissage et la sélection de l'action dans ce modèleIn this thesis work, we modelled the role of dopamine in learning and in the processes of action selection through its interaction with the basal ganglia. During the 90’s, the work of Schultz and colleagues has led to major progress in understanding the neural mechanisms underlying the influence of feedback on learning. The activity of dopaminergic neurons exhibited properties of the reward prediction error signal used in so-called Temporal Difference (TD) machine learning algorithms. Thus, DA has been thought to be the neural signal that help us to adapt our behavior. In the first part of my PhD, we analyze the information encoded by dopaminergic neurons recorded during a multi-choice task. In this purpose, we modeled the task and simulated different TD learning algorithms to quantitatively compare their ability to reproduce dopamine neurons activity. Our results show that the information carried out by dopamine neurons is only partly consistent with a reward prediction error and seems to be dissociated from behavioral adaptation.In the second part of my PhD, we study the effect of different levels of dopamine in a biologically plausible model of primates basal ganglia that considers existing connections often neglected in the literature. Indeed, most of current models of basal ganglia assume the existence of two segregated pathway: the direct pathway associated with reward and the indirect pathway associated with punishment. However, anatomical studies in primates revealed that these two pathways are not dissociated. We study the ability of such a model to reproduce beta oscillations observed in Parkinsonian and the differences in reward and punishment sensitivity, with high or low-level of dopamine

    Generic cognitive architecture for the coordination of learning strategies in robotics

    No full text
    L’objectif principal de cette thèse est de proposer une nouvelle méthode d’adaptation en ligne de l’apprentissage robotique, permettant aux robots d’adapter dynamiquement et de manière autonome leur comportement en fonction des variations de leur propre performance. La méthode élaborée est suffisamment générale et tâche-indépendante pour qu’un robot l’utilisant puisse effectuer différentes tâches dynamiques de nature variée sans ajustement des algorithmes ou des paramètres par le programmeur. Les algorithmes qui sous-tendent cette méthode consistent en un système de méta-contrôle permettant au robot de faire appel à deux experts décisionnels suivant une stratégie comportementale différente. L’expert model-based construit un modèle des effets des actions à long-terme et utilise ce modèle pour décider ; cette stratégie est coûteuse en termes de ressources calculatoires, mais converge rapidement vers la solution. L’expert model-free est quant à lui peu coûteux en termes de ressources calculatoires, mais met du temps à converger vers la solution optimale. Dans ce travail, nous avons élaboré un nouveau critère de coordination de ces deux experts permettant au robot de changer dynamiquement de stratégie au cours du temps. Nous montrons dans ce travail que notre méthode de coordination de comportements permet au robot de maintenir une performance optimale en termes de performance et de temps de calcul. Nous montrons aussi que la méthode permet de faire face à des changements brusques de l’environnement, des changements d’objectifs ou de comportements du partenaire humain dans le cas des tâches d’interaction.The main objective of this thesis is to propose a new method for online adaptation of robotic learning, allowing robots to dynamically and autonomously adapt their behavior according to variations in their own performance. The developed method is sufficiently general and task-independent that a robot using it can perform different dynamic tasks of various nature without any algorithm or parameter adjustment by the programmer. The algorithms underlying this method consist of a meta-control system that allows the robot to call upon two decision-making experts following a different behavioral strategy. The model-based expert builds a model of the effects of long-term actions and uses this model to decide; this strategy is computationally expensive, but quickly converges to the solution. The model-free expert is inexpensive in terms of computational resources, but takes time to converge to the optimal solution. In this work, we have developed a new criterion for the coordination of these two experts allowing the robot to dynamically change its strategy over time. We show in this work that our behavior coordination method allows the robot to maintain an optimal performance in terms of performance and computation time. We also show that the method can cope with abrupt changes in the environment, changes in goals or changes in the behavior of the human partner in the case of interaction tasks

    Coordination of memory systems : theoretical models of human and animals behavior

    No full text
    Durant ce doctorat financé par l'observatoire B2V des mémoires, nous avons réalisé une modélisation mathématique du comportement dans trois tâches distinctes (avec des sujets humains, des sujets singes et des rongeurs), mais qui supposent toutes une coordination entre systèmes de mémoire. Dans la première expérience, nous avons reproduit le comportement de sujets humains (choix et temps de réaction) en combinant les modèles mathématiques d'une mémoire de travail et d'une mémoire inflexible. Nous avons associé pour un sujet son comportement au meilleur modèle possible en comparant des modèles génériques de coordination de ces deux mémoires issues de la littérature actuelle ainsi que notre propre proposition d'une interaction dynamique entre les mémoires. Au final, c'est notre proposition d'une interaction au lieu d'une séparation stricte qui s'est avérée la plus efficace dans la majorité des cas pour expliquer le comportement des sujets. Dans une deuxième expérience, les mêmes modèles de coordination ont été testés dans une tâche chez le singe. Considérée comme un test de transférabilité, cette expérience démontre principalement la nécessité de coordination de mémoires pour expliquer le comportement de certains singes. Dans une troisième expérience, nous avons modélisé le comportement d'un groupe de souris confronté à l'apprentissage d'une séquence d'action motrice dans un labyrinthe sans indices externes. En comparant avec deux autres stratégies d'apprentissages (intégration de chemin et planification dans un graphe), la combinaison d'une mémoire épisodique avec une mémoire inflexible s'est révélée être le meilleur modèle pour reproduire le comportement des souris.During this PhD funded by the B2V Memories Observatory, we performed a mathematical modeling of behavior in three distinct tasks (with human subjects, monkeys and rodents), all involving coordination between memory systems. In the first experiment, we reproduced the behavior of human subjects (choice and reaction time) by combining the mathematical models of working memory and procedural memory. For each subject, we associated their behavior to the best possible model by comparing generic models of coordination of these two memories from the current literature as well as our own proposal of a dynamic interaction between memories. In the end, it was our proposal of an interaction instead of a strict separation which proved most effective in the majority of cases to explain the behavior of the subjects. In a second experiment, the same coordination models were tested in a monkey task. Considered as a transferability test, this experiment mainly demonstrates the need for coordination of memories to explain the behavior of certain monkeys. In a third experiment, we modeled the behavior of a group of mice confronted with the learning of a motor action sequence in a labyrinth without visual cues. Comparing with two other learning strategies (path integration and graph planning), the combination of an episodic memory with a procedural memory proved to be the best model to reproduce the behavior of mice
    corecore