1,720,965 research outputs found
Multi-Agent Area Coverage Control Using Reinforcement Learning Techniques
An area coverage control law in cooperation with reinforcement learning techniques is proposed for deploying multiple autonomous agents in a two-dimensional planar area. A scalar field characterizes the risk density in the area to be covered yielding nonuniform distribution of agents while providing optimal coverage. This problem has traditionally been addressed in the literature to date using locational optimization and gradient descent techniques, as well as proportional and proportional-derivative controllers. In most cases, agents' actuator energy required to drive them in optimal configurations in the workspace is not considered. Here the maximum coverage is achieved with minimum actuator energy required by each agent.
Similar to existing coverage control techniques, the proposed algorithm takes into consideration time-varying risk density. These density functions represent the probability of an event occurring (e.g., the presence of an intruding target) at a certain location or point in the workspace indicating where the agents should be located. To this end, a coverage control algorithm using reinforcement learning that moves the team of mobile agents so as to provide optimal coverage given the density functions as they evolve over time is being proposed. Area coverage is modeled using Centroidal Voronoi
Tessellation (CVT) governed by agents. Based on [1,2] and [3], the application of Centroidal Voronoi tessellation is extended to a dynamic changing harbour-like environment.
The proposed multi-agent area coverage control law in conjunction with reinforcement learning techniques is implemented in a distributed manner whereby the multi-agent team only need to access information from adjacent agents while simultaneously providing dynamic target surveillance for single and multiple targets and feedback control of the environment. This distributed approach describes how automatic flocking behaviour of a team of mobile agents can be achieved by leveraging the geometrical properties of centroidal Voronoi
tessellation in area coverage control while enabling multiple targets tracking without the need of consensus between individual agents.
Agent deployment using a time-varying density model is being introduced which is a function of the position of some unknown targets in the environment. A nonlinear derivative of the error coverage function is formulated based on the single-integrator agent dynamics. The agent, aware of its local coverage control condition, learns a value function online while leveraging the same from its neighbours. Moreover, a novel computational adaptive optimal control methodology based on work by [4] is proposed that employs the approximate dynamic programming technique online to iteratively solve the algebraic Riccati equation with completely unknown system dynamics as a solution to linear quadratic regulator problem. Furthermore, an online tuning adaptive optimal control algorithm is implemented using an actor-critic neural network recursive least-squares solution framework. The work in this thesis illustrates that reinforcement learning-based techniques can be successfully applied to non-uniform coverage control. Research combining non-uniform coverage control with reinforcement learning techniques is still at an embryonic stage and several limitations exist. Theoretical results are benchmarked and validated with related works in area coverage control through a set of computer simulations where multiple agents are able to deploy themselves, thus paving the way for efficient distributed Voronoi coverage control problems
Coordinated Deployment of Multiple Autonomous Agents in Area Coverage Problems with Evolving Risk
Coordinated missions with platoons of autonomous agents are rapidly becoming popular because of technological advances in computing, networking, miniaturization and combination of electromechanical systems. These multi-agents networks coordinate their actions to perform challenging spatially-distributed tasks such as search, survey, exploration, and mapping. Environmental monitoring and locational optimization are among the main applications of the emerging technology of wireless sensor networks where the optimality refers to the assignment of sub-regions to each agent, in such a way that a suitable coverage metric is maximized. Usually the coverage metric encodes a distribution of risk defined on the area, and a measure of the performance of individual robots with respect to points inside the region of interest. The risk density can be used to quantify spatial distributions of risk in the domain.
The solution of the optimal control problem in which the risk measure is not time varying is well known in the literature, with the optimal con figuration of the robots given by the centroids of the Voronoi regions forming a centroidal Voronoi tessellation of the area. In other words, when the set of mobile robots converge to the corresponding centroids of the Voronoi tessellation dictated by the coverage metric, the coverage itself is maximized.
In this work, it is considered a time-varying risk density evolving according to a diffusion equation with varying boundary conditions that quantify a time-varying risk on the border of the workspace. Boundary conditions model a time varying flux of external threats coming into the area, averaged over the boundary length, so that rather than considering individual kinematics of incoming threats it is considered an averaged, distributed effect. This approach is similar to the one commonly adopted in continuum physics, in which kinematic descriptors are averaged over spatial domain and suitable continuum fields are introduced to describe their evolution. By adopting a first gradient constitutive relation between the flux and the density, a simple diffusion equation is obtained. Asymptotic convergence and optimality of the non-autonomous system are studied by means of Barbalat's lemma and connections with varying boundary conditions are established. Some criteria on time-varying boundary conditions and evolution are established to guarantee
the stabilities of agents' trajectories. A set of numerical simulations illustrate theoretical results
Multi-Agent Area Coverage Control Using Reinforcement Learning Techniques
An area coverage control law in cooperation with reinforcement learning techniques is proposed for deploying multiple autonomous agents in a two-dimensional planar area. A scalar field characterizes the risk density in the area to be covered yielding nonuniform distribution of agents while providing optimal coverage. This problem has traditionally been addressed in the literature to date using locational optimization and gradient descent techniques, as well as proportional and proportional-derivative controllers. In most cases, agents' actuator energy required to drive them in optimal configurations in the workspace is not considered. Here the maximum coverage is achieved with minimum actuator energy required by each agent. \ud
Similar to existing coverage control techniques, the proposed algorithm takes into consideration time-varying risk density. These density functions represent the probability of an event occurring (e.g., the presence of an intruding target) at a certain location or point in the workspace indicating where the agents should be located. To this end, a coverage control algorithm using reinforcement learning that moves the team of mobile agents so as to provide optimal coverage given the density functions as they evolve over time is being proposed. Area coverage is modeled using Centroidal Voronoi\ud
Tessellation (CVT) governed by agents. Based on [1,2] and [3], the application of Centroidal Voronoi tessellation is extended to a dynamic changing harbour-like environment.\ud
The proposed multi-agent area coverage control law in conjunction with reinforcement learning techniques is implemented in a distributed manner whereby the multi-agent team only need to access information from adjacent agents while simultaneously providing dynamic target surveillance for single and multiple targets and feedback control of the environment. This distributed approach describes how automatic flocking behaviour of a team of mobile agents can be achieved by leveraging the geometrical properties of centroidal Voronoi\ud
tessellation in area coverage control while enabling multiple targets tracking without the need of consensus between individual agents.\ud
Agent deployment using a time-varying density model is being introduced which is a function of the position of some unknown targets in the environment. A nonlinear derivative of the error coverage function is formulated based on the single-integrator agent dynamics. The agent, aware of its local coverage control condition, learns a value function online while leveraging the same from its neighbours. Moreover, a novel computational adaptive optimal control methodology based on work by [4] is proposed that employs the approximate dynamic programming technique online to iteratively solve the algebraic Riccati equation with completely unknown system dynamics as a solution to linear quadratic regulator problem. Furthermore, an online tuning adaptive optimal control algorithm is implemented using an actor-critic neural network recursive least-squares solution framework. The work in this thesis illustrates that reinforcement learning-based techniques can be successfully applied to non-uniform coverage control. Research combining non-uniform coverage control with reinforcement learning techniques is still at an embryonic stage and several limitations exist. Theoretical results are benchmarked and validated with related works in area coverage control through a set of computer simulations where multiple agents are able to deploy themselves, thus paving the way for efficient distributed Voronoi coverage control problems
Nonuniform Coverage with Time-Varying Risk Density Function
Multi-agent systems are extensively used in several applications. An important class of applications involves the optimal spatial distribution of a group of mobile robots on a given area, where the optimality refers to the assignment of subregions to the robots, in such a way that a suitable coverage metric is maximized. Typically the coverage metric encodes a risk distribution defined on the area, and a measure of the performance of individual robots with respect to points inside the region of interest. The coverage metric will be maximized when the set of mobile robots configure themselves as the centroids of the Voronoi tessellation dictated by the risk density. In this work we advance on this result by considering a generalized area control problem in which the coverage metric is non-autonomous, that coverage metric is time varying independently of the states of the robots. This generalization is motivated by the study of coverage control problems in which the coordinated motion of a set of mobile robots accounts for the kinematics of objects penetrating from the outside. Asymptotic convergence and optimality of the non-autonmous system are studied by means of Barbalat's Lemma, and connections with the kinematics of the moving intruders is established. Several numerical simulation results are used to illustrate theoretical predictions
Coordinated Deployment of Multiple Autonomous Agents in Area Coverage Problems with Evolving Risk
Coordinated missions with platoons of autonomous agents are rapidly becoming popular because of technological advances in computing, networking, miniaturization and combination of electromechanical systems. These multi-agents networks coordinate their actions to perform challenging spatially-distributed tasks such as search, survey, exploration, and mapping. Environmental monitoring and locational optimization are among the main applications of the emerging technology of wireless sensor networks where the optimality refers to the assignment of sub-regions to each agent, in such a way that a suitable coverage metric is maximized. Usually the coverage metric encodes a distribution of risk defined on the area, and a measure of the performance of individual robots with respect to points inside the region of interest. The risk density can be used to quantify spatial distributions of risk in the domain.
The solution of the optimal control problem in which the risk measure is not time varying is well known in the literature, with the optimal con figuration of the robots given by the centroids of the Voronoi regions forming a centroidal Voronoi tessellation of the area. In other words, when the set of mobile robots converge to the corresponding centroids of the Voronoi tessellation dictated by the coverage metric, the coverage itself is maximized.
In this work, it is considered a time-varying risk density evolving according to a diffusion equation with varying boundary conditions that quantify a time-varying risk on the border of the workspace. Boundary conditions model a time varying flux of external threats coming into the area, averaged over the boundary length, so that rather than considering individual kinematics of incoming threats it is considered an averaged, distributed effect. This approach is similar to the one commonly adopted in continuum physics, in which kinematic descriptors are averaged over spatial domain and suitable continuum fields are introduced to describe their evolution. By adopting a first gradient constitutive relation between the flux and the density, a simple diffusion equation is obtained. Asymptotic convergence and optimality of the non-autonomous system are studied by means of Barbalat's lemma and connections with varying boundary conditions are established. Some criteria on time-varying boundary conditions and evolution are established to guarantee
the stabilities of agents' trajectories. A set of numerical simulations illustrate theoretical results
Design and Implementation of Control Techniques for Differential Drive Mobile Robots: An RFID Approach
Localization and motion control (navigation) are two major tasks for a successful mobile robot navigation. The motion controller determines the appropriate action for the robot’s actuator based on its current state in an operating environment. A robot recognizes its environment through some sensors and executes physical actions through actuation mechanisms. However, sensory information is noisy and hence actions generated based on this information may be non-deterministic. Therefore, a mobile robot provides actions to its actuators with a certain degree of uncertainty. Moreover, when no prior knowledge of the environment is available, the problem becomes even more difficult, as the robot has to build a map of its surroundings as it moves to determine the position. Skilled navigation of a differential drive mobile robot (DDMR) requires solving these tasks in conjunction, since they are inter-dependent. Having resolved these tasks, mobile robots can be employed in many contexts in indoor and outdoor environments such as delivering payloads in a dynamic environment, building safety, security, building measurement, research, and driving on highways. This dissertation exploits the use of the emerging Radio Frequency IDentification (RFID) technology for the design and implementation of cost-effective and modular control techniques for navigating a mobile robot in an indoor environment. A successful realization of this process has been addressed with three separate navigation modules. The first module is devoted to the development of an indoor navigation system with a customized RFID reader. This navigation system is mainly pioneered by mounting a multiple antenna RFID reader on the robot and placing the RFID tags in three dimensional workspace, where the tags’ orthogonal position on the ground define the desired positions that the robot is supposed to reach. The robot generates control actions based on the information provided by the RFID reader for it to navigate those pre-defined points. On the contrary, the second and third navigation modules employ custom-made RFID tags (instead of the RFID reader) which are attached at different locations in the navigation environment (on the ceiling of an indoor office, or on posts, for instance). The robot’s controller generates appropriate control actions for it’s actuators based on the information provided by the RFID tags in order to reach target positions or to track pre-defined trajectory in the environment. All three navigation modules were shown to have the ability to guide a mobile robot in a highly reverberant environment with variant degrees of accuracy
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
- …
