1,721,045 research outputs found
Data-driven Safe Control of Linear Systems Under Epistemic and Aleatory Uncertainties
Safe control of constrained linear systems under both epistemic and aleatory
uncertainties is considered. The aleatory uncertainty characterizes random
noises and is modeled by a probability distribution function (PDF) and the
epistemic uncertainty characterizes the lack of knowledge on the system
dynamics. Data-based probabilistic safe controllers are designed for the cases
where the noise PDF is 1) zero-mean Gaussian with a known covariance, 2)
zero-mean Gaussian with an uncertain covariance, and 3) zero-mean non-Gaussian
with an unknown distribution. Easy-to-check model-based conditions for
guaranteeing probabilistic safety are provided for the first case by
introducing probabilistic contractive sets. These results are then extended to
the second and third cases by leveraging distributionally-robust probabilistic
safe control and conditional value-at-risk (CVaR) based probabilistic safe
control, respectively. Data-based implementations of these probabilistic safe
controllers are then considered. It is shown that data-richness requirements
for directly learning a safe controller is considerably weaker than
data-richness requirements for model-based safe control approaches that
undertake a model identification. Moreover, an upper bound on the minimal risk
level, under which the existence of a safe controller is guaranteed, is learned
using collected data. A simulation example is provided to show the
effectiveness of the proposed approach
Conflict-Aware Risk-Averse and Safe Reinforcement Learning: A Meta-Cognitive Learning Framework
Presented online April 14, 2021 at 12:15 p.m.Hamidreza Modares is an Assistant Professor in the Department of Mechanical Engineering at Michigan State University. Prior to joining Michigan State University, he was an Assistant Professor in the Department of Electrical Engineering, Missouri University of Science and Technology. His current research interests include control and security of cyber–physical
systems, machine learning in control, distributed control of multi-agent systems, and robotics. He is an Associate Editor of IEEE Transactions on Neural Networks and Learning Systems.Runtime: 60:34 minutesWhile the success of reinforcement learning (RL) in computer games has shown impressive engineering feat, unlike the computer games, safety-critical settings such as unmanned vehicles must thrash around in the real world, which makes the entire enterprise unpredictable. Standard RL practice generally implants pre-specified performance metrics or objectives into the RL agent to encode the designers’ intention and preferences in achieving different and sometimes conflicting goals (e.g., cost efficiency, safety, speed of response, accuracy, etc.). Optimizing pre-specified performance metrics, however, cannot provide safety and performance guarantees across a vast variety of circumstances that the system might encounter in non-stationary and hostile environments. In this talk, I will discuss novel metacognitive RL algorithms to learn not only a control policy that optimizes accumulated reward values, but also what reward functions to optimize in the first place to formally assure safety with a good enough performance. I will present safe RL algorithms that adapt the focus of attention of RL algorithm to its variety of performance and safety objectives to resolve conflict and thus assure the feasibility of the reward function in a new circumstance. Moreover, model-free RL algorithms will be presented to solve the risk-averse optimal control (RAOC) problem to optimize the expected utility of outcomes while reducing the variance of cost under aleatory uncertainties (i.e., randomness). This is because, performance-critical systems must not only optimize the expected performance, but also reduce its variance to avoid performance fluctuation during RL’s course of operation. To solve the RAOC problem, I will present the three variants of RL algorithms and analyze their advantages and preferences for different situations/systems: 1) a one-shot static convex program based RL, 2) an iterative value iteration algorithm that solves a linear programming optimization at each iteration, and 3) an iterative policy iteration algorithm that solves a convex optimization at each iteration and guarantees the stability of the consecutive control policies
Optimal Tracking Control Of Uncertain Systems: On-policy And Off-policy Reinforcement Learning Approaches
Over the last few decades, strong connections between reinforcement learning (RL) and optimal control have prompted a major effort towards developing online and model-free RL algorithms to learn the solution to optimal control problems. Although RL algorithms have been widely used to solve the optimal regulation problems, few results considered solving the optimal tracking control problem (OTCP), despite the fact that most real-world control applications are tracking problems. On the other hand, existing methods for solving OTCP require complete knowledge of the system dynamics. This research begins with developing an adaptive optimal algorithm for linear quadratic tracking problem (LQT). A discounted performance function is introduced for the LQT problem. A discounted algebraic Riccati equation (ARE) is then derived which gives the solution to the LQT problem. The integral reinforcement learning (IRL) technique and off-policy RL technique are used to learn the solution to the discounted ARE online and without requiring complete knowledge of the system dynamics. The proposed idea is then extended to solve optimal tracking control for nonlinear systems. The input constraints are also taken into account for nonlinear systems.In the next step, the proposed method is extended to solve the CT two-player zero-sum game arising in the H8 tracking control problem. An off-policy RL algorithm is developed which enables us to find the solution to the H8 tracking control problem online in real time and without requiring the disturbance being adjustable, which is usually impractical for most of real systems. The next results show how to design dynamic OPFP controllers for CT linear systems with unknown dynamics. To this end, it is first shown that the system state can be constructed using some limited observations on the system output over a period of the history of the system. A Bellman equation is then developed to evaluate a control policy and find an improved policy simultaneously using only some limited observations on the system output. Then, using this Bellman equation, a model-free IRL-based OPFB controller is developed. Next, a model-free approach is developed for solving output synchronization of heterogeneous multi-agent systems. Both the leader’s and the follower’s dynamics is assumed to be unknown. First, a distributed adaptive observer is designed to estimate the leader’s state for each agent. A model-free off-policy RL algorithm is then developed to solve the optimal output synchronization problem online in real time. It is shown that this distributed RL approach implicitly solves the output regulation equations without actually doing so and without requiring knowledge of the leader or of agent’s dynamics. Finally, a model-free RL based method is design for the human-robot interaction system to help the robot adapt itself to the level of the human skills. This assists the human operator to perform a given task with minimum workload demands and optimize the overall human-robot system performance. First, a robot-specific neuro-adaptive controller is designed to make the unknown nonlinear robot behave like a prescribed robot impedance model. Then, a task-specific outer-loop controller is designed to find the optimal parameters of the prescribed robot impedance model, online in real time
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
Deterministic and Stochastic Fixed-time Stability of Discrete-time Autonomous Systems
This paper studies deterministic and stochastic fixed-time stability of
autonomous nonlinear discrete-time (DT) systems. Lyapunov conditions are first
presented under which the fixed-time stability of deterministic DT system is
certified. Extensions to systems under deterministic perturbations as well as
stochastic noise are then considered. For the former, the sensitivity to
perturbations for fixed-time stable DT systems is analyzed, and it is shown
that fixed-time attractiveness is resulted from the presented Lyapunov
conditions. For the latter, sufficient Lyapunov conditions for fixed-time
stability in probability of nonlinear stochastic DT systems are presented. The
fixed upper bound of the settling-time function is derived for both fixed-time
stable and fixed-time attractive systems, and the stochastic settling-time
function fixed upper bound is derived for stochastic DT systems. Illustrative
examples are given along with simulation results to verify the introduced
results
- …
