1,721,138 research outputs found

    Effects of structural failure on the safe flight envelope of aircraft

    Get PDF
    The research presented in this paper focuses on the effects of structural failures on the safe flight envelope of an aircraft, which is the set of all the states in which safe maneuver of the aircraft can be assured. Nonlinear reachability analysis basedonan optimal control formulation is performed to estimate the safe flight envelope using actual aircraft control surface inputs. This approach uses the physical model of an aircraft, where the aerodynamic stability and control derivatives are calculated using Digital Datcom. Symmetrical damages to a Cessna Citation II are considered with 25, 50, 75, and 100% spanwise vertical tail tip losses, leading to gradual shrinkage in the safe flight envelope. Based on the estimated safe flight envelopes, a discussion on the effects of structural damages and different flight conditions on the safe flight envelope is presented. In particular, the interpolatibility of the resulting safe flight envelopes is demonstrated. This property is essential for a novel database-driven flight envelope prediction method, where a database of safe flight envelopes is created offline to be accessed later in real time.Control & SimulationWind Energ

    Guaranteed globally optimal continuous reinforcement learning

    No full text
    Self-learning controllers offer various strong benefits over conventional controllers, the most important one being their ability to adapt to unexpected circumstances. Their application is however limited, the most important reason being that, for self-learning controllers to work on continuous domains, nonlinear function approximators are required, and as soon as nonlinear function approximators are involved, it is uncertain whether convergence will occur. This project has as goal to contribute towards achieving convergence guarantees. The first focus lies on using reinforcement learning, combined with neural network function approximation, to create self-learning controllers. A reinforcement learning controller architecture has been set up which is capable of controlling systems with continuous states and actions. Also an extension has been made that enables the controller to freely vary its timestep without any significant consequences. A literature research has shown that there are no convergence proofs yet of practically feasible reinforcement learning controllers with nonlinear function approximators. Several proofs of convergence for the case of linear function approximators have been provided. Also, a proof exists that a certain reinforcement learning controller algorithm with nonlinear function approximation converges. However, the corresponding learning algorithm requires an infinite amount of iterations for the value function to converge even before the policy is updated, making it practically unfeasible. The reasons why convergence often does not occur for reinforcement learning controllers with neural network function approximators include overestimated value functions, incorrect generalization, function approximators incapable of approximating the actual value function and error amplification due to bootstrapping. Furthermore, when convergence does occur, it’s not necessarily to a global optimum. Ideas to solve these issues have been offered, but it is still far from certain whether these ideas will work. Furthermore, they will make the algorithm very complex. Proving that the resulting algorithm will converge, even if it’s only convergence to a local optimum, will prove to be extremely difficult, if it is at all possible. Hence reinforcement learning controllers with neural network function approximators do not seem to be appropriate if convergence to the global optimum needs to be ensured. The literature research has also revealed that so far no attempt has been made in combining reinforcement learning with interval analysis. To investigate the possibilities of combining these two fields, the interval Q-learning algorithm has been designed. This algorithm combines a discrete version of Q-learning with interval analysis techniques. This algorithm has proven convergence for the discrete case. Subsequently, the discrete interval Q-learning algorithm has been expanded to the continuous domain. This was done using the main RL value assumption, which assumes that the derivatives of the value function with respect to every state parameter and action parameter have known bounds. The resulting continuous interval Q-learning algorithm was shown to have proven convergence to the optimal value function. Furthermore, bounds on how fast the algorithm converges were given. The most important downside of the first version of the continuous interval Q-learning algorithm was its slow run-time. This made the algorithm practically infeasible. To still meet the goals of the project, a different function approximator was designed. This function approximator used many small blocks to bound the optimal value function in every part of the state-action-space. Furthermore, the number of blocks was increased dynamically as the algorithm learned, thus giving the function approximator a theoretically unlimited accuracy. Though the resulting algorithm used information slightly less efficiently than its precursor, its run-time was significantly improved. It became practically feasible to apply this algorithm. In the end of the report, the algorithm has been applied to a few simple test problems. An important parameter here was the dimension D of the problem, which equals the sum of the number of state and action parameters. For two- and three-dimensional problems, the algorithm was able to sufficiently bound the optimal value function quite quickly, resulting in a controller with satisfactory performance. For the cart and pendulum system, which is a five-dimensional problem, this turned out to be different. A long training (in the order of several hours or more) will be required before a satisfactory performance can be obtained. However, since the algorithm has proven convergence to the globally optimal policy, this does not necessarily have to be a big problem. Finally, it is mentioned that this thesis report has introduced the world’s first combination of reinforcement learning and interval analysis. It also introduced the world’s first practically feasible continuous RL controller with proven convergence to the global optimum. The key to accomplishing such a controller turned out to be (A) letting go of conventional ways of designing continuous RL controllers, and (B) quantifying the assumption that the value Q of two nearby states are similar through the main RL value assumption.Control & SimulationAerospace Engineerin

    Selective Velocity Obstacle method for Autonomous Collision Avoidance: An experimental validation

    No full text
    Unmanned aerial vehicle (UAV) applications are increasing and there is a need for safe operations in terms of avoidance. The Velocity Obstacle (VO) method uses position and velocity vectors to determine if a collision is going to happen; an adaptation of the VO- method is called the Selective Velocity Obstacle (SVO) method and adds navigation modes and right of way rules for cooperative flights. The characteristics of the SVO-method have been evaluated before in a simulated environment, but the contribution of this paper is an experimental validation of the SVO-method by including the factors that are neglected in simulation such as noise, delay and unmodeled dynamics. Additionally, it is shown how adaptations need to be made when actual drones are used for avoidance. Multiple situations are tested where two UAVs are flown on trajectories to create colliding situations. For the experimental setup, the SVO-method is implemented on a Parrot® AR.Drone 2.0 while an OptiTrack system provides position and velocity data. The Paparazzi autopilot system uses this data for its flight plans to fly autonomously between waypoints. The results of the experiment show that the SVO-method is a safe cooperative avoidance method.Aerospace EngineeringControl & SimulationControl & Operations: ATM, Airports and SafetyAE531

    Autonomous navigation using feature based hierarchical reinforcement learning

    No full text
    The ease of availability of low cost aerial platforms has given rise to extensive research in the field of autonomous navigation. There are strong indications in existing research that UAV autonomy leads to significant gains in terms of safety as well as performance in a number of scenarios, including but not limited to search and rescue missions in disaster areas. This paper tackles the problem of autonomous indoor navigation by applying reinforcement learning. In specific terms, this paper employs hierarchical reinforcement learning methods in order to overcome the challenges posed by the complexity and size of the problem space when dealing with navigational problems. In addition, a feature-based relative state is used so as to contain the number of states to a manageable level, while ensuring that the learnt policies are perfectly valid, independent of the size of the problem. The findings of this paper successfully demonstrate the ability of the Options method to solve a complex navigation problem involving goal finding, room exiting and battery recharging in an efficient manner, and further to solve a problem that cannot be solved using traditional flat Q-learning. In addition, this paper provides evidence of the positive influence of prior high level knowledge on the learning rate for a navigation problem as a) the problem size scales up and b) the problem complexity scales up.Aerospace EngineeringControl & OperationsControl & Simulatio

    Intelligent Controller Selection for Aggressive Quadrotor Manoeuvring: A reinforcement learning approach

    No full text
    A novel intelligent controller selection method for quadrotor attitude and altitude control is presented that maintains performance in different regimes of the flight envelope. Conventional quadrotor controllers can behave insufficiently during aggressive manoeuvring, in extreme angles the quadrotor is unable to maintain height which may result in loss of performance. By implementing several controllers designed specifically for these more extreme manoeuvres it is possible to maintain performance while expanding the flight envelope beyond conventional limitations. The method proposed uses Q-Learning to learn which controller performs best in different scenarios. The controllers that can be used consist of a low angle Nonlinear Dynamic Inversion (NDI) controller and a specialised high angle NDI controller that tries to prevent actuator saturation by controlling height through the pitch and roll angles. The algorithm is split into 2 parts. During the off-line training phase the quadrotor learns a baseline policy that can be used on-line to skip the initial exploration phase. The resulting policy is a clear representation of controller selection throughout the flight envelope. The success of on-line implementation is highly dependent on the quality of the simulation model. In the on-line phase the policy is mostly exploited and learning is done at a much lower rate. Finally it is shown that reinforcement learning is a very effective way to tune complex systems. The complex system dynamics do not need to be known, only important performance parameters need to be correctly formulated in a reward function.Control & SimulationControl & OperationsAerospace Engineerin

    Safe Hierarchical Reinforcement Learning: Hierarchical Reinforcement Learning with Safe Space Exploration

    No full text
    Reinforcement Learning is a much researched topic for autonomous machine behavior and is often applied to navigation problems. In order to deal with growing environments and larger state/action spaces, Hierarchical Reinforcement Learning has been introduced. Unfortunately learning from experience, which is central to Reinforcement Learning, makes guaranteeing safety a complex problem. This paper demonstrates an approach, named the actor-judge approach, to make the exploration safer while imposing as few as possible restrictions on the agent. The approach combines ideas from the fields of Hierarchical Reinforcement Learning and Safe Reinforcement Learning to develop a Safe Hierarchical Reinforcement Learning algorithm. The algorithm is tested in a simulated environment where the agent represents an Unmanned Aerial Vehicle able to move laterally in four directions using quadridirectional range sensors to establish a relative position. Although this approach does not guarantee the agent to never explore unsafe areas of the state domain, results show the actor-judge method increases agent safety and can be used on multiple levels an HRL agent hierarchy.Aerospace EngineeringControl & OperationsControl & Simulatio

    Incremental Nonlinear Dynamic Inversion Flight Control: Stability and Robustness Analysis and Improvements

    No full text
    Incremental Nonlinear Dynamic Inversion (INDI) is a variation on Nonlinear Dynamic Inversion (NDI) retaining the high-performance advantages of NDI, while increasing controller robustness to model uncertainties and decreasing the dependency on the vehicle model. After a successful flight test with a multirotor Micro Aerial Vehicle (MAV), the question arises whether this technique can be used to successfully design a Flight Control System (FCS) for aircraft in general. This requires additional research on aircraft characteristics that could cause issues related to the stability and performance of the INDI controller. Typical characteristics are additional time delays due to data buses and measurement systems, slower actuator and sensor dynamics, and a lower control frequency. The main contributions of this article are 1) an analytical stability analysis showing that implementing discrete-time INDI with a sampling time smaller than 0.02s results in large stability margins regarding system characteristics and controller gains; 2) a simulation study showing significant performance degradation requiring controller adaptation due to actuator measurement bias, angular rate measurement noise, angular rate measurement delay and actuator measurement delay; 3) the use of a real-time time delay identification algorithm based on latency to successfully synchronize the angular rate and actuator measurement delay together with pseudo control hedging (PCH) to prevent oscillatory behavior; and 4) recommendations regarding control modes, assessment criteria and PH-LAB Cessna Citation specific issues to be used by future contributors to a flight test with INDI on the PH-LAB aircraft.Aerospace EngineeringControl & OperationsControl & Simulatio

    Incremental Generalized Policy Iteration for Adaptive Attitude Tracking Control of a Spacecraft

    No full text
    This paper proposes a novel dynamic programming algorithm for nonlinear system optimal control problem, namely Incremental Generalized Policy Iteration (IGPI). The proposed IGPI algorithm combines the advantages of Incremental Control(IC) and Generalized Policy Iteration(GPI). Incremental control can handle the nonlinearity and uncertainty in nonlinear systems without knowing the nonlinear system information, GPI can learn an optimal control law for dynamical systems. Based on the proposed IGPI algorithm, a data-driven adaptive attitude controller is designed for a spacecraft with sloshing liquid fuel. Simulation results demonstrate the effectiveness of the spacecraft attitude controller.Green Open Access added to TU Delft Institutional Repository ‘You share, we take care!’ – Taverne project https://www.openaccess.nl/en/you-share-we-take-care Otherwise as indicated in the copyright section: the publisher is the copyright holder of this work and the author uses the Dutch legislation to make this work public.Control & Simulatio
    corecore