328 research outputs found

    Digging through the dirt: a general method for abstract discrete state estimation with limited prior knowledge

    No full text
    Autonomous robots are often successfully deployed in controlled environments. Operation in uncontrolled situations remains challenging; it is hypothesized that the detection of abstract discrete states (ADS) can improve operation in these circumstances. ADS are high-level system states that are not directly detectable and influence system dynamics. An example of a typical ADS problem that is used in this thesis is that of a wheeled robot driving through puddles of mud that, when entered, alters the velocity of the robot. When the robot is in such a puddle, it is in an ADS 'mud', and when it is not, it is in an ADS 'free'. ADS can be indirectly inferred through the analysis of lower-level data such as the velocity of the robot. The goal of this thesis is to design a general abstract discrete state estimator (ADSE) operating with limited prior knowledge. An ADSE is a hierarchical system for detecting changes in ADS. The ADSE should be general; applicable to multiple ADSE problems. The ADSE should further operate under limited prior knowledge: only assuming that the amount of ADS and the ADS that describes the regular operation are known. The basis for the ADSE designed in this thesis is a Gaussian hidden Markov model (GHMM), a hidden Markov model enhanced with Gaussian emissions. Randomly generated experiments are done on a simple but general ADSE problem. Two unsupervised learning methods derived from Expectation Maximization are evaluated, namely Baum-Welch (BW) and forward extraction (FWE). FWE is introduced in this thesis and is a simpler implementation of Viterbi extraction, leveraging assumptions of ADSE to in theory gain computational efficiency. We found that both BW and FWE exhibit superior performance compared to a likelihood-based baseline estimator when the maximum score of the learning curve is considered. When the final score is considered, in some cases, FWE displays a deteriorating learning curve, resulting in worse final scores compared to the baseline. Furthermore, it was found that the lower the overlap coefficient (therefore the less similar the ADS), the higher the maximum reached score. It was further shown that BW exhibits better convergence than FWE to the true model parameters. Besides this, FWE obtained comparable or in some cases even superior scores compared to BW. In general, from the results, the diversity of the experiments conducted, and the assumptions made we can conclude that the GHMM can be a general method for an ADSE with limited prior knowledge. To quantify the suitability of the GHMM for ADSE, further research should include the evaluation of different ADSE methods on the same problem. There exists a tradeoff between the lower computational cost FWE and the more stable but more computationally intensive BW learning. Therefore, future research can include a combination of these methods. Other extensions include extending the GHMM to a Gaussian mixture hidden Markov model to allow for the modeling of more complex distributions, or the application to multiple states or a changing environment.https://github.com/Wouter-deBoer/adseMechanical Engineering | Vehicle Engineering | Cognitive Robotic

    Back to the future:sequential alignment of text representations

    Get PDF
    Language evolves over time in many ways relevant to natural language processing tasks. For example, recent occurrences of tokens 'BERT' and 'ELMO' in publications refer to neural network architectures rather than persons. This type of temporal signal is typically overlooked, but is important if one aims to deploy a machine learning model over an extended period of time. In particular, language evolution causes data drift between time-steps in sequential decision-making tasks. Examples of such tasks include prediction of paper acceptance for yearly conferences (regular intervals) or author stance prediction for rumours on Twitter (irregular intervals). Inspired by successes in computer vision, we tackle data drift by sequentially aligning learned representations. We evaluate on three challenging tasks varying in terms of time-scales, linguistic units, and domains. These tasks show our method outperforming several strong baselines, including using all available data. We argue that, due to its low computational expense, sequential alignment is a practical solution to dealing with language evolution

    Online System Identification in a Duffing Oscillator by Free Energy Minimisation

    No full text
    Online system identification is the estimation of parameters of a dynamical system, such as mass or friction coefficients, for each measurement of the input and output signals. Here, the nonlinear stochastic differential equation of a Duffing oscillator is cast to a generative model and dynamical parameters are inferred using variational message passing on a factor graph of the model. The approach is validated with an experiment on data from an electronic implementation of a Duffing oscillator.The proposed inference procedure performs as well as offline prediction error minimisation in a state-of-the-art nonlinear model

    Robust domain-adaptive discriminant analysis

    Get PDF
    Consider a domain-adaptive supervised learning setting, where a classifier learns from labeled data in a source domain and unlabeled data in a target domain to predict the corresponding target labels. If the classifier’s assumption on the relationship between domains (e.g. covariate shift, common subspace, etc.) is valid, then it will usually outperform a non-adaptive source classifier. If its assumption is invalid, it can perform substantially worse. Validating assumptions on domain relationships is not possible without target labels. We argue that, in order to make domain-adaptive classifiers more practical, it is necessary to focus on robustness; robust in the sense that an adaptive classifier will still perform at least as well as a non-adaptive classifier without having to rely on the validity of strong assumptions. With this objective in mind, we derive a conservative parameter estimation technique, which is transductive in the sense of Vapnik and Chervonenkis, and show for discriminant analysis that the new estimator is guaranteed to achieve a lower risk on the given target samples compared to the source classifier. Experiments on problems with geographical sampling bias indicate that our parameter estimator performs well

    On regularization parameter estimation under covariate shift

    No full text
    This paper identifies a problem with the usual procedure for L2-regularization parameter estimation in a domain adaptation setting. In such a setting, there are differences between the distributions generating the training data (source domain) and the test data (target domain). The usual cross-validation procedure requires validation data, which can not be obtained from the unlabeled target data. The problem is that if one decides to use source validation data, the regularization parameter is underestimated. One possible solution is to scale the source validation data through importance weighting, but we show that this correction is not sufficient. We conclude the paper with an empirical analysis of the effect of several importance weight estimators on the estimation of the regularization parameter.</p

    Planning to avoid ambiguous states through Gaussian approximations to non-linear sensors in active inference agents

    No full text
    In nature, active inference agents must learn how observations of the world represent the state of the agent. In engineering, the physics behind sensors is often known reasonably accurately and measurement functions can be incorporated into generative models. When a measurement function is non-linear, the transformed variable is typically approximated with a Gaussian distribution to ensure tractable inference. We show that Gaussian approximations that are sensitive to the curvature of the measurement function, such as a second-order Taylor approximation, produce a state-dependent ambiguity term. This induces a preference over states, based on how accurately the state can be inferred from the observation. We demonstrate this preference with a robot navigation experiment where agents plan trajectories

    Corrole basicity in the excited states: Insights on structure-property relationships

    No full text
    Steady-state fluorescence measurements and quantum-chemical DFT geometry optimizations are applied to extend the structure-property relationships between the free-base corrole macrocycle conformation and its basicity to the lowest excited S-1 and T-1 states. Direct basicity estimation in the lowest excited S-1 state is demonstrated by means of fluorescence quantum yield measurements. The long wavelength T1 tautomer is found to retain its basicity in the S(1 )state, whereas the short wavelength T2 tautomer shows a noticeable decrease in basicity in the S-1 state, which is related to the in-plane tilting of the pyrrole ring to be protonated. The conformational changes upon going from the ground to the lowest excited T-1 state and the influence of the meso-aryl substitution pattern on the overall degree of distortions and tilting of the pyrrole ring to be protonated are also discussed from the point of view of macrocycle basicity.Prof. W. Maes acknowledges the Research Foundation-Flanders (FWO) and Hasselt University for financial support. Prof. M. Kruk, Prof. L. Gladkov and Ass. Prof. D. Klenitsky acknowledge the State Program of Scientific Researches of the Republic of Belarus "Photonics, opto-and microelectronics", grant no.1.3.01, for financial support.Maes, W (corresponding author), Hasselt Univ, Inst Mat Res IMO IMOMEC, Agoralaan 1, B-3590 Diepenbeek, Belgium. Kruk, MM (corresponding author), Belarusian State Technol Univ, Phys Dept, Sverdlov Str 13a, Minsk 220006, BELARUS. [email protected]; [email protected]

    Planning to avoid ambiguous states through Gaussian approximations to non-linear sensors in active inference agents

    Get PDF
    In nature, active inference agents must learn how observations of the world represent the state of the agent. In engineering, the physics behind sensors is often known reasonably accurately and measurement functions can be incorporated into generative models. When a measurement function is non-linear, the transformed variable is typically approximated with a Gaussian distribution to ensure tractable inference. We show that Gaussian approximations that are sensitive to the curvature of the measurement function, such as a second-order Taylor approximation, produce a state-dependent ambiguity term. This induces a preference over states, based on how accurately the state can be inferred from the observation. We demonstrate this preference with a robot navigation experiment where agents plan trajectories.13 pages, 3 figures. Accepted to the International Workshop on Active Inference 202

    Bayesian autoregression to optimize temporal Matérn kernel Gaussian process hyperparameters

    No full text
    Gaussian processes are important models in the field of probabilistic numerics. We present a procedure for optimizing Matérn kernel temporal Gaussian processes with respect to the kernel covariance function's hyperparameters. It is based on casting the optimization problem as a recursive Bayesian estimation procedure for the parameters of an autoregressive model. We demonstrate that the proposed procedure outperforms maximizing the marginal likelihood as well as Hamiltonian Monte Carlo sampling, both in terms of runtime and ultimate root mean square error in Gaussian process regression

    Information-seeking polynomial NARX model-predictive control through expected free energy minimization

    Get PDF
    We propose an adaptive model-predictive controller that balances driving the system to a goal state and seeking system observations that are informative with respect to the parameters of a nonlinear autoregressive exogenous model. The controller's objective function is derived from an expected free energy functional and contains information-theoretic terms expressing uncertainty over model parameters and output predictions. Experiments demonstrate that the proposed controller dynamically balances information-seeking and goal-seeking behaviour based on parameter uncertainty
    corecore