1,720,970 research outputs found
Evaluating transformers as memory systems in reinforcement learning
Memory is an important component of effective learning systems and is crucial in non-Markovian as well as partially observable environments. In recent years, Long Short-Term Memory (LSTM) networks have been the dominant mechanism for providing memory in reinforcement learning, however, the success of transformers in natural language processing tasks has highlighted a promising and viable alternative. Memory in reinforcement learning is particularly difficult as rewards are often sparse and distributed over many time steps. Early research into transformers as memory mechanisms for reinforcement learning indicated that the canonical model is not suitable, and that additional gated recurrent units and architectural modifications are necessary to stabilize these models. Several additional improvements to the canonical model have further extended its capabilities, such as increasing the attention span, dynamically selecting the number of per-symbol processing steps and accelerating convergence. It remains unclear, however, whether combining these improvements could provide meaningful performance gains overall. This dissertation examines several extensions to the canonical Transformer as memory mechanisms in reinforcement learning and empirically studies their combination, which we term the Integrated Transformer. Our findings support prior work that suggests gating variants of the Transformer architecture may outperform LSTMs as memory networks in reinforcement learning. However, our results indicate that while gated variants of the Transformer architecture may be able to model dependencies over a longer temporal horizon, these models do not necessarily outperform LSTMs when tasked with retaining increasing quantities of information
Learning to Coordinate Efficiently through Multiagent Soft Q-Learning in the presence of Game-Theoretic Pathologies
In this work we investigate the convergence of multiagent soft Q-learning in continuous games where learning is most likely to be affected by relative overgeneralisation. While this will occur more often in multiagent independent learner problems, it is present in joint-learner problems when information is not used efficiently in the learning process. We first investigate the effect of different samplers and modern strategies of training and evaluating energy-based models on learning to get a sense of whether the pitfall is due to sampling inefficiencies or underlying assumptions of the multiagent soft Q-learning extension (MASQL). We use the word sampler to refer to mechanisms that allow one to get samples from a given (target) distribution. After having understood this pitfall better, we develop opponent modelling approaches with mutual information regularisation. We find that while the former (the use of efficient samplers) is not as helpful as one would wish, the latter (opponent modelling with mutual information regularisation) offers new insights into the required mechanism to solve our problem. The domain in which we work is called the Max of Two Quadratics differential game where two agents need to coordinate in a non-convex landscape, and where learning is impacted by the mentioned pathology, relative overgeneralisation. We close this research investigation by offering a principled prescription on how to best extend single-agent energy-based approaches to multiple agents, which is a novel direction
On noise regularised neural networks: initialisation, learning and inference
Thesis (PhD)--Stellenbosch University, 2019.ENGLISH ABSTRACT: Innovation in regularisation techniques for deep neural networks has been a key factor in the rising success of deep learning. However, there is often limited guidance from theory in the development of these techniques and our understanding of the functioning of various successful regularisation techniques remains impoverished.
In this work, we seek to contribute to an improved understanding of regularisation in deep learning. We specifically focus on a particular approach to regularisation that injects noise into a neural network. An example of such a technique which is often used
is dropout (Srivastava et al., 2014).
Our contributions in noise regularisation span three key areas of modeling: (1) learning,
(2) initialisation and (3) inference. We first analyse the learning dynamics of a simple class of shallow noise regularised neural networks called denoising autoencoders (DAEs) (Vincent et al., 2008), to gain an improved understanding of how noise affects the learning process. In this first part, we observe a dependence o f learning behaviour on initialisation, which leads us to study how noise interacts with the initialisation of a deep neural network in terms of signal propagation dynamics during the forward and
backward pass. Finally, we consider how noise affects inference in a Bayesian context.
We mainly focus on fully-connected feedforward neural networks with rectifier linear unit (ReLU) activation functions throughout this study.
To analyse the learning dynamics of DAEs, we derive closed form solutions to a system of decoupled differential equations that describe the change in scalar weights during the course of training as they approach the eigenvalues of the input covariance matrix
(under a convenient change of basis). In terms of initialisation, we use mean field theory to approximate the distribution of the pre-activations of individual neurons, and use this to derive recursive equations that characterise the signal propagation behaviour of the noise regularised network during the first forward and backward pass o f training. Using these equations, we derive new initialisation schemes for noise regularised neural networks that ensure stable signal propagation. Since this analysis is only valid at initialisation, we next conduct a large-scale controlled experiment, training thousands of networks under a theoretically guided experimental design, for further testing the effects of initialisation on training speed and generalisation. To shed light on the influence of noise on inference, we develop a connection between randomly initialised deep noise regularised neural networks and Gaussian processes (GPs)—non-parametric models that perform exact Bayesian inference—and establish new connections between a particular initialisation of such a network and the behaviour of its corresponding GP. Our work ends with an application of signal propagation theory to approximate Bayesian inference in deep learning
where we develop a new technique that uses self-stabilising priors for training deep Bayesian neural networks (BNNs).
Our core findings are as follows: noise regularisation helps a model to focus on the more prominent statistical regularities in the training data distribution during learning which should be useful for later generalisation. However, if the network is deep and not
properly initialised, noise can push network signal propagation dynamics into regimes of poor stability. We correct this behaviour with proper “noise-aware” weight initialisation.
Despite this, noise also limits the depth to which networks are able to train successfully, and networks that do not exceed this depth limit demonstrate a surprising insensitivity to initialisation with regards to training speed and generalisation. In terms of inference, noisy neural network GPs perform best when their kernel parameters correspond to the new initialisation derived for noise regularised networks, and increasing the amount of injected noise leads to more constrained (simple) models with larger uncertainty (away from the
training data). Lastly, we find our new technique that uses self-stabilising priors makes training deep BNNs more robust and leads to improved performance when compared to other state-of-the-art approaches.AFRIKAANSE OPSOMMING: Innovasie in regulariseringstegnieke vir diep neurale netwerke is ’n belangrike aspek van die toenemende sukses van diepleer. Daar is egter beperkte teoretiese leiding in die ontwikkeling van hierdie tegnieke, en ons begrip van hoe verskeie suksesvolle regulariseringstegnieke
funksioneer is steeds onvolledig.
Hierdie proefskrif poog om ’n bydra te maak tot ’n beter begrip van regularisering in diepleer. Ons fokus spesifiek op benaderings wat van ruis gebruik maak om neurale
netwerke te regulariseer. Een voorbeeld van so ’n tegniek wat gereeld gebruik word, is weglating (“dropout”) (Srivastava et al., 2014).
Ons bydrae tot ruisregularisering span drie sleutelareas van modellering: (1) leer, (2) inisialisering en (3) inferensie. Ons ondersoek eers die leerdinamika van ’n eenvoudige klas van vlak ruisgeregulariseerde neurale netwerke, genaamd ontruisende outo-enkodeerders
(DAEs) (Vincent et al., 2008), om ’n beter begrip te kry van hoe ruis die leerproses beïnvloed. In die eerste deel, neem ons waar dat leergedrag afhanklik is van inisialisering.
Dit motiveer ’n studie waar ons die wisselwerking tussen ruis en die inisialisering van diep neurale netwerke bestudeer in terme van die dinamika van seinvloei gedurende die aanvanklike voorwaartse en terugwaartse deurvloei. Laastens word daar gekyk na
die invloed van ruis in die Bayes-konteks. Hierdie studie fokus hoofsaaklik op vollediggekoppelde vorentoevoer neurale netwerke met die gerektifiseerde lineêre eenheid (ReLU) aktiveringsfunksie.
Om die leerdinamika van DAEs te ontleed, lei ons geslotevorm oplossings af vir ’n stelsel ontkoppelde differensiaalvergelykings wat die verandering in skalaargewigte tydens die afrigproses beskryf, soos wat hulle na die eiewaardes van die toevoerkovariansiematriks neig (onder ’n gerieflike basisverandering). Wat inisialisering betref, gebruik ons
gemiddelde-veldteorie om die verdeling van die voor-aktiverings van individuele neurone te benader, en gebruik ons dan hierdie resultate om rekursiewe vergelykings af te lei wat die seinvloei gedrag van ruisgeregulariseerde netwerke tydens die eerste voorwaartse en terugwaartse deurvloei beskryf. Ons gebruik hierdie vergelykings om nuwe inisialiseringskemas vir ruisgeregulariseerde neurale netwerke wat stabiele seinvloei verseker te verkry. Aangesien hierdie analise slegs geldig is tydens inisialisering, voer ons ’n grootskaalse
gekontroleerde eksperiment uit deur duisende netwerke af te rig volgens ’n geskikte eksperimentele ontwerp, om sodoende die effek van inisialisering op die afrigspoed en veralgemening van die afrigte netwerke te toets. Om lig te werp op die effek van ruis op
inferensie, ontwikkel ons ’n verband tussen ewekansigge geïnitialiseerde diep ruisgeregulariseerde
neurale netwerke en Gaussiese prosesse (GPs)—nie-parametriese modelle wat presiese Bayesiaanse inferensie uitvoer—sowel as nuwe verbande tussen ’n spesifieke inisialisering van die netwerk en die optrede van die ooreenstemmende GP. Die proefstuk word
afgesluit met ’n toepassing van seinvloeiteorie op Bayesiaanse inferensie in diep neurale netwerke, waar ons ’n nuwe tegniek, wat gebruik maak van “selfstabiliserende priors”, ontwikkel.
Ons kernbevindinge is as volg: ruisregularisering help ’n model om te fokus op die meer prominente statistiese reëlmatighede in die verdeling van die afrigdata, wat vir latere veralgemening nuttig behoort te wees. As die netwerk egter diep is en nie behoorlik geïnisialiseer is nie, kan ruis veroorsaak dat die dinamika van die netwerk se seinvloei onstabiel raak. Ons kan egter hierdie optrede teenwerk met behoorlike “ruisbewuste” gewigsinisialisering.
Nietemin, beperk ruis die diepte waartoe netwerke suksesvol afgerig kan word, en netwerke wat nie hierdie dieptebeperking oorskry nie, toon ’n verbasende
onsensitiwiteit tot inisialisering met betrekking tot hul afrigspoed en veralgemening. Wat inferensie betref, presteer ruisige neurale netwerk GPs die beste wanneer hul kernparameters ooreenstem met die nuwe inisialisering wat afgelei is vir ruisgeregulariseerde netwerke, en ’n toename in die hoeveelheid ruis wat bygevoeg word lei tot meer beperkte (eenvoudige)
modelle met groter onsekerheid (weg van die afrigtingsdata). Laastens vind ons dat ons nuwe tegniek, wat gebruik maak van “selfstabiliserende priors”, die afrigting van Bayesiaanse neurale netwerke meer robuust maak en tot verbeterde prestasie lei in vergelyking met ander moderne benaderings.Doctora
Advances in random forests with application to classification
e, a novel random forest framework, viz. oblique random rotation
forests, is proposed. Although not entirely satisfactory, the framework
serves as an example of a heuristic approach towards novel proposals based on
bias-variance analyses, instead of an ad hoc approach, as is often found in the
literature.
The analysis of comparative studies regarding advances in random forest algo rithms is also considered. It is of interest to critically evaluate the conclusions
that can be drawn from these studies, and to infer whether novel random forest
algorithms are found to significantly outperform Forest-RI. For this purpose,
a meta-analysis is conducted in which an evaluation is given of the state of research
on random forests based on all (34) papers that could be found in which
a novel random forest algorithm was proposed and compared to already existing
random forest algorithms. Using the reported performances in each paper,
a novel two-step procedure is proposed, which allows for multiple algorithms
to be compared over multiple data sets, and across different papers. The metaanalysis
results indicate weighted voting strategies and variable weighting in
high-dimensional settings to provide significantly improved performances over
the performance of Breiman’s popular Forest-RI algorithm.National Research Foundatio
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
