1,721,047 research outputs found
Toward Supervised Anomaly Detection
Anomaly detection is being regarded as an unsupervised learning task as anomalies stem from adversarial or unlikely events with unknown distributions. However, the predictive performance of purely unsupervised anomaly detection often fails to match the required detection rates in many tasks and there exists a need for labeled data to guide the model generation. Our first contribution shows that classical semi-supervised approaches, originating from a supervised classifier, are inappropriate and hardly detect new and unknown anomalies. We argue that semi-supervised anomaly detection needs to ground on the unsupervised learning paradigm and devise a novel algorithm that meets this requirement. Although being intrinsically non-convex, we further show that the optimization problem has a convex equivalent under relatively mild assumptions. Additionally, we propose an active learning strategy to automatically filter candidates for labeling. In an empirical study on network intrusion detection data, we observe that the proposed learning methodology requires much less labeled data than the state-of-the-art, while achieving higher detection accuracies
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Development and application of new statistical methods for the analysis of multiple phenotypes to investigate genetic associations with cardiometabolic traits
Die biotechnologischen Entwicklungen der letzten Jahre ermöglichen eine immer detailliertere Untersuchung von genetischen und molekularen Markern mit multiplen komplexen Traits. Allerdings liefern vorhandene statistische Methoden für diese komplexen Analysen oft keine valide Inferenz.
Das erste Ziel der vorliegenden Arbeit ist, zwei neue statistische Methoden für Assoziationsstudien von genetischen Markern mit multiplen Phänotypen zu entwickeln, effizient und robust zu implementieren, und im Vergleich zu existierenden statistischen Methoden zu evaluieren. Der erste Ansatz, C-JAMP (Copula-based Joint Analysis of Multiple Phenotypes), ermöglicht die Assoziation von genetischen Varianten mit multiplen Traits in einem gemeinsamen Copula Modell zu untersuchen. Der zweite Ansatz, CIEE (Causal Inference using Estimating Equations), ermöglicht direkte genetische Effekte zu schätzen und testen.
C-JAMP wird in dieser Arbeit für Assoziationsstudien von seltenen genetischen Varianten mit quantitativen Traits evaluiert, und CIEE für Assoziationsstudien von häufigen genetischen Varianten mit quantitativen Traits und Ereigniszeiten. Die Ergebnisse von umfangreichen Simulationsstudien zeigen, dass beide Methoden unverzerrte und effiziente Parameterschätzer liefern und die statistische Power von Assoziationstests im Vergleich zu existierenden Methoden erhöhen können - welche ihrerseits oft keine valide Inferenz liefern.
Für das zweite Ziel dieser Arbeit, neue genetische und transkriptomische Marker für kardiometabolische Traits zu identifizieren, werden zwei Studien mit genom- und transkriptomweiten Daten mit C-JAMP und CIEE analysiert. In den Analysen werden mehrere neue Kandidatenmarker und -gene für Blutdruck und Adipositas identifiziert. Dies unterstreicht den Wert, neue statistische Methoden zu entwickeln, evaluieren, und implementieren. Für beide entwickelten Methoden sind R Pakete verfügbar, die ihre Anwendung in zukünftigen Studien ermöglichen.In recent years, the biotechnological advancements have allowed to investigate associations of genetic and molecular markers with multiple complex phenotypes in much greater depth. However, for the analysis of such complex datasets, available statistical methods often don’t yield valid inference.
The first aim of this thesis is to develop two novel statistical methods for association analyses of genetic markers with multiple phenotypes, to implement them in a computationally efficient and robust manner so that they can be used for large-scale analyses, and evaluate them in comparison to existing statistical approaches under realistic scenarios. The first approach, called the copula-based joint analysis of multiple phenotypes (C-JAMP) method, allows investigating genetic associations with multiple traits in a joint copula model and is evaluated for genetic association analyses of rare genetic variants with quantitative traits. The second approach, called the causal inference using estimating equations (CIEE) method, allows estimating and testing direct genetic effects in directed acyclic graphs, and is evaluated for association analyses of common genetic variants with quantitative and time-to-event traits.
The results of extensive simulation studies show that both approaches yield unbiased and efficient parameter estimators and can improve the power of association tests in comparison to existing approaches, which yield invalid inference in many scenarios.
For the second goal of this thesis, to identify novel genetic and transcriptomic markers associated with cardiometabolic traits, C-JAMP and CIEE are applied in two large-scale studies including genome- and transcriptome-wide data. In the analyses, several novel candidate markers and genes are identified, which highlights the merit of developing, evaluating, and implementing novel statistical approaches. R packages are available for both methods and enable their application in future studies
Detecção e diagnóstico de oscilações em malhas de controle
Em indústrias de processos, a oscilação é um problema de grande incidência que degrada a rentabilidade da planta. Essas indústrias possuem tipicamente entre 500 e 5000 malhas de controle. O número elevado dificulta a inspeção individual de cada malha para a detecção de falhas o que cria a necessidade por técnicas automáticas. Nos últimos 30 anos, dezenas dessas técnicas foram propostas para a detecção e diagnóstico da oscilação, mesmo assim, suas eficiências são abaixo da desejável quando aplicadas em plantas reais. Este trabalho tem por objetivo o refinamento da literatura na detecção e diagnóstico da oscilação em indústrias de processos. Para isso, o trabalho inicia com uma revisão completa da literatura na área da detecção da oscilação com o propósito de apresentar, classificar e discutir prós e contras de cada técnica de modo a facilitar sua escolha pelo engenheiro. A seguir, dez métodos de detecção da oscilação são avaliados a dados industriais. Nota-se que a acuracidade das técnicas ainda é insatisfatória e a principal causa é a variedade de características encontradas em dados industriais. Assim, o trabalho prossegue com a proposta de três técnicas. A primeira é uma técnica de detecção da oscilação baseada em inteligência artificial onde um modelo é treinado a partir de exemplos, diferente das técnicas tradicionais baseadas algoritmos. Essa abordagem garante que a técnica cubra um maior número de características dos dados industriais. A segunda técnica objetiva a detecção e o diagnóstico simultâneo da oscilação. Para isso, o padrão dos sinais de saída do controlador e saída do processo é classificado por modelo baseado em redes neurais. A última técnica resolve um problema específico na área: o diagnóstico em sinais de baixa amostragem. O desenvolvimento da técnica foi baseado na resolução de um problema real em uma refinaria brasileira. As três técnicas propostas foram testadas a dados industriais e retornaram melhores resultados quando comparadas a técnicas tradicionais.In process industries, oscillation is a problem of high incidence that degrades plant profitability. These industries typically have between 500 and 5000 control loops. The high number turns difficult the individual inspect of each loop for fault detection, which makes necessary the use of automatic techniques. In the last 30 years, dozens of these techniques have been proposed for oscillation detection and diagnosis, however, their efficiencies are below desirable when applied on real plants. This work aims the refinement of the literature on oscillation detection and diagnosis in process industries. For this, the work begins with a complete literature review in the area of oscillation detection with the purpose of presenting, classifying and discussing pros and cons of each technique in order to facilitate the choice by engineers. Following, ten oscillation detection methods are evaluated on industrial data. It is verified that the accuracy of the techniques is still unsatisfactory, and the main cause is the variety of characteristics found in industrial data. Thus, the work proceeds with the proposal of three techniques. The first is an oscillation detection technique based on artificial intelligence, in which a model is trained by examples, different from traditional techniques based on algorithms. This approach ensures that the technique covers a greater number of characteristics found on industrial data. The second technique aims simultaneous oscillation detection and diagnosis. For this, the pattern of the output signals of the controller and output of the process is classified by a model based on neural networks. The latter technique solves a specific problem in the area: the diagnosis in signals with low sampling. The development of the technique was based on the resolution of a real problem in a Brazilian refinery. The three proposed techniques were tested on industrial data and returned better results when compared to traditional techniques
Detecção e diagnóstio de oscilações em malhas de controle
Em indústrias de processos, a oscilação é um problema de grande incidência que degrada a rentabilidade da planta. Essas indústrias possuem tipicamente entre 500 e 5000 malhas de controle. O número elevado dificulta a inspeção individual de cada malha para a detecção de falhas o que cria a necessidade por técnicas automáticas. Nos últimos 30 anos, dezenas dessas técnicas foram propostas para a detecção e diagnóstico da oscilação, mesmo assim, suas eficiências são abaixo da desejável quando aplicadas em plantas reais. Este trabalho tem por objetivo o refinamento da literatura na detecção e diagnóstico da oscilação em indústrias de processos. Para isso, o trabalho inicia com uma revisão completa da literatura na área da detecção da oscilação com o propósito de apresentar, classificar e discutir prós e contras de cada técnica de modo a facilitar sua escolha pelo engenheiro. A seguir, dez métodos de detecção da oscilação são avaliados a dados industriais. Nota-se que a acuracidade das técnicas ainda é insatisfatória e a principal causa é a variedade de características encontradas em dados industriais. Assim, o trabalho prossegue com a proposta de três técnicas. A primeira é uma técnica de detecção da oscilação baseada em inteligência artificial onde um modelo é treinado a partir de exemplos, diferente das técnicas tradicionais baseadas algoritmos. Essa abordagem garante que a técnica cubra um maior número de características dos dados industriais. A segunda técnica objetiva a detecção e o diagnóstico simultâneo da oscilação. Para isso, o padrão dos sinais de saída do controlador e saída do processo é classificado por modelo baseado em redes neurais. A última técnica resolve um problema específico na área: o diagnóstico em sinais de baixa amostragem. O desenvolvimento da técnica foi baseado na resolução de um problema real em uma refinaria brasileira. As três técnicas propostas foram testadas a dados industriais e retornaram melhores resultados quando comparadas a técnicas tradicionais.In process industries, oscillation is a problem of high incidence that degrades plant profitability. These industries typically have between 500 and 5000 control loops. The high number turns difficult the individual inspect of each loop for fault detection, which makes necessary the use of automatic techniques. In the last 30 years, dozens of these techniques have been proposed for oscillation detection and diagnosis, however, their efficiencies are below desirable when applied on real plants. This work aims the refinement of the literature on oscillation detection and diagnosis in process industries. For this, the work begins with a complete literature review in the area of oscillation detection with the purpose of presenting, classifying and discussing pros and cons of each technique in order to facilitate the choice by engineers. Following, ten oscillation detection methods are evaluated on industrial data. It is verified that the accuracy of the techniques is still unsatisfactory, and the main cause is the variety of characteristics found in industrial data. Thus, the work proceeds with the proposal of three techniques. The first is an oscillation detection technique based on artificial intelligence, in which a model is trained by examples, different from traditional techniques based on algorithms. This approach ensures that the technique covers a greater number of characteristics found on industrial data. The second technique aims simultaneous oscillation detection and diagnosis. For this, the pattern of the output signals of the controller and output of the process is classified by a model based on neural networks. The latter technique solves a specific problem in the area: the diagnosis in signals with low sampling. The development of the technique was based on the resolution of a real problem in a Brazilian refinery. The three proposed techniques were tested on industrial data and returned better results when compared to traditional techniques
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Scalable Inference in Latent Gaussian Process Models
Latente Gauß-Prozess-Modelle (latent Gaussian process models) werden von Wissenschaftlern benutzt, um verborgenen Muster in Daten zu er- kennen, Expertenwissen in probabilistische Modelle einfließen zu lassen und um Vorhersagen über die Zukunft zu treffen. Diese Modelle wurden erfolgreich in vielen Gebieten wie Robotik, Geologie, Genetik und Medizin angewendet. Gauß-Prozesse definieren Verteilungen über Funktionen und können als flexible Bausteine verwendet werden, um aussagekräftige probabilistische Modelle zu entwickeln. Dabei ist die größte Herausforderung, eine geeignete Inferenzmethode zu implementieren. Inferenz in probabilistischen Modellen bedeutet die A-Posteriori-Verteilung der latenten Variablen, gegeben der Daten, zu berechnen. Die meisten interessanten latenten Gauß-Prozess-Modelle haben zurzeit nur begrenzte Anwendungsmöglichkeiten auf großen Datensätzen.
In dieser Doktorarbeit stellen wir eine neue effiziente Inferenzmethode für latente Gauß-Prozess-Modelle vor. Unser neuer Ansatz, den wir augmented variational inference nennen, basiert auf der Idee, eine erweiterte (augmented) Version des Gauß-Prozess-Modells zu betrachten, welche bedingt konjugiert (conditionally conjugate) ist. Wir zeigen, dass Inferenz in dem erweiterten Modell effektiver ist und dass alle Schritte des variational inference Algorithmus in geschlossener Form berechnet werden können, was mit früheren Ansätzen nicht möglich war. Unser neues Inferenzkonzept ermöglicht es, neue latente Gauß-Prozess- Modelle zu studieren, die zu innovativen Ergebnissen im Bereich der Sprachmodellierung, genetischen Assoziationsstudien und Quantifizierung der Unsicherheit in Klassifikationsproblemen führen.Latent Gaussian process (GP) models help scientists to uncover hidden structure in data, express domain knowledge and form predictions about the future. These models have been successfully applied in many domains including robotics, geology, genetics and medicine. A GP defines a distribution over functions and can be used as a flexible building block to develop expressive probabilistic models. The main computational challenge of these models is to make inference about the unobserved latent random variables, that is, computing the posterior distribution given the data. Currently, most interesting Gaussian process models have limited applicability to big data.
This thesis develops a new efficient inference approach for latent GP models. Our new inference framework, which we call augmented variational inference, is based on the idea of considering an augmented version of the intractable GP model that renders the model conditionally conjugate. We show that inference in the augmented model is more efficient and, unlike in previous approaches, all updates can be computed in closed form. The ideas around our inference framework facilitate novel latent GP models that lead to new results in language modeling, genetic association studies and uncertainty quantification in classification tasks
- …
