1,721,135 research outputs found

    An outlier detection approach in the presence of uncertainty

    No full text
    Un des aspects de complexité des nouvelles données, issues des différents systèmes de traitement,sont l’imprécision, l’incertitude, et l’incomplétude. Ces aspects ont aggravés la multiplicité etdissémination des sources productrices de données, qu’on observe facilement dans les systèmesde contrôle et de monitoring. Si les outils de la fouille de données sont devenus assez performants avec des données dont on dispose de connaissances a priori fiables, ils ne peuvent pas êtreappliqués aux données où les connaissances elles mêmes peuvent être entachées d’incertitude etd’imprécision. De ce fait, de nouvelles approches qui prennent en compte cet aspect vont certainement améliorer les performances des systèmes de fouille de données, dont la détection desoutliers, objet de notre recherche dans le cadre de cette thèse. Cette thèse s’inscrit dans cette optique, à savoir la proposition d’une nouvelle méthode pourla détection d’outliers dans les données incertaines et/ou imprécises. En effet, l’imprécision etl’incertitude des expertises relatives aux données d’apprentissage, est un aspect de complexitédes données. Pour pallier à ce problème particulier d’imprécision et d’incertitude des donnéesexpertisées, nous avons combinés des techniques issues de l’apprentissage automatique, et plusparticulièrement le clustering, et des techniques issues de la logique floue, en particulier les ensembles flous, et ce, pour pouvoir projeter de nouvelles observations, sur les clusters des donnéesd’apprentissage, et après seuillage, pouvoir définir les observations à considérer comme aberrantes(outliers) dans le jeu de données considéré.Concrètement, en utilisant les tables de décision ambigües (TDA), nous sommes partis des indices d’ambigüité des données d’apprentissage pour calculer les indices d’ambigüités des nouvellesobservations (données de test), et ce en faisant recours à l’inférence floue. Après un clustering del’ensemble des indices d’ambigüité, une opération α-coupe, nous a permis de définir une frontièrede décision au sein des clusters, et qui a été utilisée à son tour pour catégoriser les observations,en normales (inliers) ou aberrantes (outliers). La force de la méthode proposée réside dans sonpouvoir à traiter avec des données d’apprentissage imprécises et/ou incertaines en utilisant uniquement les indices d’ambigüité, palliant ainsi aux différents problèmes d’incomplétude des jeuxde données. Les métriques de faux positifs et de rappel, nous ont permis d’une part d’évaluer lesperformances de notre méthode, et aussi de la paramétrer selon les choix de l’utilisateur.One of the complexity aspects of the new data produced by the different processing systems is the inaccuracy, the uncertainty, and the incompleteness. These aspects are aggravated by the multiplicity and the dissemination of data-generating sources, that can be easily observed within various control and monitoring systems. While the tools of data mining have become fairly efficient with data that have reliable prior knowledge, they cannot be applied to data where the knowledge itself may be tainted with uncertainty and inaccuracy. As a result, new approaches that take into account this aspect will certainly improve the performance of data mining systems, including the detection of outliers,which is the subject of our research in this thesis.This thesis deals therefore with a particular aspect of uncertainty and accuracy, namely the proposal of a new method to detect outliers in uncertain and / or inaccurate data. Indeed, the inaccuracy of the expertise related to the learning data, is an aspect of complexity. To overcome this particular problem of inaccuracy and uncertainty of the expertise data, we have combined techniques resulting from machine learning, especially clustering, and techniques derived from fuzzy logic, especially fuzzy sets. So we will be able to project the new observations, on the clusters of the learning data, and after thresholding, defining the observations to consider as aberrant (outliers) in the considered dataset.Specifically, using ambiguous decision tables (ADTs), we proceeded from the ambiguity indices of the learning data to compute the ambiguity indices of the new observations (test data), using the Fuzzy Inference. After clustering, the set of ambiguity indices, an α-cut operation allowed us to define a decision boundary within the clusters, which was used in turn to categorize the observations as normal (inliers ) or aberrant (outliers). The strength of the proposed method lies in its ability to deal with inaccurate and / or uncertain learning data using only the indices of ambiguity, thus overcoming the various problems of incompleteness of the datasets. The metrics of false positives and recall, allowed us on one hand to evaluate the performances of our method, and also to parameterize it according to the choices of the user

    imperfect and multi-sources data merging : application to spatial qualified agricultural practices

    No full text
    Notre thèse s'inscrit dans le cadre de la mise en place d'un observatoire des pratiques agricoles dans le bassin versant de la Vesle. L'objectif de ce système d'information agri-environnemental est de comprendre les pratiques responsables de la pollution de la ressource en eau par les pesticides d'origine agricole sur le territoire étudié et de fournir des outils pertinents et pérennes pour estimer leurs impacts. Notre problématique concerne la prise en compte de l'imperfection dans le processus de la fusion de données multi-sources et imparfaites. En effet, l'information sur les pratiques n'est pas exhaustive et ne fait pas l'objet d'une déclaration, il nous faut donc construire cette connaissance par l'utilisation conjointe de sources multiples et de qualités diverses en intégrant dans le système d'information la gestion de l'information imparfaite. Dans ce contexte, nous proposons des méthodes pour une reconstruction spatialisée des informations liées aux pratiques agricoles à partir de la télédétection, du RPG, d'enquêtes terrain et de dires d'experts, reconstruction qualifiée par une évaluation de la qualité de l'information. Par ailleurs, nous proposons une modélisation conceptuelle des entités agronomiques imparfaites du système d'information en nous appuyant sur UML et PERCEPTORY. Nous proposons ainsi des modèles de représentation de l'information imparfaite issues des différentes sources d'information à l'aide soit des ensembles flous, soit de la théorie des fonctions de croyance et nous intégrons ces modèles dans le calcul d'indicateurs agri-environnementaux tels que l'IFT et le QSA.Our thesis is part of a regional project aiming the development of a community environmental information system for agricultural practices in the watershed of the Vesle. The objective of this observatory is 1) to understand the practices of responsible of the water resource pollution by pesticides from agriculture in the study area and 2) to provide relevant and sustainable tools to estimate their impacts. Our open issue deals with the consideration of imperfection in the process of merging multiple sources and imperfect data. Indeed, information on practices is not exhaustive and is not subject to return, so we need to build this knowledge through the combination of multiple sources and of varying quality by integrating imperfect information management information in the system. In this context, we propose methods for spatial reconstruction of information related to agricultural practices from the RPG remote sensing, field surveys and expert opinions, skilled reconstruction with an assessment of the quality of the information. Furthermore, we propose a conceptual modeling of agronomic entities' imperfect information system building on UML and PERCEPTORY.We provide tools and models of representation of imperfect information from the various sources of information using fuzzy sets and the belief function theory and integrate these models into the computation of agri-environmental indicators such as TFI and ASQ

    Collaborative approaches for complex data classification

    No full text
    La présente thèse s'intéresse à la classification collaborative dans un contexte de données complexes, notamment dans le cadre du Big Data, nous nous sommes penchés sur certains paradigmes computationels pour proposer de nouvelles approches en exploitant des technologies de calcul intensif et large echelle. Dans ce cadre, nous avons mis en oeuvre des classifieurs massifs, au sens où le nombre de classifieurs qui composent le multi-classifieur peut être tres élevé. Dans ce cas, les méthodes classiques d'interaction entre classifieurs ne demeurent plus valables et nous devions proposer de nouvelles formes d'interactions, qui ne se contraignent pas de prendre la totalité des prédictions des classifieurs pour construire une prédiction globale. Selon cette optique, nous nous sommes trouvés confrontés à deux problèmes : le premier est le potientiel de nos approches à passer à l'echelle. Le second, relève de la diversité qui doit être créée et maintenue au sein du système, afin d'assurer sa performance. De ce fait, nous nous sommes intéressés à la distribution de classifieurs dans un environnement de Cloud-computing, ce système multi-classifieurs est peut etre massif et ses propréités sont celles d'un système complexe. En terme de diversité des données, nous avons proposé une approche d'enrichissement de données d'apprentissage par la génération de données de synthèse, à partir de modèles analytiques qui décrivent une partie du phenomène étudié. Aisni, la mixture des données, permet de renforcer l'apprentissage des classifieurs. Les expérientations menées ont montré un grand potentiel pour l'amélioration substantielle des résultats de classification.This thesis focuses on the collaborative classification in the context of complex data, in particular the context of Big Data, we used some computational paradigms to propose new approaches based on HPC technologies. In this context, we aim at offering massive classifiers in the sense that the number of elementary classifiers that make up the multiple classifiers system can be very high. In this case, conventional methods of interaction between classifiers is no longer valid and we had to propose new forms of interaction, where it is not constrain to take all classifiers predictions to build an overall prediction. According to this, we found ourselves faced with two problems: the first is the potential of our approaches to scale up. The second, is the diversity that must be created and maintained within the system, to ensure its performance. Therefore, we studied the distribution of classifiers in a cloud-computing environment, this multiple classifiers system can be massive and their properties are those of a complex system. In terms of diversity of data, we proposed a training data enrichment approach for the generation of synthetic data from analytical models that describe a part of the phenomenon studied. so, the mixture of data reinforces learning classifiers. The experimentation made have shown the great potential for the substantial improvement of classification results

    Interest of using phylogenesis in building digital neuro-genetic neural networks entity

    No full text
    Les réseaux de neurones artificiels sont entraînés en imitant la plasticité synaptique biologique, mais ces algorithmes ont des limites : une méthode de construction des modèles d’IA considérée comme trop empirique, des problèmes d’apprentissage liés à la propagation d’erreurs lorsque la topologie devient complexe. Pour tenter d’améliorer ces modes d'apprentissage, une nouvelle méthode appelée Phylogenetic Replay Learning (PRL) est présentée dans cette thèse. L'objectif est d'en évaluer l'efficacité d’apprentissage en utilisant un algorithme de planification génétique de l’évolution de la topologie lors de l’apprentissage. …/…Les expériences démontrent que l’approche phylogénétique, PRL, est capable de produire un modèle plus performant qu’avec la méthode classique dite DDL. Second point, l'entropie de Shannon des poids des modèles générés par la PRL montre que les couches contiennent une meilleure répartition statistique des informations que lorsqu'un modèle est entraîné avec la DDL. …/…Les différentes étapes constituant la recherche : …/…1. Évaluation de la méthode PRL : en s'appuyant sur le chemin évolutif (préalablement enregistré). …/…2. Évaluation de la méthode d’apprentissage DDL : en partant du modèle Champion trouvé précédemment, on le réentraîne X fois avec DDL. …/…3. Comparaison des résultats obtenus via les deux méthodes. S’il y a une différence, on identifie laquelle et on en recherche la ou les raison(s). …/…4. Tests de transférabilité : on valide les précédentes expériences en utilisant de nouveaux jeux de données. …/ 5. Tests de reproductibilité : on refait les 4 premières étapes en partant d’un nouveau modèle initial, d’un nouveau chemin phylogénétique.Artificial neural networks are trained by imitating biological synaptic plasticity, but these algorithms have limitations: a topology construction method deemed too empirical, error propagation issues when the topology becomes complex, or the initial assignment of random weights. To improve these learning methods, a new method called Phylogenetic Replay Learning (PRL) has been developed and presented during this thesis. It takes inspiration from natural evolutionary mechanisms. The goal of this thesis is to evaluate the effectiveness of this new method using a genetic planning simulation algorithm for the evolution of a brain. …/…These experiments demonstrate that the phylogenetic approach, PRL, can produce a model more effective than with the DDL method. Secondly, the Shannon entropy of the weights of the models generated by PRL shows that the deeper layers statistically contain more information than when a model is trained traditionally (with the DDL). …/…The different steps constituting the research are: …/…1. Evaluation of the "phylogenetic replay learning" method, PRL: the initial model is retrained X times following the same recorded phylogenetic path. …./…2. Evaluation of the direct learning method (classical), DDL: retrained X times using the DDL method. …/…3. Comparison of the results obtained via the two methods. If there is a difference, it is identified and the reason(s) for it are sought. …/…4. Transferability tests: the previous experiments are validated using new datasets. …/…5. Reproducibility tests: new experiments are conducted starting from new initial model, new phylogenetic path, and thus a new Champion

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis

    Dispelling the Myths Behind First-author Citation Counts

    Get PDF
    We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more sophisticated methods

    Author Index

    No full text
    Nao informado
    corecore