1,720,959 research outputs found

    Modèles linéaires pour données fonctionnelles multivariées

    No full text
    In this thesis, we were interested in the problem of predicting a real or categorical variable using multivariate functional variables. In the existing literature, the proposed methods often assume the case of a single domain. This means that each dimension of the multivariate functional variable has the same domain of definition. This assumption limits their use to a limited number of application fields. Indeed, technological advances in data collection and storage have made it possible to observe several functional characteristics, sometimes of different natures, for the same statistical individual. To solve the prediction problem with this type of variables, we proposed two methods inspired by the PLS regression: MFPLS and TMFPLS. The first one is an extension of the PLS algorithm to the case of explanatory multivariate functional data, where the dimensions are potentially defined on different domains. This method can be used for regression and binary classification. The second method: TMFPLS, is a decision tree which can be used for more complex classification tasks (non-linear relationship between the target variable and the predictors, multiclass classification). These methods can be used for a wide range of applications; however, their interpretation becomes difficult when the predictors have numerous dimensions. This is typically the case when many sensors are used to measure a functional variable in several locations. Or more generally, when it comes to repeated functional data. In this case, we presented parsimonious methods based on the fusion penalty, to obtain more interpretable models. Applications on simulated data and real data (EEG, ECG, etc.) have demonstrated the good performance of our methods.Le cadre méthodologique de cette thèse est l'analyse de données fonctionnelles. Nous nous intéressons particulièrement au problème de la prédiction d'une variable réelle ou catégorielle à l'aide de variables fonctionnelles multivariées. Dans la littérature existante, les méthodes ont souvent recours au cadre restrictif du domaine unique. Il signifie que chaque dimension de la variable fonctionnelle multivariée a le même domaine de définition. Cette hypothèse limite leurs utilisations pour un certain nombre de domaines d'application. En effet, l'émergence des nouvelles technologies de collecte et de stockage de données a permis l'observation de plusieurs caractéristiques fonctionnelles, parfois de type différent, pour un même individu statistique. Pour répondre à la problématique de prédiction avec ce type de variables, nous proposons des méthodes basées sur la régression PLS : MFPLS et TMFPLS. Le premier est une extension de l'algorithme PLS au cas des données fonctionnelles multivariées explicatives, où les dimensions sont potentiellement définies sur différents domaines. Cette méthode peut être utilisée pour la régression et la classification binaire. La deuxième méthode : TMFPLS, est un arbre de décision qui permet de répondre à des tâches de classification plus complexes (relation non-linéaire entre la variable à prédire et les variables explicatives, plusieurs classes tolérées). Ces méthodes peuvent couvrir un éventail de problèmes dans les applications, cependant, les interpréter devient difficile lorsque les données explicatives ont de nombreuses dimensions. C'est le cas typiquement lorsque plusieurs capteurs sont utilisés pour mesurer une variable fonctionnelle suivant plusieurs localisations. Ou plus généralement, lorsque l'on a à faire à des données fonctionnelles répétées. Dans ce cas, nous présentons des méthodes parcimonieuses basées sur la pénalité fusion permettant d'obtenir une meilleure interprétation des modèles. Les applications sur des données simulées et données réelles (EEG, ECG, etc.) ont permis de démontrer la bonne performance de nos méthodes

    Modèles linéaires pour données fonctionnelles multivariées

    No full text
    In this thesis, we were interested in the problem of predicting a real or categorical variable using multivariate functional variables. In the existing literature, the proposed methods often assume the case of a single domain. This means that each dimension of the multivariate functional variable has the same domain of definition. This assumption limits their use to a limited number of application fields. Indeed, technological advances in data collection and storage have made it possible to observe several functional characteristics, sometimes of different natures, for the same statistical individual. To solve the prediction problem with this type of variables, we proposed two methods inspired by the PLS regression: MFPLS and TMFPLS. The first one is an extension of the PLS algorithm to the case of explanatory multivariate functional data, where the dimensions are potentially defined on different domains. This method can be used for regression and binary classification. The second method: TMFPLS, is a decision tree which can be used for more complex classification tasks (non-linear relationship between the target variable and the predictors, multiclass classification). These methods can be used for a wide range of applications; however, their interpretation becomes difficult when the predictors have numerous dimensions. This is typically the case when many sensors are used to measure a functional variable in several locations. Or more generally, when it comes to repeated functional data. In this case, we presented parsimonious methods based on the fusion penalty, to obtain more interpretable models. Applications on simulated data and real data (EEG, ECG, etc.) have demonstrated the good performance of our methods.Le cadre méthodologique de cette thèse est l'analyse de données fonctionnelles. Nous nous intéressons particulièrement au problème de la prédiction d'une variable réelle ou catégorielle à l'aide de variables fonctionnelles multivariées. Dans la littérature existante, les méthodes ont souvent recours au cadre restrictif du domaine unique. Il signifie que chaque dimension de la variable fonctionnelle multivariée a le même domaine de définition. Cette hypothèse limite leurs utilisations pour un certain nombre de domaines d'application. En effet, l'émergence des nouvelles technologies de collecte et de stockage de données a permis l'observation de plusieurs caractéristiques fonctionnelles, parfois de type différent, pour un même individu statistique. Pour répondre à la problématique de prédiction avec ce type de variables, nous proposons des méthodes basées sur la régression PLS : MFPLS et TMFPLS. Le premier est une extension de l'algorithme PLS au cas des données fonctionnelles multivariées explicatives, où les dimensions sont potentiellement définies sur différents domaines. Cette méthode peut être utilisée pour la régression et la classification binaire. La deuxième méthode : TMFPLS, est un arbre de décision qui permet de répondre à des tâches de classification plus complexes (relation non-linéaire entre la variable à prédire et les variables explicatives, plusieurs classes tolérées). Ces méthodes peuvent couvrir un éventail de problèmes dans les applications, cependant, les interpréter devient difficile lorsque les données explicatives ont de nombreuses dimensions. C'est le cas typiquement lorsque plusieurs capteurs sont utilisés pour mesurer une variable fonctionnelle suivant plusieurs localisations. Ou plus généralement, lorsque l'on a à faire à des données fonctionnelles répétées. Dans ce cas, nous présentons des méthodes parcimonieuses basées sur la pénalité fusion permettant d'obtenir une meilleure interprétation des modèles. Les applications sur des données simulées et données réelles (EEG, ECG, etc.) ont permis de démontrer la bonne performance de nos méthodes

    Comparing shape-based and pixel-based approaches for melanoma detection

    No full text
    In recent years, both the number and the size of image datasets have grown at an uncontrollable rate. This creates a serious challenge for the analysis of image databases, especially for researchers who may not have access to expensive and powerful supercomputers. In this research report, we study an alternative to pixel based analysis: a functional representation of the shapes of objects within images. This representation offers several advantages. By being of much lower dimension, it greatly reduces the computational cost of subsequent analyses; it is also far more interpretable and can leverage the extensive set of tools already developed for the analysis of multivariate functional data. We investigate this shape-based approach in the context of classification using a real dataset, the HAM10000 dataset, and our results demonstrate a clear computational benefit with similar predictive power

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis

    Classification of multivariate functional data on different domains with Partial Least Squares approaches

    No full text
    International audienceClassification of multivariate functional data is explored in this paper, particularly for functional data defined on different domains. Using the partial least squares (PLS) regression, we propose two classification methods. The first one uses the equivalence between linear discriminant analysis and linear regression. The second is a decision tree based on the first technique. Moreover, we prove that multivariate PLS components can be estimated using univariate PLS components. This offers an alternative way to calculate PLS for multivariate functional data. Finite sample studies on simulated data and real data applications show that our algorithms are competitive with linear discriminant on principal components scores and black-boxes models

    Fusion regression methods with repeated functional data

    Get PDF
    Linear regression and classification methods with repeated functional data are considered. For each statistical unit in the sample, a real-valued parameter is observed over time under different conditions related by some neighborhood structure (spatial, group, etc.). Two regression methods based on fusion penalties are proposed to consider the dependence induced by this structure. These methods aim to obtain parsimonious coefficient regression functions, by determining if close conditions are associated with common regression coefficient functions. The first method is a generalization to functional data of the variable fusion methodology based on the 1-nearest neighbor. The second one relies on the group fusion lasso penalty which assumes some grouping structure of conditions and allows for homogeneity among the regression coefficient functions within groups. Numerical simulations and an application of electroencephalography data are presented

    Interactions entre les services écosystémiques et l'utilisation des terres en France : une analyse statistique spatiale

    No full text
    International audienceThe provision of ecosystem services (ESs) is driven by land use and biophysical conditions and is thus intrinsically linked to space. Large-scale ES models, developed to inform policy makers on ES drivers, do not usually consider spatial autocorrelation that could be inherent to the distribution of these ESs or to the modeling process. The objective of this study is to estimate the drivers of ecosystem services in France using statistical models and show how taking into account spatial autocorrelation improves the predictive quality of these models. We study six regulating ESs (habitat quality index, water retention index, topsoil organic matter, carbon storage, soil erosion control, and nitrogen oxide deposition velocity) and three provisioning ESs (crop production, grazing livestock density, and timber removal). For each of these ESs, we estimated and compared five spatial statistical models to investigate the best specification (using statistical tests and goodness-of-fit metrics). Our results show that (1) taking into account spatial autocorrelation improves the predictive accuracy of all ES models (Δ R 2 ranging from 0.13 to 0.58); (2) land use and biophysical variables (weather and soil texture) are significant drivers of most ESs; (3) forest was the most balanced land use for provision of a diversity of ESs compared to other land uses (agriculture, pasture, urban, and others); (4) Urban area is the worst land use for provision of most ESs. Our findings imply that further studies need to consider spatial autocorrelation of ESs in land use change and optimization scenario simulations.La fourniture de services écosystémiques (SE) est déterminée par l'utilisation des terres et les conditions biophysiques et est donc intrinsèquement liée à l'espace. Les modèles d'ES à grande échelle, développés pour informer les décideurs politiques sur les moteurs des ES, ne prennent généralement pas en compte l'autocorrélation spatiale qui pourrait être inhérente à la distribution de ces ES ou au processus de modélisation. L'objectif de cette étude est d'estimer les moteurs des services écosystémiques en France en utilisant des modèles statistiques et de montrer comment la prise en compte de l'autocorrélation spatiale améliore la qualité prédictive de ces modèles. Nous étudions six SE régulateurs (indice de qualité de l'habitat, indice de rétention d'eau, matière organique de la couche arable, stockage du carbone, contrôle de l'érosion du sol et vitesse de dépôt des oxydes d'azote) et trois SE fournisseurs (production végétale, densité du bétail de pâturage et prélèvement de bois). Pour chacun de ces SE, nous avons estimé et comparé cinq modèles statistiques spatiaux afin d'étudier la meilleure spécification (en utilisant des tests statistiques et des mesures d'adéquation). Nos résultats montrent que (1) la prise en compte de l'autocorrélation spatiale améliore la précision prédictive de tous les modèles d'ES (Δ R 2 allant de 0,13 à 0,58) ; (2) l'utilisation des terres et les variables biophysiques (météo et texture du sol) sont des facteurs significatifs de la plupart des ES ; (3) la forêt était l'utilisation des terres la plus équilibrée pour la fourniture d'une diversité d'ES par rapport aux autres utilisations des terres (agriculture, pâturage, urbain et autres) ; (4) la zone urbaine est la pire utilisation des terres pour la fourniture de la plupart des ES. Nos résultats impliquent que d'autres études doivent prendre en compte l'autocorrélation spatiale des SE dans les simulations de scénarios de changement d'affectation des sols et d'optimisation
    corecore