1,720,958 research outputs found
Multi-view data integration by linear and non-linear dimensionality reduction
Technological advancements and global data sharing allow for the collection of information from multiple sources on the same samples. Such data are usually referred to as multi-view data, and the dataset from each source as data-view. Statistical properties, such as heterogeneity and noise, make the analysis of multi-view data challenging. An integrative analysis of the different data-views, provides an improved and more accurate understanding of the data. In this thesis, both linear and non-linear solutions are investigated for the visualisation, clustering and classification of multi-view data. In particular, various solutions that perform data integration through dimensionality reduction are explored. Sparse solutions of Canonical Correlation Analysis, a well-known linear integration approach on two data-views, are described, and extensions to the analysis of multiple data-views are proposed. Further, adaptations of non-linear dimensionality reduction (or manifold learning) methods for multi-view data are presented. The proposed algorithms are based on t-distributed Stochastic Neighbour Embedding (t-SNE), Locally Linear Embedding (LLE) and Isometric Feature Mapping (ISOMAP). Manifold learning approach multi-SNE, the multi-view extension based on t-SNE was found to be the best performing solution, providing accurate visualisations of the samples, confirmed both qualitatively and quantitatively. An extension of the algorithm that allows the classification of the samples in a semi-supervised manner is introduced. The uncommon notion of incorporating the response variables as an additional data-view is explored on both linear and non-linear solutions. This thesis ends with the analysis of single-cell multi-omics data, a new and challenging type of biological data. The proposed linear and non-linear integrative algorithms were implemented for the estimation of cell subtypes, cell identification, visualisation and other tasks. This thesis investigates the limitations and strengths of the proposed algorithms through various experiments on numerous real and synthetic multi-view data.Open Acces
Semi-supervised classification and visualisation of multi-view data
An increasing number of multi-view data are being published by studies in several fields. This type of data corresponds to multiple data-views, each representing a different aspect of the same set of samples. We have recently proposed multi-SNE, an extension of t-SNE, that produces a single visualisation of multi-view data. The multi-SNE approach provides low-dimensional embeddings of the samples, produced by being updated iteratively through the different data-views. Here, we further extend multi-SNE to a semi-supervised approach, that classifies unlabelled samples by
regarding the labelling information as an extra data-view. We look deeper into the performance, limitations and strengths of multi-SNE and its extension, S-multi-SNE, by applying the two methods on various multi-view datasets with different challenges. We show that by including the labelling information, the projection of the samples improves drastically and it is accompanied by a strong classification performance
Multi-view data visualisation via manifold learning
Non-linear dimensionality reduction can be performed by manifold learning approaches, such as Stochastic Neighbour Embedding (SNE), Locally Linear Embedding (LLE) and Isometric Feature Mapping (ISOMAP). These methods aim to produce two or three latent embeddings, primarily to visualise the data in intelligible representations. This manuscript proposes extensions of Student’s t-distributed SNE (t-SNE), LLE and ISOMAP, for dimensionality reduction and visualisation of multi-view
data. Multi-view data refers to multiple types of data generated from the same samples.
The proposed multi-view approaches provide more comprehensible projections of the samples compared
to the ones obtained by visualising each data-view separately. Commonly, visualisation is used for identifying underlying patterns within the samples. By incorporating the obtained low-dimensional embeddings from the multi-view manifold approaches into the K-means clustering algorithm, it is shown that clusters of the samples are accurately identified. Through extensive comparisons of novel and existing multi-view manifold learning algorithms on real and synthetic data, the proposed multi-view extension of t-SNE, named multi-SNE, is found to have the best performance, quantified both qualitatively and quantitatively by assessing the clusterings obtained.
The applicability of multi-SNE is illustrated by its implementation in the newly developed and challenging
multi-omics single-cell data. The aim is to visualise and identify cell heterogeneity and cell types in biological tissues relevant to health and disease. In this application, multi-SNE provides an improved performance over single-view manifold learning approaches and a promising solution for unified clustering of multi-omics single-cell data
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Subtypes estimation in single-cell multi-omics data via S-multi-SNE
Single-cell RNA sequencing (scRNA-seq) has rapidly become an essential method in modern biology, as it allows global characterisation of transcriptomics in individual cells from different tissues, or organisms. scRNA-seq data are collected from a single individual organism and they are composed as mRNA count matrices with rows representing the cells and columns the genes. New technologies further allow multi-omics data to be obtained on the same set of single cells.
In this paper, multi-SNE and S-multi-SNE, two multi-view dimensionality reduction approaches, are adapted for visualization and classification of cellular subtypes, two important challenges in the analysis of scRNA-seq data. S-multi-SNE has been adapted for cell subtype estimation of scRNA-seq data. In a series of experiments we illustrate that it is possible to identify cell subtypes by leveraging information from individuals from the same species, from different species, and when data are generated by different sequencing technologies. In the conducted analyses, we show that S-multi-SNE consistently outperformed several other machine learning techniques and, in most comparisons, surpassed existing reference-based single-cell classification algorithms.
We further illustrate how two single-cell multi-omics datasets, scATAC-seq and scRNA-seq, can be integrated together through S-multi-SNE for improved cell subtype identification
Multi-view biclustering via non-negative matrix tri-factorisation
Multi-view data is ever more apparent as methods for production, collection and storage of data become more feasible both practically and fiscally. However, not all features are relevant to describe the patterns for all individuals. Multi-view biclustering aims to simultaneously cluster both rows and columns, discovering clusters of rows as well as their view-specific identifying features. A novel multi-view biclustering approach based on non-negative matrix factorisation is proposed named ResNMTF. Demonstrated through extensive experiments on both synthetic and real datasets, ResNMTF successfully identifies both overlapping and non-exhaustive biclusters, without pre-existing knowledge of the number of biclusters present, and is able to incorporate any combination of shared dimensions across views. Further, to address the lack of a suitable bicluster-specific intrinsic measure, the popular silhouette score is extended to the bisilhouette score. The bisilhouette score is demonstrated to align well with known extrinsic measures, and proves useful as a tool for hyperparameter tuning as well as visualisation
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
