1,720,971 research outputs found
Non-negative matrix factorisation: Algorithms and applications
Non-negative matrix factorisation (NMF) is attractive in data analysis because it can produce a sparse and parts based representation of the data. In this thesis we investigate and demonstrate solutions to several aspects of NMF. In particular, we consider the oft overlooked issue of model selection utilising a principled approach using information theory, provide a method for including external data into the NMF formulation, implement an autoencoder framework that can perform variants of NMF, and extend this autoencoder approach using variational methods to produce a probabilistic form of NMF. Two problems of model selection are explored in this thesis by taking a minimum description length (MDL) approach. First we produce and demonstrate a method using MDL to provide an estimate of the optimal subspace size to project onto. Secondly we extend this work by utilising MDL within the objective function for NMF to provide an automatic method of regularising the factorised matrices. The final representation should then be produced from a principled trade-off between accuracy and complexity reducing the problem of over-fitting. In standard NMF we factorise our data into two matrices without consideration of any external sources of information. If we could include these exogenous drivers into the NMF formulation it might enable an improved factorisation to be formed using information not directly available to the internal data. Our solution to this problem, called XNMF, finds a combined representation which includes these external drivers by utilising an extended version of the popular multiplicative update method. We prove theoretically that our method is guaranteed to reduce the objective function monotonically and that it also does so empirically by testing on financial data. In addition, we demonstrate that XNMF produces an improved representation and that it may produce better clusters in the data than standard NMF. We also investigate the broader utility of XNMF by application to a problem in biology - spatial proteomics. We demonstrate that there are useful data analysis advantages to using XNMF and that there may be many other biological applications of this technique. Our next contribution is to explore the use of autoencoders in producing the NMF factorisation (AE-NMF). We provide a set of methods that can produce the two factorised matrices and investigate advantages of using AE-NMF. In particular, we show that AENMF allows for the easy extension of NMF to perform a wider variety of tasks and with different objective functions. Finally, we produce a method of combining AE-NMF with variational autoencoders to produce a probabilistic version of NMF. This method, unlike standard NMF, allows us to: generate new data; provides a probabilistic representation between the input and latent space; gives some sense of the uncertainty in our representation; and produces a representation that is regularised in a principled manner. Unlike some other forms of probabilistic NMF this approach does not require the use of sampling such as the Markov Chain Monte Carlo technique
A method of integrating spatial proteomics and protein-protein interaction network data
The increase in quantity of spatial proteomics data requires a range of analytical techniques to effectively analyse the data. We provide a method of integrating spatial proteomics data together with protein-protein interaction (PPI) networks to enable the extraction of more information. A strong relationship between spatial proteomics and PPI network data was demonstrated. Then a method of converting the PPI network into vectors using spatial proteomics data was explained which allows the integration of the two datasets. The resulting vectors were tested using machine learning techniques and reasonable predictive accuracy was found
Rank selection in non-negative matrix factorization using minimum description Length
Nonnegative matrix factorization (NMF) is primarily a linear dimensionality reduction technique that factorizes a nonnegative data matrix into two smaller nonnegative matrices: one that represents the basis of the new subspace and the second that holds the coefficients of all the data points in that new space. In principle, the nonnegativity constraint forces the representation to be sparse and parts based. Instead of extracting holistic features from the data, real parts are extracted that should be significantly easier to interpret and analyze. The size of the new subspace selects how many features will be extracted from the data. An effective choice should minimize the noise while extracting the key features. We propose a mechanism for selecting the subspace size by using a minimum description length technique. We demonstrate that our technique provides plausible estimates for real data as well as accurately predicting the known size of synthetic data. We provide an implementation of our code in a Matlab format
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Non-negative matrix factorization with exogenous inputs for modeling financial data
Non-negative matrix factorization (NMF) is an effective dimensionality reduction technique that extracts useful latent spaces from positive value data matrices. Constraining the factors to be positive values, and via additional regularizations, sparse representations, sometimes interpretable as part-based representations have been derived in a wide range of applications. Here we propose a model suitable for the analysis of multi-variate financial time series data in which the variation in data is explained by latent subspace factors and contributions from a set of observed macro-economic variables. The macro-economic variables being external inputs, the model is termed XNMF (eXogenous inputs NMF). We derive a multiplicative update algorithm to learn the factorization, empirically demonstrate that it converges to useful solutions on real data and prove that it is theoretically guaranteed to monotonically reduce the objective function. On share prices from the FTSE 100 index time series, we show that the proposed model is effective in clustering stocks in similar trading sectors together via the latent representations learned
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
