1,721,248 research outputs found
Advances in model based clustering for the social sciences
This dissertation attempts to gather the main research topics I engaged during my PhD, in collaboration with several national and international researchers. The primary focus of this work is to highlight the power of model based clustering for identifying latent structures in complex data and its usefulness in the social sciences. This methods have become increasingly popular in social science research as they allow for more accurate and nuanced understanding of complex data structures. In the thesis are presented 3 papers that contribute to the development and application of model-based clustering in social science research, covering a range of scenario. The thesis pays particular attention to the practical applications of the treated methods, providing insights that can improve our understanding of complex social phenomena.
The first chapter of this dissertation introduces the usefulness of clustering model to deal with the complexity of society, and aware of some of the main issues when analysing socio-economic data. Following this conceptual introduction, the second chapter delves more into the technical aspects of model based clustering and estimation. These first two chapters pave the road for the three developments presented thereafter. The third chapter includes the application of a Mixture of Matrix-Normals classification model to the Migrant Integration Policy Index (MIPEX), that measures and evaluates countries policies toward migrants’ integration over time. The used model is suitable for longitudinal data and allows for the identification of clusters of countries with similar patterns of migrant integration policies over time. The work is published in Alaimo et al. [2021a]. The fourth chapter uses MIPEX data too, but for a single year, and a finite mixtures of multivariate Gaussian is applied to identify groups of countries with a similar level of integration. Then, the relative proportion of immigrants held in prison among clusters is estimated, exploiting Fisher’s noncentral hypergeometric model. The aim of this work is test the existence of an association between countries’ level of integration of immigrants and the proportion of immigrants in prison. The work is currently in referral process. The fifth chapter introduce the work developed during my visiting research period at University of Lyon, Lyon 2. It specify the Bayesian partial membership model for soft clustering of multivariate data, namely when units have fractional membership to multiple groups. The model is specified for count data, and it is applied on the data of the bike sharing company of Washington DC and on the data of Serie A football players. The last chapter summarizes the main points of the dissertation, underlining the most relevant findings, the contributions, and stressing out how clustering models altogether yield a cohesive treatment of socio-economic data
Supercontinents, orogenesis and magmatism: a tribute to the career of J. Brendan Murphy
[EN] Special Publication 542 is a tribute to the remarkable career of J. Brendan Murphy and features 32 articles by 134 authors from 17 different countries; a testament to the high-profile and far-reaching influence of Brendan’s work. The topics are wide-ranging in accord with Brendan’s diverse research interests, but fall into three broad categories that encompass Brendan’s main fields of influence: (i) supercontinents and the supercontinent cycle, including reconstructions and modelling, (ii) orogenesis and terranes, with a focus on the Appalachian–Variscan and Central Asian orogenic belts and the oceans with which they are associated and (iii) magmatism and magmatic processes, with an emphasis on the geochemistry and isotopic compositions of magmas in arc and rift settings. Like Brendan’s own research, the scope of the papers spans the globe from Canada to China and ranges from regional field-based studies to conceptual global analyses. All of the articles, however, are focused on unravelling some critical aspect of geology or aimed at clarifying some crucial geologic process. Hence they also share a theme common to Brendan’s many contributions in emphasizing the importance of process-oriented research.Peer reviewe
Partial membership models for soft clustering of multivariate count data
The standard mixture modelling framework has been widely used to study heteroge
neous populations, by modelling them as being composed of a finite number of homoge
neous sub-populations. However, the standard mixture model assumes that each data point
belongs to one and only one mixture component, or cluster, but when data points have
fractional membership in multiple clusters this assumption is unrealistic. It is in fact con
ceptually very different to represent an observation as partly belonging to multiple groups
instead of belonging to one group with uncertainty. For this purpose, various soft clustering
approaches, or individual-level mixture models, have been developed. In this context, (6)
formulated the Bayesian partial membership model (BPM) as an alternative structure for
individual-level mixtures, which also captures partial membership in the form of attribute
specific mixtures, but does not assume a factorization over attributes. Our work proposes
using the BPM for soft clustering of count data. Learning and inference are carried out
using Markov chain Monte Carlo methods. The method is applied on Capital Bike share
data of Washington DC from 15 of June to 15 of July 2022
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
