1,720,986 research outputs found

    Customer churn prediction - A case study in retail banking

    No full text
    This work focuses on one of the central topics in customer relationship management (CRM): transfer of valuable customers to a competitor. Customer retention rate has a strong impact on customer lifetime value, and understanding the true value of a possible customer churn will help the company in its customer relationship management. Customer value analysis along with customer churn predictions will help marketing programs target more specific groups of customers. We predict customer churn with logistic regression techniques and analyze the churning and nonchurning customers by using data from a consumer retail banking company. The result of the case study show that using conventional statistical methods to identify possible churners can be successful

    Improving Worker Performance with Human-Centered Data Science

    Get PDF
    Advances in information technologies not only provide novel tools to support work in the traditional sectors; they also create additional employment opportunities in the modern workforce where work contexts have been largely changed. All these changes call for new efforts to study worker performance. Indeed, information technologies, especially data science techniques, render unprecedented large-scale rich data and sophisticated analytic tools to investigate worker performance. However, it remains unclear how we can combine the strengths of big data analytics in data science and our existing knowledge in social science to enhance worker performance. In this dissertation, we propose a human-centered data science framework that integrates machine learning, causal inference, field experiments, and social science theories: First, machine learning (with counterfactual reasoning) enables the prediction (and explanation) of human behavior in work practice via large-scale data analysis. Existing insights from social theories can further enhance its predictive power by informing feature construction, model architecture, and model explanation. Field experiments can help to evaluate the effectiveness of these models in real-world practices. Second, field experiments perform precise interventions and establish causality with randomized controlled trials. Yet, the experimental analysis mainly supports the understanding of treatment effects at aggregate levels, such as average treatment effect. Machine learning empowers more sophisticated analyses of experimental data by revealing heterogeneous effects at a finer granularity, such as individual treatment effects. Third, while these data-driven discoveries complement social science theories and provide rich insights for describing, explaining, and predicting human behavior, they require rigorous analytic tools, such as experiments and machine learning, to validate or disconfirm their applicability in specific contexts. In addition to testing theories, causal insights derived from field experiments and counterfactual machine learning models could support the development of new theories that better reflect reality. To exemplify the various applications of this framework in both traditional sectors and in the modern workforce, we present three empirical studies: developing machine learning models to improve the outreach performance for government specialists, leveraging a field experiment to enhance the performance of the gig economy workers, and using counterfactual machine learning to unpack individual treatment effects of field experiments on worker performance in the gig economy. These studies illustrate that the framework of human-centered data science is effective and flexible in increasing worker performance.PhDInformationUniversity of Michigan, Horace H. Rackham School of Graduate Studieshttp://deepblue.lib.umich.edu/bitstream/2027.42/169835/1/tengye_1.pd

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis

    A Machine Learning Based System for Semi-Automatically Redacting Documents

    No full text
    Redacting text documents has traditionally been a mostly manual activity, making it expensive and prone to disclosure risks. This paper describes a semi-automated system to en- sure a specified level of privacy in text data sets. Recent work has attempted to quantify the likelihood of privacy breaches for text data. We build on these notions to provide a means of obstructing such breaches by framing it as a multi-class classification problem. Our system gives users fine-grained control over the level of privacy needed to obstruct sensi- tive concepts present in that data. Additionally, our system is designed to respect a user-defined utility metric on the data (such as disclosure of a particular concept), which our methods try to maximize while anonymizing. We describe our redaction framework, algorithms, as well as a prototype tool built in to Microsoft Word that allows enterprise users to redact documents before sharing them internally and obscure client specific information. In addition we show experimen- tal evaluation using publicly available data sets that show the effectiveness of our approach against both automated attack- ers and human subjects.The results show that we are able to preserve the utility of a text corpus while reducing disclosure risk of the sensitive concept

    Machine learning

    No full text
    This chapter introduces you to the value of machine learning in the social sciences, particularly focusing on the overall machine learn- ing process as well as clustering and classification methods. You will get an overview of the machine learning pipeline and methods and how those methods are applied to solve social science problems. The goal is to give an intuitive explanation for the methods and to provide practical tips on how to use them in practice

    Dispelling the Myths Behind First-author Citation Counts

    Get PDF
    We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more sophisticated methods

    Author Index

    No full text
    Nao informado
    corecore