1,720,960 research outputs found
Feature Selection in High Dimension Sample Spaces
Classification and multivariate class prediction problems arise in many scientific disciplines often relying on large data sets in the form of sample vectors of high dimension. In order to reduce dimensionality and effectively treat the problem, we may try to pre-process data and extract a minimum subset of feature variables that is sufficient for items identification, a problem addressed in literature as Minimum Test Set and known to be NP-hard by a reduction to Set Cover. We use a procedure based on computing mutual information between sets of symbolic feature variables and we present results obtained by applying the procedure over a previously discretised data set from integrated circuit industry collecting wafer data by sampling chips over a large number of real valued attribute variables (typically about a thousand). The experimentation gives evidence that the procedure effectively and feasibly yields a small number of features that provide sufficient information for
chip failure prediction
Customer Return Detection with Features Selection
We address the semiconductor industry problem of detecting microchips that escape production tests but are returned by customers as non-functional. This problem deals with analyzing high dimensional unbalanced databases collecting only a very small number of customer return samples. We show how to construct a model for effectively discriminating, based on wafer probe test data, potential customer returns from other good chips at the cost of a low overkill, where a model is a pair consisting of a selected set of wafer probe tests with minimal redundancy and a 1-class-SVM (Support Vector Machine) with optimal kernel parameters. We report about an experimentation on real data from EWS (Electronic Wafer Sort) test and customer returns showing the capability of predicting customer returns at cost of a relatively low overkill
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Fibrillazione ventricolare: sopravvivenza raddoppiata se il first responders è un laico: esperienza di Piacenza
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
PredictMed-epilepsy: A multi-agent based system for epilepsy detection and prediction in neuropediatrics
Background and objective: Epileptic seizures are associated with a higher incidence of Developmental Disabilities and Cerebral Palsy. Early evaluation and management of epilepsy is strongly recommended. We propose and discuss an application to predict epilespy (PredictMed-Epilepsy) and seizures via a deep-learning module (PredictMed-Seizures) encompassed within a multi-agent based healthcare system (PredictMed-MHS); this system is meant, in perspective, to be integrated into a clinical decision support system (PredictMed-CDSS). PredictMed-Epilespy, in particular, aims to identify factors associated with epilepsy in children with Developmental Disabilities and Cerebral Palsy by using a prediction-learning model named PredictMed. PredictMed-epilespy methods: We performed a longitudinal, multicenter, double-blinded, descriptive study of one hundred and two children with Developmental Disabilities and Cerebral Palsy (58 males, 44 females; 65 inpatients, 37 outpatients; 72 had epilepsy - 22 of intractable epilepsy, age: 16.6±1.2y, range: 12–18y). Data from 2005 to 2021 on Cerebral Palsy etiology, diagnosis, type of epilepsy and spasticity, clinical history, communication abilities, behaviors, intellectual disability, motor skills, and eating and drinking abilities were collected. The machine-learning model PredictMed was exploited to identify factors associated with epilepsy. The guidelines of the “Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis” Statement (TRIPOD) were followed. PredictMed-epilepsy results: Cerebral Palsy etiology [(prenatal > perinatal > postnatal causes) p=0.036], scoliosis (p=0.048), communication (p=0.018) and feeding disorders (p=0.002), poor motor function (p<0.001), intellectual disabilities (p=0.007), and type of spasticity [(quadriplegia/triplegia > diplegia > hemiplegia), p=0.002)] were associated with having epilepsy. The prediction model scored an average of 82% of accuracy, sensitivity, and specificity. Thus, PredictMed defined the computational phenotype of children with Developmental Disabilities/Cerebral Palsy at risk of epilepsy. Novel contribution of the work: We have been developing and we have prototypically implemented a Multi-Agent Systems (MAS) that encapsulates the PredictMed-Epilepsy module. More specifically, we have implemented the Patient Observing MAS (PoMAS), which, as a novelty w.r.t. the existing literature, includes a complex event processing module that provides real-time detention of short- and long-term events related to the patient's condition
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
