Research Data UNIPD
Not a member yet
461 research outputs found
Sort by
Feature selection and molecular classification of cancer phenotypes: a comparative study
Classification of high dimensional gene expression data is key to the development of effective di-agnostic and prognostic tools. Feature selection involves finding the best subset with the highest power in predicting class labels. We here conducted a comparative study focused on different combinations of feature selectors (Chi-Squared, mRMR, Relief-F, Genetic Algorithms) and classi-fication learning algorithms (Random Forests, PLS-DA, SVM, Regularized Logistic/Multinomial Regression, kNN) to identify those with the best predictive capacity. The performance of each combination is evaluated through an empirical study on three benchmark cancer-related micro-array datasets. Our results first suggest that the quality of the data relevant to the target classes is key for the successful classification of cancer phenotypes. We also proved that, for a given classi-fication learning algorithm and dataset, all filters have a similar performance. Interestingly, fil-ters achieve comparable or even better results with respect to the GA-based wrappers, while also being easier to implement and faster. Taken together, our findings suggest that simple, well-established feature selectors in combination with optimized classifiers guarantee good per-formances, with no need for complicated and computationally demanding methodologie
Coronary Sinus Diameter to Predict Residual Congestion and Survival in Hemodialysis Patients (CORSICO Study): A Pilot Study
Database reporting clinical data collected from patient
Dataset on Italian Regional Ministers: socio-economic characteristics and political careers in different territorial levels (1995-2020)
The Dataset collects all information about socio-economic characteristics and political institutional offices covered by Regional Ministers in Ordinary Statute Regions from 1995 to 2020
Data availability material of "Lazzarin et al. - Influence of bed roughness on flow and turbulence structure around a partially-buried, isolated freshwater mussel"
This folder contains data of the figure shown in the paper
Mapping of local gambling initiatives in Italy
The dataset provides the complete enumeration of gambling policies implemented by Italian Municipalities between 2003 and 2021. The dataset comprises information on all municipalities existing in 2017 and following years (thus considering also merging). The following variables are available: Municipality ISTAT Code (ID), Municipality Name (Name) Province, Region, Researcher, and a series of variables identifying the number and the type of rulings adopted (“regolamento”, “ordinanza” and “delibera”). The rulings are distinguished between identified and downloaded or only identified (because the document is no longer available). Additional variables describe municipal activism with other administrative acts (Anti-gambling Manifesto, events, projects or tax reductions).
Overall, the dataset comprises 8031 units
Statistical characterization of erosion mechanics in shallow tidal environments
Dataset and results of the statistical characterization of erosion events in shallow tidal environment
Alarm logs of industrial packaging machines
The advent of the Industrial Internet of Things (IIoT) has led to the availability of huge amounts of data, that can be used to train advanced Machine Learning algorithms to perform tasks such as Anomaly Detection, Fault Classification and Predictive Maintenance. Even though not all pieces of equipment are equipped with sensors yet, usually most of them are already capable of logging warnings and alarms occurring during operation. Turning this data, which is easy to collect, into meaningful information about the health state of machinery can have a disruptive impact on the improvement of efficiency and up-time. The provided dataset consists of a sequence of alarms logged by packaging equipment in an industrial environment. The collection includes data logged by 20 machines, deployed in different plants around the world, from 2019-02-21 to 2020-06-17. There are 154 distinct alarm codes, whose distribution is highly unbalanced. This data can be used to address the following tasks: 1. Next alarm forecasting: this problem can be framed as a supervised multi-class classification task, or a binary classification task when a specific alarm code is considered. 2. Predicting alarms occurring in a future time frame: here the goal is to forecast the occurrence of certain alarm types in a future time window. Since many alarms can occur, this is a supervised multi-label classification. 3. Future alarm sequence prediction: here the goal is predicting an ordered sequence of future alarms, in a sequence-to-sequence forecasting scenario. 4. Anomaly Detection: the task is to detect abnormal equipment conditions, based on the pattern of alarms sequence. This task can be either unsupervised, if only the input sequence is considered, or supervised if future alarms are taken into account to assess whether or not there is an anomaly. All of the above tasks can also be studied from a continual learning perspective. Indeed, information about the serial code of the specific piece of equipment can be used to train the model; however, a scalable model should also be easy to apply to new machines, without the need of a new training from scratch
Distensibility of Deformable Aortic Replicas Assessed by an Integrated In-Vitro and In-Silico Approach - DATA
Raw data of the numerical simulation and video analysis of the paper "Distensibility of Deformable Aortic Replicas Assessed by an Integrated In-Vitro and In-Silico Approach"
Social networking behaviors differences between Italian gay and heterosexual men
The data were collected for a study that investigates differences between gay and heterosexual men regarding both social networking behaviors and addiction. Furthermore, it explores the possible effects of grandiose and vulnerable narcissism, fear of missing out, and physical appearance on social networking behaviors and addiction. A total of 586 Italian men (334 gay and 252 heterosexual) were recruited with snowball sampling and they completed an online questionnaire