1,721,008 research outputs found
Approximate Data Mining Techniques on Clinical Data
The past two decades have witnessed an explosion in the number of medical and healthcare datasets available to researchers and healthcare professionals. Data collection efforts are highly required, and this prompts the development of appropriate data mining techniques and tools that can automatically extract relevant information from data. Consequently, they provide insights into various clinical behaviors or processes captured by the data. Since these tools should support decision-making activities of medical experts, all the extracted information must be represented in a human-friendly way, that is, in a concise and easy-to-understand form. To this purpose, here we propose a new framework that collects different new mining techniques and tools proposed. These techniques mainly focus on two aspects: the temporal one and the predictive one. All of these techniques were then applied to clinical data and, in particular, ICU data from MIMIC III database. It showed the flexibility of the framework, which is able to retrieve different outcomes from the overall dataset. The first two techniques rely on the concept of Approximate Temporal Functional Dependencies (ATFDs). ATFDs have been proposed, with their suitable treatment of temporal information, as a methodological tool for mining clinical data. An example of the knowledge derivable through dependencies may be "within 15 days, patients with the same diagnosis and the same therapy usually receive the same daily amount of drug". However, current ATFD models are not analyzing the temporal evolution of the data, such as "For most patients with the same diagnosis, the same drug is prescribed after the same symptom". To this extent, we propose a new kind of ATFD called Approximate Pure Temporally Evolving Functional Dependencies (APEFDs). Another limitation of such kind of dependencies is that they cannot deal with quantitative data when some tolerance can be allowed for numerical values. In particular, this limitation arises in clinical data warehouses, where analysis and mining have to consider one or more measures related to quantitative data (such as lab test results and vital signs), concerning multiple dimensional (alphanumeric) attributes (such as patient, hospital, physician, diagnosis) and some time dimensions (such as the day since hospitalization and the calendar date). According to this scenario, we introduce a new kind of ATFD, named Multi-Approximate Temporal Functional Dependency (MATFD), which considers dependencies between dimensions and quantitative measures from temporal clinical data. These new dependencies may provide new knowledge as "within 15 days, patients with the same diagnosis and the same therapy receive a daily amount of drug within a fixed range". The other techniques are based on pattern mining, which has also been proposed as a methodological tool for mining clinical data. However, many methods proposed so far focus on mining of temporal rules which describe relationships between data sequences or instantaneous events, without considering the presence of more complex temporal patterns into the dataset. These patterns, such as trends of a particular vital sign, are often very relevant for clinicians. Moreover, it is really interesting to discover if some sort of event, such as a drug administration, is capable of changing these trends and how. To this extent, we propose a new kind of temporal patterns, called Trend-Event Patterns (TEPs), that focuses on events and their influence on trends that can be retrieved from some measures, such as vital signs. With TEPs we can express concepts such as "The administration of paracetamol on a patient with an increasing temperature leads to a decreasing trend in temperature after such administration occurs". We also decided to analyze another interesting pattern mining technique that includes prediction. This technique discovers a compact set of patterns that aim to describe the condition (or class) of interest. Our framework relies on a classification model that considers and combines various predictive pattern candidates and selects only those that are important to improve the overall class prediction performance. We show that our classification approach achieves a significant reduction in the number of extracted patterns, compared to the state-of-the-art methods based on minimum predictive pattern mining approach, while preserving the overall classification accuracy of the model. For each technique described above, we developed a tool to retrieve its kind of rule. All the results are obtained by pre-processing and mining clinical data and, as mentioned before, in particular ICU data from MIMIC III database
Approximate Temporal Functional Dependencies on Clinical Data
The aim of this PhD project is to provide a framework for temporal data mining
Fermentazione con Bacillus Stearothermophilus: produzione di 2,3-butandiolo, studio della via metabolica ed applicazioni biocatalitiche
The here presented PhD regarded fermentation with the bacteria Bacillus stearothermophilus ATCC
2027 evaluating its potential production of 2,3-butanediol, a metabolic product with wide
applications, evaluating also the metabolic pathway and the involved enzyme.
B. steraothermophilus was previously successfully used in our Laboratory in biocatalysis obtaining
kinetic resolution of chiral alcohols via oxidation e stereo selective reduction of 1,2-diketons to S,Sdiols.
Growing substrate analysis evidenced bacteria capacity to produce significant amounts of 2,3-
butanediol (2,3-BD) and its precursor acetoin (AC) using sucrose as carbon source.
2,3-butanediol is widely used in many fields, from alimentary to polymers industries, beside its use
as bio-fuel or additive in fuels, increasing the relative importance of its production.
The first aim of my PhD activity was to verify real quality of 2,3-BD and AC produced by B.
stearothermophilus sucrose’s fermentation, and to flowingly screen other mono- and disaccharides
as carbon sources.
Experiments were conducted at fixed sucrose concentration (40, 30, 20, 10 gr/l) and results
indicated fermentation with 30 gr/l of sucrose as the best in terms of yield (about 100%) and carbon
consume (residual about 1 gr/l). This result was compared with those obtained using other carbon
sources as glucose, fructose, xylose, maltose, lactose and cellobiose, and cane molasses, chosen in
relation to their high production in agroindustrial processes.
Obtained results evidenced B. stearothermophilus complete sucrose fermentation and an only
partial fermentation of its constituents monosaccharides, fructose and glucose. The other sugars
give small amounts of the two metabolites.
An important point of the present research activity was the comprehension of biochemical
mechanisms allowing microorganism production of the cited metabolites. Gaschromatographic endproducts
analysis with a chiral column demonstrated B. stearothermophilus production of high
quantity of (R)-acetoin, (2R,3R)-butanediol and meso-butanediol.
Consequently, a study on the cited bacteria metabolic pathway was developed adopting methods
previously used with other microorganisms.
Similarly to other cogeneric bacteria, Bacillus sterarothermophilus presents two metabolic
pathways to produce butanediol. The firs, more diffused, “catabolic way” starts from pyruvate
derived from glicolysis and by the way of three enzymatic steps (condensation, decarboxylation and
reduction) produces 2,3-butanediol.
The second, less diffused, indicated as “butanediol cycle” starts from diacetil and produces
butanediol by the way of three enzymatic reactions (condensation, reduction, hydrolysis).
The present research evidenced a new S-stereo specific acetoin-reductase (AC-reductase) that in
association with the enzyme diacetil acetoin reductase (BSDR) previously used in biocatalysis,
produces meso-butanediol.
“Butanediol cycle” was confirmed by the presence of an acetil acetoin synthetase (AAC-synthetase)
capable of acetilacetoin (AAC) production from diacetil.
While AC-reductase was partially purified, AAC-synthetase was used raw in biocatalysis.
Other 1,2 dichetons, beside diacetil, were considered as possible starting-products to obtain α-
hydroxy-dichetons variably substituted. 3,4-hexanedione and 1-phenyl-2,3-propanedione were used
obtaining respectively 4-hydroxy-4-ethyl-3,5-heptanedione and 1-phenil-2-hydroxy-2-methyl-1,3-
butanedione. Obtained results evidenced AAC-synthetase efficiency in this catalysis showing an
almost total conversion in the first case (82%) and a lower product yield in the second one (45%)
with optical purity of chiral product about 40%.
The present PhD research activity allowed also publications on fermentation1, metabolic pathway2
and the formation of C-C linkage mediated by AAC-synthetase.3
References
1. P. P. GIOVANNINI, M. MANTOVANI, A. MEDICI, P. PEDRINI– Productions of 2,3-butandiol by Bacillus
sterothermophilus: fermentation and metabolic pathway. Proceedings of IBIC 2008, Chem. Eng. Transactions,
14, 281-286 (2008).
2. P.P. GIOVANNINI, M. MANTOVANI, M. FOGAGNOLO, S. MAIETTI, A. MEDICI, P. PEDRINI – Bacillus
stearothermophilus fermentation: the enzymatic route to 3R-hydroxy-2-butanone and meso-and 2R,3Rbutanediol.
Journal of Molecular Catalysis B: enzymatic, in stampa.
3. P.P. GIOVANNINI, M. MANTOVANI, A. MEDICI, P. PEDRINI - Enzymatic Carbon-Carbon Bond Formation:
Synthesis of a-Hydroxy-1,3-Diketones from the Corresponding 1,2-Diketones. Organic Letters, in stampa
FARmAPP: a process-driven solution to prevent and oppose illegal recruitment in agriculture in Northern Italy
Illegal recruitment in agriculture is an issue that affects many different aspects, from the workers’ physical and psychological health conditions to the overall economy. This phenomenon is particularly complex, and many disciplines are trying to face it. In the domain of computer science, one of the possibilities is to improve the systems used by the recruitment/temp agencies. In this study, we propose FARmAPP, a process-driven tool used by the recruitment agencies and farms. FARmAPP was developed using an agile approach with a direct contribution from three different recruitment agencies that operates in three Italian regions. FARmAPP collects and analyzes usage data to monitor “suspect” behaviors from the farms that could lead back to illegal recruitment or workers exploitation. We also created a new custom algorithm to analyze the CVs of the unemployed people to suggest the best candidates for each different job. After the development of FARmAPP, we trained over 80 agencies employees to use and manage FARmAPP autonomously. Their feedback was overall positive, and they stated that FARmAPP is a helpful tool to be included in their system when dealing with agricultural jobs
A Migration Framework for Active BPMN Processes in Healthcare
The Business Process Modeling Notation (BPMN) is a diagrammatical notation to describe complex process models, and it can be used as a common language among stakeholders. These stakeholders could intervene in the development of the processes, also in an agile environment. Thus, the same process could evolve through different versions over time. BPMN has been already adopted in the healthcare domain, where designing health care processes is fundamental for delivering optimal and efficient services to patients without overburdening health care professionals. Adapting and updating BPMN processes within an agile development is paramount in the rapidly evolving healthcare domain. The challenges of migrating a process to its revised versions have been discussed since the introduction of BPMN. However, there is a lack of migration policies that consider compensation strategies when migrating to a revised version at runtime, combined with healthcare- related migration risk classes. In this paper, we propose a methodological framework that includes migration strategies to adopt when the user is already running the process in a previous version. We propose compen-satory strategies to integrate the novelties of the revised process based on the actual users' completion status, with a specific focus on the addition, modification, and removal of tasks. Moreover, the proposed framework includes a color-coded risk classification system encapsulating migration risks and potential impact on patients. This system provides a visual and intuitive way to understand potential risks, enabling clinicians to make informed decisions about the migration strategy. Finally, we showed an application of the proposed framework in a real-world scenario, that is, through an ERAS-inspired prehabilitation program for pancreatic surgery currently developed at the Verona Pancreas Institute
Discovering Evolving Temporal Information: Theory and Application to Clinical Databases
Functional dependencies (FDs) allow us to represent database constraints, corresponding to requirements as “patients having the same symptoms undergo the same medical tests.” Some research eforts have focused on extending such dependencies to consider also temporal constraints such as “patients having the same symptoms undergo in the next period the same medical tests.” Temporal functional dependencies are able to represent such kind of temporal constraints in relational databases. Another extension for FDs allows one to represent approximate functional dependencies (AFDs), as “patients with the same symptoms generally undergo the same medical tests.” It enables data to deviate from the defned constraints according to a user-defned percentage. Approximate temporal functional dependencies (ATFDs) merge the concepts of temporal functional dependency and of approximate functional dependency. Among the diferent kinds of ATFD, the Approximate Pure Temporally Evolving Functional Dependencies (APE-FDs for short) allow one to detect patterns on the evolution of data in the database and to discover dependencies as “For most patients with the same initial diagnosis, the same medical test is prescribed after the occurrence of same symptom.” Mining ATFDs from large databases may be computationally expensive. In this paper, we focus on APE-FDs and prove that, unfortunately, verifying a single APE-FD over a given database instance is in general NP-complete. In order to cope with this problem, we propose a framework for mining complex APE-FDs in real-world data collections. In the framework, we designed and applied sound and advanced model-checking techniques. To prove the feasibility of our proposal, we used real-world databases from two medical domains (namely, psychiatry and pharmacovigilance) and tested the running prototype we developed on such databases
Acute Kidney Injury Prediction with Gradient Boosting Decision Trees enriched with Temporal Features
This paper aims to predict the risk of Acute Kidney Injury (AKI) in intensive care units (ICUs) using machine learning techniques and statistical approaches. The data used in the study are derived from Medical Information Mart for Intensive Care (MIMIC) III, which is a freely accessible database of de-identified ICU-related data. The paper focuses on two different phases. The first one consists of a scrupulous phase of extraction and transformation of MIMIC data to retrieve all the criteria specified in the Kidney Disease Improving Global Outcomes (KDIGO) clinical practice guideline definition of AKI. The main features included are demographics, medications, comorbidities, charted vital signs, and laboratory events. In the second phase, we used several different techniques already used to predict AKI, and we also added a complex temporal feature, called Trend-Event Feature (TE-F). The prediction was performed using a rolling observational window design that includes the data collection window (length: 1 to 6 days) and the prediction window (7 days). The Gradient Boosting Decision Trees (GBDT) method was used, with different groups of features, to predict the risk of AKI. To evaluate the GBDT performances, we used the area under the ROC curve (AUROC), Recall, Precision, and F-score measures. We observed that lab parameters contribute the most to the prediction of the AKI risk. Moreover, adding TE-Fs to different groups of features leads to better performance results
A BPMN-based Framework to Manage ERAS-inspired Pathway for Patients Undergoing Pancreatic Surgery
It has been proved that patients on the way to pancreatic surgery could benefit from a set of activities to improve their psycho-physical status. The Enhanced Recovery After Surgery (ERAS) program is a proven set of evidence-based activities to support the perioperative pathway of patients. On the other hand, the Business Process Model and Notation (BPMN) is able to describe complex process semantics in an intuitive notation that constitutes a common language between different stakeholders. Finally, web-based applications could help clinicians to monitor, support, and manage patients in the perioperative stage, enabling communications with surgeons, psychologists, nutritionists, and physiatrists.In this paper, we propose a framework to model an ERAS-inspired prehabilitation program as BPMN processes, which includes activities in four macro areas: surgery, nutrition, physical activity, and psychology. We also propose Fit to Fight, a process-driven web application (built from scratch) that executes and manages the processes in conjunction with a BPMN process engine. Activities are scheduled during a 30 days time-frame, and the underlying processes implement a fine set of temporal constraints to ensure it. The prehabilitation program considered is the one currently adopted by the Unit of Pancreatic Surgery at the Verona University Hospital, which is part of the Verona Pancreas Institute, the first interdisciplinary center of excellence for diagnosing and treating pancreatic diseases in Italy. Thanks to this close collaboration, we adopted the agile methodology where the prehabilitation program modeling and the web application development were made with the clinicians’ continuous involvement. Finally, we developed an elaborated yet user-friendly and responsive interface to present the activities to patients
A Key Performance Indicator to Analyze Swarm Learning Performances with EHR
Swarm Learning (SL) has been recently proposed for distributed learning, where a group of individual centers perform a synchronized training. Unlike traditional machine learning models that rely on a central server, swarm learning distributes the learning process across multiple nodes. Each node independently processes data and contributes to the overall learning task. This collaboration allows the swarm to benefit from individual nodes' different data. Unlike federated learning, here model parameters are not handled by a central server but are randomly handled across each individual node. The intrinsic attention of swarm learning to data privacy makes it suitable for distributed health care analysis, where a clinical center wants to benefit from all the other ones in the swarm network. However, the benefit for a single center or for the whole network could vary depending on data distribution. In this paper, we want to analyze the performance of the swarm learning in a network with multiple nodes, where different data distribution scenarios are taken into account. This analysis will show the gain of the whole swarm network and a specific (reference) node, focusing on scenarios where this node has a different amount of data with respect to the other nodes. To perform a more analytical analysis, we introduce a new Key Performance Indicator (KPI) to measure such gain. We then applied this method using I CU data extracted from the MIMIC EHR database and discussed the results obtained by analyzing 5 nodes with different data distribution scenarios
- …
