34 research outputs found
XAIMetabolomeDiet-v1.0
<p><strong>About</strong></p>
<p>This published repository contains the scripts used for</p>
<ul>
<li>pre-processing the two datasets: UK Biobank, and HMP2 in R (Data Preprocessing.R),</li>
<li>performing metabolomics-based diagnostic prediction of IBD, applying SHAP XAI and performing diet-metabolite correlations in Python (ML-XAI-Diet.ipynb), and</li>
<li>performing LASSO regression in R (ClassificationScript for LASSO-R.R, and LASSO-R.R)</li>
</ul>
<p><strong>Authors</strong></p>
<ul>
<li>Serena Onwuka: Data Preprocessing.R, and ML-XAI-Diet.ipynb</li>
<li>Laura Bravo-Merodio: ClassificationScript for LASSO-R.R, and LASSO-R.R</li>
</ul>
<p><strong>Acknowledgement</strong></p>
<p>In creating the Data Processing.R and ML-XAI-Diet.ipynb scripts, we gratefully acknowledge the contributors of StackOverflow for their thorough answers, the developers on GitHub for their publicly available code, and ChatGPT for its assistance.</p>
Computational biology applications in the study of complex systems
In biomedicine, the advent of digitalization and big improvements in computing power and high throughput technologies has yielded an unprecedented amount of data. To harness this data’s full potential, research has become increasingly computational, with core tools of data science such as machine learning required to decipher patterns, help extract meaning and uncover underlying connections. In this work, we have explored key leading areas of computational research, from discovery science to translational medicine and complex system studies. First, a machine learning pipeline was developed to inform bioinformatic research. After validation, it was applied on a novel dataset generating insight onto the predictive power of immune features in ultra-early trauma injury, generating leads of possible biomarkers associated with the development of MultiOrgan Dysfunction. Then, in order to make our pipeline accessible, an interactive, simple, free and open-source supervised machine learning webtool was developed. Adapting this same framework to the realm of clinical translation, we then deployed two prognostic models for decision support in surgeries during the COVID-19 pandemic. This was done in collaboration with the NIHR surgical team. Lastly, we explored systems biology approaches by studying the complexity behind ageing and health to disease transitions. For this, we leveraged the UK Biobank data as a rich source of deeply phenotyped information. By calculating biological age (PhenoAge) longitudinally for nearly 400,000 participants, we identified four distinct ageing categories ranging from healthy to unhealthy ageing. These different trajectories were characterised by their chronic diseases and genetic makeup, revealing a strong association of metabolic dysfunction with unhealthier phenotypes and immune-related signals for healthier ones. Also, strong opposite-effect associations of longevity-related variants were found, with novel regulatory elements postulated as possible drivers of unhealthy phenotypes, opening new avenues for future study. Overall, our findings highlight the crucial role of computational methods in biomedicine and their potential to transform clinical practice
-Omics biomarker identification pipeline for translational medicine
BACKGROUND: Translational medicine (TM) is an emerging domain that aims to facilitate medical or biological advances efficiently from the scientist to the clinician. Central to the TM vision is to narrow the gap between basic science and applied science in terms of time, cost and early diagnosis of the disease state. Biomarker identification is one of the main challenges within TM. The identification of disease biomarkers from -omics data will not only help the stratification of diverse patient cohorts but will also provide early diagnostic information which could improve patient management and potentially prevent adverse outcomes. However, biomarker identification needs to be robust and reproducible. Hence a robust unbiased computational framework that can help clinicians identify those biomarkers is necessary.METHODS: We developed a pipeline (workflow) that includes two different supervised classification techniques based on regularization methods to identify biomarkers from -omics or other high dimension clinical datasets. The pipeline includes several important steps such as quality control and stability of selected biomarkers. The process takes input files (outcome and independent variables or -omics data) and pre-processes (normalization, missing values) them. After a random division of samples into training and test sets, Least Absolute Shrinkage and Selection Operator and Elastic Net feature selection methods are applied to identify the most important features representing potential biomarker candidates. The penalization parameters are optimised using 10-fold cross validation and the process undergoes 100 iterations and a combinatorial analysis to select the best performing multivariate model. An empirical unbiased assessment of their quality as biomarkers for clinical use is performed through a Receiver Operating Characteristic curve and its Area Under the Curve analysis on both permuted and real data for 1000 different randomized training and test sets. We validated this pipeline against previously published biomarkers.RESULTS: We applied this pipeline to three different datasets with previously published biomarkers: lipidomics data by Acharjee et al. (Metabolomics 13:25, 2017) and transcriptomics data by Rajamani and Bhasin (Genome Med 8:38, 2016) and Mills et al. (Blood 114:1063-1072, 2009). Our results demonstrate that our method was able to identify both previously published biomarkers as well as new variables that add value to the published results.CONCLUSIONS: We developed a robust pipeline to identify clinically relevant biomarkers that can be applied to different -omics datasets. Such identification reveals potentially novel drug targets and can be used as a part of a machine-learning based patient stratification framework in the translational medicine settings.</p
Explainable AI-prioritized plasma and fecal metabolites in inflammatory bowel disease and their dietary associations
Fecal metabolites effectively discriminate inflammatory bowel disease (IBD) and show differential associations with diet. Metabolomics and AI-based models, including explainable AI (XAI), play crucial roles in understanding IBD. Using datasets from the UK Biobank and the Human Microbiome Project Phase II IBD Multi’omics Database (HMP2 IBDMDB), this study uses multiple machine learning (ML) classifiers and Shapley additive explanations (SHAP)-based XAI to prioritize plasma and fecal metabolites and analyze their diet correlations. Key findings include the identification of discriminative metabolites like glycoprotein acetyl and albumin in plasma, as well as nicotinic acid metabolites andurobilin in feces. Fecal metabolites provided a more robust disease predictor model (AUC [95%]: 0.93 [0.87–0.99]) compared to plasma metabolites (AUC [95%]: 0.74 [0.69–0.79]), with stronger and more group-differential diet-metabolite associations in feces. The study validates known metabolite associations and highlights the impact of IBD on the interplay between gut microbial metabolites and diet
MOESM3 of -Omics biomarker identification pipeline for translational medicine
Additional file 3. R Markdown analysis results from the workflow developed on GSE15061
MOESM1 of -Omics biomarker identification pipeline for translational medicine
Additional file 1. R Markdown analysis results from the workflow developed on GSE15471
MOESM2 of -Omics biomarker identification pipeline for translational medicine
Additional file 2. Lipids identified in three cohorts are listed with different category. Category A: HM vs. Mixed (FM and HM combined) feeding; Category B: FM vs. Mixed (FM and HM combined ) feeding; Category C: HM vs. FM
Biomarker Prioritisation and Power Estimation Using Ensemble Gene Regulatory Network Inference
Inferring the topology of a gene regulatory network (GRN) from gene expression data is a challenging but important undertaking for gaining a better understanding of gene regulation. Key challenges include working with noisy data and dealing with a higher number of genes than samples. Although a number of different methods have been proposed to infer the structure of a GRN, there are large discrepancies among the different inference algorithms they adopt, rendering their meaningful comparison challenging. In this study, we used two methods, namely the MIDER (Mutual Information Distance and Entropy Reduction) and the PLSNET (Partial least square based feature selection) methods, to infer the structure of a GRN directly from data and computationally validated our results. Both methods were applied to different gene expression datasets resulting from inflammatory bowel disease (IBD), pancreatic ductal adenocarcinoma (PDAC), and acute myeloid leukaemia (AML) studies. For each case, gene regulators were successfully identified. For example, for the case of the IBD dataset, the UGT1A family genes were identified as key regulators while upon analysing the PDAC dataset, the SULF1 and THBS2 genes were depicted. We further demonstrate that an ensemble-based approach, that combines the output of the MIDER and PLSNET algorithms, can infer the structure of a GRN from data with higher accuracy. We have also estimated the number of the samples required for potential future validation studies. Here, we presented our proposed analysis framework that caters not only to candidate regulator genes prediction for potential validation experiments but also an estimation of the number of samples required for these experiments
Translational biomarkers in the era of precision medicine
In this chapter we discuss the past, present and future of clinical biomarker development. We explore the advent of new technologies, paving the way in which health, medicine and disease is understood. This review includes the identification of physicochemical assays, current regulations, the development and reproducibility of clinical trials, as well as, the revolution of omics technologies and state-of-the-art integration and analysis approaches
Link prediction in complex network using information flow
Link prediction in complex networks has recently attracted a great deal of attraction in diverse scientific domains, including social and biological sciences. Given a snapshot of a network, the goal is to predict links that are missing in the network or that are likely to occur in the near future. This problem has both theoretical and practical significance; it not only helps us to identify missing links in a network more efficiently by avoiding the expensive and time consuming experimental processes, but also allows us to study the evolution of a network with time. To address the problem of link prediction, numerous attempts have been made over the recent years that exploit the local and the global topological properties of the network to predict missing links in the network. In this paper, we use parametrised matrix forest index (PMFI) to predict missing links in a network. We show that, for small parameter values, this index is linked to a heat diffusion process on a graph and therefore encodes geometric properties of the network. We then develop a framework that combines the PMFI with a local similarity index to predict missing links in the network. The framework is applied to numerous networks obtained from diverse domains such as social network, biological network, and transport network. The results show that the proposed method can predict missing links with higher accuracy when compared to other state-of-the-art link prediction methods.</p
