1,721,068 research outputs found
Recommended from our members
Harnessing AI/ML for Proteomics: Post-Translational Modification Prediction and Proteome Turnover Imputation
The field of proteomics, encompassing the exhaustive study of proteins, their structures, functions, post-translational modifications, dynamics, and interactions, stands as a crucial domain in the quest to understand biological systems and disease mechanisms. The rise of high-throughput technologies, notably mass spectrometry, has exponentially increased the volume and complexity of proteomic data, posing both opportunities and challenges in large-scale data analysis and interpretation. In this context, the integration of Artificial Intelligence (AI) and Machine Learning (ML) methodologies presents a transformative strategy, promising to significantly enhance various facets of data analysis in proteomics. This dissertation is dedicated to exploring the application of AI/ML in the domain of proteomics.The first theme of this dissertation introduces MIND-S, a deep-learning platform designed to predict protein post-translational modifications (PTMs). MIND-S utilized protein sequence and structure, modeling through combination of a transformer model and a graph neural network to efficiently predict multiple PTMs. It features an interpretation module that discerns the relevance of amino acids and uncovers PTM patterns without direct supervision. Additionally, it assesses the effects of mutations on PTMs and has been validated using biological data. This work demonstrates MIND-S's accuracy and efficiency in analyzing PTM processes in both health and disease.The second theme delves into gene representation through a comprehensive, task-agnostic approach, aiming for a holistic understanding of molecular events. Traditional gene embeddings often have a narrow focus on specific tasks, missing the broader picture. This study evaluates nine gene embeddings across three categories: experimental, literature, and knowledge graph data. Using Singular Vector Canonical Correlation Analysis (SVCCA), it reveals that the representations contain unique, minimally overlapping information, fostering rich, multifaceted embeddings. This method outperforms task-specific approaches in various benchmark tests and successfully imputes missing data, enhancing individual embeddings. It offers a robust framework for comprehensive biomolecule characterization, with significant benefits for biomedical AI applications.The third theme addresses the challenge of missing values in temporal proteomics datasets, which can obscure critical measurements and impair the understanding of biomedical processes. To address this, a Data Multiple Imputation (DMI) pipeline was developed to facilitate robust analysis of protein turnover rates in time-series data. This approach was applied to murine cardiac and human plasma datasets, greatly improving the detection of protein turnover rates and uncovering new biological insights. The imputed data provided a more comprehensive depiction of proteins, enhancing the understanding of biological pathways and disease associations. Notably, DMI outperformed single imputation methods in benchmark evaluations, demonstrating its effectiveness in managing missing data challenges in temporal proteomics
Recommended from our members
An Informatics Roadmap Toward a FAIR Understanding of Mitochondrial Biology and Rare Mitochondrial Disease
Mitochondrial biology is integral to our fundamental understanding of human health and many diseases. They exist in every human cell type except for red blood cells and have critical functions in metabolism, oxidative phosphorylation, oxidation-reduction, and as signaling hubs responsible for mediating protective mechanisms. Rare mitochondrial diseases (RMDs) are devastating and complex, affect multiple organ systems, and disproportionately impact young children. Despite copious existing knowledge and increased public interest, the knowledge is fragmented and difficult to access. Clinical case reports (CCRs) on RMDs contain valuable clinical insights, but they are scarce and lack the metadata necessary to facilitate their discovery among the two million CCRs on PubMed. The unstructured text data of CCRs is also ill-suited to computational approaches, limiting our ability to derive the knowledge contained within.To address these issues, I assembled all available informatics tools and resources with mitochondrial components and used them to contribute to Gene Wiki pages that enable easy access to mitochondrial knowledge for researchers, students, clinicians, and patients. Through these efforts, I made mitochondrial gene, protein, and disease knowledge widely accessible with contributions of over 4MB of content across 541 Gene Wiki pages. Concurrently, I used Gene Wiki as an educational platform to train over 50 students in the biosciences and pre-medical studies in mitochondrial biology and disease, as well as instilling effective research and writing methods in biomedicine.To impose structure on CCRs and render them FAIR (Findable, Accessible, Interoperable, Reusable), I developed and applied a standardized metadata template to RMD CCRs and codified patient symptomology with the International Statistical Classification of Disease and Related Health Problems (ICD) system. I created the open-source, cloud-based MitoCases RMD Knowledge Platform (http://mitocases.org/) to house data on 384 RMD CCRs, including 4,561 instances of 952 unique ICD codes. Supplementing CCRs with structured metadata amplifies machine-readable information content and provides a distinct improvement in searching for CCRs as compared to indexing by title and abstract. Finally, I employed these resources to conduct a thorough review of Barth syndrome and characterized the diversity of presentations, range of genetic etiologies, and treatment paradigms
Recommended from our members
Cloud-based Analysis and Integration of Proteomics and Metabolomics Datasets
Our capabilities to define cardiovascular health and disease using highly multivariate “omics” datasets have substantially increased in recent years. Advances in acquisition technologies as well as bioinformatics methods have paved the way for ultimately resolving every biomolecule comprising various human “omes”. Understanding how different “omes” change and interact with one another temporally will ultimately unveil multi-omic molecular signatures that inform pathologic mechanisms, indicate disease phenotypes, and identify new therapeutic targets. Herein we describe a thesis project that creates novel, contemporary data science methods and workflows to extract temporal molecular signatures of disease from multi-omics analyses, and develops integrated omics knowledgebases for the cardiovascular community at-large. Chapter 1 provides an overview of cardiac physiology and pathophysiology involved in cardiac remodeling and heart failure (HF). A description of the systematic characterization of cardiac proteomes and metabolomes is included, including methodologies for multi-omics phenotyping. Finally, an overview of bioinformatics methods for driver molecule discovery is provided, discussing strategies for characterizing temporal patterns and conducting functional enrichment. Chapter 2 describes computational approaches to discern the oxidative posttranslational modification (O-PTM) proteome, an important factor in cardiac remodeling. We developed a novel platform involving a customized, quantitative biotin switch pipeline and advanced analytic workflow to profile O-PTMs in an isoproterenol (ISO)-induced cardiac remodeling mouse model. We identified 1,655 proteins containing 3,324 oxidized sites, and unveiled temporal progression of O-PTM in disease. Chapter 3 discusses computational approaches for identifying temporal metabolomics fingerprints in HF treatment. Pathologic remodeling from a healthy to diseased heart involves a series of alterations over time. Mechanical circulatory support devices (MCSD) are a promising strategy for unloading the heart and reversing this process. We sought to identify molecular drivers of pathologic remodeling and reverse remodeling in the plasma metabolome; however, machine learning (ML)-empowered technological platforms required for these analyses are lacking. Thus, we established a Multiple Reaction Monitoring (MRM)-based MS quantitative platform and ML-based computational workflow to discern metabolomics fingerprints. We quantified 610 plasma metabolites and identified those exhibiting high correlation to cardiac phenotype, demonstrating a novel platform for biomarker discovery. Finally, Chapter 4 integrates all aforementioned innovations into one unified, cloud-based computational knowledgebase, MetProt, equipped to analyze, annotate, and integrate metabolomics and proteomics information. This pipeline fully characterizes the plasma metabolome in HF, unveils the interplay of proteomes and metabolomes, and derives new knowledge in cardiovascular medicine. Innovations include engineered features for addressing large-scale clinical datasets as well as algorithms to connect various types of molecules (e.g., proteins and metabolites). Chapter 4 is subdivided into 3 projects: Project 1 describes a computational pipeline to characterize the plasma metabolome using datasets from the ISO mouse model of HF and human HF; Project 2 develops bioinformatics strategies to integrate proteome and metabolome datasets from six genetically distinct mouse strains; and Project 3 establishes the cloud-based MetProt to disseminate the above computational pipelines for the cardiovascular community at-large. Taken together, these innovations offer new approaches and workflows for integrated omics investigations that enable novel discovery and ultimately advance precision cardiovascular science and medicine
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Recommended from our members
Biomedical Knowledge Systems for AI-driven Knowledge Discovery in Cardiovascular Medicine
Recently, AI-driven discovery in biomedical research is rapidly evolving; an area benefiting from this approach is the creation of a dynamic knowledge system from the integration, processing, organization, and management of biomedical knowledge extracted from text datasets and curated knowledgebases. This thesis presents workflows toward integrating these resources to discover new knowledge in cardiovascular medicine using automated text mining, structured knowledge graphs, explainable biomedical predictions, and generative modeling. The framework cohesively links biomedical entities, identifies hidden relationships, synthesizes new knowledge, and provides actionable insights. Key contributions include uncovering novel protein-disease associations, developing a scalable knowledge graph for multi-modal machine learning applications, and grounding large language model outputs through retrieval-augmented generation. A case study on cardiovascular diseases highlights the applications of this framework to evaluate therapeutics, reveal unexplored connections, generate novel hypotheses, advancing both research and clinical applications. Furthermore, this approach focus on validation and evidence-supported predictions, explores strategies for AI education in biomedical research, bridging technical expertise and domain specific knowledge for effective and responsible AI adoption in biomedicine
Recommended from our members
Protein Half-life of Degradation Machineries in Healthy and Stressed Myocardium
Protein synthesis and degradation function in concert to maintain myocardial proteome homeostasis and render proteome dynamics triggered by pathological stresses; mounting evidence documents their dysfunction as a cardiac disease driver. Despite our knowledge of hypertrophic signaling cascades in heart, little is known about the concomitant self-regulation of protein synthesis and degradation machineries during this proteome-wide remodeling. For 6 genetic mouse strains that developed distinguishable scales of ISO-induced hypertrophy, the basal and altered protein half-lives were quantified for proteasomal subunits as well as other degradation and synthesis machineries. Contractile proteins were quantified as half-life references to hypertrophy. The workflow developed in this study revealed the unique turnover features among six genetic strains, the distinct turnover modulation of individual proteasomal subunits under β-adrenergic stimulation, and a signature of contractile protein turnover alteration. Together with functional assessments, our discovery supports future investigations to dissect these regulatory processes in maladaptive cardiac remodeling
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Recommended from our members
Mitochondrial Protein Dynamics in Cardiac Remodeling
The cardiac mitochondrial proteome contains ~1,500 distinct proteins that carry out necessary metabolic and energetic processes in the heart. To sustain cardiac function, the mitochondrial proteome must be maintained in constant renewal, or turnover, especially under stress conditions. Disruptions of protein turnover can lead to protein damage and proteotoxicity, a hallmark of many heart disease etiologies. Current quantitative proteomics experiments largely focus on the measurement of the steady-state abundance, or changes therein, of proteins that are present in a system, and give little insights into the underlying regulations of protein synthesis, degradation, and homeostasis. Protein turnover rates provide this missing temporal dimension of information, and can inform on the potential mechanism through which protein abundance may permute during the development of disease (e.g., via increased synthesis or decreased degradation). Currently, such investigations are hampered by the fact that the technology to measure protein turnover in animals on a large scale has not been well developed. This dissertation outlines a new method to measure protein turnover half-life in the cardiac mitochondrion. Basic features of the regulation of protein turnover in the mitochondrion are discussed, and how protein dynamics permutes in early-stage heart failure after hypertrophic stimuli is described. In total, we measured the turnover rates of 2,986 proteins in the mouse heart under basal conditions, isoproterenol stimulus, and post-stimulus recovery, including 1,078 proteins from isolated mitochondria. The data revealed widespread, bidirectional changes in protein turnover in 35 functional categories, and further identified a number of novel candidate disease proteins with significantly up-regulated turnover rates in disease, including HK1, ALDH1B1, and PHB, which have been obscured from previous investigations due to their inconspicuous changes in steady-state abundance. Combinatorial analysis of protein expression and protein turnover data indicates that the remodeling heart is characterized by decreased turnover but increased expression of a cohort of mitochondrial proteins including FXN, LETM1, and CYC1, suggesting a potential class of candidate disease proteins whose impaired degradation is associated with remodeling. I further discuss the implications of the data to the cardiac remodeling process at large and how such investigations may be translated to human studies in the future. Taken together, the results suggest that comparisons of protein turnover rates can be a powerful new tool to understand the temporal dynamics of disease progression in the heart
- …
