1,720,968 research outputs found
Visualization and Simulation of Variants in Personal Genomes With an Application to Premarital Testing (VSIM)
Interpretation and simulation of the large-scale genomics data are very challenging, and currently, many web tools have been developed to analyze genomic variation which supports automated visualization of a variety of high throughput genomics data. We have developed VSIM an automated and easy to use web application for interpretation and visualization of a variety of genomics data, it identifies the candidate diseases variants by referencing to four databases Clinvar, GWAS, DIDA, and PharmGKB, and predicted the pathogenic variants. Moreover, it investigates the attitude towards premarital genetic screening by simulating a population of children and analyze the diseases they might be carrying, based on the genetic factors of their parents taking into consideration the recombination hotspots. VSIM supports output formats based on Ideograms that are easy to interpret and understand, which makes it a biologist-friendly powerful tool for data visualization, and interpretation of personal genomic data. Our results show that VSIM can efficiently identify the causative variants by referencing well-known databases for variants in whole genomes associated with different kind of diseases. Moreover, it can be used for premarital genetic screening by simulating a population of offspring and analyze the disorders they might be carrying. The output format provides a better understanding of such large genomics data. VSIM thus helps biologists and marriage counsellor to visualize a variety of genomic variants associated with diseases seamlessly
Prioritizing Causative Genomic Variants by Integrating Molecular and Functional Annotations from Multiple Biomedical Ontologies
Whole-exome and genome sequencing are widely used to diagnose individual patients. However, despite its success, this approach leaves many patients undiagnosed. This could be due to the need to discover more disease genes and variants or because disease phenotypes are novel and arise from a combination of variants of multiple known genes related to the disease. Recent rapid increases in available genomic, biomedical, and phenotypic data enable computational analyses, reducing the search space for disease-causing genes or variants and facilitating the prediction of causal variants. Therefore, artificial intelligence, data mining, machine learning, and deep learning are essential tools that have been used to identify biological interactions, including protein-protein interactions, gene-disease predictions, and variant--disease associations. Predicting these biological associations is a critical step in diagnosing patients with rare or complex diseases.
In recent years, computational methods have emerged to improve gene-disease prioritization by incorporating phenotype information. These methods evaluate a patient's phenotype against a database of gene-phenotype associations to identify the closest match. However, inadequate knowledge of phenotypes linked with specific genes in humans and model organisms limits the effectiveness of the prediction. Information about gene product functions and anatomical locations of gene expression is accessible for many genes and can be associated with phenotypes through ontologies and machine-learning models. Incorporating this information can enhance gene-disease prioritization methods and more accurately identify potential disease-causing genes.
This dissertation aims to address key limitations in gene-disease prediction and variant prioritization by developing computational methods that systematically relate human phenotypes that arise as a consequence of the loss or change of gene function to gene functions and anatomical and cellular locations of activity. To achieve this objective, this work focuses on crucial problems in the causative variant prioritization pipeline and presents novel computational methods that significantly improve prediction performance by leveraging large background knowledge data and integrating multiple techniques.
Therefore, this dissertation presents novel approaches that utilize graph-based machine-learning techniques to leverage biomedical ontologies and linked biological data as background knowledge graphs. The methods employ representation learning with knowledge graphs and introduce generic models that address computational problems in gene-disease associations and variant prioritization. I demonstrate that my approach is capable of compensating for incomplete information in public databases and efficiently integrating with other biomedical data for similar prediction tasks. Moreover, my methods outperform other relevant approaches that rely on manually crafted features and laborious pre-processing. I systematically evaluate our methods and illustrate their potential applications for data analytics in biomedicine. Finally, I demonstrate how our prediction tools can be used in the clinic to assist geneticists in decision-making. In summary, this dissertation contributes to the development of more effective methods for predicting disease-causing variants and advancing precision medicine
bio-ontology-research-group/VSIM: Visualization and simulation of variants in personal genomes with an application to premarital testing
Visualization and simulation of variants in personal genomes with an application to premarital testin
VSIM: Visualization and simulation of variants in personal genomes with an application to premarital testing
Background: Interpretation of personal genomics data, for example in genetic counseling, is challenging due to the complexity of the data and the amount of background knowledge required for its interpretation. This background knowledge is distributed across several databases. Further information about genomic features can also be predicted through machine learning methods. Making this information accessible more easily has the potential to improve interpretation of variants in personal genomes. Results: We have developed VSIM, a web application for the interpretation and visualization of variants in personal genome sequences. VSIM identifies disease variants related to Mendelian, complex, and digenic disease as well as pharmacogenomic variants in personal genomes and visualizes them using a web server. VSIM can further be used to simulate populations of children based on two parent genomes, and can be applied to support premarital genetic counseling. We make VSIM available as source code as well as through a container that can be installed easily in network environments in which genomic data is specially protected. VSIM and related documentation is freely available at https://github.com/bio-ontology-research-group/VSIM. Conclusions: VSIM is a software that provides a web-based interface to variant interpretation in genetic counseling. VSIM can also be used for premarital genetic screening by simulating a population of children and analyze the disorder they might be carrying.This work was supported by funding from King Abdullah University of Science and Technology (KAUST) Office of Sponsored Research (OSR) under Award NoURF/1/3454-01-01, FCC/1/1976-08-01, and FCS/1/3657-02-01
VSIM: Visualization and Simulation of Variants in Personal Genomes with an Application to Premarital Testing
VSIM: Visualization and Simulation of Variants in Personal Genomes with an Application to Premarital Testin
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Prioritizing Copy Number Variants using Phenotype and Gene Functional Similarity
Abstract
Background: There are many types of genetic variation in the human genome, ranging from large chromosome anomalies to Single Nucleotide Variant (SNV). It is becoming necessary to develop methods for distinguishing disease-causing variants from a large number of neutral genetic variation in an individual. This problem is also relevant to Copy Number Variants (CNVs), which is a class of genetic variation where large segments of the genome differ in copy number amongst various individuals.
Results:. We have built a method that incorporates biological background knowledge about the relation between phenotypes resulting from a loss of function in mouse genes, gene functions as described using the Gene Ontology (GO), as well as the anatomical site of gene expression along with a score that predicts the pathogenicity of CNV SVScore. We use this information to build a machine learning model that ranks CNVs based on their predicted pathogenicity and the relation between genes affected by the CNV and the phenotype we observe in affected individuals. Our method achieves an F-score of 99.23%, with 99.18% precision in our evaluation set.
Introduction
Over the past several years, much progress has been made in the area of CNVs detection and understanding their role in human diseases 1,2,3. We now understand that CNVs account for much of human variability. Correspondingly, there have been several methods introduced to find disease-associated genes and SNVs 4,5,6. Constructing similar methods for CNV is challenging due to the heterogeneity in variant size, type and the possibility of multiple genes being affected by large CNVs. CNV impact prediction methods should consider these factors in order to robustly prioritize pathogenic variants.
Results
The performance of our methods is based on a dataset of CNVs detected in structure variants with known phenotypes. These CNVs were evaluated as harmful or benign. Our results show that incorporating this information leads to improvement over a baseline model (Fig 2) which uses only similarity scores between gene phenotype associations and disease associated phenotypes, as well as improvement over using only pathogenicity prediction methods for CNVs. Our method achieves an F-score of 99.23%, with 99.18% precision.
Future work
Future work is required to evaluate and improve our model using patient-derived WGS data. Moreover, establishing a workflow that incorporating existing tools for CNV calling from BAM/Fastq file to SV. Then we can test the method using real samples with known CNV disease.
- …
