1,720,994 research outputs found
Bayesian Networks for Omics Data Analysis in Hepatocellular Carcinoma Single-Cell Sequencing
Single cell multi omics techniques have shown an advancement in unrevealing complex diseases like cancer heterogeneity by providing multi-faceted insight into their individual cellular regulations. In this study, a machine learning approach, Bayesian network (BN), has been applied to integrate genomic, epigenomic, and transcriptomic data in hepatocellular carcinoma at single cell resolution. Hepatocellular carcinoma (HCC) is the most common type of liver cancer with a high metastatic rate and reckoned for poor prognosis. Heterogeneity of tumor cells is concerned with cancer progression, metastasis, therapeutic resistance, and mortality. For this purpose, a dataset from a published study of 25 single cell sequencing of hepatocellular carcinoma were used. First, DNA methylome and transcriptome data were analyzed on their own. Copy number variations were estimated from DNA methylome data by using the Hidden Markov Model method. To reveal the causal relationship between the omics, three BN models were constructed. The models were fitted to their parameters by using maximum likelihood estimation. For model evaluation, score-based criteria, Akaike information criterion and Bayesian information criterion, were used. 207 genes with significant models have been detected. The heterogeneity of the omics and their regulation mechanisms with each other have been shown, by pointing to genes that follow different BN models that take place in major pathways in HCC.Tek hücreli çoklu omik teknikleri, kendi bireysel hücresel düzenlemelerine çok yönlü bir bakış açısı sağlayarak kanser heterojenliği gibi kompleks hastalıkların ortaya çıkarmada bir ilerleme göstermiştir. Bu çalışmada, tek hücre temelli hepatoselüler karsinomda genomik, epigenomik ve transkriptomik verileri entegre etmek için bir makine öğrenimi yaklaşımı olan Bayesian ağları (BN) uygulanmıştır. Bu amaçla, yayınlanmış bir çalışmadan hepatoselüler karsinomun 25 tekil hücre dizileme veri seti kullanılmıştır. Hepatosellüler karsinom (HSK), yüksek metastatik oranla en yaygın karaciğer kanseri türüdür ve kötü prognoza sahip olduğu düşünülmektedir. Tümör hücrelerinin heterojenliği, kanserin ilerlemesi, metastaz, terapötik direnç ve mortalite ile ilgilidir. Önce, DNA metilom ve transkriptom verileri tek başlarına analiz edilmiştir. Kopya sayısı varyasyonu, Gizli Markov Modeli yöntemi kullanılarak DNA metilom verilerinden tahmin edilmiştir. Omikler arasındaki nedensel ilişkiyi incelemek için üç BN modeli oluşturulmuştur. Modeller, en çok olabilirlik kestirimi (MLE) kullanılarak parametrelerine uydurulmuştur. Model değerlendirme için puana dayalı kriterler, Akaike bilgi kriteri (AIC) ve Bayes bilgi kriteri (BIC) kullanılmıştır. Anlamlı modele sahip 207 gen tespit edilmiştir. Farklı BN model izleyen genlerin HCC’de aynı yolakta yer aldığına işaret ederek, omiklerin ve birbirleriyle regülasyon mekanizmalarının heterojenliği gösterilmiştir
Ekzom Veri Setinden Hastalığa Özgü Varyant Veri Tabanı Oluşturulması
Adabalı, Y., Building a Disease-specific Variant Database from Exome Datasets,
Hacettepe University Graduate School Health Sciences Department of Bioinformatics
Master’s Thesis, Ankara, 2019. After the completion of human reference genome
sequence with the Human Genome Project, an increasing nıumber of individuals have
been sequenced with next-generation sequencing technologies. This has shown that
millions of genetic differences, called genetic variants, exists between individuals. Some
of these genetic variants are known to cause genetic disorders. However, it is difficult to
pinpoint a disease-causing variant among the many variants present in an individual. The
aim of this thesis is to collect genetic variant data from individuals with suspected genetic
disorders to establish a variant database. This database will provide analysis of specific
variants for Turkish population and providing a platform that allows comparative analysis
of individual data. For this purpose, MS-SQL as the database mangement system; ASP,
.NET, MVC in the back-end; Entity Framework as object-relational mapping tool; HTML5
and CSS technologies and Bootsrap, Javascript ve Jquery libraries in the front-end were
used. The built system establishes an expandable database by incorporating tabular file
formats such as .tsv, .csv which are annotated by the Ion Reporter software. The
database allows query options with respect to several variant properties. In addition,
variants for an individual in the database can be compared against the other variants
which can be filtered by 3 modes of disease type and inheritance pattern, frequency of
variants in in-house and international variant databases and the effect or position of
variants. In conclusion, the established database provides a quick and effective analysis
of genetic variants that can be related to several diseases by using specific data for
Turkish population and facilitates the analysis of this big genetic data.
Key Words: Whole exome sequencing, Genetic database, Genetic variation, GenotypeONAY SAYFASI iii
YAYIMLAMA VE FİKRİ MÜLKİYET HAKLARI BEYANI iv
ETİK BEYAN SAYFASI v
TEŞEKKÜR vi
ÖZET vii
ABSTRACT viii
İÇİNDEKİLER ix
SİMGELER ve KISALTMALAR xi
ŞEKİLLER xvi
TABLOLAR xviii
1.GİRİŞ 1
2. GENEL BİLGİLER 3
2.1. İnsan Genomu ve Genom Mimarisi 3
2.1.1. İnsan Genom Projesi 4
2.1.2. Topluma ve Bireye Özgü Genetik Değişiklikler (Varyantlar) 5
2.1.3. Genom ve Varyant Veri Tabanları 9
2.2. Yeni Nesil DNA Dizileme Teknolojileri (NGS) 12
2.2.1. Farklı Platformlara Göre Ekzom Veri Eldesi 13
2.2.2. Ekzom Verisinden Hastalığa Özgü Varyantların Tespit Edilmesi 16
2.3. Veri Tabanları 22
2.3.1. Veri Tabanı Mimarileri 22
2.3.2. Veri Tabanlarında Optimizasyon İşlemleri 25
2.3.3. Dinamik Web Uygulamaları 25
2.3.4. Ön yüz mimarileri 26
2.3.5. Arka yüz mimarileri 28
2.3.6. Yazılım Geliştirme Mimarileri 30
3. GEREÇ, YÖNTEM VE BİREYLER 31
3.1. Bireyler 31
x
3.2. Çalışmada Kullanılan Yöntemler 31
3.2.1. Çalışmada Kullanılan Uygulama Geliştirme Ortamları 32
3.2.2. Veri tabanı İşlemleri 33
3.2.3. Veri aktarım uygulaması 35
3.2.4. Web Uygulamasının Geliştirilmesi 41
4. BULGULAR 45
4.1. Yazılımların Genel Özellikleri 45
4.1.1 Veri Yüklenmesi için Geliştirilen Masaüstü Uygulaması 45
4.1.2 Veri tabanı 48
4.1.3 Web Arayüzü Uygulaması 58
4.2. Filtreleme Modlarının Değerlendirilmesi 64
4.2.1. Veri tabanının Varyant Filtreleme Becerisinin Değerlendirilmesi 64
4.2.2. Farklı Filtreleme Modlarına Göre Patojenik Varyant Saptanması 70
5. TARTIŞMA 78
6. SONUÇ VE ÖNERİLER 88
7. KAYNAKLAR 89
8. EKLER
EK-1: Tez Çalışması ile İlgili Etik Kurul İzni
EK-2: Tez Çalışması Orijinallik Raporu
EK-3: MS-SQL Management Studio üzerinden csv dosyası aktarımı
EK-4: “tmpGrch38” tablosundan gen ile ilişkili tablolara aktarım yapan T-SQL
sorguları
EK-5: “GetVariants” isimli fonksiyona çağrı yapan kod bloğu
9. ÖZGEÇMİŞAdabalı, Y., Ekzom Veri Setinden Hastalığa Özgü Varyant Veri Tabanı Oluşturulması,
Hacettepe Üniversitesi Sağlık Bilimleri Enstitüsü Biyoinformatik Programı Yüksek
Lisans Tezi, Ankara, 2019. İnsan Genom Projesi ile insan referans genom dizisinin
oluşturulması ardından ileri nesil dizileme teknolojileri ile katlanarak artan sayıda bireyin
genomik dizileri ortaya çıkarılmaya devam etmiştir. Böylece, insanlar arasında genomik
dizide milyonlarca farklılığın (varyantın) olduğunu görülmüştür. Bu genetik varyantların
bir kısmının genetik hastalıklara sebep olduğu bilinmektedir; ancak bir dizileme
analizinde ortaya çıkan pek çok varyanttan hangisinin hastalıkla ilişkili olduğunu
saptamak zordur. Bu tezin amacı, genetik hastalık şüphesi olan kişilerden “Ekzom”
dizilemesi yapılarak elde edilen genetik varyant verilerinin Türk populasyonuna özgü
verileri içeren bir varyant veri tabanı oluşturmakta kullanılması ve karşılaştırmalı varyant
analizi için bir platform sağlanmasıdır. Bu amaçla, veri tabanı yönetim sistemi olarak MSSQL; uygulama arka yüz teknolojisi olarak ASP, .NET, MVC; nesne ile ilişkisel
haritalama/eşleme aracı olarak Entity Framework; ön yüz teknolojisi olarak ise HTML5,
CSS teknolojileriyle Bootstrap, Javascript ve Jquery kütüphaneleri kullanılmıştır. Kurulan
sistem, Ion Reporter yazılımı aracılığıyla anote edilen tabuler dosya formatlarını (.tsv, .csv
gibi) kullanan genişletilebilir bir veri tabanıdır. Veri tabanında çeşitli varyant özelliklerine
göre sorgu yapılabilmektedir. Ayrıca, veri tabanındaki hastaya özgü varyantlar, hastalık
tipine ve kalıtım modeline göre 3 farklı modda, kurumsal ve uluslararası varyant veri
tabanlarındaki varyant sıklıklarına ve varyantların pozisyonu/etkisine göre karşılaştırmalı
varyant filtrelenebilmektedir. Sonuçta, oluşturulan veri tabanı hastalıkla ilişkili olabilecek
genetik varyantların Türk toplumuna özgü veriler kullanılarak hızlı ve etkin analizini
sağlamakta ve büyük genetik verilerin analizini kolaylaştırmaktadır.
Anahtar kelimeler: Tüm Ekzom dizileme, Genetik veri tabanı, Genetik varyasyon,
Genoti
Identification of bio-markers for insulin resistance and sensitivity through multi-omics analysis
Aims: This study aims to identify multi-omics bio-markers for insulin resistance and sensitivity using machine learning approaches on a dataset integrated from several omics. Methods: The study included 362 patients with Insulin Resistance and Insulin Sensitivity from the Integrative Personal Omics Profiling (iPOP) database. Combining the multi-omics data from the Integrative Human Microbiome Project, this study used machine learning to reveal the relationship between insulin resistance and insulin sensitivity. Results: Of 362 patients 186 were insulin resistance and 176 were insulin sensitivity. 11,585 features were used, including clinical features, RNA transcripts, gut microbiota, cytokines, proteins, and metabolomics. We found 21 features capable of distinguishing insulin resistance from insulin sensitivity using a well-known artificial neural network (ANN) method. The model had an area under the receiver operating characteristic (AUC) of 0.97 in the validation dataset and 0.89 in the test dataset. The ANN model’s performance was compared with Random Forest model. Of the 21 new findings, two metabolites (methyl-uric acid and methylxanthine) are xenobiotics, and three RNA transcripts (SERPINF1, SLC2A2, and CHL1). Conclusion: A small number of multi-omics features identified from 11,585 potential candidates for a machine learning model can accurately predict insulin resistance and sensitivity
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
