1,721,039 research outputs found

    Disease Comorbidity Analysis

    Get PDF
    Lühikokkuvõte: Personaalmeditsiin on uus lähenemine tervisekaitsele, milles tuuakse esile patsientide individuaalsus ja asetatakse rõhku haiguste ennetamisele nendest tekkinud tagajärgedele reageerimise asemel. Sealjuures võetakse arvesse võimalikult palju nii patsientide kui haiguste kohta teadaolevast ja ka muust meditsiinilisest teabest ning üritatakse nende vahel seoseid leida. Personaalmeditsiini üldeesmärgid on pakkuda tulevikus kõigile senisest lühema aja jooksul efektiivsemat ravi madalate kuludega. Käesoleva töö eesmärgiks on uurida haiguste komorbiidsust Eesti populatsioonis. Töös koostatakse Eesti E-tervise Sihtasutuse 2012.-2013. aasta epikriiside andmete põhjal kõigi sama patsiendi puhul koosesinevate RHK-10 registri haiguste paaride kohta 2x2 sõltuvustabelid. Haigustevahelist võimalikku seost hinnatakse Fisheri täpse testiga, filtreeritakse välja tugevamini assotsieeritud paarid ja visualiseeritakse tulemusi nn kuumuskaartide abil. Haiguste koosesinemise uurimine on eelduseks tulevases teadustöös haigusepisoodide kaevandamisele. Võtmesõnad: bioinformaatika, personaalmeditsiin, epidemioloogia, haiguste komorbiidsus, RHK 10, 2x2 sõltuvustabelid, Fisheri täpne testAbstract: Personalised medicine is a new approach to health care, in which the focus is on the individuality of patients, and disease prediction and prevention are emphasised, as opposed to only reacting to the consequences of medical disorders. As much data about the patients and diseases as possible, as well as other medical information, is taken into account while attempting to find if and how they are linked to each other. The main objective of personalised medicine is to offer more effective treatment to every patient in a shorter period of time at a lower cost in the future. The aim of this thesis is to study and analyse disease comorbidity in the Estonian population. 2x2 contingency tables are constructed about every pair of co-occurring ICD-10 diagnose codes in epicrises gathered by the Estonian E-Health Foundation in the years 2012-2013. The potential correlation between diseases is measured with Fisher’s exact test and diagnose pairs with a stronger association are filtered. The results are visualised using heat maps. Disease comorbidity analysis is a prerequisite for future research about disease episode mining. Keywords: bioinformatics, personalised medicine, epidemiology, disease comorbidity, ICD 10, 2x2 contingency tables, Fisher’s exact tes

    Software for Clustering Using k-means Algorithms

    Get PDF
    Klasteranalüüsis on laialt levinud k-keskmiste meetod, mis võimaldab andmeid grupeerida nende tunnuste järgi, seejuures minimeerides ruutvigade summat klastrites olevate andmeobjektide ja vastava klastri keskpunktide vahel. Kuna k-keskmiste meetodi kui optimeerimisülesandele täpse lahenduse leidmine on NP-raske, siis on probleemi lahendamiseks võetud kasutusele mitmeid lähendeid otsivaid algoritme. Bakalaureusetöö eesmärgina valmis rakendus, mis lubab kasutada viit k-keskmiste klasterdusalgoritmi ja nelja algsete keskpunktide valimise meetodit. Kasutades nii reaalelulisi kui ka sünteetilisi andmestikke antakse ülevaade rakenduses implementeeritud algoritmide jõudlusest, mälukasutusest ja edukusest leida hea lähend k-keskmiste optimeerimisülesandele.In cluster analysis k-means method is a method popularly used for grouping data by their features. The method aims to minimize within-cluster sum of squared errors between data objects in clusters and their corresponding center means. Because solving k-means optimization task exactly is NP-hard there have been introduced several heuristic algorithms for finding approximations. As the goal of the thesis a software was made, which enables use of nine different algorithms, which are 5 k-means clustering algorithms and 4 methods for choosing initial centers. Using real life and synthetic datasets an overview of the application’s capabilities is given by measuring algorithms performance, memory use and approximation capabilities

    A Study of Clustering Methods Using Visual Data

    Get PDF
    Klasteranalüüs on laia kasutusvaldkonnaga andmeanalüüsi tehnika, mille rakendamiseks on olemas mitu erinevat algoritmi. Käesoleva töö eesmärk on anda ülevaade kolme levinuma klasteranalüüsi meetodi tööpõhimõtetest ja eripäradest, rakendades hierarhilise klasterda-mise, k-keskmiste klasterdamise ja Kohoneni võrgu algoritme näidisandmestiku peal. Li-saks algoritmide tööpõhimõtetele on kirjeldatud ka põhjus, miks näidisandmestikuks on valitud visuaalsed andmed ehk pildid ning kuidas on implementeeritud klasteranalüüsi meetodite rakendamiseks kasutatav skript. Töö sisaldab ka skripti rakendamisel saadud klasterduste analüüsi.Cluster analysis is a widely used data analysis technique that can be applied by using sev-eral different algorithms. This thesis aims to give an overview of the working principles and specifics of the three most commonly used cluster analysis methods by applying hier-archical clustering, k-means clustering and self-organizing map algorithms on sample data. In addition to the description of working principles of the clustering algorithms, there is also a description of how the script used for applying the clustering methods is implement-ed and an explanation for why visual data or pictures are chosen as the sample data. The thesis also includes an analysis of the clustering results produced by the script

    Price Elasticity Based Recommender System

    Get PDF
    Soovitussüsteeme on palju uuritud ja edukalt rakendatud paljudes valdkondades, et suurendada läbimüüki tehes klienditele asjakohaseid soovitusi. Käesoleva magistritöö eesmärgiks on välja töötada uudne soovitussüsteem, mis teeb klientidele personaalseid pakkumisi toote soodushinna huvipakkuvuse põhjal. Seda saab rakendada olukordades, kus soovitusi tehakse allahinnatud toodete seast. Näiteks valides kampaaniatooteid kliendile saadetavasse personaalsesse uudiskirja. Me kaasame tootepõhise kaasfiltreerimise algoritmi täiendusena majanduse valdkonnas kasutatavat nõudluse hinnaelastsust, et võtta arvesse, et tootel on kampaaniaperioodil tavapärasest madalam hind. Hinnates mudeli abil omaelastsuse väärtuse, saame kliendi tootereitingu, mis näitab, kuidas hinna muutumine mõjutab ostetavat kogust. Toodete sarnasuste leidmiseks kasutame ristelastsust, mis liigitab tooted asendus- ja täiendkaupadeks. Selle suuruse leidmine ei nõua tihti esinevat kaugusmõõtude tingimust, et kaks toodet peavad olema ostetud samade klientide poolt. Kirjeldatud soovitussüsteem on rakendatud reaalelulistele supermarketi tehingute andmetele. Süsteemi headuse testimiseks kasutame kahte kampaaniaperioodi, mille allahinnatuid tooteid kasutame võimalike soovitustena. Me saavutame märgatavalt paremad tulemused kasutades ainult kampaaniatoote elastsusi ja mitte asendustoodete vastavaid väärtusi. Kaasates ainult kliendid, kellele leidsime vähemalt 5 pakkumist, saavutame tunduvalt paremad tulemused. Täpsemalt, tehes neile klientidele 12 soovitust (vähem kui 1% kampaaniatoodete arvust), tabame kõik klientide kampaaniatoodete ostud. Parima meetodi korral saavutame kordustäpsuse 0,24, mis on üle 10 korra parem võrreldes meetodiga, mida ettevõte hetkel kasutab, kus soovitused on manuaalselt valitud kliendi segmentide omaduste põhjal. Supermarketi kett on kinnitanud oma soovi, et esitletud meetodit testida, seega antud soovitussüsteemi rakendatakse reaalselt klientidele huvipakkuvate soovituste tegemiseks.Recommender systems have been widely studied and successfully applied in a variety of areas to increase sales by guiding people toward items they are more likely to find interesting. The aim of this thesis is to develop a novel recommender system that suggests items to a client based on the appeal of a product discount. This can be applied to situations where recommendations are made from a list of discounted items such as campaign products selected into personalized sales promotion letters. To take into consideration that the products have cheaper price than usual during the campaign period, we propose an extension to an item based collaborative filtering algorithm, namely the price elasticity of demand known from the field of economics. We represent a client's rating about an item by estimating with a model own elasticity which measures the sensitivity of quantity demanded to the changes in the prices. The similarities of items are computed using cross elasticity which exhibits the substitutional and complementary effects among the products. Unlike traditional similarity metrics, this measure does not assume that the two items have to be purchased by the same clients. The proposed recommender system based on the price elasticity is applied to a real world supermarket transactions dataset. The performance of the system is evaluated on two campaign periods where recommendations are made from the discounted products. The experiments show that it is better to make recommendations based on only campaign product elasticities, without considering the elasticities of campaign products' substitutes. Furthermore, when including only the customers for whom we have found at least 5 recommendations, the performance is considerably better. In particular, when making 10 suggestions (less than 1% of all campaign products), we detect all campaign products that the clients indeed purchased. Using the best method, our approach achieves precision of 0.24, which is over 10 times better in comparison to the method currently used by the company where employees manually select recommendations based on the characteristics of customer segments. The supermarket chain has confirmed their interest in testing the proposed method in practice, hence it will be applied in real world to make more relevant recommendations to the customers

    Clustering-based motif discovery from short peptides

    Get PDF
    Uute sekveneerimistehnoloogiate abil genereeritakse palju erineva taustaga bioloogilisi andmeid. Olulise info leidmiseks tuleb neid andmeid analüüsida. Antud töös koostame meetodi, mis suudab tuvastada motiive suurest hulgast lühikestest aminohapete järjestustest ehk peptiididest, mis sisaldavad infot konkreetse inimese organismis olevate antikehade kohta. On alust arvata, et leitud motiivide abil võib olla võimalik tuvastada, milliseid haiguseid inimene on põdenud. Kuna ükski uuritud olemasolevatest tööriistadest selle probleemi lahendamiseks ei sobinud, koostasime motiivide tuvastamiseks uue meetodi. Meetodi esimene osa, sarnaste peptiidigruppide tuvastamine, põhineb hierarhilisel klasterdamisel ning sisaldab kahte erinevat võimalust hierarhilise klasterduse puust automaatselt klastrite eraldamiseks. Meetodi teine osa on sarnaste peptiidide klastritest motiivide tuvastamine. Kuna pärisandmetes olevad motiivid ei ole teada, genereerisime sünteetilised andmed, mille peal koostatud meetodit valideerida. Koostatud meetod suutis vastavalt sünteetiliste andmete omadustele tuvastada 50% kuni 100% sinna sisestatud motiividest, pärisandmetele eeldatavalt kõige sarnasema andmestiku peal 86%. Motiivide lugemise meetod töötas samamoodi hästi, etteantud mürata klastrite pealt suudetakse tuvastada 100% motiividest ning müraga klastrite pealt 90% motiividest. Koostatud meetodit on võimalik rakendada ka teistest bioloogilistest andmetest motiivide otsimiseks. Sel juhul peaks muutma teatud parameetreid, mis selles töös kasutatava andmestiku jaoks on seatud. Edaspidiseks tööks võiks olla meetodi töötamise valideerimine teiste omadustega andmete peal.With the help of new sequencing technologies we can generate a lot of biological data of different backgrounds. These data need to be analysed in order to extract the most important information from them. In this work we develop a method for extracting motifs from a large amount of short amino acid sequences called peptides that contain information about antibodies in that organism. Motifs found from these peptides could be linked to diseases that a person has had. Since none of the tested existing methods were suitable for solving this problem, we developed our own method that consists of two parts. First part, finding groups of similar peptides, is based on hierarchical clustering and has two different options for automatically extracting clusters from the hierarchical clustering tree. Second part is reading motifs from groups of similar peptides. Since we cannot validate the method on real data due to the lack of knowledge about the true motifs in them, we generate synthetic datasets that we validate the developed method on. The percentage of motifs the developed method could identify from synthetic data with different properties ranged from 50% to 100%, with 86% on the data that should be most similar to the real data. Method that reads motifs from group of similar peptides worked also very well. It could identify 100% of motifs from groups of peptides where no noise was added and 90% of motifs from noisier peptide groups. The developed method could be also used for motif discovery on different biological datasets. In that case we would have to change some parameters that were specifically chosen for this problem. Future work could be to test how well this method performs on different biological datasets

    Pattern Discovery from Biosequences

    No full text
    ei saavutettav

    A Linear Model of Genetic Transcription Regulation that Combines Microarray and Genome Sequence Data

    Get PDF
    The thesis proposes a novel method for the analysis of microarray data based on fitting a specific linear model that combines microarray data with DNA sequence information. The model is both descriptive and predictive: its coefficients provide insight into the structure of the genetic regulatory networks, and its predictive performance may be used to find a set of genes that play important role in transcription regulation (transcription factors). An efficient algorithm is proposed for calculating the least-squares fit for the parameters of the model. The proposed method is tested on a synthetic dataset and the results indicate that the approach is capable of detecting interesting relations in the data

    Web Application to Calculate Genetic Risk Scores Based on Imputed Data.

    Get PDF
    Viimastel aastatel on genotüpiseerimise hinna langus teinud võimalikuks geneetilise informatsiooni lisamise tervishoiusüsteemi. Eesti on üks vähestest riikidest, kellel on võimekus muuta see informatsioon arstide jaoks igapäevaseks tööriistaks, kes saaksid seeläbi teha paremini informeeritud otsuseid oma patsientide kohta. Tartu Ülikooli Eesti Geenivaramu (EGV) on üks asutustest, kes töötavad geneetilise info põhjal uute haigusriskide ennustamise mudelite kallal. Teadustöö EGV-s on loonud erinevaid mudeleid polügeensete haiguste riskide hindamiseks. Selle magistritöö käigus esitame tarkvara, mis võimaldab ennustusmudelite kiiremat väljatöötamist ja arenduse käigus tehtud katsetusi.The falling cost of genotyping has made feasible to include genetic information to national healthcare system. Estonia is one of the few countries that have great potential of converting this information into a everyday tool for clinicians, who then would be able to make more informed decisions related to their patient’s health. Estonian Genome Center, University of Tartu (EGCUT) is one of the institutions that is working on creating new risk prediction models based on genetic information. Researchers in EGCUT have created models to evaluate the risk for polygenic disorders. Current thesis focuses on development of a software that would enable fast application of these risk prediction models to collected genetic data and visualization of the results together with clinical data
    corecore