imagine (Institute of molecular genetics and genetic engineering)
Not a member yet
    3088 research outputs found

    Leveraging Open Source Hardware and Physics-Informed Machine Learning for Accurate Experimental Identification of Bioink Thermophysical Properties in 3D Bioprinting

    No full text
    Three-dimensional bioprinting serves as a foundation for a number of modern tissue engineering technologies, for example organ-on-a-chip models and nerve conduits. The ability of 3D bioprinting to create functional tissue structures depends on the properties of the used bioink – the material that consists of cells and supporting structure. Most frequently used supporting component of a bioink for extrusion-based 3D bioprinting is a temperature-activated hydrogel (e.g. kappa carrageenan or sodium alginate), a hydrogel that solidifies when it reaches a specific activation temperature. Use of composite hydrogels that consist of different components (e.g. a combination of hydroxyapatite nanorods and gelatin) allows for precise control of bioprinted construct properties. Knowledge of the hydrogel’s activation temperature and the associated thermophysical properties is essential for optimizing the bioprinting process since they influence the settings that one needs to select in order to maximize the probability of successful experiment – printhead temperature, printing speed and extrusion multiplier. Additionally, thermophysical properties have an effect on mechanical properties of the final printed model and also affect the viability of the cells within it. Thermophysical properties, as well as mechanical properties, can be controlled by changing the composition of hydrogel. This study provides a design of experiment for determining thermophysical properties of a composite hydrogel bioink sample using temperature sensors of a 3D bioprinter and a Physics-Informed Machine Learning algorithm. The algorithm combines temperature sensor data, physical simulation of the heat exchange process within a sample and accompanying Machine Learning model that selects the most promising combination of hydrogel components for experimental testing. The experiment is based on a hardware solution of custom open design – the experimental setup can be reproduced using a consumer-grade 3D printer and electronic components readily available on the market. We demonstrate that the combination of open source hardware, artificial intelligence control system and physical simulation allows for accurate assessment of the bioink properties thus making 3D bioprinting more reproducible, robust and accessible.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024

    Estimating the dimensionality of omics network embedding space

    No full text
    Thanks to the advances in capturing technology, huge amounts of large-scale biological, omics data have been accumulated. These data are naturally modeled as networks in which nodes represent entities (e.g., patients, genes, metabolites) and edges represent interactions between them. Because of the computational complexity of directly mining networks, current approaches first embed these networks in low-dimensional vector space, and then mine the resulting node embedding vectors for new biomedical knowledge. However, despite successful applications of network-embedding methodologies for mining biological data, there is still no gold-standard approach for determining its key parameter; the number of dimensions of the embedding space. Thus, to set this parameter, most studies rely on computationally inefficient grid-searches. Recently, Two Nearest Neighbors (2NN), a methodology that estimates the intrinsic dimensionality of data-points in high dimensional space, has been successfully applied to estimate the number of dimensions needed to embed synthetic and toy example networks. In this work, we investigate the applicability of 2NN for determining the dimensionality of biological, omics network embedding spaces. On the protein-protein interaction networks and the gene co-expression networks of budding yeast and of homo sapiens, we relate the obtained dimensionality estimations with various network topological properties and with biomedical downstream analysis tasks.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024

    STIM: Multipurpose method for spatial transcriptomics data integration across different technologies

    No full text
    We developed an innovative, statistically based data integration method specifically tailored for spatial transcriptomics data. Our method successfully performed all data integration tasks, while removing batch effects by correcting the entire gene expression matrix, ensuring superior preservation of biological information. The outstanding preservation of biological information is significantly enhanced by employing piece-wise affine transformations for aligning gene expression distributions across samples. Our technique robustly demonstrated exceptional batch-effects correction performances across various experimental technologies, datasets, and integration tasks by outperforming all existing methods, especially in preservation of biological information. As an integral feature, we developed an entirely new spatially aware clustering method capable of accurate identification of tissues and spatial domains. Together with a novel cross-sample clusters mapping methodology, the method ensured robust cross-sample clustering applicable to spatial domains clusters, as well as cell-type clusters. Due to all these features, our method demonstrated remarkable versatility, enabling batch-effectsfree integration of multiple samples, 3D clustering, and the seamless incorporation of healthy and diseased samples. Moreover, our method is the one and only that directly corrects a gene expression matrix by applying transformations which keep the gene regulatory information preserved, allowing studying of gene regulations and gene co-expressions after data integration. This makes our method unique and the only choice for any downstream analysis task which requires full gene expression matrix as an input.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024

    Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping

    No full text
    Unsupervised learning, particularly clustering, is crucial for disease subtyping and patient stratification. With the availability of large-scale multi-omics data, clustering algorithms can be empowered by deep learning models, e.g. variational autoencoder (VAE), to exploit the betweenindividual heterogeneity. However, the impact of confounders—external factors unrelated to the condition, e.g. batch effect and age—on clustering is often overlooked, introducing bias and spurious biological conclusions. We proposed four VAE-based deconfounding approaches utilizing multi-omics data: i) removal of latent features correlated with confounders ii) a conditional variational autoencoder (cXVAE), iii) adversarial training, and iv) adding a regularization term to the loss function. Based on reallife multi-omics data from The Cancer Genome Atlas, we simulated various confounding effects (linear, non-linear, categorical, combination) and evaluated each model’s performance across 50 repetitions based on reconstruction error, clustering stability, and deconfounding effectiveness, measured via the adjusted rand index (ARI). We demonstrated a substantial impact of the artificially introduced confounder effect on patient clustering (ARI: 0.33±0.12), yet various proposed models effectively mitigated this effect, with cXVAE clearly outperforming other frameworks (ARI: 0.66±0.07). cXVAE not only accurately recovers true patient labels but also reveals meaningful pathological associations among cancer types, reinforcing deconfounded representation validity. Conversely, our study proved that some of the proposed strategies, such as adversarial training, are incapable of sufficiently removing confounders. Our study contributes to delivering accurate patient subgrouping by not only (i) proposing novel frameworks for simultaneous multi-omics data integration, dimensionality reduction, and deconfounding of clustering, but also by (ii) benchmarking respective frameworks on open-access data to aid fellow researchers in selecting an appropriate framework that they can readily apply in a health-related settings.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024

    Advancing Supervised Machine Learning for scRNA-seq Data Analysis

    No full text
    The exponential growth of single-cell transcriptomic data presents a significant challenge for the analysis of single-cell transcriptomic data. Current best practices rely on unsupervised clustering. The applications of supervised machine learning (ML) for the analysis of singlecell transcriptomic (scRNA-seq) data have increased in recent years. The main advantages of supervised ML are higher classification accuracy, and reproducibility and reliability of results, compared to unsupervised clustering. However, single-cell transcriptomic technologies are evolving rapidly, resulting in limited reproducibility of results due to changes in biological sample processing and technical differences between subsequent experimental measurements. A lack of high-quality standardized reference datasets increases the risk of model overfitting and reduces model generalization properties. Benchmarking supervised machine learning algorithms is challenging because of the lack of reference datasets. For the advancement of scRNA-seq applications, we need high-quality annotated standardized datasets. To address the need for the deployment of supervised ML in this field, we developed a single-cell transcriptomic database of reference datasets for healthy human peripheral blood mononuclear cells (PBMC). We collected over two million single-cell data from multiple public data sources and applied advanced cell annotation methods to create multi-annotation labels in the healthy PBMC reference dataset. Each cell has labels designating cell type and subtypes, cell cycle, and cell state, along with assigned degree of belief. The annotations are based on the multi-dimensional cell ontology that we have designed. scRNA-seq data in our database were converted into a standardized format using a defined protocol that enables the direct use of data for supervised ML tasks. The data standardization pipeline and cell annotation tools are deployed within the database. The database is deployed as a publicly accessible web server for the study of single-cell PBMC.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024

    Development of automated pharmacogenetic report for evaluating possible side effects of acute lymphoblastic leukemia therapy

    No full text
    Pharmacogenetics is the study of the genetic basis of individual responses to drugs. Adverse drug reactions (ADR) effects occur when an individual with a certain risk genotype gets a certain drug. The majority of risk genotype carriers are unaware of their genetic predisposition. Information about an association between genotypes and drug responses helps drug prescription and dosage. It is particularly important when patients receive many drugs during the treatment courses, such as treatment from acute lymphoblastic leukemia (ALL) in Dmitry Rogachev Center of Pediatric Hematology, Oncology and Immunology. The goal of our work was to develop a personal pharmacogenetic report for childhood ALL patients based on WGS data. In collaboration with the doctors we have compiled a list of drugs and the most relevant side effects of the treatment. Based on literature and public resources analysis we collected a database of genomic variants and haplotypes associated with ADR of ALL therapy. Our pharmacogenetic report lists risks of side effects of 25 drugs based on haplotypes of 13 pharmacogenes and almost 100 short variants. To define haplotypes of the pharmacogenes we have developed a bioinformatics pipeline that includes public tools and our own code. The pipeline was validated on reference samples from the GetRM dataset and 1000 WGS samples from the “100,000+Me” Initiative. A Django-based system was developed for the ADR risks calculation and reports generation. The system takes WGS data as an input, creates predictions for each drug and generates a final pdf report for the patient. Genotypes, haplotypes and predictions are stored in a database for further analysis. Over 50 reports were already issued to the chemotherapy specialists to assist them during the drug and dosage selection process.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024

    Pharmacogenetics-based Dosing Algorithm for Acenocoumarol in the Serbian Population

    No full text
    Pharmacogenetics, as a discipline which correlates genetics of an individual and the effects of drugs, has given new possibilities for personalized approaches in medicine. It is possible to design algorithms to predict the effects of a certain therapeutic by analysing relevant genetic variants as well as non-genetic factors which may influence therapy. It has been shown that algorithms designed in this way allow for better prediction in comparison to traditional trial and error method and represent a more cost-effective approach for health systems. Additionally, the contribution of factors affecting therapy may vary markedly between different ethnic groups. One of the most considered drugs in pharmacogenetics are coumarins (warfarin, acenocoumarol, phenprocoumon), anticoagulation drugs, used in treating and preventing thromboses. This work was aimed to design pharmacogenetics-based algorithm for acenocoumarol. Assuming that population-specific algorithm may take advantage over models used in a generalized manner, we aimed to design a mathematical model for predicting individual drug dosage in the Serbian population based on clinical-demographic and genetic data. Patients with stable acenocoumarol maintenance dose (N = 200) were divided into two cohort – derivation cohort (N = 100) and testing cohort (N = 100) – on a random basis. On the derivation cohort multiple regression analysis was applied in order to select predictors to be used for estimating the individual dose of acenocoumarol and to derive a model for dose prediction. The testing cohort was used for assessing the quality of the derived model. Mathematical model for predicting individual acenocoumarol dose was designed and its unadjusted R2 was 61.8. In addition to genetic factors (VKORC1*2, CYP2C9*2, CYP2C9*3), we identified age, weight and gender of the patients as significant predictors of drug dosage. In comparison with the model given by other authors our model showed better prediction of individual acenocoumarol dose for patients in the Serbian population.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024

    Identifying the cluster of differentiation markers deregulated in colon cancer through analysis of Gene Expression Omnibus database

    No full text
    While the incidence of late-onset (≥50 years) colon cancer (LOCC) has been decreased worldwide, the number of early-onset (<50 years) colon cancer (EOCC) is increasing. Cluster of differentiation (CD) markers are widely and successfully used for identification of leukocyte population using flow cytometry in order to differentiate phenotypes among blood cancers and to address appropriate treatments. However, the alternations in their expression levels have been observed in non-blood cancers as well, including colon cancer, which enhance their diagnostic, prognostic and therapeutic biomarker potential. The aim of this research was to identify potential CD biomarkers for EOCC and LOCC using open access to expression profiling by high throughput sequencing in order to conduct further experimental exploration. Transcriptomic data of dataset GSE240623, obtained from formalin-fixed paraffinembedded tumor tissue samples from 13 EOCC and 13 LOCC patients, and their pairedadjacent normal colon tissues, from Gene Expression Omnibus (GEO) database was used. Differentially expressed genes for EOCC and LOCC paired groups were obtained using DESeq2 package in R software with log2 fold change threshold set to 1. Up and downregulated genes were subsequently filtered only by CD markers for both comparisons. In EOCC patients, CD79B, CD22, CD19, CD79A, CD37 and CD48 were downregulated, while CD276 was upregulated in tumor tissue compared to control tissue. On the other hand, only CD33 was downregulated in LOCC tumor tissue compared to normal colon tissue. Finally, CD27 was downregulated in tumor tissues of both groups, EOCC and LOCC, in comparison to their counterparts. In this study, we have identified nine CD markers that have potential diagnostic, prognostic and/or therapeutic significance for colon cancer and will be further examined in in silico and in vitro study.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024

    Echinococcus spp.in golden jackals (Canis aureus)

    No full text
    The golden jackal (Canis aureus) is a confirmed definitive host for tapeworms of the genus Echinococcus. As Serbia has one of Europe’s largest resident populations of golden jackals, investigating their role in the transmission and distribution of Echinococcus spp. is of interest to public health. To analyze the population genetics of Echinococcus spp. circulating in golden jackals, the project WORM_PROFILER, which started in March of 2024, is collecting gastrointestinal (GI) tracts of legally hunted animals from different areas of Serbia. For this study, GI tracts of 33 animals were processed. After defrosting, the intestines were cut open longitudinally and feces were transferred into 50 mL centrifuge tubes. The mucosa was scraped and analyzed microscopically to isolate parasites. Approximately 3 g of feces was processed by ZnCl2 flotation and sequential mesh filtration to collect taeniid eggs. The DNA from adult Echinococcus spp. and taeniid eggs was extracted using quick boiling in 0.02 M NaOH and screened by multiplex PCR designed to detect E. multilocularis, E. granulosus and E. canadensis. Examination of the mucosa yielded several gravid Echinococcus tapeworms from the small intestine of one animal from western Serbia and eggs were collected from the feces. PCR identified the tapeworm species as E. multilocularis. Single egg picking and sequencing of the cox1 and nad1 genes is underway. Taeniid eggs were additionally collected from another jackal from the same area, but could not be identified as Echinococcus spp. by PCR. These early findings suggest that E. multilocularis is present in western Serbia and sampling in the same geographical area will be intensified. Processing of additional collected samples is ongoing.Book of Abstracts:The XIV European Multicolloquium of Parasitology Wrocław, Poland August 26–30, 202

    Establishment of a platform for genetic polymorphism research in the Balkan population

    No full text
    There are about 30 million people affected by a rare disease in Europe and 80% of them have a genetic component. The development of the whole-genome sequencing (WGS) method enabled the simultaneous identification of genetic variants that potentially cause diseases in affected individuals and by this contributes to enhance diagnostic accuracy, understand disease pathogenesis, and develop therapies. The completion of the Human Genome Project in 2003 led in a revolution in genomic medicine and multiple countries initiated genome projects in order to build reference genome datasets needed for characterization of population-specific variations. Currently, there is no reference genome specific for populations residing in the Balkan Peninsula making interpretation of data achieved by WGS in context of health and disease challenging for these populations. The Balkan Genome Project is designed as a joint scientific collaboration of institutions from countries residing in the Balkan Peninsula and the aim of the Project is to genetically characterize populations of the Balkan Peninsula in order to create a comprehensive and detailed catalogue of human genetic variations specific for populations residing in the Balkan Peninsula. The proposed bilateral project aims to establish at the Institute of Molecular Genetics and Genetic Engineering, University of Belgrade (IMGGE) a platform for genetic polymorphism research of habitants of the Republic of Serbia, as a part of the Balkan Peninsula, which will make significant contribution to initiating the Balkan Genome Project.Principal Investigator: Dr Danijela Drakulić, IMGGEDuration period: 2024-202

    1,327

    full texts

    3,088

    metadata records
    Updated in last 30 days.
    imagine (Institute of molecular genetics and genetic engineering)
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇