imagine (Institute of molecular genetics and genetic engineering)
Not a member yet
3088 research outputs found
Sort by
Leveraging Open Source Hardware and Physics-Informed Machine Learning for Accurate Experimental Identification of Bioink Thermophysical Properties in 3D Bioprinting
Three-dimensional bioprinting serves as a foundation for a number of modern tissue
engineering technologies, for example organ-on-a-chip models and nerve conduits. The
ability of 3D bioprinting to create functional tissue structures depends on the properties
of the used bioink – the material that consists of cells and supporting structure. Most
frequently used supporting component of a bioink for extrusion-based 3D bioprinting is a
temperature-activated hydrogel (e.g. kappa carrageenan or sodium alginate), a hydrogel that
solidifies when it reaches a specific activation temperature. Use of composite hydrogels that
consist of different components (e.g. a combination of hydroxyapatite nanorods and gelatin)
allows for precise control of bioprinted construct properties. Knowledge of the hydrogel’s
activation temperature and the associated thermophysical properties is essential for
optimizing the bioprinting process since they influence the settings that one needs to select
in order to maximize the probability of successful experiment – printhead temperature,
printing speed and extrusion multiplier. Additionally, thermophysical properties have an
effect on mechanical properties of the final printed model and also affect the viability of the
cells within it. Thermophysical properties, as well as mechanical properties, can be controlled
by changing the composition of hydrogel. This study provides a design of experiment for
determining thermophysical properties of a composite hydrogel bioink sample using
temperature sensors of a 3D bioprinter and a Physics-Informed Machine Learning algorithm.
The algorithm combines temperature sensor data, physical simulation of the heat exchange
process within a sample and accompanying Machine Learning model that selects the most
promising combination of hydrogel components for experimental testing. The experiment
is based on a hardware solution of custom open design – the experimental setup can be
reproduced using a consumer-grade 3D printer and electronic components readily available
on the market. We demonstrate that the combination of open source hardware, artificial
intelligence control system and physical simulation allows for accurate assessment of the
bioink properties thus making 3D bioprinting more reproducible, robust and accessible.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024
Estimating the dimensionality of omics network embedding space
Thanks to the advances in capturing technology, huge amounts of large-scale biological,
omics data have been accumulated. These data are naturally modeled as networks in which
nodes represent entities (e.g., patients, genes, metabolites) and edges represent interactions
between them. Because of the computational complexity of directly mining networks, current
approaches first embed these networks in low-dimensional vector space, and then mine the
resulting node embedding vectors for new biomedical knowledge. However, despite successful
applications of network-embedding methodologies for mining biological data, there is still no
gold-standard approach for determining its key parameter; the number of dimensions of the
embedding space. Thus, to set this parameter, most studies rely on computationally inefficient
grid-searches. Recently, Two Nearest Neighbors (2NN), a methodology that estimates the
intrinsic dimensionality of data-points in high dimensional space, has been successfully
applied to estimate the number of dimensions needed to embed synthetic and toy example
networks.
In this work, we investigate the applicability of 2NN for determining the dimensionality of
biological, omics network embedding spaces. On the protein-protein interaction networks
and the gene co-expression networks of budding yeast and of homo sapiens, we relate the
obtained dimensionality estimations with various network topological properties and with
biomedical downstream analysis tasks.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024
STIM: Multipurpose method for spatial transcriptomics data integration across different technologies
We developed an innovative, statistically based data integration method specifically tailored
for spatial transcriptomics data. Our method successfully performed all data integration
tasks, while removing batch effects by correcting the entire gene expression matrix,
ensuring superior preservation of biological information. The outstanding preservation
of biological information is significantly enhanced by employing piece-wise affine
transformations for aligning gene expression distributions across samples. Our technique
robustly demonstrated exceptional batch-effects correction performances across various
experimental technologies, datasets, and integration tasks by outperforming all existing
methods, especially in preservation of biological information.
As an integral feature, we developed an entirely new spatially aware clustering method
capable of accurate identification of tissues and spatial domains. Together with a novel
cross-sample clusters mapping methodology, the method ensured robust cross-sample
clustering applicable to spatial domains clusters, as well as cell-type clusters. Due to all
these features, our method demonstrated remarkable versatility, enabling batch-effectsfree
integration of multiple samples, 3D clustering, and the seamless incorporation of
healthy and diseased samples.
Moreover, our method is the one and only that directly corrects a gene expression matrix by
applying transformations which keep the gene regulatory information preserved, allowing
studying of gene regulations and gene co-expressions after data integration. This makes
our method unique and the only choice for any downstream analysis task which requires full
gene expression matrix as an input.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024
Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping
Unsupervised learning, particularly clustering, is crucial for disease subtyping and patient
stratification. With the availability of large-scale multi-omics data, clustering algorithms can be
empowered by deep learning models, e.g. variational autoencoder (VAE), to exploit the betweenindividual
heterogeneity. However, the impact of confounders—external factors unrelated to
the condition, e.g. batch effect and age—on clustering is often overlooked, introducing bias and
spurious biological conclusions.
We proposed four VAE-based deconfounding approaches utilizing multi-omics data: i) removal
of latent features correlated with confounders ii) a conditional variational autoencoder (cXVAE),
iii) adversarial training, and iv) adding a regularization term to the loss function. Based on reallife
multi-omics data from The Cancer Genome Atlas, we simulated various confounding effects
(linear, non-linear, categorical, combination) and evaluated each model’s performance across 50
repetitions based on reconstruction error, clustering stability, and deconfounding effectiveness,
measured via the adjusted rand index (ARI).
We demonstrated a substantial impact of the artificially introduced confounder effect on patient
clustering (ARI: 0.33±0.12), yet various proposed models effectively mitigated this effect, with cXVAE
clearly outperforming other frameworks (ARI: 0.66±0.07). cXVAE not only accurately recovers true
patient labels but also reveals meaningful pathological associations among cancer types, reinforcing
deconfounded representation validity. Conversely, our study proved that some of the proposed
strategies, such as adversarial training, are incapable of sufficiently removing confounders.
Our study contributes to delivering accurate patient subgrouping by not only (i) proposing novel
frameworks for simultaneous multi-omics data integration, dimensionality reduction, and
deconfounding of clustering, but also by (ii) benchmarking respective frameworks on open-access
data to aid fellow researchers in selecting an appropriate framework that they can readily apply
in a health-related settings.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024
Advancing Supervised Machine Learning for scRNA-seq Data Analysis
The exponential growth of single-cell transcriptomic data presents a significant challenge for
the analysis of single-cell transcriptomic data. Current best practices rely on unsupervised
clustering. The applications of supervised machine learning (ML) for the analysis of singlecell
transcriptomic (scRNA-seq) data have increased in recent years. The main advantages of
supervised ML are higher classification accuracy, and reproducibility and reliability of results,
compared to unsupervised clustering. However, single-cell transcriptomic technologies
are evolving rapidly, resulting in limited reproducibility of results due to changes in
biological sample processing and technical differences between subsequent experimental
measurements. A lack of high-quality standardized reference datasets increases the risk of
model overfitting and reduces model generalization properties. Benchmarking supervised
machine learning algorithms is challenging because of the lack of reference datasets.
For the advancement of scRNA-seq applications, we need high-quality annotated
standardized datasets. To address the need for the deployment of supervised ML in this
field, we developed a single-cell transcriptomic database of reference datasets for healthy
human peripheral blood mononuclear cells (PBMC). We collected over two million single-cell
data from multiple public data sources and applied advanced cell annotation methods to
create multi-annotation labels in the healthy PBMC reference dataset. Each cell has labels
designating cell type and subtypes, cell cycle, and cell state, along with assigned degree
of belief. The annotations are based on the multi-dimensional cell ontology that we have
designed. scRNA-seq data in our database were converted into a standardized format using
a defined protocol that enables the direct use of data for supervised ML tasks. The data
standardization pipeline and cell annotation tools are deployed within the database. The
database is deployed as a publicly accessible web server for the study of single-cell PBMC.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024
Development of automated pharmacogenetic report for evaluating possible side effects of acute lymphoblastic leukemia therapy
Pharmacogenetics is the study of the genetic basis of individual responses to drugs.
Adverse drug reactions (ADR) effects occur when an individual with a certain risk genotype
gets a certain drug. The majority of risk genotype carriers are unaware of their genetic
predisposition. Information about an association between genotypes and drug responses
helps drug prescription and dosage. It is particularly important when patients receive many
drugs during the treatment courses, such as treatment from acute lymphoblastic leukemia
(ALL) in Dmitry Rogachev Center of Pediatric Hematology, Oncology and Immunology. The
goal of our work was to develop a personal pharmacogenetic report for childhood ALL
patients based on WGS data.
In collaboration with the doctors we have compiled a list of drugs and the most relevant
side effects of the treatment. Based on literature and public resources analysis we collected
a database of genomic variants and haplotypes associated with ADR of ALL therapy. Our
pharmacogenetic report lists risks of side effects of 25 drugs based on haplotypes of 13
pharmacogenes and almost 100 short variants.
To define haplotypes of the pharmacogenes we have developed a bioinformatics pipeline
that includes public tools and our own code. The pipeline was validated on reference samples
from the GetRM dataset and 1000 WGS samples from the “100,000+Me” Initiative. A
Django-based system was developed for the ADR risks calculation and reports generation.
The system takes WGS data as an input, creates predictions for each drug and generates
a final pdf report for the patient. Genotypes, haplotypes and predictions are stored in a
database for further analysis. Over 50 reports were already issued to the chemotherapy
specialists to assist them during the drug and dosage selection process.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024
Pharmacogenetics-based Dosing Algorithm for Acenocoumarol in the Serbian Population
Pharmacogenetics, as a discipline which correlates genetics of an individual and the effects
of drugs, has given new possibilities for personalized approaches in medicine. It is possible
to design algorithms to predict the effects of a certain therapeutic by analysing relevant
genetic variants as well as non-genetic factors which may influence therapy. It has been
shown that algorithms designed in this way allow for better prediction in comparison to
traditional trial and error method and represent a more cost-effective approach for health
systems. Additionally, the contribution of factors affecting therapy may vary markedly
between different ethnic groups. One of the most considered drugs in pharmacogenetics
are coumarins (warfarin, acenocoumarol, phenprocoumon), anticoagulation drugs, used in
treating and preventing thromboses.
This work was aimed to design pharmacogenetics-based algorithm for acenocoumarol.
Assuming that population-specific algorithm may take advantage over models used in a
generalized manner, we aimed to design a mathematical model for predicting individual
drug dosage in the Serbian population based on clinical-demographic and genetic data.
Patients with stable acenocoumarol maintenance dose (N = 200) were divided into two
cohort – derivation cohort (N = 100) and testing cohort (N = 100) – on a random basis. On
the derivation cohort multiple regression analysis was applied in order to select predictors
to be used for estimating the individual dose of acenocoumarol and to derive a model
for dose prediction. The testing cohort was used for assessing the quality of the derived
model.
Mathematical model for predicting individual acenocoumarol dose was designed and its
unadjusted R2 was 61.8. In addition to genetic factors (VKORC1*2, CYP2C9*2, CYP2C9*3),
we identified age, weight and gender of the patients as significant predictors of drug
dosage. In comparison with the model given by other authors our model showed better
prediction of individual acenocoumarol dose for patients in the Serbian population.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024
Identifying the cluster of differentiation markers deregulated in colon cancer through analysis of Gene Expression Omnibus database
While the incidence of late-onset (≥50 years) colon cancer (LOCC) has been decreased
worldwide, the number of early-onset (<50 years) colon cancer (EOCC) is increasing.
Cluster of differentiation (CD) markers are widely and successfully used for identification
of leukocyte population using flow cytometry in order to differentiate phenotypes among
blood cancers and to address appropriate treatments. However, the alternations in their
expression levels have been observed in non-blood cancers as well, including colon cancer,
which enhance their diagnostic, prognostic and therapeutic biomarker potential.
The aim of this research was to identify potential CD biomarkers for EOCC and LOCC using
open access to expression profiling by high throughput sequencing in order to conduct
further experimental exploration.
Transcriptomic data of dataset GSE240623, obtained from formalin-fixed paraffinembedded
tumor tissue samples from 13 EOCC and 13 LOCC patients, and their pairedadjacent
normal colon tissues, from Gene Expression Omnibus (GEO) database was
used. Differentially expressed genes for EOCC and LOCC paired groups were obtained
using DESeq2 package in R software with log2 fold change threshold set to 1. Up and
downregulated genes were subsequently filtered only by CD markers for both comparisons.
In EOCC patients, CD79B, CD22, CD19, CD79A, CD37 and CD48 were downregulated,
while CD276 was upregulated in tumor tissue compared to control tissue. On the other
hand, only CD33 was downregulated in LOCC tumor tissue compared to normal colon
tissue. Finally, CD27 was downregulated in tumor tissues of both groups, EOCC and LOCC,
in comparison to their counterparts. In this study, we have identified nine CD markers that
have potential diagnostic, prognostic and/or therapeutic significance for colon cancer and
will be further examined in in silico and in vitro study.Book of abstracts: 5th Belgrade Bioinformatics Conference, Serbia, Belgrade,17-20 june 2024
Echinococcus spp.in golden jackals (Canis aureus)
The golden jackal (Canis aureus) is a confirmed definitive host for
tapeworms of the genus Echinococcus. As Serbia has one of Europe’s
largest resident populations of golden jackals, investigating
their role in the transmission and distribution of Echinococcus spp.
is of interest to public health. To analyze the population genetics
of Echinococcus spp. circulating in golden jackals, the project
WORM_PROFILER, which started in March of 2024, is collecting
gastrointestinal (GI) tracts of legally hunted animals from
different areas of Serbia. For this study, GI tracts of 33 animals
were processed. After defrosting, the intestines were cut open
longitudinally and feces were transferred into 50 mL centrifuge
tubes. The mucosa was scraped and analyzed microscopically to
isolate parasites. Approximately 3 g of feces was processed by
ZnCl2 flotation and sequential mesh filtration to collect taeniid
eggs. The DNA from adult Echinococcus spp. and taeniid eggs was
extracted using quick boiling in 0.02 M NaOH and screened by
multiplex PCR designed to detect E. multilocularis, E. granulosus
and E. canadensis. Examination of the mucosa yielded several
gravid Echinococcus tapeworms from the small intestine of one
animal from western Serbia and eggs were collected from the
feces. PCR identified the tapeworm species as E. multilocularis.
Single egg picking and sequencing of the cox1 and nad1 genes is
underway. Taeniid eggs were additionally collected from another
jackal from the same area, but could not be identified as Echinococcus
spp. by PCR. These early findings suggest that E. multilocularis
is present in western Serbia and sampling in the same geographical
area will be intensified. Processing of additional collected
samples is ongoing.Book of Abstracts:The XIV European Multicolloquium of Parasitology Wrocław, Poland August 26–30, 202
Establishment of a platform for genetic polymorphism research in the Balkan population
There are about 30 million people affected by a rare disease in Europe and 80% of them have a genetic
component. The development of the whole-genome sequencing (WGS) method enabled the
simultaneous identification of genetic variants that potentially cause diseases in affected individuals and
by this contributes to enhance diagnostic accuracy, understand disease pathogenesis, and develop
therapies. The completion of the Human Genome Project in 2003 led in a revolution in genomic
medicine and multiple countries initiated genome projects in order to build reference genome datasets
needed for characterization of population-specific variations. Currently, there is no reference genome
specific for populations residing in the Balkan Peninsula making interpretation of data achieved by WGS
in context of health and disease challenging for these populations. The Balkan Genome Project is
designed as a joint scientific collaboration of institutions from countries residing in the Balkan Peninsula
and the aim of the Project is to genetically characterize populations of the Balkan Peninsula in order to
create a comprehensive and detailed catalogue of human genetic variations specific for populations
residing in the Balkan Peninsula. The proposed bilateral project aims to establish at the Institute of
Molecular Genetics and Genetic Engineering, University of Belgrade (IMGGE) a platform for genetic
polymorphism research of habitants of the Republic of Serbia, as a part of the Balkan Peninsula, which
will make significant contribution to initiating the Balkan Genome Project.Principal Investigator: Dr Danijela Drakulić, IMGGEDuration period: 2024-202