imagine (Institute of molecular genetics and genetic engineering)
Not a member yet
3088 research outputs found
Sort by
Pharmacogenomics and pharmacotranscriptomics of acute leukemia in children: a path to personalized medicine
Personalized medicine is focused on research disciplines which contribute to the individualization of
therapy, like pharmacogenomics and pharmacotranscriptomics. Acute lymphoblastic leukemia (ALL) is
the most common malignancy of childhood. It is one of the pediatric malignancies with the highest cure
rate, but still a lethal outcome due to therapy accounts for 1- 3% of deaths. Further improvement of
treatment protocols is needed through implementation of pharmacogenomics and
pharmacotranscriptomics. Emerging high-throughput technologies, microarrays and next-generation
sequencing, have provided an enormous amount of molecular data with potential to be implemented in
childhood ALL treatment protocols. In the current review, we summarized the contribution of these
novel technologies to pharmacogenomics and pharmacotranscriptomics of childhood ALL. We have
presented data on molecular markers responsible for efficacy, side effects and toxicity of the drugs
commonly used for childhood ALL treatment, i.e., glucocorticoid drugs, vincristine, asparaginase,
anthracyclines, thiopurines and methotrexate. Big data was generated using high-throughput
technologies, but their implementation in clinical practice is poor. Research efforts have to be focused
on data analysis and designing prediction model using machine learning algorithms. Bioinformatics
tools and implementation of artificial intelligence are expected to open the door wide for personalized
medicine in clinical practice of childhood ALL.Book of abstracts: International Conference of Biochemists and Molecular Biologists in Bosnia and Herzegovina - ABMBBIH May, 202
Β-glucosidase b from microbacterium sp. Bg28 as a biofilm control agent In food processing environment
About one in ten people contract a foodborne illness within a year. Children under the age of five are
the most affected, with 125,000 deaths each year. Many of the foodborne illness outbreaks can be
linked to the presence of biofilms in the food industry, and Salmonella enteritidis is an extremely
important foodborne pathogen that thrives in these conditions. It has been shown that biofilms can be
resistant to physical and chemical treatments used in cleaning and disinfection procedures in food
processing. The problem with using more aggressive disinfectants is that they often violate food safety
regulations. The use of enzymes which degrade biofilm matrix structural components should facilitate
current disinfection procedures and not compromise food safety. In this study, the anti-biofilm activity
of recombinantly expressed β-glucosidase B and its potential use as a protective agent to control
Salmonella biofilm formation is investigated. The putative target of this enzyme is cellulose, the
structural component of the Salmonella biofilm matrix. β-Glucosidase B deriving from the
environmental strain Microbacterium sp. BG28 was heterologously expressed in Escherichia coli and
successfully purified by affinity chromatography. The anti-biofilm activity of the enzyme was
evaluated in in vitro assays using various clinical isolates of S. enteritidis. The toxicity of the enzyme
was studied in Caenorhabditis elegans. β-glucosidase B effectively inhibited the formation of
Salmonella biofilms grown in a temperature range of 8°C to 37°C, achieving 50% inhibition at
concentrations of 100μg/ml. Biochemical characterization showed that the optimal pH activity of the
enzyme is between 6 and 7, with the highest activity observed at temperatures between 37°C and 47°C.
The absence of toxicity and other presented results indicate that beta-glucosidase B can be used in
biofilm control in the food industry.Book of abstracts: International Conference of Biochemists and Molecular Biologists in Bosnia and Herzegovina - ABMBBIH May, 202
Crossing roads of plastic degradation and biomaterial production
Iako je već čitav vek prisutna u životima ljudi, plastika je i dalje jedan od
najraznovrsnijih, najčešće proizvođenih i korišćenih materijala. Nekada najveća prednost
plastike – izdržljivost – danas predstavlja veliki problem, jer je čini teško razgradivim
materijalom koji se gomila u životnoj sredini [1]. Najčešće korišćen pristup za odlaganje
ovog polimera je deponovanje, koje je pored ekološke pretnje ujedno i ekonomski izazov,
jer se ovakvim odlaganjem plastike gubi uložena energija i mogućnost za ponovnu upotrebu
materijala. Sa druge strane, biološki proces, zasnovan na enzimskoj razgradnji, pruža
nekoliko prednosti: blage uslove, nizak utrošak energije, i odsustvo opasnih hemikalija [2].
Poseban značaj enzimske razgradnje plastike je što obezbeđuje niz metabolita - polaznih
jedinjenja za proizvodnju novih vrednih polimera [3], čime se doprinosi uspostavljanju
cirkularne ekonomije kada su u pitanju plastični materijali.
Paralelno sa razvijanjem i unapređivanjem bioloških pristupa u razgradnji plastičnog
otpada, istražuju se i ekološki prihvatljivi materijali koji bi mogli zameniti plastiku, kao što
je biorazgradiv biopolimer - bakterijska nanoceluloza (Slika 1). Zahvaljujući svojim
izvanrednim svojstvima kao što su mehanička čvrstoća, hidrofilnost, biokompatibilnost,
obnovljivost i netoksičnost, bakterijska nanocelloza ima potencijal za primenu u različitim
granama industrije [4]. Međutim, produkcija bakterijske nanoceluloze na industrijskoj skali
je otežana visokom cenom medijuma za rast bakterija proizvođača. Iz tog razloga su svetska
istraživanja poslednjih godina usmerena na optimizaciju produkcije bakterijske
nanoceluloze korišćenjem različitih vrsta otpada [5].Knjiga izvoda: 9. simpozijum Hemija i zaštita životne sredine Kladovo, 4-7. jun 2023. BOOK OF ABSTRACTS : 9th Symposium Chemistry and Environmental Protection Kladovo, 4-7th June 202
Metagenomic Analysis of Bacterial Community and Isolation of Representative Strains from Vranjska Banja Hot Spring, Serbia
The hot spring Vranjska Banja is the hottest spring on the Balkan Peninsula with a water temperature of 63–95 °C and a pH value of 7.1, in situ. According to the physicochemical analysis, Vranjska Banja hot spring belongs to the bicarbonated and sulfated hyperthermal waters. The structures of microbial community of this geothermal spring are still largely unexplored. In order to determine and monitor the diversity of microbiota of the Vranjska Banja hot spring, a comprehensive culture-independent metagenomic analysis was conducted in parallel with a culture-dependent approach for the first time. Microbial profiling using amplicon sequencing analysis revealed the presence of phylogenetically novel taxa, ranging from species to phyla. Cultivation-based methods resulted in the isolation of 17 strains belonging to the genera Anoxybacillus, Bacillus, Geobacillus, and Hydrogenophillus. Whole-genome sequencing of five representative strains was then performed. The genomic characterization and OrthoANI analysis revealed that the Vranjska Banja hot spring harbors phylogenetically novel species of the genus Anoxybacillus, proving its uniqueness. Moreover, these isolates contain stress response genes that enable them to survive in the harsh conditions of the hot springs. The results of the in silico analysis show that most of the sequenced strains have the potential to produce thermostable enzymes (proteases, lipases, amylases, phytase, chitinase, and glucanase) and various antimicrobial molecules that can be of great importance for industrial, agricultural, and biotechnological applications. Finally, this study provides a basis for further research and understanding of the metabolic potential of these microorganisms.This is the peer reviewed version of the paper: Malesević, M., Stanisavljević, N., Matijašević, D., Ćurčić, J., Tasić, V., Tasić, S.,& Kojić, M.. (2023). Metagenomic Analysis of Bacterial Community and Isolation of Representative Strains from Vranjska Banja Hot Spring, Serbia. in Microbial Ecology. [https://doi.org/10.1007/s00248-023-02242-6]Related to published version:[ https://imagine.imgge.bg.ac.rs/handle/123456789/1893
Inverting convolutional neural networks for super-resolution identification of regime changes in epidemiological time series
Inferring the timing and amplitude of perturbations in epidemiological systems from
their stochastically spread low-resolution outcomes is as relevant as challenging. It
is a requirement for current approaches to overcome the need to know the details
of the perturbations to proceed with the analyses. However, the general problem of
connecting epidemiological curves with the underlying incidence lacks the highly effective
methodology present in other inverse problems, such as super-resolution and dehazing
from machine vision. I will present an unsupervised physics-informed convolutional neural
network approach in reverse to connect death records with an incidence that allows the
identification of regime changes at a single-day resolution. Applied to COVID-19 data
with proper regularization and model-selection criteria, the approach can identify the
implementation and removal of lockdowns and other nonpharmaceutical interventions
with ± 0.9-day accuracy over the span of a year.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Computational bioengineering for heart disease
In silico clinical trials are a new paradigm for development of a new drug and medical device.
SILICOFCM project is multiscale modeling of familial cardiomyopathy which considers
a comprehensive list of patient specific features as genetic, biological, pharmacologic,
clinical, imaging and cellular aspects.
The 3D deformable-body represents the left and right ventricle of the heart. Blood flow
is modeled during the filling phase by applying the fluid-solid interaction method. The
ventricle wall is modeled by 3D brick 8-node solid elements, with fibers that have threedimensional
direction. The Navier-Stokes equations are solved using the ALE formulation
for fluid with large displacements of the boundary. The ventricle wall model is simulated
by the muscle material model. Muscle fiber orientation is defined by direction vector in 3D
prescribed through input data. The outlet blood pressure is used as the boundary condition.
At the same time, the wall muscle fibers are activated according to the activation function
taken from specific patient measurements.
Computational Platform for Multiscale Modelling in biomedical engineering is results of
SGABU project that is served as an educational tool for students and researchers. The
platform integrates already developed solutions and various datasets related to cancer,
cardiovascular, bone disorders and tissue engineering into one multiscale platform. This
will enable further validation and parameterization of models, creation of environment for
future trends, e.g. in silico clinical trials, virtual surgery, development of prediction models.
InSilc project is devoted to in silico mechanical stent testing within ISO 25539 standards
and in silico stent deployment for metallic and biodegradable material.
In-silico projects will connect basic experimental research with clinical study and
bioinformatics, data mining and image processing tools using very advanced computer
models for drug, stent and patient database in order to reduce animal and clinical studies.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
An agnostic analysis of the human AlphaFold2 proteome using local protein conformations
For more than 30 years, different computational approaches have been implemented to
propose 3D structural models of proteins from their amino acid sequence. Using deep
Learning, AlphaFold 2 obtained particularly remarkable results; some models were within
the uncertainties of the experimental resolution (Jumper et al., Nature 2021). AlphaFold 2
code is freely avalaible and EBI provides structural model databases (Tunyasuvunakool
et al., Nature 2021), i.e. 98.5% of the human proteome is given. 36% of these models are
predicted with atomistic quality.
The human protein models provided by AlphaFold were analyzed using its confidence
index (pLDDT score), with classic secondary structure and finer analysis of local protein
conformation, e.g. γ-turns, β-turns and bends, β-turn types, PolyProline II (PPII), helix
curvatures, β-bulges, and a structural alphabet, namely Protein Blocks (PB).
As expected, the large majority of α-helices are well predicted with high pLDDT scores.
However, some points are intriguing and could potentially lead to improvements in
the future: (i) PPII helices are too often encountered with a low confidence index. They
represent 4-5% of all residues and are important in protein-protein interactions; it could
so be an issue to be poorly approximated. (ii) In a very surprising way, while β-turns
(turns of 4 residues) are well predicted, 55% of γ-turns (3 residues) have very low pLDDT
scores. (iii) Even more strikingly, 94.8% of cis ω angles associated with low pLDDT scores,
i.e. AlphaFold is clearly unable to propose proper cis ω angles. (iv) β-sheet occurrence is
lower than expected, while PB d (i.e. β-sheet core geometry) occurrence is completely in
accordance with the expected frequencies. There are so potentially β-sheets that were
not founded until the end, which would explain this low frequency (de Brevern, Biochimie
2023). AlphaFold 2 had impacted the structural modeling area but works remained
(Tourlet et al., BioMedInformatics 2023)Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Zero- and Few-Shot Machine Learning for Named Entity Recognition in Biomedical Texts
Named entity recognition (NER) is an NLP that involves identifying and classifying named
entities in text. Token classification is a crucial subtask of NER that assumes assigning
labels to individual tokens within a text, indicating the named entity category to which
they belong. Fine-tuning large language models (LLMs) on labeled domain datasets has
emerged as a powerful technique for improving NER performance. By training a pretrained
LLM such as BERT on domain-specific labeled data, the model learns to recognize
named entities specific to that domain with high accuracy. This approach has been applied
to a wide range of domains including biomedical and has demonstrated significant
improvements in NER accuracy.
Still, data for fine-tuning pre-trained LLMs is large and labeling is a time-consuming
and expensive process that requires expert domain knowledge. Also, domains with an
open set of classes yield difficulties in traditional machine learning approaches since the
number of classes to predict needs to be pre-defined.
Our solution to the two mentioned problems is based on data transformation for
factorizing the initial multiple classification problem into a binary one and applying crossencoder-
based BERT architecture for zero- and few-shot learning.
To create our dataset, we transformed six widely used biomedical datasets that contain
various biomedical entities such as genes, drugs, diseases, adverse events, chemicals,
etc., into a uniform format. This transformation process enabled us to merge the datasets
into a single cohesive dataset of 26 different named entity classes.
We then fine-tuned two pre-trained language models: BioBERT and PubMedBERT for the
NER task in zero- and few-shot settings. The results of the experiment for 9 classes in
zero-shot mode are promising for semantically similar classes and improve significantly
after providing only a few supporting examples for almost all classes. The best results
were obtained using a fine-tuned PubMedBERT model, with average F1 scores of
35.44%, 50.10%, 69.94%, and 79.51% for zero-shot, one-shot, 10-shot, and 100-shot NER
respectively.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Using AI to design antibodies
Advancements in antibody engineering are crucial for developing effective and safe therapeutic
candidates. Traditional approaches often involve limited screening of sequence space, resulting
in drug candidates with suboptimal binding affinity, developability, or immunogenicity. However,
recent breakthroughs in deep learning and generative artificial intelligence (AI) offer promising
solutions to overcome these challenges.
In our work, we utilized deep contextual language models trained on high-throughput affinity data
to quantitatively predict binding of unseen antibody sequence variants. Our approach spans a wide
range of binding affinities, demonstrating the potential to optimize antibody engineering. Additionally,
we introduced a “naturalness” metric that measures similarity to natural immunoglobulins. We
found that naturalness is associated with measures of drug developability and immunogenicity,
allowing us to optimize it alongside binding affinity using a genetic algorithm.
Additionally, we explored generative AI-based antibody design, and achieved successful design of
all complementarity-determining regions (CDRs) in the heavy chain of the antibodies. Our designed
antibodies exhibit high binding rates, surpassing randomly sampled antibodies from the Observed
Antibody Space. Moreover, these AI-designed binders display high diversity, low sequence identity
to known antibodies, and favorable naturalness scores, indicating desirable developability profiles
and reduced immunogenicity.
Collectively, our findings demonstrate the immense potential of deep learning and generative
AI in revolutionizing antibody optimization and design. By leveraging large-scale data, predictive
models, and high-throughput experimentation, we can accelerate and improve our antibody
engineering capabilities. The integration of deep contextual language models and the incorporation
of naturalness into the design process provide intelligent screening approaches. Similarly, the
application of generative AI enables us to efficiently and precisely design antibodies from scratch,
outperforming traditional methods in terms of speed and quality.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Computer analysis of glioma gene network structure
Computer analysis of disease susceptibility genes using online bioinformatics tools and
open databases allows the identification of potential target genes for therapy. In the
course of this study we reconstructed the gene network for genes associated with glioma.
The relevance of the work is due to the fact that gliomas are the most common primary
brain tumors. Gliomas originate from glial cells that support and protect nerve cells in
the brain and spinal cord. Despite surgical removal, gliomas are still prone to recurrence
because they grow rapidly in the brain, are resistant to chemotherapy, and are very
aggressive (Byun Y.H. et al, 2022).
The task was to collect a list of glioma genes, analyze gene ontologies, reconstruct the
gene network, and analyze the spatial structures of the associated proteins.
The following online bioinformatics tools were used: STRING-DB (https://string-db.org/)
for gene network construction, MalaCards (https://www.malacards.org/), OMIM database
(https://omim.org/). The search was performed using the keyword “glioma”. AlphaFold
(https://alphafold.ebi.ac.uk/), PDB (https://www.rcsb.org/) resources were used to model
and visualize 3D protein structures. PANTER (http://www.pantherdb.org/) and DAVID
(https://david.ncifcrf.gov/summary.jsp) resources were used to analyze gene ontologies.
The list of genes for analysis consisted of 176 genes.
The most significant categories for glioma genes according to DAVID are: binding of
identical proteins, negative regulation of biological processes, regulation of programmed
cell death, regulation of cell death, and cell population proliferation.
The gene network was reconstructed using the STRING-DB resource (https://string-db.
org/). MicroRNA genes were not recognized by the program. The graph included 150
genes. The study of the gene network structure showed high connectivity of genes
within certain clusters. The EGFR and TP53 genes, which are known and well-studied
oncogenes, had the greatest number of connections, as well as STAT3, KRAS, PIK3CA,
IDH1, KDR. Construction of the glioma gene network showed that some elements of the
graph are sufficiently linked, while others are only partially linked so that the search for
target proteins for glioma treatment can be facilitated.
Three-dimensional structures of KRAS and PIK3CA proteins were constructed using
AlphaFold software (https://alphafold.ebi.ac.uk/). PAE viewer (http://www.subtiwiki.
uni-goettingen.de/v4/paeViewerDemo) was used to check the validity of the predicted
protein structure. The structure of KRAS protein was found to be similar to that of 7ROV
protein obtained from PDB (https://www.rcsb.org/) and the structure of PIK3CA protein
was found to be similar to that of 4YKN protein.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202