imagine (Institute of molecular genetics and genetic engineering)
Not a member yet
3088 research outputs found
Sort by
Mapping of Disease Names to Disease Codes based on Natural Language Processing Techniques
Information aggregation from various gen, disease, and gen-disease databases such
as DisGeNet, COSMIC, HumsaVar, Orphanet, ClinVar, HPO, and Diseases into a unique
database would enable researchers to analyze and compare valuable domain findings
in a more convenient and systematic way. However, the aggregation poses numerous
challenges due to non-uniform information annotation across the databases. In this work,
we address the problem of mapping a disease name, when needed, into a standardized
disease code (DOID) based on Natural Language Processing text representation
techniques. We examine the benefits and limitations of using off-the-shelf embeddings
such as Med2vec, and language models such as BioBERT, UmlsBERT, and PubMedBERT
in retrieval scenarios with respect to standard full-text search. In addition to qualitative
improvements, we elaborate on the technical requirements and computational
complexities that come with the embracement of language models and semantic search.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
The use of Active Machine Learning for Protospacer-Adjacent Motif recovery in Class 2 CRISPR-Cas systems
The recognition of target DNA sequences during the interference phase of prokaryotic
CRISPR-Cas immunity relies on Protospacer-Adjacent Motif (PAM) sequences, specific
for each Cas effector. PAM identification is a laborious and time consuming process
that requires multiple stages including in vitro and in vivo cleavage assays followed by
Next Generation Sequencing of targets that withstood cleavage. Determining PAM is
an essential step of characterisation of any novel Cas9 ortholog and determines the
likelihood of its potential use. This study investigates the potential of machine learning
to predict PAM sequences for a given Cas9 ortholog based on the results of cleavage
experiments and employing an Active Learning process akin to Reinforcement Learning
with Human Feedback. Machine learning-facilitated PAM identification would streamline
and accelerate existing pipelines for describing novel Cas proteins. We demonstrate that
simple models with a small amount of data are sufficient for confident PAM predictions
when training is effectively orchestrated.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Computational Modelling of Drug Effects on Cardiomyopathy and Analysis of Myocardial Work
Analysis of myocardial work is essential in determination of left ventricle ejection fraction (LVEF)
and non-invasive assessment of different types of cardiomyopathies. Two major classifications of
cardiomyopathy are: dilated (DCM) and hypertrophic (HCM) cardiomyopathy. Although there are
clinical improvements in cardiomyopathy risk assessment, patients are still under high risk of severe
events. Computational modeling of and computer-aided drug design can significantly advance the
understanding of cardiac muscle activity in DCM and HCM cardiomyopathies, speed up the drug
discovery and reduce the risk of severe events, aiming to improve the treatment of cardiomyopathy.
The main advantage and novelty of presented study are coupled macro and micro simulations into the
integrated Fluid Solid Interaction (FSI) system and its application for examination of heart behavior and
drug interactions. In contrary to detailed and patient-specific models where FSI analyses are very timeconsuming,
our models are parametric and based on dimensions of specific LV components. FSI algorithm
within the PAK software is used for modeling the LV with nonlinear material model, together with stretches
integration along muscle fibers. The methods are integrated within the SILICOFCM platform, and aim to
propose an advanced approach for the assessment of work indices and biomechanical characteristics of
cardiomyopathies and drugs effects, based on computational modelling.
In this study, simulations of the effect of drugs on improving performance of DCM LV parametric model
include the drugs that affect calcium transients (Dygoxin) and changes in kinetic parameters (2-deoxy
adenosine triphosphate - dATP). Myocardial word is presented through changes of pressures and
volumes (P-V diagrams) for DCM LV model at basic condition (without administered drug) and with using
Dygoxin and dATP. Due to increased LV size, the P-V loop for the DCM model without administered drug
is shifted toward lower ventricular pressure and lager ventricular volume, with LVEF = 56.83%. Effects of
drugs on DCM show an increase in ventricular peak pressures and LVEFs, while the P-V loops are shifted
toward decreased volumes, corresponding to healthy hearts.
Computational modeling and drug design approaches can speed up the drug discovery and
significantly reduce expenses aiming to improve the treatment of cardiomyopathyBook of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Newest Advances on the FeatureCloud Platform for Federated Learning in Biomedicine
AI in biomedicine has been a central research topic in recent years. Although there are many
different techniques and strategies, the majority rely on data that is of both high quality
and quantity. Despite the steady growth in the amount of data generated for patients, it is
frequently difficult to make that data useful for research because of strong restrictions through
privacy regulations such as the GDPR. Through federated learning (FL), we are able to use
distributed data for machine learning while keeping patient data inside the respective hospital.
Instead of sharing the patient data, like in traditional machine learning, each participant trains
an individual machine learning model and shares the model parameters and weights. Existing
FL frameworks, however, frequently have restrictions on certain algorithms or application
domains, and they frequently call for programming knowledge.
With FeatureCloud, we addressed these limitations and provided a user-friendly solution for
both developers and end-users. FeatureCloud greatly simplifies the complexity of developing
federated applications and executing FL analyses in multi-institutional settings. Additionally,
it provides an app store that makes it easy for the community to publish and reuse federated
algorithms. Apps can be chained together to form pipelines and executed without programming
knowledge, making them ideal for flexible clinical applications. Apps on FeatureCloud can receive
certification from both internal and external reviewers to guarantee safety. FeatureCloud
effectively separates local components from sensitive data systems by utilizing containerization
technology, making it robust to execute in any system environment and guaranteeing data
security. To further ensure the privacy of data, FeatureCloud incorporates privacy-enhancing
technologies and complies with strict data privacy regulations, such as GDPR.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Synthesis, characterization, and in vivo evaluation of the anticancer activity of a series of 5- and 6-(halomethyl)-2,2′-bipyridine rhenium tricarbonyl complexes
We report the synthesis, characterization, and in vivo evaluation of the anticancer activity of a series of 5- and 6-(halomethyl)-2,2′-bipyridine rhenium tricarbonyl complexes. The study was promoted in order to understand if the presence and position of a reactive halomethyl substituent on the diimine ligand system of fac-[Re(CO)3]+ species may be a key molecular feature for the design of active and non-toxic anticancer agents. Only compounds potentially able to undergo ligand-based alkylating reactions show significant antiproliferative activity against colorectal and pancreatic cell lines. Of the new species presented in this study, one compound (5-(chloromethyl)-2,2′-bipyridine derivative) shows significant inhibition of pancreatic tumour growth in vivo in zebrafish-Panc-1 xenografts. The complex is noticeably effective at 8 μM concentration, lower than its in vitro IC50 values, being also capable of inhibiting in vivo cancer cells dissemination
Numerical and Biological Modeling Approach in the Analysis of the Cancer Viability and Apoptosis
Biomedicine is a multidisciplinary branch of science that requires a clear approach to the
study and analysis of various life processes necessary for a deeper understanding of
human health. This research focuses on the use of numerical simulations with the aim of an
improved comprehension of cancer viability and apoptosis during treatment with commercial
chemotherapeutic agents. In recent times, the usage of numerical models was successfully
applied to predict the behavior of tumors. This study includes a wide range of numerical results
that have been obtained by examining cell viability in real-time, determining the type of cell
death and the genetic factors that control these processes. The results of the in vitro test were
used to develop a numerical model that provides a new perspective on the proposed problem.
In this study, colon, and breast cancer cell lines (HCT-116 and MDA-MB-231), as well as healthy
lung fibroblast cell line (MRC-5) were treated with commercial chemotherapeutic agents. The
obtained results showed a decrease in viability and the occurrence of predominantly late
apoptosis upon treatment, as well as a strong correlation between parameters. A mathematical
model was developed and used to gain a better understanding of the investigated processes.
This method can accurately simulate the behavior of cancer cells and reliably predict their
growth.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Efficient bioinformatics workflow for de novo transcriptome assembly of Pelargonium zonale
Variegated Pelargonium zonale is a widely cultivated ornamental plant characterized by
green, photosynthetically active tissue (GL) and white, non-photosynthetic tissue (WL).
The aim of this study was to investigate the transcriptomic differences between these
two tissue types.
We performed RNA-seq analysis of GL and WL on Illumina HiSeq 2500 platform. The
raw reads were processed using in-house scripts to remove low-quality reads, adapter
sequences, poly-N sequences, and contaminants. High-quality clean reads were subjected
to de novo transcriptome assembly using Trinity (min_kmer_cov = 2, min_glue = 2). The
redundancy was removed and longest transcripts per cluster were selected as unigenes.
Gene expression levels were estimated using RSEM by mapping clean data back to the
assembled transcriptome (Bowtie2 with mismatch = 0). Differential expression analysis
between GL and WL (three biological replicates per each) was performed with DESeq2
R package (p values adjusted according to Benjamini and Hochberg for controlling False
Discovery Rate). Genes with abs (log2 FC) ≥ 2 and adjusted p value < 0.05 were assigned
as statistically significant differentially expressed. Functional enrichment analysis was
performed using GOseq R package and KOBAS software (corrected p < 0.05).
We annotated 85,374 unigenes (61.17%), providing a valuable resource for future
functional genomics studies. Out of 8896 gene clusters that were statistically significantly
differentially expressed between the green and white leaf tissues (p value < 0.05 and
abs(log2 fold change) ≥ 2), 5585 were upregulated in the WL, while 3311 were upregulated
in the GL. These findings shed light on the transcriptomic differences between the
two leaf tissue types in P. zonale and provide a foundation for further research on the
functional significance of these differences. Also, this study demonstrated utility of the
Trinity pipeline for de novo transcriptomic analysis of organism whose genomes are yet
not sequenced.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Integrated relational database of human protein-protein interactions
Protein-protein interactions’ data are stored in various publicly available databases of
different types and formats. In this work, a new database for protein-protein interactions
is created by integrating data from multiple existing databases. This task is not trivial
since different databases use distinct gene or protein identifiers for protein annotation.
Additionally, they use different methods to determine interaction scores, and the
interactions are obtained through diverse experimental or predictive methods. As a result,
two databases may store different data about the same interaction.
To integrate data from various databases, namely BioGRID, STRING, HIPPIE, IntAct, and
Reactome, into a single PPI database, the following process is undertaken. Initially, data
is downloaded from these databases in the MITAB format, encompassing all pertinent
interaction information such as protein identifiers, publication sources and other. In order
to obtain unique protein identifiers in all PPIs in the database, the UniProt ID mapping tool
was used to determine UniProt IDs. Next, since scoring systems differ among databases,
for every interaction a new score is calculated using MISCORE tool as an additional metrics
unique for all the PPIs in the database. The resulting database contains tens of millions of
human PPIs from five different sources.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
In Silico analysis and prediction of novel pharmacogenomic markers of pediatric ALL treatment
Acute lymphoblastic leukemia (ALL) is the most common childhood neoplasm. Side
effects of therapy occur in 75% of patients and 1-3% of patients have a lethal outcome due
to treatment. More efficient treatment of pediatric ALL has been developed by avoiding
drug adverse effects included in the treatment protocols. Therefore, implementation of
pharmacogenomics is paramount in pediatric ALL treatment. Next generation sequencing
(NGS) contributed to discovery of novel genetic markers, potential candidates for targeted
therapy and predictors of efficacy and toxicity of drugs.
We aimed to discover novel potential pharmacogenomic markers in pediatric ALL.
DNA samples from bone marrow of 17 pediatric ALL patients were analyzed using
the platform TruSeq Amplicon – Cancer Panel (Ilumina) for somatic mutations in 48
oncogenes. DNA samples from blood of 100 individuals, using the platform TruSightOne
(Ilumina), were analyzed for germinative mutations. An in-house virtual panel for GC
response markers was designed. Predicting the effects of novel variants was performed
using the SIFT, PolyPhen-2 and PROVEAN software tools. For protein structure stability
and modeling we used STRUM method and i-TASSER server.
In the NGS study of somatic mutations in pediatric ALL, 9 novel variants have been
identified. Bioinformatic analysis has shown that STK11 c.1023G>T and ERBB2 c.2341C>T
possess potential as pharmacogenomic markers, therefore, they are candidates for
molecular targeted therapy. In the exome sequencing study, according to the prediction
algorithms, 3 new potential markers in pharmacogenes related to GC response have been
identified, ABCB1 c.947A>G, NCOA3 rs138733364 and TBX21 rs14059812.
Using NGS analysis and prediction algorithms, we have detected 2 novel somatic mutations,
candidates for targeted molecular therapy, as well as 3 novel germinative variants,
potential pharmacogenomic markers of GC response in pediatric ALL. Pharmacogenomic
profiling of each pediatric ALL patient is indispensable for new therapy approaches and it
could lead to better outcomes.Book of abstract: 4th Belgrade Bioinformatics Conference, June 19-23, 202
Supplementary data for article: Solarz, D., Witko, T., Karcz, R., Malagurski, I., Ponjavić, M., Levic, S., Nešić, A., Guzik, M., Savić, S.,& Nikodinović-Runić, J.. (2023). Biological and physiochemical studies of electrospun polylactid/polyhydroxyoctanoate PLA/ P(3HO) scaffolds for tissue engineering applications. in RSC Advances, 13(34), 24112-24128. https://doi.org/10.1039/D3RA03021K
Polyhydroxyoctanoate, as a biocompatible and biodegradable biopolymer, represents an ideal candidate for biomedical applications. However, physical properties make it unsuitable for electrospinning, currently the most widely used technique for fabrication of fibrous scaffolds. To overcome this, it was blended with polylactic acid and polymer blend fibrous biomaterials were produced by electrospinning. The obtained PLA/PHO fibers were cylindrical, smaller in size, more hydrophilic and had a higher degree of biopolymer crystallinity and more favorable mechanical properties in comparison to the pure PLA sample. Cytotoxicity evaluation with human lung fibroblasts (MRC5 cells) combined with confocal microscopy were used to visualize mouse embryonic fibroblasts (MEF 3T3 cell line) migration and distribution showed that PLA/PHO samples support exceptional cell adhesion and viability, indicating excellent biocompatibility. The obtained results suggest that PLA/PHO fibrous biomaterials can be potentially used as biocompatible, biomimetic scaffolds for tissue engineering applications.Supplementary material for:[https://doi.org/10.1039/D3RA03021K]Related to published version: [https://imagine.imgge.bg.ac.rs/handle/123456789/2059