1,720,969 research outputs found
Caratterizzazione di strutture proteiche ripetute nella rivoluzione di AlphaFold
For the past fifty years, one of the greatest challenges in bioinformatics has been answering the question: "How can we predict the three-dimensional structure of a protein from its amino acid sequence?”. Experimentally determining the three-dimensional structure of a protein is often a slow and challenging process. Instead, sequencing its amino acids has become a high-throughput task thanks to advancements in technology and reduced costs. This disparity has led structural biology and bioinformatics to focus primarily on globular proteins, which reliably fold into a consistent three-dimensional structure and are therefore more accessible to computational and experimental studies. This focus aligns with the sequence-structure-function paradigm that has long guided our understanding of protein function. However, many proteins belong to the lesser-studied category of Non-Globular Proteins (NGPs), which display more diverse structural and functional characteristics, making them harder to observe.
The introduction of cutting-edge protein structure prediction algorithms like AlphaFold2 and RoseTTAFold has revolutionized the field. These algorithms demonstrated remarkable accuracy in the 14th edition of the CASP competition held in 2020, sparking what is often referred to as the AlphaFold revolution. Today, computational models of protein structures are available for nearly every known protein, significantly accelerating research in structural biology.
Despite this progress certain classes of NGPs, including Tandem Repeat proteins, pose unique challenges in terms of detection and classification. This work focuses on STRPs, a subset of TR proteins with well-defined structural features. The AlphaFold revolution has prompted major updates to databases that store structural data, such as RepeatsDB, which specializes in STRPs. This thesis outlines the evolution of RepeatsDB from its 3rd version in 2021 to its 4th version in 2024, showcasing improvements in manual curation, automated prediction, and scalability in response to the surge of available structural data.
RepeatsDB 4 introduces enhancements to the manual curation process through the development of the RepeatsDB Bio-curation Tool, which has helped refine the definition and classification of STRPs. In collaboration with Pfam, manually collected STRPs have been compared between the two databases, validating and improving the information held in both. The addition of automated prediction methods, such as the newly developed STRPsearch algorithm, represents another significant step forward. STRPsearch integrates curated STRP data with the fast structural search capabilities of FoldSeek to improve and speed up predictions, allowing RepeatsDB to scale and eventually cover the entire AlphaFoldDB.
Furthermore, reengineering RepeatsDB 4 has generated various side products, including the ngx-mol-viewers Angular library, which enhances biological molecule visualization and has already been included in other databases like MobiDB. Moreover, RepeatsDB 4 utilizes a specialized pipeline to enrich biological data by integrating external resources and software. The pipeline has been designed to be executed on a High Performance Computing (HPC) cluster environment via the DRMAAtic library for efficient cluster communication. The architecture of RepeatsDB 4 is designed to be reusable across other projects, such as the DOME registry, underscoring its broader applications in bioinformatics
DRMAAtic: dramatically improve your cluster potential
Motivation The accessibility and usability of high-performance computing (HPC) resources remain significant challenges in bioinformatics, particularly for researchers lacking extensive technical expertise. While Distributed Resource Managers (DRMs) optimize resource utilization, the complexities of interfacing with these systems often hinder broader adoption. DRMAAtic addresses these challenges by integrating the Distributed Resource Management Application API (DRMAA) with a user-friendly RESTful interface, simplifying job management across diverse HPC environments. This framework empowers researchers to submit, monitor, and retrieve computational jobs securely and efficiently, without requiring deep knowledge of underlying cluster configurations. Results We present DRMAAtic, a flexible and scalable tool that bridges the gap between web interfaces and HPC infrastructures. Built on the Django REST Framework, DRMAAtic supports seamless job submission and management via HTTP calls. Its modular architecture enables integration with any DRM supporting DRMAA APIs and offers robust features such as role-based access control, throttling mechanisms, and dependency management. Successful applications of DRMAAtic include the RING web server for protein structure analysis, the CAID Prediction Portal for disorder and binding predictions, and the Protein Ensemble Database deposition server. These deployments demonstrate DRMAAtic's potential to enhance computational workflows, improve resource efficiency, and facilitate open science in life sciences
RING 4.0: faster residue interaction networks with novel interaction types across over 35,000 different chemical structures
Residue interaction networks (RINs) are a valuable approach for representing contacts in protein structures. RINs have been widely used in various research areas, including the analysis of mutation effects, domain-domain communication, catalytic activity, and molecular dynamics simulations. The RING server is a powerful tool to calculate non-covalent molecular interactions based on geometrical parameters, providing high-quality and reliable results. Here, we introduce RING 4.0, which includes significant enhancements for identifying both covalent and non-covalent bonds in protein structures. It now encompasses seven different interaction types, with the addition of pi-hydrogen, halogen bonds and metal ion coordination sites. The definitions of all available bond types have also been refined and RING can now process the complete PDB chemical component dictionary (over 35000 different molecules) which provides atom names and covalent connectivity information for all known ligands. Optimization of the software has improved execution time by an order of magnitude. The RING web server has been redesigned to provide a more engaging and interactive user experience, incorporating new visualization tools. Users can now visualize all types of interactions simultaneously in the structure viewer and network component. The web server, including extensive help and tutorials, is available from URL: https://ring.biocomputingup.it/.Graphical Abstrac
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Contacts prediction of linear peptides from genomic data
The rise of metagenomics and the technological improvements in the fields of bioinformatics and computational biology led to an exponential increase in the amount of biological data available to be studied. However, the rate at which biological data are studied is much slower than the rate at which they are stored. This issue pushed the development of programs capable of extracting significant information from newly sourced data without the need of human intervention. More specifically, some of these programs have been developed to infer structural information from protein sequences. Since the structure of a protein is strictly bound to its function, it is easy to understand the importance of such task. Among the structural information which can be inferred looking at a protein sequence, there are contact maps. Contact maps define whether two residues are functionally linked within the same protein chain or two different ones. Despite much work has been carried out for intra-chain contact maps prediction using sequence information, less can be found about inter-chain contact
maps. Moreover, methods are usually presented and tested on benchmark dataset generated for such purpose. In this, a whole pipeline for both intra-chain and inter-chain contact predictions is presented. Instead of using a generic benchmark set of protein sequences as input, the pipeline starts from predictions of linear interacting peptides at residues level. Linear interacting peptides are regions in a protein sequence which are thought to not have a fixed folding, but to adapt their structure to the functional needs of the protein itself. Needles to say, fewer studies have been conducted about this specific issue in literature. Finally, an analysis of the results is carried out. The analysis focuses on the evaluation of methods implied for contact predictions over the given dataset. Particular attention is paid to the comparison of the performances on inter-chain alignments with respect to the ones achieved on intra-chain alignments. Furthermore, the effect of linear interacting peptides is taken into account
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
RepeatsDB in 2025: expanding annotations of structured tandem repeats proteins on AlphaFoldDB
: RepeatsDB (URL: https://repeatsdb.org) stands as a key resource for the classification and annotation of Structured Tandem Repeat Proteins (STRPs), incorporating data from both the Protein Data Bank (PDB) and AlphaFoldDB. This latest release features substantial advancements, including annotations for over 34 000 unique protein sequences from >2000 organisms, representing a fifteenfold increase in coverage. Leveraging state-of-the-art structural alignment tools, RepeatsDB now offers faster and more precise detection of STRPs across both experimental and predicted structures. Key improvements also include a redesigned user interface and enhanced web server, providing an intuitive browsing experience with improved data searchability and accessibility. A new statistics page allows users to explore database metrics based on repeat classifications, while API enhancements support scalability to manage the growing volume of data. These advancements not only refine the understanding of STRPs but also streamline annotation processes, further strengthening RepeatsDB's role in advancing our understanding of STRP functions
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
