1,720,980 research outputs found
Caraterrizzazione computazionale di melo. Funzione proteica, disordine e variabilità.
Melo domestico è la coltivazione da frutta più diffusa nelle zone temperate ed è stata coltivata in Asia e in Europa dall'antichità. Si osserva un'ampia variabilità fenotipica all'interno di una stessa. Si ritiene che questa variabilità sia alla base della grande diversità della resa fruttifera di melo che si osserva tra le cultivar. Tuttavia, mentre il genoma di melo è disponibile dal 2012, la sua annotazione è carente nelle risorse standard che raccolgono dati sul genoma e sulle proteine, nonostante la sua importanza economica. Ensembl non include il genoma di melo o di qualsiasi altra specie del genere Malus e non presenta nemmeno un modo per codificare le informazioni relative alla variabilità tra cultivar/ecotipi. UniProt annota solo circa il 10% dei geni di melo. Io sostengo che all'origine della mancanza di annotazioni si trova un fenomeno noto come Big Data, che in questo caso è prodotto dalle tecnologie NGS. A causa di una quantità sempre crescente di genomi sequenziati, sono necessari metodi rapidi e precisi per mantenere il passo con il sequenziamento di nuovi genomi, che producano annotazioni da archiviare in database biologici (DB), dove possono essere facilmente recuperate. Tuttavia, i dati relativi alle piante sono poco integrati e questo ha portato a frammentare le informazioni relative alle piante in diversi DB specifici. Questo è il motivo per cui la mela addomesticata è assente o per lo più assente da Ensembl e UniProt. Per colmare il vuoto lasciato dalle risorse esistenti, abbiamo sviluppato PhytoTypeDB, un database che contiene la variabilità inter-cultivar di proteine vegetali annotate funzionalmente. PytoTypeDB è una risorsa facile da usare sviluppata per i ricercatori a recuperare informazioni aggiornate sulla funzione e variabilità dei geni. Per generare l'annotazione PhytoTypeDB, il concetto di famiglia genica e dominio è stato ampiamente sfruttato. Questi concetti ruotano attorno alla premessa che la conservazione delle sequenze proteiche guidi la conservazione delle funzioni. Sebbene questa idea sia senza dubbio giusta per le proteine globulari, non è sufficiente a coprire l'intero spazio delle funzioni proteiche. Inoltre, limitando l'annotazione ai domini conservati automaticamente la copertura delle annotazioni sono polarizzate verso le regioni altamente conservate. Molte proteine tuttavia non hanno una struttura tridimensionale stabile e sono piuttosto intrinsecamente disordinate (ID) in condizioni native. Queste proteine svolgono ruoli critici nella cellula e ospitano la maggior parte della variabilità dal momento che sono in rapida evoluzione. Per indagare su questa classe di proteine abbiamo trasferito dagli Stati Uniti in Italia la risorsa centrale di annotazioni del disordine manualmente curate, DisProt. Dopo aver ri-annotato le voci già presenti e aggiunto duecento nuove annotazioni come sforzo comunitario europeo, abbiamo confrontato DisProt con altre risorse di ID gestite manualmente. Inoltre, abbiamo utilizzato annotazioni manuali per valutare i predittori esistenti, ponendo le basi per una futura valutazione periodica della predizione di ID simile al CASP - chiamata Critical Assessment of Intrinsic protein Disorder (CAID) - che stiamo attualmente eseguendo. Nel campo della predizione dell'ID abbiamo pubblicato un nuovo metodo - MobiDB-lite - che è stato incluso come il primo predittore di ID nella famosa risorsa di annotazione del dominio InterPro. MobiDB-lite è stato usato per predire l'ID sulla più grande scala possibile, tutta lo spazio delle sequenze proteiche contenute in UniProt. Queste annotazioni sono state raccolte nella versione più recente di MobiDB, insieme alle annotazioni da molte altre fonti. Infine, ho approfondito l'analisi della gamma di fenomeni che ricade sotto il nome di ID cercando di classificare ulteriormente e estrapolare modelli su un set di dati su larga scala.Domesticated apple is the most important temperate fruit crop and has been cultivated in Asia and Europe from antiquity. As a consequence of its self-incompatibility a wide variability within a same population is observed for phenotype characters of domesticated apple. This retained variability is thought to be the base of the great diversity of domesticated apple yield among different cultivars. However, while the genome of domesticated apple has been available since 2012, annotation for domesticated apple is lacking in standard resources gathering genome and protein data despite its economic importance. Ensembl, the repository for nucleic acid sequences does not include the genome of domesticated apple or any other species from the genus Malus and does not even feature a way to encode information relative to variability among cultivars/accessions. UniProt only annotates around 10% of domesticated apple genes. I argue that at the origin of lacking annotation stands a recent phenomenon known as Big Data, which in this case is produced by NGS technologies. Due to an ever-increasing amount of sequenced genomes, fast and accurate methods are required to keep the pace with the sequencing of new genomes. Produced annotations must then be stored in biological database (DB)s, where final users can easily access and retrieve information of interest. However, plant-specific data is less integrated than human data and this brought to plant-related information to be fragmented in many different plant-specific DBs. This is why domesticated apple is absent or mostly absent from Ensembl and UniProt. To fill the gap left by existing resources, we developed PhytoTypeDB (http://phytotypedb.bio.unipd.it), a database containing the inter-cultivar variability of functionally annotated plant proteins.PhytoTypeDB is a user-friendly resource developed to help plant scientist to retrieve updated information about gene function and variability. To generate PhytoTypeDB annotation, the concept of gene family and domain were vastly exploited. This Concepts revolve around the premise of protein sequence conservation being the drive for function conservation. While this idea is undoubtedly right for globular proteins, it is not enough to cover the whole protein function space. Furthermore, limiting the annotation to conserved domains automatically bias annotation coverage towards highly conserved regions. Many proteins however lack such stable three-dimensional structure and are rather intrinsically disordered under native conditions. These proteins play critical roles in the cell and host the vast majority of variability since they are quickly evolving. To investigate this class of proteins we transferred from the U.S.A. to Italy the central resource of high quality manually curated Intrinsic Disorder (ID) annotation, Database Of Protein Disorder (DisProt). After re-annotating all legacy entries and adding a supplementary two hundred annotations as a community effort of the European Cost Action NGP-Net, we compared DisProt annotation to other manually curated resources of ID. Furthermore, we used manually curated annotations to evaluate existing automatic detection methods, posingthe foundations for a future periodic assessment of ID prediction similar to CASP - called Critical Assessment of Intrinsic protein Disorder (CAID) - that we are currently running. In the field of prediction of ID we published a novel method - MobiDB-lite - that was included as the first ID predictor in the famous domain annotation resource InterPro.MobiDB-lite was used to predict ID on the largest scale possible, all protein sequence of the UniProt sequence space. These annotations were collected in the newest version of MobiDB, along with annotation from many different sources. Finally, I delved in the analysis of the very nuanced array of phenomena that fall under the name of ID trying to further classify and extrapolate patterns on a large scale dataset
MobiDB-lite: Fast and highly specific consensus prediction of intrinsic disorder in proteins
Intrinsic disorder (ID) is established as an important feature of protein sequences. Its use in proteome annotation is however hampered by the availability of many methods with similar performance at the single residue level, which have mostly not been optimized to predict long ID regions of size comparable to domains. Here, we have focused on providing a single consensus-based prediction, MobiDB-lite, optimized for highly specific (i.e. few false positive) predictions of long disorder. The method uses eight different predictors to derive a consensus which is then filtered for spurious short predictions. Consensus prediction is shown to outperform the single methods when annotating long ID regions. MobiDB-lite can be useful in large-scale annotation scenarios and has indeed already been integrated in the MobiDB, DisProt and InterPro databases
Large-scale analysis of intrinsic disorder flavors and associated functions in the protein sequence universe
Where differences resemble: sequence-feature analysis in curated databases of intrinsically disordered proteins
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
A comprehensive assessment of long intrinsic protein disorder from the DisProt database
Motivation: Intrinsic disorder (ID), i.e. the lack of a unique folded conformation at physiological conditions, is a common feature for many proteins, which requires specialized biochemical experiments that are not high-throughput. Missing X-ray residues from the PDB have been widely used as a proxy for ID when developing computational methods. This may lead to a systematic bias, where predictors deviate from biologically relevant ID. Large benchmarking sets on experimentally validated ID are scarce. Recently, the DisProt database has been renewed and expanded to include manually curated ID annotations for several hundred new proteins. This provides a large benchmark set which has not yet been used for training ID predictors.Results: Here, we describe the first systematic benchmarking of ID predictors on the new DisProt dataset. In contrast to previous assessments based onmissing X-ray data, this dataset contains mostly long ID regions and a significant amount of fully ID proteins. The benchmarking shows that ID predictors work quite well on the new dataset, especially for long ID segments. However, a large fraction of ID still goes virtually undetected and the ranking of methods is different than for PDB data. In particular, many predictors appear to confound ID and regions outside X-ray structures. This suggests that the ID prediction methods capture different flavors of disorder and can benefit from highly accurate curated examples.Availability and implementation: The raw data used for the evaluation are available from URL: http://www.disprot.org/assessment/.Contact: [email protected] information: Supplementary data are available at Bioinformatics online
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
