1,721,130 research outputs found
The evolutionarily conserved Dim1 protein defines a novel branch of the thioredoxin fold superfamily
Zhang, Yu-Zhu, Kathleen L. Gould, Roland L. Dunbrack, Jr., Hong Cheng, Heinrich Roder, and Erica A. Golemis. The evolutionarily conserved Dim1 protein defines a novel branch of the thioredoxin fold superfamily. Physiol. Genomics 1: 109–118, 1999.—Dim1 is a small evolutionarily conserved protein essential for G2/M transition that has recently been implicated as a component of the mRNA splicing machinery. To date, the mechanism of Dim1 function remains poorly defined, in part because of the absence of informative sequence homologies between Dim1 and other functionally defined proteins or protein domains. We have used a combination of molecular modeling and NMR structural analysis to demonstrate that ∼125 of the 142 amino acids of human Dim1 (hDim1) define a novel branch of the thioredoxin fold superfamily. Mutational analysis of Dim1 based on the predicted fold indicates that alterations in the region corresponding to the thioredoxin active site do not affect Dim1 activity. However, removal of a very short carboxy-terminal extension generates a dominant negative form of the protein [hDim1-(1–128)] that when overproduced induces cell cycle arrest in G2, via a mechanism likely to involve alteration of Dim1 association with partner molecules. In sum, this study identifies the Dim1 proteins as a novel sixth branch of the thioredoxin superfamily involved in cell cycle.</jats:p
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
The role of balanced training and testing data sets for binary classifiers in bioinformatics.
Training and testing of conventional machine learning models on binary classification problems depend on the proportions of the two outcomes in the relevant data sets. This may be especially important in practical terms when real-world applications of the classifier are either highly imbalanced or occur in unknown proportions. Intuitively, it may seem sensible to train machine learning models on data similar to the target data in terms of proportions of the two binary outcomes. However, we show that this is not the case using the example of prediction of deleterious and neutral phenotypes of human missense mutations in human genome data, for which the proportion of the binary outcome is unknown. Our results indicate that using balanced training data (50% neutral and 50% deleterious) results in the highest balanced accuracy (the average of True Positive Rate and True Negative Rate), Matthews correlation coefficient, and area under ROC curves, no matter what the proportions of the two phenotypes are in the testing data. Besides balancing the data by undersampling the majority class, other techniques in machine learning include oversampling the minority class, interpolating minority-class data points and various penalties for misclassifying the minority class. However, these techniques are not commonly used in either the missense phenotype prediction problem or in the prediction of disordered residues in proteins, where the imbalance problem is substantial. The appropriate approach depends on the amount of available data and the specific problem at hand
- …
