1,720,958 research outputs found
PREreview of "Defining amino acid pairs as structural units suggests mutation sensitivity to adjacent residues"
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/8237332.
Note: We reviewed an updated version of this preprint for a journal. The comments are posted with this preprint in the hopes the authors post the updated version that we commented on here as the journal we are reviewing for does not place limits on updating preprints during the peer review process.
Ashraya Ravikumar and James Fraser
Summary:
The traditional Ramachandan plot uses the ϕ and ψ torsion angles about the N-C bond and C-C bond respectively to represent aspects of the three dimensional protein backbone structure in two dimensions. Some of the atoms involved in the calculation of ϕ and ψ torsion angles of a residue come from the adjacent residues in the protein chain. In this work, the authors consider the ψ angle of residue i and ϕ angle of residue i+1 as an entity and analyze the distribution of these amino acid pairs. Their approach has the advantage of the torsion angle pair being fully contained in an amino acid pair and the ease of representation of these pairs in the familiar Ramachandran like plot. The authors show that their cross peptide bond plot covers more area than the traditional ϕ,ψ plot and identifies certain structural elements that are "recurring outliers" using the traditional plot. They also show some differences in conformational preference between thermophilic and mesophilic proteins. There is an initial attempt at experimental validation, with small stability changes (measured by melting temperature) upon point mutations to amino acids more favored for that specific region of the cross plot; however, this validation is limited and would benefit from examples intended to be neutral and destabilizing. The major strength of the paper is a new concept that is very simple yet also powerful for identifying regions of conformational space that should be considered "valid", not outliers. In doing so, their method provides a lot of scope for some interesting future work and new ways of validating protein structures, refinement procedures, and structure predictions. The major recurring issue in the manuscript however has been lack of clarity and lack of attention to detail, which can be improved in a future (third?) iteration of the manuscript. The major and minor points of concern are expanded below:
Major points:
1. In Figure 1, the authors claim that "Standard secondary structures (such as left and right -helices, -strands and turns) are clearly recognizable" from the gray stick plots. While the -helices are recognizable especially in clusters 12 and 19, it is not possible to identify -strands and turns from these images. More amino acids may have to be added to the visualization on either side of the pair to make this clear.
2. Their analysis on the correlation of cross bond angles using MD simulations need more details and discussion. The text points to Figure 2 and Table S1 to show the correlation but Figure 2 simply shows the five structures which were simulated along with the amino acid pairs and the cross peptide bond plot. It is unclear how to interpret the correlation between ϕ(k) and ψ(k+1) from this figure. Also, their choice of proteins for this analysis seems arbitrary. What do they mean by "small" protein? What is the dataset of structures they started out with before randomly picking these five structures?
3. How do the authors claim that cluster 9 and 10 shown in Figure 4 represent a transition into helix? The ϕ(k+1),ψ(k+1) distribution is not indicative of being part of a helix.
4. The authors refer to Figure 5 in their discussion about the cluster 15 representing the Type II turn. The representatives of the cluster do appear in different structural contexts. But, are they all indeed turns? Do they satisfy other criteria to be called a beta turn - either distance between C atoms of i and i+3 residue or H bond between carbonyl oxygen of i and amide hydrogen of i+3. This information is needed to support the statement that type II turns are also common in random coil regions. Also, Figure 5 caption says the representatives are from cluster 6, but we are assuming the authors mean cluster 15.
5. The authors have not mentioned what is the color scale used to color the nodes in Figure 6. We are assuming that warmer color means higher probability of the cluster being occupied by thermophiles. The clusters that are part of the most prominent transitions are very close to each other. Maybe many of the amino acid pairs from thermophiles that are classified into cluster 12 could be part of cluster 19 with some minor change in ψ or ϕ angle. Are the alphafold predicted structures accurate enough to distinguish between such close ψ,ϕ angles? These observations could also be a result of inherent biases in alphafold. Perhaps the authors could also analyze experimental structures of mesophiles and thermophiles to see if these trends hold.
6. The statement "Visual inspection of the most prominent cases suggests that the preferred clusters in thermophiles, presumably the more thermostable ones, are those which appear more ordered" is not well supported. If this statement is being made purely based on the cartoon representation of the two cluster transitions shown as inset in Figure 6, then it does not look convincing, especially 1 to 11. Perhaps the authors could analyze the extent of disorder in a more systematic way by comparing preferred clusters of thermophiles and mesophiles and quantitatively look at the difference in disorder, if any. The authors can also show specific examples of transition matrix along with pictorial representations of differences in angles/planes so that the reader can understand them better.
7. In the context specific mutation analysis, the differences in ΔTm observed are very small and the raw DSF data is not shown in the supplemental figure. Are these changes significant enough to conclude about the effect of these context specific mutations on protein stability? Additional experiments are likely needed to place these ΔTm changes in context. What is the typical change in ΔTm for a predicted neutral or deleterious mutation? Can anything be inferred based on prior deep mutational scans for GFP or another protein to help give this analysis more power?
8. Over the years, Ramachandran angle restraints have become part of structure refinement protocols, which when applied inappropriately, could lead to over-optimization of ϕ-ψ angles. For such cases a simple Ramachandran validation will fail to identify issues in the structure and needs a more global approach such as the Ramachandran Z-score (https://www.sciencedirect.c.... From the way the authors' approach is designed, it could have potential for a similar application. They have shown in Figure 3 how outliers of Ramachandran plot fall into acceptable regions in their cross-peptide bond plot. Along these lines, the authors should discuss about global validation metrics derivable from their method (perhaps related to the normalized marginal distribution distance shown in Figure 7)
Minor points:
1. In the introduction section, the authors mention the disadvantage of the (ϕ,ψ)2 method. But apart from mentioning the protein blocks (PB) method, they don't explain how their approach improves over the PB method. It's important to place this work in context to PB since PB's have been used for several structure related applications. Another related work is TERMs (https://www.sciencedirect.c... where the protein structure is broken down into smaller structural entities which are then used to assess model quality, sequence/structure compatibility, conformational transitions, etc. The authors could discuss about the relationship of between their work and TERMs
2. The authors should add labels to panels within figures instead of addressing them as top, middle, etc.
3. What is the correct resolution cut off used to extract structures from PDB? Results and discussion says 1.8Å but methods say 1.5Å. Also, the authors need to provide the list of PDB structures used.
4. The authors have probably cited the wrong reference for the proteomes of mesophile/thermophile bacterial pairs in materials and methods
5. Page 5, para 2, text says Figure 7 while it's actually referring to Figure 6.
6. Page 5, para 3, Fig S3 should be changed to Fig S4
7. Page 6, para 1, unclear if authors are referring to Fig S2 or Figure 7 since Figure S2 does not have the FH and YH distribution. Same para, reference to Figure 6 instead of Figure 7.
8. In Materials and Methods, page 7, under "Calculation of correlation coefficients between dihedral angles", it should be Figure 2
Ashraya Ravikumar and James Fraser
Competing interests
The author declares that they have no competing interests
PREreview of "Approximating conformational Boltzmann distributions with AlphaFold2 predictions"
<p><strong>This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at <a href="https://prereview.org/reviews/10048116">https://prereview.org/reviews/10048116</a>.</strong></p>
<p>In this manuscript the authors have tested the hypothesis that the MSA constructed by AlphaFold2 (AF2) contains information about the distribution of different conformational states of a protein. Whereas the current way of thinking about AF2's MSA-predicted Cβ–Cβ distance maps focuses on their power to provide binary classifications of inter-residue contacts, the authors propose that Cβ–Cβ distances should instead be thought of as a set of collective variables that approximate a Boltzmann distribution. This is a novel hypothesis that lends AF2 the ability to decipher the conformational Boltzmann distributions of proteins. The authors test this in the contexts of protein dynamics, mutation impacts, and protein-protein interactions. They start with analyzing the correlation between AF2 contact distance and spin label distance distributions obtained from EPR spectroscopy using T4 lysozyme as a model, finding a general agreement despite broader AF2 distributions. Following this, they explore if AF2 can approximate free energy changes in systems that contain multiple biologically important minima, using EGFR KD studies for this purpose. AF2 accurately identifies altered contact distance distributions corresponding to active or inactive conformations in several mutations, indicating a sensitivity to alterations that stabilize particular conformational states. Next, they assess sensitivity to thermodynamically destabilizing mutations. AF2 was able to predict different contact distance probabilities for disruptive mutations like L198R in UBA1, but was less sensitive for milder mutations like L198A. Lastly, AF2's sensitivity to protein-protein interactions was explored using the μ-opioid receptor (μOR). Although the helix displacement distances observed in the predicted structure of isolated and complexed μOR do not exactly match with expected values, AF2 did successfully predict differences in select contact distance distributions of active/inactive-state μOR. Demonstrating that Cβ–Cβ distance probabilities from the same AF2-learned distribution reflect distances observed in differentially behaving domains of a protein lends strong support to the hypothesis that AF2 contact distance distributions can approximate conformational distributions. </p><p>The manuscript explores the correlations and sensitivities of AF2 predicted Cβ–Cβ distances across a variety of protein contexts, giving a broad view of its capabilities and limitations. Transitions between the various sections flowed well, and overall the writing was well worded and easily comprehensible. In addition, the presentation was balanced. It doesn't just focus on the success of AF2, but also highlights where its sensitivities might vary or fall short, providing a balanced view of its capabilities. Given limited computational resources, the conformational space explored by MD and MCMC simulations is limited by their initial states. AI methods are instead limited by how informative their system definitions (MSAs and pre-set theoretical or experimental contact distance distributions) are, allowing AI methods, such as the AF2 method outlined by the authors, to more effectively sample conformational space. This is a very fascinating implication of their work which the authors have briefly mentioned in the discussion. This (and the connection to Figure 7 in the paper) warrants a deeper discussion, but the main conclusions the authors come to are within the scope of the manuscript, and are backed up by the evidence presented. </p><p>There are a few points we would like to bring to the attention of the authors to strengthen the manuscript further.</p><p><b>Major points:</b></p><ol><li><p>There are some difficulties interpreting Figure 2. </p><p>(a) It is important to mark the distances between the two chosen pairs of atoms in the active and inactive state. Without this information, the purpose of Figure 2D is unclear and Figure 2D, F and G are difficult to understand. </p><p>(b) Also, what is the threshold distance to classify a state as active or inactive?</p><p>(c) Figure 2E seems confusing with different axis and ranges.</p></li><li><p>In case of DDR1, does the MD simulations reflect the peak distances (between 7.5 and 10.0 Å for DFG-in and between 16.0 and 18.0 Å for DFG-out) observed for AF2 distance distributions? Also, the probability distribution shift towards shorter distances for Y755A does not seem particularly strong at first glance. Is this why the double alanine mutant was included? Are there also MD simulations of the double mutant that show a reduced preference for the DFG-out conformation?</p></li><li><p>The overall results on EGFR mutants are striking. Many of these mutants (most notably L858R have structures deposited in the PDB (ID:2ITT and many others) that are potentially part of the overall training of AF2/OpenFold. Can you comment on how this might affect the results? </p></li></ol><p><b>Minor Points:</b></p><ol><li><p>There is some ambiguity in the statement, "The central hypothesis of this manuscript is that the collective contact distance distributions predicted by AF2 contain relevant information that can approximate Boltzmann distributions provided the relevant conformational states can be adequately described by these contact distances." We suggest adding to this such that a stronger connection is formed between the theory section and the remainder of the paper. For example, the authors could explain that the contact distances specified in each section are the set of CVs you describe earlier, "we identify a set of CVs, <b>ξ</b> = (ξ1, ξ2, …, ξ<i>m</i>)...". It would also be helpful to clarify that the distributions predicted by AF2 represent the ensemble averaged observable, as described by equation 4. Lastly, the authors mention that these distributions can approximate Boltzmann distributions, but this is somewhat vague. This could be reworded to say that AF2 distributions can approximate experimentally derived Boltzmann distributions of the same distance.</p></li><li><p>The authors are comparing Cβ–Cβ distances determined by AF2 to spin label distances from EPR. This is explained in the methods section, but the procedure for adjusting the spin label distances to facilitate a meaningful comparison between them and the AF2 distances is somewhat unclear. To make a stronger justification for why these are comparable, the authors could clarify the procedure. For example, some context from the authors' previous paper, <i>De Novo High-Resolution Protein Structure Determination from Sparse Spin labeling EPR Data</i>: "[distance from spin label] d<sub>SL</sub> is a starting point for the upper estimate of d<sub>Cβ</sub>, and subtracting the effective distance of 6Å twice from d<sub>SL</sub> gives a starting point for the lower estimate of d<sub>Cβ</sub>" could be beneficial. Including a rank correlation coefficient, as hinted above, could also help emphasize that the results demonstrate "similar <i>relative</i> probabilities among the contact distances for AF2 and EPR"</p></li><li><p>In the comparison of distance distributions between AF2 predictions and EPR measures, the major peaks of the two distributions are similar but in certain cases (127CB - 154CB, 120CB - 131CB), some additional peaks are found beyond 10A. A statistical comparison of the distributions, perhaps using a KS test, will help in evaluating the significance of the similarities.</p></li><li><p>Typo in Hamiltonian Equation 1 (should be momentum squared)</p></li><li><p>In the T4 Lysozyme example, how were the six contacts between the 12 unique residues found?</p></li><li><p>In Figure 5, the fourth row could have more discussion/explanation. What does the colorbar represent? There is no label.</p></li><li><p>As mentioned earlier, the connection between the Discussion and Figure 7 is not well established. The authors could expand on their writing and/or make the figure more simplified to match the discussion better.</p></li></ol><p>Reviewed by:</p><p>Jessica Flowers, Angelica Lam, Ashraya Ravikumar, James Fraser</p>
<h2>Competing interests</h2>
<p>
The author declares that they have no competing interests.
</p>
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
koamabayili/VECTRON-author-checklist: VECTRON author checklist
We have done our best to complete the author checklist relating to the use of animals in the hut study. Note that the objective for the hut study was to evaluate the IRS treatment applications for residual efficacy against Anopheles mosquitoes, including the local An. coluzzii mosquito population. Cows were only used to attract mosquitoes into the huts and no tests were carried out directly on the cows. The author checklist is intended for use with studies where experiments are carried out on animals, which is why we have had such difficulty in completing this for the hut study, as many of the questions do not relate to how the cows were used
Author-wise bibliometric analysis based on entropy.
Author-wise bibliometric analysis based on entropy.</p
- …
