1,720,984 research outputs found
Annotation Regression for Genome-Wide Association Studies with an Application to Psychiatric Genomic Consortium Data
Although genome-wide association studies (GWAS) have been successful at finding thousands of disease-associated genetic variants (GVs), identifying causal variants and elucidating the mechanisms by which genotypes influence phenotypes are critical open questions. A key challenge is that a large percentage of disease-associated GVs are potential regulatory variants located in noncoding regions, making them difficult to interpret. Recent research efforts focus on going beyond annotating GVs by integrating functional annotation data with GWAS to prioritize GVs. However, applicability of these approaches is challenged by high dimensionality and heterogeneity of functional annotation data. Furthermore, existing methods often assume global associations of GVs with annotation data. This strong assumption is susceptible to violations for GVs involved in many complex diseases. To address these issues, we develop a general regression framework, named Annotation Regression for GWAS (ARoG). ARoG is based on a finite mixture of linear regressions model where GWAS association measures are viewed as responses and functional annotations as predictors. This mixture framework addresses heterogeneity of effects of GVs by grouping them into clusters and high dimensionality of the functional annotations by enabling annotation selection within each cluster. ARoG further employs permutation testing to evaluate the significance of selected annotations. Computational experiments indicate that ARoG can discover distinct associations between disease risk and functional annotations. Application of ARoG to autism and schizophrenia data from Psychiatric Genomics Consortium led to identification of GVs that significantly affect interactions of several transcription factors with DNA as potential mechanisms contributing to these disorders.11Nscopu
Adaptive estimation with partially overlapping models
In many problems, one has several models of interest that capture key parameters describing the distribution of the data. Partially overlapping models are taken as models in which at least one covariate effect is common to the models. A priori knowledge of such structure enables efficient estimation of all model parameters. However, in practice, this structure may be unknown. We propose adaptive composite M-estimation (ACME) for partially overlapping models using a composite loss function, which is a linear combination of loss functions defining the individual models. Penalization is applied to pairwise differences of parameters across models, resulting in data driven identification of the overlap structure. Further penalization is imposed on the individual parameters, enabling sparse estimation in the regression setting. The recovery of the overlap structure enables more efficient parameter estimation. An oracle result is established. Simulation studies illustrate the advantages of ACME over existing methods that fit individual models separately or make strong a priori assumption about the overlap structure
atSNP: transcription factor binding affinity testing for regulatory SNP detection
Motivation: Genome-wide association studies revealed that most disease-associated single nucleotide polymorphisms (SNPs) are located in regulatory regions within introns or in regions between genes. Regulatory SNPs (rSNPs) are such SNPs that affect gene regulation by changing transcription factor (TF) binding affinities to genomic sequences. Identifying potential rSNPs is crucial for understanding disease mechanisms. In silico methods that evaluate the impact of SNPs on TF binding affinities are not scalable for large-scale analysis.
Results: We describe affinity testing for regulatory SNPs (atSNP), a computationally efficient R package for identifying rSNPs in silico. atSNP implements an importance sampling algorithm coupled with a first-order Markov model for the background nucleotide sequences to test the significance of affinity scores and SNP-driven changes in these scores. Application of atSNP with >20 K SNPs indicates that atSNP is the only available tool for such a large-scale task. atSNP provides user-friendly output in the form of both tables and composite logo plots for visualizing SNP-motif interactions. Evaluations of atSNP with known rSNP-TF interactions indicate that atSNP is able to prioritize motifs for a given set of SNPs with high accuracy.33
Ensemble estimation and variable selection with semiparametric regression models
SummaryWe consider scenarios in which the likelihood function for a semiparametric regression model factors into separate components, with an efficient estimator of the regression parameter available for each component. An optimal weighted combination of the component estimators, named an ensemble estimator, may be employed as an overall estimate of the regression parameter, and may be fully efficient under uncorrelatedness conditions. This approach is useful when the full likelihood function may be difficult to maximize, but the components are easy to maximize. It covers settings where the nuisance parameter may be estimated at different rates in the component likelihoods. As a motivating example we consider proportional hazards regression with prospective doubly censored data, in which the likelihood factors into a current status data likelihood and a left-truncated right-censored data likelihood. Variable selection is important in such regression modelling, but the applicability of existing techniques is unclear in the ensemble approach. We propose ensemble variable selection using the least squares approximation technique on the unpenalized ensemble estimator, followed by ensemble re-estimation under the selected model. The resulting estimator has the oracle property such that the set of nonzero parameters is successfully recovered and the semiparametric efficiency bound is achieved for this parameter set. Simulations show that the proposed method performs well relative to alternative approaches. Analysis of an AIDS cohort study illustrates the practical utility of the method.11Nsciescopu
Iteratively Reweighted Group Lasso Based on Log-Composite Regularization
The paper considers supervised learning problems of labeled data with grouped input features. The groups are nonoverlapped such that the model coefficients corresponding to the input features form disjoint groups. The coefficients have group sparsity structure in the sense that coefficients corresponding to each group shall be simultaneously either zero or nonzero. To make effective use of such group sparsity structure given a priori, we introduce a novel log-composite regularizer, which can be minimized by an iterative algorithm. In particular, our algorithm iteratively solves for a traditional group least absolute shrinkage and selection operator (LASSO) problem that involves summing up the l(2) norm of each group until convergent. By updating group weights, our approach enforces a group of smaller coefficients from the previous iterate to be more likely to set to zero compared to the group LASSO. Theoretical results include a minimizing property of the proposed model as well as the convergence of the iterative algorithm to a stationary solution under mild conditions. We conduct extensive experiments on synthetic and real datasets, indicating that our method yields a performance that is superior to that of the state-of-the-art methods in linear regression and binary classification.11Nsciescopu
Nanoscale Visualization of the Electron Conduction Channel in the SiO/Graphite Composite Anode
Conductive atomic force microscopy (C-AFM) is widely used to determine the electronic conductivity of a sample surface with nanoscale spatial resolution. However, the origin of possible artifacts has not been widely researched, hindering the accurate and reliable interpretation of C-AFM imaging results. Herein, artifact-free C-AFM is used to observe the electron conduction channels in Si-based composite anodes. The origin of a typical C-AFM artifact induced by surface morphology is investigated using a relevant statistical method that enables visualization of the contribution of artifacts in each C-AFM image. The artifact is suppressed by polishing the sample surface using a cooling cross-section polisher, which is confirmed by Pearson correlation analysis. The artifact-free C-AFM image was used to compare the current signals (before and after cycling) from two different composite anodes comprising single-walled carbon nanotubes (SWCNTs) and carbon black as conductive additives. The relationship between the electrical degradation and morphological evolution of the active materials depending on the conductive additive is discussed to explain the improved electrical and electrochemical properties of the electrode containing SWCNTs.
atSNP Search: a web resource for statistically evaluating influence of human genetic variation on transcription factor binding
AbstractSummaryUnderstanding the regulatory roles of non-coding genetic variants has become a central goal for interpreting results of genome-wide association studies. The regulatory significance of the variants may be interrogated by assessing their influence on transcription factor binding. We have developed atSNP Search, a comprehensive web database for evaluating motif matches to the human genome with both reference and variant alleles and assessing the overall significance of the variant alterations on the motif matches. Convenient search features, comprehensive search outputs and a useful help menu are key components of atSNP Search. atSNP Search enables convenient interpretation of regulatory variants by statistical significance testing and composite logo plots, which are graphical representations of motif matches with the reference and variant alleles. Existing motif-based regulatory variant discovery tools only consider a limited pool of variants due to storage or other limitations. In contrast, atSNP Search users can test more than 37 billion variant-motif pairs with marginal significance in motif matches or match alteration. Computational evidence from atSNP Search, when combined with experimental validation, may help with the discovery of underlying disease mechanisms.Availability and implementationatSNP Search is freely available at http://atsnp.biostat.wisc.edu.Supplementary informationSupplementary data are available at Bioinformatics online.11Nsciescopu
atSNPInfrastructure, a Case Study for Searching Billions of Records While Providing Significant Cost Savings over Cloud Providers
1
Ultra-widefield fundus autofluorescence in age-related macular degeneration.
Establish accuracy and reproducibility of subjective grading in ultra-widefield fundus autofluorescence (FAF) imaging in patients with age-related macular degeneration (AMD), and determine if an association exists between peripheral FAF abnormalities and AMD.This was a prospective, single-blinded case-control study. Patients were consecutively recruited for the study. Patients were excluded if there was a history of prior or active ocular pathology other than AMD or image quality was insufficient for analysis as determined by two independent graders. Control patients were those without any evidence of AMD or other ophthalmic disease apart from cataract. Using the Optos 200Tx (Optos, Marlborough, MA, USA), a ResMax central macula and an ultra-widefield peripheral retina image was taken for each eye in both normal color and short wavelength FAF. Ultra-widefield photographs were modified to mask the macula. Each ResMax and ultra-widefield image was independently graded by two blinded investigators.There were 28 AMD patients and 11 controls. There was a significant difference in the average age between AMD patients and control groups (80 versus 64, respectively P<0.001). There was moderate, statistically significant agreement between observers regarding image interpretation (78.4%, K = 0.524, P<0.001), and 69.0% (K = 0.49, P<0.001) agreement between graders for FAF abnormality patterns. Patients with AMD were at greater risk for peripheral FAF abnormalities (OR: 3.43, P = 0.019) and patients with FAF abnormalities on central macular ResMax images were at greater risk of peripheral FAF findings (OR: 5.19, P = 0.017).Subjective interpretation of FAF images has moderate reproducibility and validity in assessment of peripheral FAF abnormalities. Peripheral FAF abnormalities are seen in both AMD and control patients. Those with AMD, poor visual acuity, and macular FAF abnormalities are at greater risk
INFIMA leverages multi-omics model organism data to identify effector genes of human GWAS variants.
Genome-wide association studies reveal many non-coding variants associated with complex traits. However, model organism studies largely remain as an untapped resource for unveiling the effector genes of non-coding variants. We develop INFIMA, Integrative Fine-Mapping, to pinpoint causal SNPs for diversity outbred (DO) mice eQTL by integrating founder mice multi-omics data including ATAC-seq, RNA-seq, footprinting, and in silico mutation analysis. We demonstrate INFIMA\u27s superior performance compared to alternatives with human and mouse chromatin conformation capture datasets. We apply INFIMA to identify novel effector genes for GWAS variants associated with diabetes. The results of the application are available at http://www.statlab.wisc.edu/shiny/INFIMA/
- …
