1,720,977 research outputs found

    A comparison of statistical methods for identifying differentially expressed genes using Affymetrix oligonucleotide arrays

    No full text
    生物晶片能同時平行檢測成千上萬基因的 mRNA 含量,間接說明基因的表現程度。Affymetrix GeneChipTM 公司的專利產品—Affymetrix 高密度寡聚核苷酸晶片 (high-density oligonucleotide array),為一種精準度及再現性 (reproducibility) 較高的 DNA 生物晶片。 Affymetrix GeneChipTM 公司 (2002) 的 Affymetrix Microarray Suite 5.0 (MAS 5.0)、Li and Wong (2001) 的 Model Based Expression Index (MBEI)、Irizarry et al. (2003) 的 Robust Multi-array Average (RMA) 為目前常用之三種表現量轉換方法。在我們的研究中,利用學生氏t檢定、Efron et al. (2001) 的 penalized t-statistic、無母數檢定 (Mann-Whitney test 或 Wilcoxon signed rank test) (Conover, 1999) 以及結合學生氏t統計值或無母數統計值之 Pepe et al. (2003) 的 selection probability function 選拔方法,討論三種表現量轉換方法之鑑別顯著差異表現探針組 (probe sets) 的結果,發現三種表現量轉換方法是有差異的。 此外,我們修正 Hess and Iyer (2004) 的模擬方法,在 R 環境下模擬試驗數據。利用統計模擬,我們間接證明了 Affymetrix 高密度寡聚核苷酸晶片的再現性。我們使用結合學生氏t統計值或無母數統計值之 selection probability function 選拔方法,以 sensitivity、specificity 和 false discovery rate 比較三種表現量轉換方法。在我們的研究中,建議使用 RMA 表現量轉換方法,MBEI 為緊跟其後具競爭力的方法。我們建議使用 Rat 230A 晶片試驗之重複數 (sample size) 在3 ~ 7之間。學生氏t統計值為相較於 Mann-Whitney test 統計值高效且穩定之 selection probability function 選拔方法統計值的選擇。最後將討論重複數的部分整合成一非常實用的演算法 (algorithm),提供給研究人員作為決定重複數之參考,以期能在試驗成本及效率之間取得平衡。Microarray technology has made it possible to measure the abundance of mRNA transcripts for thousands of genes simultaneously. In particular, Affymetrix high-density oligonucleotide array, a patent for Affymetrix GeneChipTM, is very popular in the scientific community due to its high specificity and reproducible property. In this study, we first review three statistical methods, Affymetrix Microarray Suite 5.0 (MAS 5.0) (Affymetrix GeneChipTM, 2002), Model Based Expression Index (MBEI) (Li and Wong, 2001) and Robust Multi-array Average (RMA) (Irizarry et al., 2003), that are currently in use for background correction, normalization and expression transformation. Then we evaluate their performance based on significance tests of the resulting fold change estimates obtained from these methods. Student t-test, penalized t-statistic provided by Efron et al. (2002), and nonparametric test (Mann-Whitney test or Wilcoxon signed rank test) (Conover, 1999) are implemented for the significance test. It is shown that MAS 5.0, MBEI and RMA can lead to quite different conclusions for identification of the differentially expressed probe sets. Therefore, we develop a simulation mechanism to generate replicated experiments. The simulation study is modified from the method recently proposed by Hess and Iyer (2004). Our modified method can mimic naturally occurring data and is based on a real “temperate” array data. For each simulated data set, we directly use the selection probability function proposed by Pepe et al. (2003) with Student t statistic for ranking the expression levels of probe sets. We calculate sensitivity and false discovery rate (FDR) of the three methods based on 100 simulated data sets for various scenarios. We recommend RMA for routine applications because it appears to have higher sensitivity and smaller FDR in all the scenarios under study. Note that MBEI is competitive with RMA in most scenarios. In addition, we develop a practical algorithm to determine sample size of the experiments using Affymetrix oligonucleotide arrays.第一章 前言 1 1.1 Affymetrix高密度寡聚核苷酸晶片之簡介……………………….1 1.2 研究動機與目的……………………………………………………3 第二章 Affymetrix高密度寡聚核苷酸晶片探針組的表現量 5 2.1 晶片背景值校正……………………………………………………5 2.2 晶片正規化…………………………………………………………7 2.3 表現量轉換………………………………………………………..11 2.4 實際試驗資料分析………………………………………………..15 第三章 鑑別具表現差異的探針組 19 3.1 檢定方法…………………………………………………………..19 3.2 Pepe et al. (2003)之selection probability function………………..27 第四章 統計模擬比較及重複數之決定 32 4.1 Hess and Iyer (2004)的模擬方法………………………………….32 4.2 模擬Affymetrix高密度寡聚核苷酸晶片資料……………………34 4.3 模擬結果與討論…………………………………………………...42 4.4 選擇適當重複數之演算法………………………………………...54 第五章 結論與未來研究 55 5.1 總結………………………………………………………………..55 5.2 未來研究…………………………………………………………..57 參考文獻……………………………………………………………………...58 附錄A R和 Bioconductor 的下載及使用………………………………...61 附錄B 討論重複數之R程式………………………………………………6

    Conformance Proportions in a Normal Variance Components Model

    No full text
    良質率定義為一個感興趣的特性值(characteristic) 落於預先指定的可接受範圍(acceptance region) 內之比例。良質率不僅可應用於製造業的製程評估,而且在農業管理及環境監測等領域上,亦可用來幫助評估感興趣的特徵表現。例如一個適當的乾物質含量(dry matter content) 範圍有助於決定飼料用玉米的最佳收穫時機;水果的甜度(sweetness) 要求必須高於某個規格;亦或在農藥殘留檢測中要求毒素的濃度必須低於某個上限值等等,都是想要估計一個隨機變量(random variable) 超過某個給定界線(specification limit) 或是落於某個給定範圍(specification region) 的比例,本質上即是在估計良質率。 本論文建議在實際應用上,以良質率的區間估計作為容許區間(tolerance interval) 的替代方法。首先,我們討論良質率的區間估計與容許區間之間的關係與異同。接下來,對於雙邊良質率(bilateral conformance proportion),我們建構了兩種信賴區間估計方法,一個是以廣義樞軸量(generalized pivotal quantity) 為 基礎,另一個是利用修正大樣本法(modified large sample method) 的概念。我們同樣提出兩種方法來建構單邊良質率(unilateral conformance proportion) 的信賴區間,第一種也是利用廣義樞軸量,第二種則是建構在學生氏t 分佈(Student’s t distribution) 上。一個以拔靴法(bootstrap method) 為基礎的校正(calibration) 概念被應用到單邊及雙邊良質率的區間估計上,使其經驗覆蓋率(empirical coverage probability) 更接近給定的數值。同時,我們也考慮不均衡數據(unbalanced data)的估計。此外,我們藉由一些例子來說明本論文所提出的良質率區間估計方法及 其應用,並且透過統計模擬研究來評估這些方法的成效。結果顯示,在實際應用上,本論文所提出之良質率區間估計為可建議使用之方法。Conformance proportion is defined as the proportion of a performance characteristic of interest that falls within a prespecified acceptance region. It can be used not only in manufacture industry but also in agricultural management or environmental monitoring. For instance, determining best harvest timing for forage maize under an appropriate range of dry matter content, monitoring the sweetness of fruits to be above a lower limit, or requiring the concentration of a toxin to be below an upper limit in pesticide residue tests. It is of desire to estimate the probability that a random variable exceeds a specification limit or falls into a specification region, which is essentially the conformance proportion. In this dissertation, we propose the approach of a conformance proportion as an alternative to that of a tolerance interval for practical use. First, we discuss the connections between the two approaches. Then, two methods are developed for computing confidence limits for bilateral conformance proportions, one is based on the concept of a generalized pivotal quantity and the other is based on the modified large sample method. For unilateral conformance proportions, we also propose two methods for interval estimation, the first one is also based on the concept of a generalized pivotal quantity and the second one is based on the Student’s t distribution. A bootstrap calibration approach is adapted for both bilateral and unilateral conformance proportions to have empirical coverage probability sufficiently close to the nominal level. Furthermore, we consider the situations with unbalanced data scenarios. Some examples are given to illustrate the proposed methods. The performances of these approaches are evaluated by detailed statistical simulation studies, showing that they can be recommended for practical use

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis

    Dispelling the Myths Behind First-author Citation Counts

    Get PDF
    We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more sophisticated methods

    Author Index

    No full text
    Nao informado

    koamabayili/VECTRON-author-checklist: VECTRON author checklist

    No full text
    We have done our best to complete the author checklist relating to the use of animals in the hut study. Note that the objective for the hut study was to evaluate the IRS treatment applications for residual efficacy against Anopheles mosquitoes, including the local An. coluzzii mosquito population. Cows were only used to attract mosquitoes into the huts and no tests were carried out directly on the cows. The author checklist is intended for use with studies where experiments are carried out on animals, which is why we have had such difficulty in completing this for the hut study, as many of the questions do not relate to how the cows were used
    corecore