1,720,971 research outputs found
Synthetic - FEGA Sweden Heilsa synthetic dataset December 2023
Synthetic - This submission contains a subset of a synthetic dataset derived from the project Heilsa Tryggvedottir - a Nordic collaboration on sharing sensitive human data. Heilsa Tryggvedottir is funded by the Nordic e-Infrastructure Collaboration (NeIC), the ELIXIR nodes of Finland, Norway, and Sweden, Computerome in Denmark, and the Estonian Scientific Computing Infrastructure (ETAIS).
In the synthetic data creation process, it was attempted to strike a fine balance between the usability of the datasets (e.g. technical FEGA development, testing, user training, and basic bioinformatics) and compliance with GDPR. File names and file content (e.g. headers in fastq) are anonymized. Moreover, the X, Y, and mitochondrial sequences have been discarded from the original data since these data can be used for maternal, paternal, or ethnic origin tracing. The dataset does not follow natural haplotype distribution (inherent to imputation panels). The only inputs derived from real sequence data are variant distribution density per chromosome and learning sequencing error models.
The synthetic dataset consists of two fastq files, a cram file, a vcf file, and two index files.
This dataset is 1 of 1 included in the study titled Synthetic - FEGA Sweden Heilsa synthetic dataset December 2023, http://identifiers.org/ega.study:EGAS50000000086
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
RNAseq data from 112 samples of benign or malignant ovarian tumours
This dataset contains ~1.2 TB RNA sequencing data in fastq format from 112 samples of fresh-frozen ovarian tumours from a total of 111 women, two samples being replicates. The samples were collected from the U-CAN collection at Uppsala Biobank and include both benign (n = 18) and malignant (n = 94) tumours.
The RNA sequencing samples were sequenced using paired-end sequencing (2x150 bp) on an Illumina NovaSeq 6000 instrument at the SciLifeLab National Genomics Infrastructure (NGI) in Uppsala.
This dataset is 1 of 1 included in the study titled Total RNA expression in benign ovarian and malignant ovarian tumours, http://identifiers.org/ega.study:EGAS50000001045
Whole exome sequencing data from 120 AML samples
This dataset contains bam-files from whole exome sequencing of 120 paired tumor-normal pairs from AML. Tumor DNA was extracted from either bone marow or peripheral blood from primary AML samples. Normal DNA was extracted from cultured skin fibroblast samples. The libraries were prepared using the Nextera rapid capture exome kit and sequenced on an Illumina NextSeq 500 using 2x151bp paired end chemistry. The fastq files generated by sequencing were aligned to the human hg19 reference genome (ucsc.hg19.fasta from the GATK resource bundle) using bwa (0.7.9a-r786 or 0.7.15-r1140) and duplicate reads were identified using samblaster (0.1.24).
This dataset is 1 of 4 included in the study titled The cellular state space of AML unveils novel NPM1 subtypes with distinct clinical outcomes and immune evasion properties, http://identifiers.org/ega.study:EGAS50000001084
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
RNA sequencing of 120 AML samples
This dataset contains fastq-files from bulk RNA sequencing of 120 AML samples. RNA was extracted from either bone marow or peripheral blood from primary AML samples. The libraries were prepared using Illumina Truseq RNA library preparation kit v2 and sequenced on an Illumina NextSeq 500 using 2x151bp paired end chemistry.
This dataset is 1 of 4 included in the study titled The cellular state space of AML unveils novel NPM1 subtypes with distinct clinical outcomes and immune evasion properties, http://identifiers.org/ega.study:EGAS50000001084
- …
