1,721,009 research outputs found
genozip: a fast and efficient compression tool for VCF files
genozip is a new lossless compression tool for VCF (Variant Call Format) files. By applying field-specific algorithms and fully utilizing the available computational hardware, genozip achieves the highest compression ratios amongst existing lossless compression tools known to the authors, at speeds comparable with the fastest multi-threaded compressors. genozip is freely available to non-commercial users. It can be installed via conda-forge, Docker Hub, or downloaded from github.com/divonlan/genozip. Supplementary data are available at Bioinformatics online.Divon Lan, Raymond Tobler, Yassine Souilmi and Bastien Llama
Genozip: a universal extensible genomic data compressor
We present Genozip, a universal and fully featured compression software for genomic data. Genozip is designed to be a general-purpose software and a development framework for genomic compression by providing five core capabilities - universality (support for all common genomic file formats), high compression ratios, speed, feature-richness, and extensibility. Genozip delivers high-performance compression for widely-used genomic data formats in genomics research, namely FASTQ, SAM/BAM/CRAM, VCF, GVF, FASTA, PHYLIP, and 23andMe formats. Our test results show that Genozip is fast and achieves greatly improved compression ratios, even when the files are already compressed. Further, Genozip is architected with a separation of the Genozip Framework from file-format-specific Segmenters and data-type-specific Codecs. With this, we intend for Genozip to be a general-purpose compression platform where researchers can implement compression for additional file formats, as well as new codecs for data types or fields within files, in the future. We anticipate that this will ultimately increase the visibility and adoption of these algorithms by the user community, thereby accelerating further innovation in this space. Availability: Genozip is written in C. The code is open-source and available on GitHub (https://github.com/divonlan/genozip). The package is free for non-commercial use. It is distributed as a Docker container on DockerHub and through the conda package manager. Genozip is tested on Linux, Mac, and Windows. Supplementary information: Supplementary data are available at Bioinformatics online.Divon Lan, Ray Tobler, Yassine Souilmi, Bastien Llama
Ancient DNA studies in pre-Columbian Mesoamerica
Mesoamerica is a historically and culturally defined geographic area comprising current central and south Mexico, Belize, Guatemala, El Salvador, and border regions of Honduras, western Nicaragua, and northwestern Costa Rica. The permanent settling of Mesoamerica was accompanied by the development of agriculture and pottery manufacturing (2500 BCE–150 CE), which led to the rise of several cultures connected by commerce and farming. Hence, Mesoamericans probably carried an invaluable genetic diversity partly lost during the Spanish conquest and the subsequent colonial period. Mesoamerican ancient DNA (aDNA) research has mainly focused on the study of mitochondrial DNA in the Basin of Mexico and the Yucatán Peninsula and its nearby territories, particularly during the Postclassic period (900–1519 CE). Despite limitations associated with the poor preservation of samples in tropical areas, recent methodological improvements pave the way for a deeper analysis of Mesoamerica. Here, we review how aDNA research has helped discern population dynamics patterns in the pre-Columbian Mesoamerican context, how it supports archaeological, linguistic, and anthropological conclusions, and finally, how it offers new working hypotheses.Xavier Roca-Rada, Yassine Souilmi, João C. Teixeira and Bastien Llama
Frequency data and SweepFinder2 output
<p>This directory contains the input SFS (site frequency spectrum) data for each of the populations, as well as SweepFinder2 output CLR (composite likelihood ratio) for each population split by chromosome.</p>
yassineS/HTSkey: HTSKey first release
<p>This is an early release for a high-throughput sequencing data processing pipeline.</p>
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
- …
