Abstract
Abstract Motivation: Restriction-site–associated genomic markers are a powerful tool for investigating evolutionary questions at the population level, but are limited in their utility at deeper phylogenetic scales where fewer orthologous loci are typically recovered across disparate taxa. While this limitation stems in part from mutations to restriction recognition sites that disrupt data generation, an additional source of data loss comes from the failure to identify homology during bioinformatic analyses. Clustering methods that allow for lower similarity thresholds and the inclusion of indel variation will perform better at assembling RADseq loci at the phylogenetic scale. Results: PyRAD is a pipeline to assemble de novo RADseq loci with the aim of optimizing coverage across phylogenetic datasets. It uses a wrapper around an alignment-clustering algorithm, which allows for indel variation within and between samples, as well as for incomplete overlap among reads (e.g. paired-end). Here I compare PyRAD with the program Stacks in their performance analyzing a simulated RADseq dataset that includes indel variation. Indels disrupt clustering of homologous loci in Stacks but not in PyRAD , such that the latter recovers more shared loci across disparate taxa. I show through reanalysis of an empirical RADseq dataset that indels are a common feature of such data, even at shallow phylogenetic scales. PyRAD uses parallel processing as well as an optional hierarchical clustering method, which allows it to rapidly assemble phylogenetic datasets with hundreds of sampled individuals. Availability : Software is written in Python and freely available at http://www.dereneaton.com/software/ Contact: daeaton.chicago@gmail.com Supplementary Information: Supplementary data are available at Bioinformatics online.
Keywords
Affiliated Institutions
Related Publications
Stacks 2: Analytical methods for paired‐end sequencing improve RADseq‐based population genomics
Abstract For half a century population genetics studies have put type II restriction endonucleases to work. Now, coupled with massively‐parallel, short‐read sequencing, the fami...
PHYLUCE is a software package for the analysis of conserved genomic loci
Abstract Summary: Targeted enrichment of conserved and ultraconserved genomic elements allows universal collection of phylogenomic data from hundreds of species at multiple time...
<i>Stacks</i>: Building and Genotyping Loci <i>De Novo</i> From Short-Read Sequences
Abstract Advances in sequencing technology provide special opportunities for genotyping individuals with speed and thrift, but the lack of software to automate the calling of te...
<i>Oases:</i>robust<i>de novo</i>RNA-seq assembly across the dynamic range of expression levels
Abstract Motivation: High-throughput sequencing has made the analysis of new model organisms more affordable. Although assembling a new genome can still be costly and difficult,...
CLEVER: clique-enumerating variant finder
Abstract Motivation: Next-generation sequencing techniques have facilitated a large-scale analysis of human genetic variation. Despite the advances in sequencing speed, the comp...
Publication Info
- Year
- 2014
- Type
- article
- Volume
- 30
- Issue
- 13
- Pages
- 1844-1849
- Citations
- 741
- Access
- Closed
External Links
Social Impact
Social media, news, blog, policy document mentions
Citation Metrics
Cite This
Identifiers
- DOI
- 10.1093/bioinformatics/btu121