DADA2: High-resolution sample inference from Illumina amplicon data

  • Nat Methods. 2016 Jul;13(7):581-3. doi: 10.1038/nmeth.3869.
Benjamin J Callahan  1 Paul J McMurdie  2 Michael J Rosen  3 Andrew W Han  2 Amy Jo A Johnson  2 Susan P Holmes  1
Affiliations
  • 1. Department of Statistics, Stanford University, Stanford, California, USA.
  • 2. Second Genome, South San Francisco, California, USA.
  • 3. Department of Applied Physics, Stanford University, Stanford, California, USA.
Abstract

We present the open-source software package DADA2 for modeling and correcting Illumina-sequenced amplicon errors (https://github.com/benjjneb/dada2). DADA2 infers sample sequences exactly and resolves differences of as little as 1 nucleotide. In several mock communities, DADA2 identified more real variants and output fewer spurious sequences than other methods. We applied DADA2 to vaginal samples from a cohort of pregnant women, revealing a diversity of previously undetected Lactobacillus crispatus variants.