The literature, compressed

Things that will bite you

The rest of this site is about choices inside the pipeline. This page is about the things that were already wrong before the pipeline started, or that no pipeline choice can fix. Every one of them has a canonical paper behind it, and every one of them is still being rediscovered the hard way.

1 · The kitome

Salter et al. (BMC Biology, 2014) showed that contaminating bacterial DNA is ubiquitous in DNA extraction kits and laboratory reagents, varies between kits and between batches of the same kit, and critically distorts results from low-biomass samples. It affects 16S surveys and shotgun metagenomics alike. This is the single most common way a microbiome study produces a confident, reproducible, entirely artefactual result.

contaminant DNA per reaction — constant kit · buffers · polymerase · water high biomass — stool <2% real community low biomass — swab, tissue, air, blood the kit real community 67% contaminant fraction ∝ 1 / input biomass if biomass differs by group, so does the contaminant and it will look like a real, consistent, replicable effect extraction blanks · PCR blanks · sequenced never empty — their content is the measurement decontam: frequency vs concentration, or prevalence

2 · Controls, and what each one rules out

ControlWhat it catchesWhat happens without it
Extraction blank
kit, no sample
Contaminant DNA in the kit and reagentsReagent taxa are reported as biology. Unfalsifiable after the fact.
PCR blank
water template
Contamination introduced at amplification, and index hoppingYou cannot tell an extraction contaminant from a PCR one, so you cannot fix the right step.
Mock community
known composition
Classifier accuracy, primer bias, chimera rate, copy-number distortionNo idea whether your taxonomy is right. This is the only direct check you have.
Technical replicates
same sample, twice
How much of your between-sample variation is the protocolBiological variance and technical variance are pooled and reported as biology.
Randomised batchesConfounding by kit lot, plate, run, extraction dayNothing downstream can separate the batch from the treatment. See Forking Paths, study one.

If you have a mock community, classify it and report the result. If you do not, say so — that is a limitation, not an omission, and reviewers treat the two very differently.

3 · 16S copy number is not one

Bacterial genomes carry between one and roughly fifteen copies of the 16S rRNA gene. A taxon with seven copies contributes seven times the reads of a taxon with one at the same cell abundance. So an amplicon table is not a table of organisms — it is a table of gene copies, and the mapping between them is taxon-specific and largely unknown for uncultured lineages.

Copy-number correction tools exist and are contested: the correction depends on predicting copy number for taxa whose genomes have never been sequenced, which is exactly the set you care about in an environmental sample. The defensible position is to state that abundances are gene copies, be cautious comparing across distantly related taxa, and treat within-taxon comparisons across samples — which are unaffected — as the sturdier analysis.

4 · Every primer pair has a blind spot

No 16S primer pair amplifies all bacteria equally. Mismatches to a lineage's binding site suppress it, sometimes to zero, and different regions resolve different clades. The widely used 515F/806R V4 pair, for instance, is known to under-recover some groups relative to alternatives — which is why updated variants of those primers exist.

The consequences are structural rather than fixable: your community composition is primer-conditional, comparisons across studies using different regions are not clean, and a taxon absent from your table may be absent from your primers rather than from your sample. Name the primer pair in the abstract-adjacent part of the methods, not in a supplement.

5 · ASVs versus OTUs, settled enough

Callahan, McMurdie and Holmes (ISME Journal, 2017) argued that exact amplicon sequence variants should replace OTUs in marker-gene analysis: ASV methods infer biological sequences without an arbitrary dissimilarity threshold, resolve variants differing by a single nucleotide, match or beat OTU methods on sensitivity and specificity, and — the decisive practical advantage — are reusable and comparable between studies, because an exact sequence means the same thing everywhere.

De novo OTUs are the opposite: they are defined by clustering your dataset, so an OTU from one study has no counterpart in another without re-clustering everything together. The counter-argument that survives is a narrow one — broad ecological patterns are often robust to the choice, so re-analysing an old OTU study as ASVs will not usually overturn its headline. That is a reason not to panic about the literature, not a reason to start a new study on OTUs.

6 · The inference problem

Direction is not in the data

A cross-sectional association between a taxon and a disease is compatible with the taxon causing the disease, the disease creating conditions the taxon likes, and a third factor — diet, medication, transit time, inflammation — driving both. Sequencing depth does not help. Longitudinal sampling, intervention, or gnotobiotic transfer does.

Dispersion is a result, not a nuisance

Zaneveld et al. (Nature Microbiology, 2017) named the Anna Karenina principle: stressed or dysbiotic individuals vary more in community composition than healthy ones. A significant PERMDISP is often the finding. See the animation.

Alpha diversity is not health

"Higher diversity is healthier" is a folk theorem, not a general result. It holds in some systems and inverts in others. Say what diversity measures in your system before attaching a value judgement to a Shannon index.

Relative is not absolute

Everything an ordinary 16S survey measures is a proportion. A total-load measurement — spike-in, qPCR, flow cytometry — converts it into something you can make amount-claims about, and without one you cannot. The compositional trap.

Where these come from

ConceptSource
Reagent contamination, the kitomeSalter et al., BMC Biology 12:87, 2014 — "Reagent and laboratory contamination can critically impact sequence-based microbiome analyses"
ASVs over OTUsCallahan, McMurdie & Holmes, ISME Journal 11:2639–2643, 2017
Anna Karenina principleZaneveld, McMinds & Vega Thurber, Nature Microbiology 2:17121, 2017 — "Stress and stability"
Rarefaction, againstMcMurdie & Holmes, PLOS Computational Biology, 2014 — "Waste not, want not: why rarefying microbiome data is inadmissible"
Rarefaction, reanalysedSchloss, mSphere, 2023/24 — "Waste not, want not: revisiting the analysis that called into question the practice of rarefaction"
DADA2Callahan et al., Nature Methods 13:581–583, 2016
QIIME 2Bolyen et al., Nature Biotechnology 37:852–857, 2019
MaAsLin 2Mallick et al., PLOS Computational Biology, 2021
MaAsLin 3Nickols et al., Nature Methods, 2025/26
SILVAQuast et al., Nucleic Acids Research 41:D590, 2013, and the 2026 NAR database issue update for the DSMZ era

Volume and page numbers were taken from publisher records at the time of writing. If you are citing these, pull the record yourself rather than trusting a web page — including this one.