A method for selectively enriching microbial DNA from contaminating vertebrate host DNA.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (partial-strong), independently re-confirmed. The in-scope pipeline result (Fig 3: E. coli/human defined-mixture MBD2-Fc enrichment) was rerun end-to-end on «our HPC» SLURM «job» (node n164, 27:02, exit 0): all 12 PRJNA208794 Ion Torrent runs length-filtered >=50bp (seqkit) and aligned bowtie2 --sensitive --end-to-end to a combined hg19 (UCSC) + E. coli MG1655 (ENA U00096.3) index, then classified by reference contig via samtools idxstats. The central finding is unambiguously reproduced: input mixtures 1.6-12.4% E. coli rise to 47.5-86.9% E. coli after enrichment (7-30x fold), with host DNA depleted into the bead-bound fraction (99%+ human) - exactly the MBD2-Fc mechanism. Per-ratio percentages mostly land inside the paper's reported bands: A1 within-tol, A4 exact; A2/A3 graded partial only because the extreme 80:1 mixture (which starts with the least E. coli, 1.6%) reaches a lower enrichment ceiling (47.5%) than the 65-85% band. This re-run reproduced the earlier «job» numbers bit-for-bit, confirming the pipeline is deterministic. Deviations: bowtie2 2.5.4 vs paper's 2.0.4 (presets stable across versions); E. coli ref pinned to U00096.3/NC_000913.3 vs paper's unversioned 'MG1655'. Out-of-scope wet-lab + legacy/web-server paths not attempted (see scope.md).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-21 ⛓ 1fa30a0205ce
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors test whether a methyl-CpG binding domain fused to human IgG Fc (MBD-Fc), bound to Protein A paramagnetic beads, can selectively enrich microbial DNA from vertebrate host DNA by exploiting differences in CpG methylation density.
- ★ MBD-Fc bound to Protein A paramagnetic beads selectively binds methylated (vertebrate host) DNA, depleting it from mixed samples method
- ★ Enrichment substantially decreases host-derived sequencing reads while increasing microbial (bacterial/Plasmodium) reads finding
- ★ Efficient MBD-Fc binding requires a DNA methylation density threshold of roughly 2-3 methyl-CpG per kilobase finding
- ★ MBD-Fc enrichment preserves relative microbial species abundance and diversity compared to unenriched samples finding
- ★ MBD-Fc enrichment increases recovery of Plasmodium falciparum DNA from mock human-parasite mixtures without introducing coverage or GC bias finding
- ★ The enrichment method is broadly applicable across host types, including human saliva, human blood, and black molly fish finding
- Enrichment enables detection of low-abundance microbes and bacterial drug-resistance genes not detected in unenriched samples finding
- Bacterial genomes generally lack sufficient CpG methylation density to bind MBD-Fc, explaining the specificity of depletion mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Scintillation counting and gel densitometry | 3H-labeled E. coli K12 DNA mixed with IMR-90, HeLa, or NIH 3T3 mammalian DNA | MBD-Fc/Protein A bead enrichment | percent DNA in bound vs unbound (supernatant) fractions | scintillation counter; gel densitometry |
| Whole genome shotgun sequencing | IMR-90 human DNA mixed with E. coli DNA | MBD-Fc enrichment (unenriched, enriched supernatant, bead-bound fractions) | percent reads mapping to E. coli vs human reference genome | Ion Torrent PGM; Bowtie 2.0.4 alignment |
| Restriction digestion, in vitro methylation, gel electrophoresis/densitometry | T7 and lambda phage DNA fragments methylated with M.HhaI and/or M.HpaII | titrated methyl-CpG density | percent DNA bound to MBD-Fc beads vs methyl-CpG density | polyacrylamide gel electrophoresis |
| Whole genome shotgun sequencing / metagenomics | Human saliva DNA | MBD-Fc enrichment (unenriched, enriched, bead-bound) | reads mapping to HOMD and PhageSeed databases; species abundance via MetaPhlAn; Shannon-Wiener diversity | SOLiD 4 |
| Whole genome shotgun sequencing / metagenomics | Human blood DNA (commercial samples) | MBD-Fc enrichment | reads mapping to HOMD database; genus-level abundance | SOLiD 4 |
| Whole genome shotgun sequencing | Mock 90% human / 10% Plasmodium falciparum blood DNA mixture | MBD-Fc enrichment | percent reads mapping to P. falciparum genome; genome coverage evenness; GC content bias | Illumina HiSeq 2000 and MiSeq |
| Whole genome shotgun sequencing / metagenomic assembly and annotation | Whole black molly fish (Poecilia cf. sphenops) DNA | MBD-Fc enrichment | microbial genus abundance concordance; de novo assembly contigs; drug-resistance gene annotation | Illumina GAIIx and MiSeq; MG-RAST; CLC de novo assembler; MetaPhlAn |
- – Host genome-mapped reads decreased 50-fold after enrichment while bacterial and Plasmodium reads increased 8-11.5-fold 50-fold decrease (host); 8-11.5-fold increase (microbe)
- ▼ With 40 uL MBD-Fc beads, only 1-4% of mammalian DNA remained in the supernatant while 84-100% of E. coli DNA remained unbound 1-4% mammalian DNA remaining; 84-100% E. coli DNA remaining
- – After enrichment of human-E. coli mixtures, 65-85% of Ion Torrent reads mapped to E. coli vs only 15-35% to human, while bead-bound DNA was 97-99% human 65-85% E. coli reads (enriched) vs 97-99% human reads (bound)
- – Efficient bead binding occurred only above a threshold of 2-3 methyl-CpG per kilobase
- ▲ 94-96% of human-mapped reads in saliva were depleted after enrichment, corresponding to an 8-fold increase in reads mapping to HOMD 8-fold increase; 94-96% depletion
- – Shannon-Wiener diversity index for 147 saliva bacterial species was nearly identical before (4.72) and after (4.80) enrichment H'=4.72 vs 4.80
- ▲ Plasmodium falciparum reads increased 8-fold after enrichment, with even genome coverage and no GC bias relative to unenriched sample (60% zero-coverage) 8-fold increase
- ▲ In black molly fish, enrichment yielded 198 contigs (>10-fold coverage) matching bacterial drug-resistance genes, including QnrS2 at 49x coverage 49x coverage; 198 contigs
- fold_change 50-fold decrease in host reads; 8-11.5-fold increase in microbial reads (genome mapping across enrichment experiments)
- other Shannon-Wiener H'=4.72 (unenriched) vs H'=4.80 (enriched) (saliva bacterial species diversity)
- count 1-4% mammalian DNA remaining in supernatant; 84-100% E. coli DNA remaining in supernatant (40 uL beads) (mixed cell line/E. coli depletion experiment)
- fold_change 8-fold increase in reads mapping to HOMD database (saliva enrichment)
- count 94-96% of human reads depleted after enrichment (saliva sample)
- count 65-85% of reads mapped to E. coli reference after enrichment; 15-35% mapped to human (Ion Torrent IMR-90/E. coli mixtures)
- fold_change 8-fold increase in Plasmodium-mapped reads after enrichment (mock human-Plasmodium mixture)
- other methylation density threshold of 2-3 methyl-CpG per kilobase for efficient binding (T7/lambda methylation titration experiment)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This proof-of-concept methods paper evaluated MBD-Fc bead-based enrichment of microbial DNA from vertebrate host DNA across four sample types: defined cell-line mixtures, human saliva and blood, a mock malaria-infected blood sample, and black molly fish. Enrichment performance was assessed primarily by comparing the percentage of sequencing reads mapping to host vs. microbial reference genomes in unenriched, enriched (supernatant), and bead-bound fractions. Microbial community composition was compared between enriched and unenriched samples using Shannon-Wiener diversity indices and visual concordance plots of relative species abundances; no formal inferential hypothesis tests were reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Shannon-Wiener diversity index (descriptive comparison of point estimates, no formal significance test) | Comparison of bacterial species diversity in human saliva before vs. after MBD-Fc enrichment | 147 bacterial species observed (body text; abstract states 149); number of sequencing reads underlying this calculation not specified | not stated |
| Percentage of sequencing reads mapped to reference genome (descriptive proportion comparison) | Host vs. microbial read proportions before and after enrichment across all sequencing experiments (Figures 3, 5, 8, 9) | 1–2 million reads (Ion Torrent PGM); 174–346 million reads per blood sample and 501–537 million reads per saliva sample (SOLiD 4); Illumina read counts not specified | not stated |
| Gel densitometry (quantitative image analysis of DNA band intensities) | Quantification of mammalian vs. E. coli DNA fractions in bound/unbound fractions and methylation-density binding experiments (Figures 2, 4) | 250–500 ng input DNA per condition; number of independent experimental replicates not stated | not stated |
| 3H scintillation counting | Quantification of radiolabeled E. coli DNA recovery in supernatant fraction of MBD-Fc enrichment (Figure 2B) | 500 ng input DNA per condition; number of replicates not stated | not stated |
| Concordance plot (visual comparison of relative microbial abundances) | Comparison of microbial genus/species relative abundances between enriched and unenriched saliva, blood, and fish samples (Figures 6, 9) | — | not stated |
-
Microbial diversity preservation was assessed by reporting two Shannon-Wiener index point estimates (H′ = 4.72 unenriched vs. H′ = 4.80 enriched) without quantifying uncertainty around each estimate↳ Could also: Bootstrap confidence intervals around each H′ estimate, or a permutation-based test of the difference in diversity indices (e.g., via the vegan R package), could also be applied; rarefaction to equal sequencing depth before computing H′ is another standard step — CIs around diversity estimates would allow readers to assess whether the 0.08-unit difference is within sampling variation, and rarefaction ensures that differences in sequencing depth do not confound the diversity comparison
-
Concordance between enriched and unenriched community profiles was presented as a visual scatter/concordance plot with no summary statistic↳ Could also: A Spearman or Pearson correlation coefficient with a 95% CI on log-transformed relative abundances could also be reported alongside the concordance plot — A numeric correlation coefficient provides a reproducible, comparable summary of concordance that is less dependent on visual interpretation and allows quantitative comparison across different sample types or methods
-
Enrichment performance was reported as ranges of fold-change or percentage-point differences (e.g., 84–100% E. coli recovery, 8–11.5-fold microbial increase) without a measure of central tendency or spread across replicates↳ Could also: Reporting mean ± SD (or individual data points with a mean line) across independent enrichment replicates could also convey reproducibility — Summary statistics with dispersion measures allow readers to gauge the consistency of the method across runs; ranges alone conflate variability across conditions (e.g., different input ratios) with variability across replicate experiments
-
Relative microbial abundance data (proportions of reads summing to 1) were compared directly as ordinary proportions across conditions↳ Could also: Compositional data analysis approaches such as Aitchison log-ratio (ALR or CLR) transformations could also be applied before comparing abundances across enriched and unenriched libraries — Sequencing read proportions are inherently compositional (constrained to sum to 1), and log-ratio methods address the spurious correlations that can arise when directly comparing proportions sharing a common denominator
-
Diversity indices and relative-abundance comparisons were made between enriched and unenriched libraries without subsampling reads to equal sequencing depth↳ Could also: Rarefaction (random subsampling) to the minimum library size across compared samples could also be applied before computing diversity and abundance metrics — When library sizes differ substantially across conditions, rarefaction helps ensure that apparent differences in species detection or diversity reflect biology rather than differences in sequencing depth
-
The binding-threshold experiment (Figure 4B) displayed replicate measurements as overlaid individual data points without reporting means or a measure of spread↳ Could also: Mean ± SD or a fitted sigmoid/logistic curve with confidence band across the methylation-density axis could also be shown to summarize the threshold relationship — A fitted curve with uncertainty band would make the estimated ~2–3 methyl-CpG/kb threshold more precise and allow readers to assess its reproducibility across the replicates already collected
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 24204593
Paper: Feehery GR et al. "A method for selectively enriching microbial DNA from contaminating vertebrate host DNA." PLoS ONE 2013;8(10):e76096. DOI 10.1371/journal.pone.0076096 · PMCID PMC3810253.
Nature of paper: Primarily a wet-lab method paper. It introduces an MBD2-Fc (methyl-CpG binding domain) bead-based depletion of CpG-methylated vertebrate host DNA to enrich unmethylated microbial DNA (the chemistry that became the NEB NEBNext Microbiome DNA Enrichment Kit). The quantitative read-out of enrichment, however, is sequencing-based and therefore has a bioinformatic pipeline behind it.
In scope (pipeline-derived, attempted)
RESULT A — E. coli / human (IMR-90) defined-mixture enrichment (Figure 3). This is the cleanest, fully-public, low-cost pipeline result.
- Data: SRA BioProject PRJNA208794, 12 Ion Torrent PGM runs = 4 host:microbe ratios (10:1, 20:1, 40:1, 80:1) × 3 fractions (input/unenriched, enriched supernatant, bound-to-beads). All public, ~0.7–2.3 M reads each.
- Pipeline (verbatim from Methods): "Fastq files were processed to remove reads shorter than 50 base pairs and mapped to a combined reference sequence database containing both the hg19 human reference genome and the E. coli MG1655 genome. Reads were mapped to this reference sequence using Bowtie 2.0.4 using the sensitive, end-to-end options."
- Reproduction recipe: download the 12 runs; drop reads <50 bp; build a
combined
hg19 + E. coli K-12 MG1655 (NC_000913)Bowtie2 index; align each run withbowtie2 --sensitive --end-to-end; classify each aligned read as E. coli vs human by reference contig; report % reads → E. coli vs human per run. - Compared against: Figure 3 / Results text (see
original/claims.tsv): input mixtures 2.5–10% E. coli; after enrichment 65–85% E. coli & 15–35% human.
Out of scope (not attempted) — and why
- MBD2-Fc bead chemistry, ³H-scintillation binding assays, library prep — wet-lab, no pipeline.
- Human blood & saliva (SOLiD 4, PRJNA208062 / PRJNA208064): legacy
color-space; mapped with Bowtie 0.12.7 to HOMD + PhageSeed + hg19. HOMD/PhageSeed
versions from 2013 are not pinned in the paper and the SOLiD color-space toolchain
is effectively unmaintained →
env_unresolvable/docs_insufficientfor a faithful redo. Skipped (hard 20%). - Black molly fish (PRJNA203390): this is the only path that actually uses the linked repo sickle (quality trimming), but downstream analysis is MG-RAST + MetaPhlAn with unpinned DB versions and a web server → not reproducibly pinnable. Skipped.
- Plasmodium falciparum enrichment (Fig 8): Illumina; "8-fold increase" qualitative; reference/version detail thin. Skipped (lower priority than A).
- Saliva Shannon-Wiener diversity (H′=4.72→4.80): depends on the SOLiD/HOMD path above. Skipped.
Note on the linked "Code" artifact
The brief lists github.com/vsbuffalo/sickle as the code. Sickle is a generic
windowed adaptive quality-trimmer; per the Methods it was applied only to the
black molly Illumina reads, not to the in-scope Figure-3 Ion Torrent data
(which used only a length filter + Bowtie2). Per brief rule P16 we reproduce the
described pipeline on the paper's own public data regardless; Result A is the
faithful, low-ambiguity target.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.