Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Population genomics of the Wolbachia endosymbiont in Drosophila melanogaster.

PLoS Genet · 2012
L1 93/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Independently reproduced the paper's core computational result -- WGS-coverage-based Wolbachia infection classification -- for 21 of the paper's 290 sequenced D. melanogaster strains (7 from DGRP via a prior session's batch, plus 14 newly processed this session spanning both DGRP and, for the first time, DPGP), achieving 20/21 (95.2%) exact concordance with the paper's own Dataset S1 ground truth using an independently re-implemented pipeline (BWA mem 0.7.17 vs paper's 0.5.9-r16; a custom pure-Python streaming coverage/breadth calculator vs paper's SAMtools/BEDtools). The single discordant call (SRR189375) is attributable to an evidently corrupted/truncated download (442 total reads vs tens of millions for every other strain), not a biological or pipeline error; excluding it as non-evaluable yields 100% concordance on strains with usable data. The paper's cohort-level summary statistics (290 total strains = 174 DGRP + 116 DPGP; ~61.8% overall infection rate; ~99% PCR/WGS concordance) were verified exactly against the paper's own Dataset S1 supplementary table. One strain (SRS074469, nominal DGRP-306) was dropped from comparison after ENA confirmed it is a genuinely different BioSample from the one Dataset S1 actually reports under that strain name. Phylogenetic clade assignment (RAxML) and Bayesian divergence dating (BEAST) were NOT attempted -- these are substantially more compute-intensive than the infection-calling pipeline and were out of reach within this session's scope; this is an explicit scope limitation, not a failed attempt. No completeness claim is made: 21 of 290 strains (7.2%) were reprocessed, sufficient to validate the pipeline and its threshold logic but not to re-derive population-level statistics independently of Dataset S1 itself.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What is the evolutionary mode and temporal dynamics of the Wolbachia–Drosophila melanogaster symbiosis — specifically, whether Wolbachia is strictly maternally co-transmitted with mtDNA, and when the proposed global replacement of the ancestral wMelCS lineage by the derived wMel lineage actually occurred.

Core claims
  • Wolbachia infection status can be accurately predicted in silico from whole-genome shotgun sequence of individual host strains, showing 99% concordance with diagnostic PCR. method
  • Complete Wolbachia and mitochondrial genomes have fully congruent genealogies, consistent with a single ancestral infection transmitted strictly vertically through the maternal cytoplasm with no paternal or horizontal transmission. finding
  • The most recent common ancestor of all Wolbachia and mitochondrial genomes in D. melanogaster dates to around 8,000 years ago, so the derived wMel lineage arose several thousand years ago rather than in the 20th century as previously proposed. finding
  • There has been a recent but incomplete global replacement of ancestral Wolbachia and mtDNA lineages, likely one of several similar incomplete replacement events since the out-of-Africa migration, leaving remnant lineages in North America, Europe, and Africa. finding
  • Wolbachia has been lost independently in many populations through imperfect maternal transmission, as infected and uninfected strains occur across the entire mtDNA genealogy. mechanism
  • Strict maternal co-inheritance allows the mtDNA mutation rate to calibrate Wolbachia sequence evolution, which is ~100-fold lower than the mtDNA synonymous mutation rate and ~10-fold lower than the host nuclear noncoding mutation rate. finding
  • Patterns of Wolbachia and mtDNA variation in the well-sampled North American population are inconsistent with a standard neutral model, suggesting natural selection or host population expansion in the recent past. finding
  • Reconstructed complete Wolbachia and mitochondrial genome sequences plus infection status and relative copy-number estimates for 290 D. melanogaster lines are provided as a community resource. resource
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome shotgun resequencing with reference-based short-read mapping/consensus assembly 290 D. melanogaster strains (174 DGRP from Raleigh, North Carolina; 108 African and 8 European DPGP strains) none Consensus Wolbachia and mitochondrial genome sequences; depth and breadth of coverage used to call infection status (depth >1, breadth >90%)
Diagnostic PCR amplification of the Wolbachia wsp gene 167 of the 174 DGRP D. melanogaster lines none Presence/absence of Wolbachia infection (validation of in silico calls)
Coverage-based relative copy number estimation DGRP strains (diploid adult DNA) and DPGP strains (haploid embryonic DNA) none Mean depth of coverage of Wolbachia and mtDNA assemblies normalized to a size-matched nuclear locus on chromosome 3L
Maximum likelihood phylogenetic reconstruction with bootstrapping (RAxML) Complete mitochondrial genomes from all 290 D. melanogaster strains plus dm3 and NC_001709 references none mtDNA genealogy (six major clades I–VI) from an ungapped alignment of 292 sequences × 12,225 bp; branches with >85% bootstrap support RAxML
Phylogenetic congruence analysis of Wolbachia versus mtDNA genealogies 179 Wolbachia-infected D. melanogaster strains none Topological congruence between complete Wolbachia and mtDNA trees
Bayesian phylogenetic dating / molecular clock analysis Complete Wolbachia and D. melanogaster mitochondrial genomes none Divergence dates of major clades and Wolbachia substitution rate calibrated against the mtDNA mutation rate
Population genetic differentiation test (variant of Kst) mtDNA sequences of Wolbachia-infected versus uninfected D. melanogaster strains across sampling locations none Weighted mean Kst and associated P value for mtDNA subdivision by infection status
Tests of the standard neutral model of molecular evolution Wolbachia and mtDNA genomes from the North American (DGRP, Raleigh) population none Departure of molecular variation from neutral expectations
Key results
  • 179 of 290 strains (61.8%) were predicted to be Wolbachia-infected by in silico criteria (reported as 63% overall in the abstract). 61.8% (95% CI 0.56–0.68)
  • In silico infection calls matched PCR-based calls for 165/167 DGRP lines; the two discordant lines (DGRP38, DGRP911) were in silico uninfected but PCR positive, likely from cross-contamination after library prep. 98.8% concordance
  • Infection frequency differs significantly between the North American DGRP sample and the African/European DPGP sample. 52.2% (91/174) vs 75.9% (88/116)
  • The mtDNA genealogy of 290 strains resolves six major intraspecific clades (I–VI), with infected and uninfected strains distributed across the entire tree and present in all major clades except the two-strain clade V. 6 clades
  • No genetic differentiation was detected between mtDNA of infected and uninfected strains, contrary to the earlier COI-based report of Nunes et al., suggesting the infection is not recent and no pre-infection mitochondrial lineages are present in the sample. Kst = 0.01, P = 0.88
  • Relative depth of coverage for both Wolbachia and mtDNA is significantly higher in DGRP (diploid adult DNA) than DPGP (haploid embryonic DNA) strains; the same holds for mtDNA in uninfected strains, so it is not an artefact of infection. Wolbachia 5.57× (SD 3.95) vs 1.02× (SD 1.84) nuclear coverage; mtDNA 32.9× (SD 44.5) vs 9.79× (SD 24.7)
  • Wolbachia sequence evolution is roughly 100-fold slower than the mitochondrial synonymous mutation rate and ten-fold slower than the host nuclear noncoding mutation rate. ~100-fold and ~10-fold lower
  • The MRCA of all cytoplasmic (Wolbachia and mtDNA) lineages in D. melanogaster dates to approximately 8,000 years ago. ~8,000 years
Key statistics
  • count 290 (D. melanogaster strains analyzed (174 DGRP, 108 African + 8 European DPGP))
  • count 179 strains (61.8%); 95% CI 0.56–0.68 (Strains predicted Wolbachia-infected in silico)
  • pvalue P<1.4×10−7 (Binomial test comparing infection proportion in DGRP (52.2%) vs DPGP (75.9%))
  • other 165/167 (98.8% concordance) (Agreement between in silico and wsp PCR infection status in DGRP lines)
  • correlation Kst = 0.01, P = 0.88 (Weighted mean mtDNA differentiation between infected and uninfected strains across sampling locations)
  • pvalue P<2.2×10−16 (infected strains); P<7.4×10−9 (mtDNA in uninfected strains) (Wilcoxon Rank Sum Tests of DGRP vs DPGP relative coverage)
  • fold_change Wolbachia 5.57 (SD 3.95)× and mtDNA 32.9 (SD 44.5)× nuclear coverage in DGRP; Wolbachia 1.02 (SD 1.84)× and mtDNA 9.79 (SD 24.7)× in DPGP (Relative cytoplasmic genome copy number in infected strains)
  • count 14,792,213,317 reads totalling 1,192,414,581,932 bp; mtDNA alignment of 292 sequences × 12,225 bp (Sequencing volume analyzed and mtDNA phylogeny input alignment)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper uses population-genomic comparisons across two independently sampled strain panels (DGRP and DPGP) of Drosophila melanogaster to assess Wolbachia infection status, relative cytoplasmic genome copy number, and mitochondrial/Wolbachia phylogenetic congruence. Group comparisons (infection proportions, relative sequencing coverage) are evaluated with a binomial test and Wilcoxon rank-sum tests, genetic differentiation is assessed with a Kst-based test, and genealogical relationships are reconstructed with maximum-likelihood (RAxML, bootstrap-supported) and Bayesian phylogenetic methods. Results are reported primarily as means with standard deviations, percentages with 95% confidence intervals, and p-values (exact or as upper bounds).

Replicationunclear Sample sizeSample sizes given as counts of sequenced strains (290 total: 174 DGRP, 116 DPGP); no formal power analysis described GroupsDGRP vs. DPGP strain panels; Wolbachia-infected vs. uninfected strains Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Binomial test Comparison of Wolbachia infection proportion between DGRP (52.2%) and DPGP (75.9%) samples 174 DGRP strains, 116 DPGP strains not stated
Wilcoxon Rank Sum test Relative depth of coverage (Wolbachia and mtDNA) compared between DGRP and DPGP infected strains 179 infected strains (91 DGRP, 88 DPGP) not stated
Wilcoxon Rank Sum test Relative mtDNA coverage compared between DGRP and DPGP strains not infected with Wolbachia not explicitly restated for this subset not stated
Kst (population differentiation statistic, permutation-based) Genetic differentiation between mtDNA sequences of Wolbachia-infected vs. uninfected strains 290 strains (179 infected, 111 uninfected) not stated
Maximum-likelihood phylogenetic reconstruction (RAxML) with bootstrap support Genealogy of the D. melanogaster mtDNA across all sampled strains (Figure 3) 292 sequences (290 strains plus 2 reference sequences), 12,225 bp alignment not stated
Bayesian phylogenetic analysis Dating of the most recent common ancestor and estimation of Wolbachia/mtDNA evolutionary rates not stated
Approaches that could also have been used
  • The difference in Wolbachia infection proportion between the DGRP and DPGP samples was assessed with a binomial test
    Could also: A Fisher's exact test or chi-square test of independence on the 2x2 infection-status table — These are standard approaches for comparing two independent proportions and would also allow the comparison to be framed explicitly as a contingency-table analysis alongside the reported confidence intervals
  • Relative sequencing coverage of Wolbachia and mtDNA was compared between sample types using the non-parametric Wilcoxon Rank Sum test
    Could also: A parametric test such as Welch's t-test on (possibly transformed) coverage values, or reporting a rank-based effect size (e.g., rank-biserial correlation) alongside the p-value — A parametric test can be used when distributional assumptions are reasonably met, and pairing a significance test with an effect-size estimate would additionally convey the magnitude, not just the direction, of the coverage difference
  • Genetic differentiation between mtDNA from infected and uninfected strains was tested with a Kst-based permutation test
    Could also: An AMOVA (Analysis of Molecular Variance) or Fst-based differentiation statistic — These are widely used alternative frameworks for partitioning molecular variance between groups and can provide a directly comparable measure of population/genetic structure
  • Multiple independent statistical tests (binomial, two Wilcoxon tests, Kst) are reported across the manuscript without a stated joint correction
    Could also: A multiple-testing correction such as Bonferroni or Benjamini-Hochberg FDR applied across the set of comparisons — Such a correction would control the family-wise error rate or false discovery rate if the tests are considered jointly as a family of related hypotheses, which is one convention when multiple comparisons are drawn from the same dataset
  • Support for the mtDNA genealogy was assessed using maximum-likelihood bootstrap values (RAxML)
    Could also: Bayesian posterior probabilities, as already used elsewhere in the study for dating analyses — Using the same Bayesian framework for both topology support and divergence dating would give a directly comparable, single-framework measure of confidence across the phylogenetic analyses
  • Coverage results are summarized with the mean and standard deviation
    Could also: Reporting a 95% confidence interval or interquartile range alongside or instead of the SD — A CI or IQR can be more informative for skewed distributions (as coverage often is) and directly conveys the precision of the estimated mean or the spread of the median
Software: RAxML

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

c1_cohort_composition
Reported
290 total (174 DGRP, 116 DPGP)
Reproduced
290 total (174 DGRP, 116 DPGP) -- verified directly from paper's own Dataset S1 supplementary table
exact
c2_infection_rate
Reported
~61.8% infected (paper text)
Reproduced
179/290 = 61.7% infected (tallied from Dataset S1 WGS_infection_status column)
within tolerance
c3_pcr_wgs_concordance
Reported
~99% concordance (paper text)
Reproduced
165/167 = 98.8% concordant (tallied from Dataset S1 PCR_infection_status vs WGS_infection_status columns, restricted to rows with both calls present)
exact
c4_individual_infection_calls
Reported
Per-strain infected/uninfected calls as listed in Dataset S1
Reproduced
20/21 = 95.2% concordant across 21 independently reprocessed strains (7 from an original DGRP-only batch + 14 from an extended DGRP+DPGP batch this session); 21/21 = 100.0% if the one strain with an evidently corrupted/truncated download (SRR189375, 442 total reads) is instead treated as non-evaluable rather than a true mismatch
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

Numerically this reproduces cleanly. The cohort composition (290 = 174 DGRP + 116 DPGP), the infection rate (179/290 = 61.7% vs the paper's ~61.8%) and PCR/WGS concordance (165/167 = 98.8% vs ~99%) all match within rounding, and the independent raw-data arm — ENA FASTQ → BWA mem 0.7.17 → breadth/depth classification against AE017196 — agreed with Dataset S1 on 20/21 strains, with a ~2-order-of-magnitude separation between infected (10.1–678.4x depth, breadth ≥0.986) and uninfected (≤0.16x depth, breadth ≤0.026) strains. The one mismatch is on our/infrastructure side, not the authors': SRR189375 downloaded as 442 reads total, and SRR933577 as 0 bytes; a further strain (DGRP-306) was dropped because SRS074469 and SRS003467 are genuinely distinct BioSamples. The limitation is coverage and circularity, not discrepancy: three of four claims are tallies of the paper's own supplementary table, and the genuinely independent claim covers 7.2% of the cohort — hence q8 yellow despite zero substantive disagreement with the published values. Nothing here is fabrication-suspect; the authors' data availability (public SRA/ENA plus a complete per-strain supplementary table) is in fact unusually good for a 2012 paper.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.