Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A Meta-Analysis of Wolbachia Transcriptomics Reveals a Stage-Specific Wolbachia Transcriptional Response Shared Across Different Hosts.

G3 (Bethesda) · 2020
L1 89/100 PQI 96
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described WELL ENOUGH to reproduce 1:1. The paper re-analyses 7 published Wolbachia RNA-seq datasets through one common pipeline (SRA->HISAT2->seqtk->FADU->edgeR->FactoMineR PCA/WGCNA); the authors deposited the FADU count matrices + full R code (Zenodo 3726272 / matt-chung repo @6fcdd77). I re-ran their OWN edgeR (v3.24.3, pinned to repo sessionInfo) + PCA code on the deposited counts on «our HPC». Result: EXACT reproduction of all headline pipeline-derived numbers I could pin -- 7/7 per-study gene totals; 7/7 Table-1 reads-to-CDS ranges; CPM-filter kept-gene counts for 6/6 reproducible studies (560/867/18/781/425/548); DE-gene counts (FDR<0.05) for 5/5 reproducible studies (0/120/473/94/373); both reported PCA PC1 variances (44.5%, 70.9%). ONE result is NOT reproducible and is flagged as a possible deposit/integrity error: Luck 2015 (wDi) reported 325 kept / 36 DE, but the deposited Luck2015 count matrix has only 0-349 total reads/sample (consistent with its own Table 1), so the authors' documented depth filter ('exclude samples with <2000 counts') removes ALL 12 samples -> 325/36 cannot be derived from the deposited data+code. Practical notes for auditors: supplementary sheet NAMES are scrambled (mapped by content); study_sample_map.tsv.txt was not deposited (groups reconstructed from column headers, validated against embedded expected outputs); the deposited xlsx headers are shifted one column (values were always correct). NOT attempted: Table-1 'Wolbachia Reads Mapped' upstream raw-read alignment (SRR lists not deposited; Darby2012 wOo is SOLiD color-space incompatible with the HISAT2 base-space pipeline; high-divergence and superseded by the deposited counts), plus WGCNA/GO-enrichment/core-gene-overlap (out of scope).

💻 Code ↗ 🗄 Data: 10.5281/zenodo.3726272

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 89
    assessed: 2026-06-17 ⛓ be79d9590c7b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-17
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Do prior Wolbachia RNA-Seq transcriptomics studies, when re-analyzed with a single unified up-to-date pipeline, reveal conserved or conflicting patterns of differential gene expression across diverse Wolbachia strains and hosts?

Core claims
  • Across datasets re-analyzed with a unified workflow, there is a general lack of global Wolbachia gene regulation. finding
  • A weak, conserved transcriptional response upregulating ribosomal proteins occurs in early larval stages across diverse Wolbachia strains from both nematode and insect hosts, suggesting a pan-Wolbachia response during host development. finding
  • Data from an early filarial nematode study (Luck et al. 2014) did not support robust conclusions about Wolbachia differential expression, so original interpretations should be reconsidered. finding
  • A unified RNA-Seq meta-analysis pipeline using current best-practice algorithms can reconcile discordant results from seven prior Wolbachia transcriptomics studies. method
  • A two-step mapping approach (splice-aware mapping to combined host+Wolbachia reference, then non-spliced re-mapping of Wolbachia reads) controls for lateral gene transfer reads. method
  • Re-analysis identified different numbers of differentially expressed genes than original studies (e.g., 0 vs 26 for Darby 2012; 473 vs 80 for Gutzwiller 2015). finding
  • Core Wolbachia genes shared across wBm, wDi, wMel, and wOo genomes were identified using PanOCT for cross-study comparison. resource
Experimental setups
Assay System Perturbation Readout Platform
RNA-Seq (re-analysis) Wolbachia endosymbiont wOo / Onchocerca ochengi (adult male and female gonad) none (sex/tissue comparison) differential gene expression / TPM SOLiD v4 (original), HISAT2/FADU/edgeR (re-analysis)
RNA-Seq (re-analysis) Wolbachia endosymbiont wDi / Dirofilaria immitis (L3, L4, adult male, adult female, microfilariae life stages) none (life-cycle stage comparison) differential gene expression / TPM Illumina GAIIx 50 bp single-end (original)
RNA-Seq (re-analysis) Wolbachia endosymbiont wDi / Dirofilaria immitis (different body tissues) none (tissue comparison) differential gene expression / TPM
RNA-Seq (re-analysis) Wolbachia endosymbiont wMelPop-CLA / Aedes albopictus RML-12 cell line drug (doxycycline treatment) differential gene expression / TPM
RNA-Seq (re-analysis) Wolbachia endosymbiont wMel / Drosophila melanogaster (embryo, larvae, pupae, adult male, adult female) none (life-cycle stage comparison) differential gene expression / TPM
RNA-Seq (re-analysis) Wolbachia endosymbiont wBm / Brugia malayi (life-cycle stages) none (life-cycle stage comparison) differential gene expression / TPM
Comparative genomics / ortholog identification (PanOCT) wBm, wDi, wMel, wOo genomes none core/shared Wolbachia genes PanOCT v3.23
Co-expression clustering (WGCNA) and functional enrichment Differentially expressed Wolbachia genes from each dataset none expression modules, over-represented GO/InterPro terms WGCNA v1.6.6; InterProScan v5.34-73
Key results
  • No wOo genes were significantly differentially expressed between adult male and female O. ochengi in re-analysis (vs 26 in original). 0 DE genes (orig. 26)
  • PCA of wOo separated male and female samples on PC1, but inter-replicate variation among males was nearly as large as inter-sample variation. PC1=44.5%, PC2=36.7%
  • Luck et al. 2014 wDi re-analysis yielded very low counts and most samples failed saturation, producing artifactual expression patterns. 495-27,118 reads; 3/5 samples not saturated
  • Re-analysis of Gutzwiller 2015 wMel identified 473 DE genes (vs 80 original), with 325 core DE genes. 473 DE genes (orig. 80)
  • Re-analysis of Darby 2014 wMelPop-CLA doxycycline study identified 120 DE genes (vs 78 original), 72 core. 120 DE genes (orig. 78)
  • Re-analysis of Chung 2019 wBm identified 373 DE genes (vs 318 original), 338 core. 373 DE genes (orig. 318)
  • Re-analysis of Grote 2017 wBm identified 94 DE genes (vs 62 original), 83 core. 94 DE genes (orig. 62)
  • Ribosomal protein genes are upregulated in early larval stages across diverse Wolbachia strains, consistent with a pan-Wolbachia response.
Key statistics
  • other FDR < 0.05 (DE significance threshold) (edgeR quasi-likelihood differential expression)
  • other PC1 = 44.5%, PC2 = 36.7% (PCA of wOo (Darby 2012) male vs female)
  • count 495-27,118 Wolbachia reads mapped to CDS (Luck 2014 wDi low coverage)
  • count ~96% of wOo genes had similar expression between samples (original Darby 2012 finding)
  • fold_change 473 DE genes (325 core) in re-analysis vs 80 original (Gutzwiller 2015 wMel)
  • count seven major published Wolbachia transcriptome studies re-analyzed (scope of meta-analysis)
  • other divergence >700 million years ago (between nematode and insect Wolbachia hosts)
  • other CPM ≥ 5 minimum expression filter (edgeR gene filtering criterion)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper re-analyzed seven published Wolbachia RNA-Seq transcriptomics datasets using a unified computational pipeline, applying a two-step splice-aware/non-splice-aware alignment strategy followed by read quantification in coding sequences. Differential expression was assessed independently for each dataset with edgeR's quasi-likelihood F-test (FDR < 0.05), and differentially expressed genes were clustered by expression pattern using WGCNA. Cross-dataset patterns were examined descriptively via PCA, hierarchical clustering, and UpSet analysis of core DE genes, with functional enrichment assessed by Fisher's exact test (FDR < 0.05).

Replicationbiological Sample sizeSample sizes per dataset described in Table 1 and per-study text; no a priori power analysis reported, as this is a re-analysis of pre-existing data GroupsLife cycle stages (e.g., L3, L4, adult, microfilariae), biological sex (male vs. female), tissue types, and antibiotic treatment vs. control — varying by dataset Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFDR correction (Benjamini-Hochberg implied by edgeR and standard R defaults; method name not explicitly stated in text)
Statistical tests used
Test Applied to n Assumptions
edgeR quasi-likelihood F-test (GLM QLF) Differential expression analysis in all re-analyzed datasets (Darby 2012, Darby 2014, Luck 2015, Gutzwiller 2015, Grote 2017, Chung 2019) Varies by dataset: 4, 6, 6, 55, 14, and 32 samples respectively (Table 1; Luck 2014 excluded due to insufficient depth) not stated
Fisher's exact test Over-representation of GO terms and InterPro descriptions in WGCNA expression modules, DE gene subsets, and UpSet categories not stated
Principal component analysis (PCA) Exploratory visualization of sample-level transcriptional variation within each re-analyzed dataset na
Hierarchical clustering with bootstrap-based p-values (pvclust) Dendrograms of sample relationships using z-score of log2 TPM values not stated
Rarefaction analysis (vegan) Assessment of sequencing depth saturation per sample across all datasets prior to DE analysis na
Approaches that could also have been used
  • Each dataset was re-analyzed independently with a unified pipeline and results were compared descriptively by overlapping DE gene lists
    Could also: A formal quantitative meta-analysis framework — such as random-effects meta-analysis of log2 fold-change estimates and their standard errors across datasets — could also synthesize findings — Formal effect-size meta-analysis would produce pooled estimates with confidence intervals and heterogeneity statistics, enabling statistical statements about cross-dataset consistency rather than relying on qualitative overlap of thresholded gene lists
  • Differential expression was assessed using edgeR with a quasi-likelihood F-test
    Could also: DESeq2 with Wald or likelihood ratio tests could also be applied to the same count matrices — The paper itself cites DESeq2 as a current best-practice tool alongside edgeR; applying both and reporting concordance of DE calls would further characterize the robustness of findings to the choice of DE tool
  • Functional enrichment was tested with Fisher's exact test applied to the discrete set of genes passing FDR < 0.05
    Could also: A rank-based gene set enrichment approach (e.g., GSEA or fgsea) using the full gene ranking by fold-change or test statistic could also be applied — Rank-based methods do not require a significance threshold to define the gene set and use signal from all tested genes, which can improve sensitivity for detecting coordinated but modest expression shifts — particularly relevant given the weak overall transcriptional responses observed
  • Expression values were normalized as TPM for visualization, PCA, and clustering inputs
    Could also: TMM-normalized CPM (edgeR's native normalization) or DESeq2's variance-stabilizing transformation (VST) could also be used for exploratory analysis — TMM-CPM and VST are designed to stabilize variance across the expression range and are the normalizations matched to the underlying DE statistical models; they may better reflect the assumptions of the DE analysis used downstream
  • Cross-dataset comparison of DE genes used UpSet plots and qualitative visual overlap at a uniform FDR threshold
    Could also: Rank aggregation methods (e.g., RRA — robust rank aggregation) could also identify genes consistently DE across datasets without requiring a fixed significance cutoff in each individual study — Rank aggregation is robust to between-study differences in statistical power and sample size, and can surface genes trending DE across datasets even when they do not individually reach threshold — addressing the power variability visible in Table 1's wide range of DE gene counts
  • WGCNA was used to cluster differentially expressed genes into co-expression modules using a topological overlap matrix
    Could also: Hierarchical clustering, k-means, or non-negative matrix factorization (NMF) could also group genes by expression pattern — WGCNA is well-suited to large gene networks; for the relatively small DE gene sets reported in most individual datasets here, simpler clustering approaches may produce equally interpretable groupings with fewer parameters requiring tuning
  • The Luck et al. 2014 dataset was excluded from DE analysis after the re-analysis revealed insufficient Wolbachia read depth
    Could also: A sensitivity analysis reporting results both with and without this dataset — or applying rarefaction-based subsampling to equalize depth across datasets before DE analysis — could also be performed — Explicit depth-equalization or sensitivity analysis would allow readers to assess how conclusions about cross-dataset patterns are affected by the exclusion decision and depth heterogeneity across the seven studies
Software: R 3.5.0 · edgeR 3.24.0 · WGCNA 1.6.6 · HISAT2 2.1.0 · FADU 1.4 · vegan 2.5-4 · FactoMineR 1.41 · pvclust 2.0-0 · InterProScan 5.34-73 · PanOCT 3.23 · UpSetR 1.4.0 · SRA-Toolkit 2.9 · Seqtk 1.2 · ggplot2 3.1.1 · dendextend 1.9 · ggdendro 0.1-20 · cowplot 0.9.4 · gplots 3.0.1.1

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
14
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

10.25387/g3.12605408 DOI in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
NC_002978.6 RefSeq in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
NC_006833.1 RefSeq in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
NC_018267.1 RefSeq in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32718933

Paper: Chung et al. 2020, A Meta-Analysis of Wolbachia Transcriptomics Reveals a Stage-Specific Wolbachia Transcriptional Response Shared Across Different Hosts. G3 10(9):3243. DOI 10.1534/g3.120.401534. PMCID PMC7467002.

Re-analysis of 7 published Wolbachia RNA-seq studies through one common pipeline.

Code / data artifacts

  • BRIEF-listed code: github.com/lh3/seqtk (seqtk = one step of the pipeline; third-party).
  • Authors' own analysis repo (used here):
    • github.com/matt-chung/wolbachia_transcriptomics_metaanalysis (literate README: bash alignment workflow + full R meta-analysis with embedded expected outputs).
    • Zenodo 10.5281/zenodo.3726272 = github.com/Dunning-Hotopp-Lab/A-meta-analysis-...-resp v1.0 (same content; ships SupTable1–4 + gff maps + figures).
  • Shipped data that makes this reproducible without re-aligning raw SRA:
    • SupTable1.xlsx = FADU fractional count matrix per study (edgeR input).
    • SupTable2.xlsx = TPM matrix per study.
    • SupTable3.xlsx = DE-gene log2TPM/z-score (heatmap) data (5 DE studies).
    • SupTable4.xlsx = functional-term enrichment / WGCNA module data.

Pipeline (per Methods)

SRA-Toolkit v2.9 download → HISAT2 v2.1.0 (host+Wolbachia) → Seqtk v1.2 extract Wolbachia-mapping reads → HISAT2 realign (no spliced, max fragment 1 kbp) → FADU v1.4 quantify (CDS, ID attr, study-specific strandedness) → edgeR v3.24.x (QL F-test, CPM≥5 filter, FDR<0.05) → FactoMineR PCA / hierarchical clustering / WGCNA.

IN SCOPE (pipeline-derived, attempted — reproduced from shipped counts)

  1. Per-study protein-coding gene totals (count-matrix integrity).
  2. Table 1 "Reads to CDS" ranges = per-sample colSums of the count matrix (min–max).
  3. Genes passing the edgeR minimum-CPM filter (kept) per study.
  4. Differentially expressed genes (FDR<0.05) per study — Table 1 "DE Genes Reanalysis".
  5. PCA PC1 variance % (Darby 2012 = 44.5%; Luck 2015 = 70.9%).

These are reproduced exactly per the authors' R code (edgeR 3.24.x, FactoMineR 1.41) using SupTable1 counts + sample groups reconstructed from the column headers (the study_sample_map.tsv.txt is not shipped; groups = sample id minus replicate suffix; verified against embedded expected outputs).

HARDER (optionally attempted)

  1. Upstream Table 1 "Wolbachia Reads Mapped" ranges — requires SRA→HISAT2→seqtk→FADU on raw reads (heavy). seqtk is the BRIEF-named code; running it on the paper's SRA data for one small study (Darby 2012 wOo) exercises that artifact directly.

OUT OF SCOPE (not pipeline / not attempted)

  • WGCNA module biology, functional/GO enrichment narrative, manuscript figures' aesthetic layout, cross-study core-genome ortholog calls (depend on external OrthoFinder/InterProScan runs and manual interpretation).
  • "DE Genes Original" column (values taken from the original papers, not this pipeline).
Figures / tables: Table1Table
gene_totals_x7
Reported
651,1086,871,871,1086,839,754
Reproduced
651,1086,871,871,1086,839,754
exact
cds_reads_ranges_x7
Reported
Table 1 'Reads to CDS' (7 studies)
Reproduced
identical to rounding (e.g. wOo 26067-146642; chung 8525-35276890)
exact
cpm_kept_wOo_darby_2012
Reported
560
Reproduced
560
exact
cpm_kept_wMel_darby_2014
Reported
867
Reproduced
867
exact
cpm_kept_wDi_luck_2014
Reported
18
Reproduced
18
exact
cpm_kept_wMel_gutzwiller_2015
Reported
781
Reproduced
781
exact
cpm_kept_wBm_grote_2017
Reported
425
Reproduced
425
exact
cpm_kept_wBm_chung_2019
Reported
548
Reproduced
548
exact
de_wOo_darby_2012
Reported
Reproduced
exact
de_wMel_darby_2014
Reported
120
Reproduced
120
exact
de_wMel_gutzwiller_2015
Reported
473
Reproduced
473
exact
de_wBm_grote_2017
Reported
94
Reproduced
94
exact
de_wBm_chung_2019
Reported
373
Reproduced
373
exact
pca_pc1_wOo_darby_2012
Reported
44.5
Reproduced
44.5
exact
pca_pc1_wDi_luck_2015
Reported
70.9
Reproduced
70.9
exact
cpm_kept_wDi_luck_2015
Reported
325
Reproduced
0 (all 12 samples removed by documented >2000-count depth filter)
did not match
de_wDi_luck_2015
Reported
36
Reproduced
not computable (0 samples after filter)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴

Running the authors' own edgeR + FactoMineR code on their deposited FADU count matrices reproduces 25/27 pinned numbers exactly (gene totals, CPM-kept, DE counts, both PCA PC1 variances), with only fractional FADU rounding as a difference — an exemplary downstream reproduction. The single exception is Luck2015 (wDi): reported 325 kept / 36 DE genes are not derivable from the deposited matrix, whose counts (0–349/sample) cause the authors' documented depth filter to drop all samples. This sits on the authors'/deposit side (likely a wrong or truncated deposited matrix, not our method), and is severe for that one dataset but does not affect the central shared-stage-response conclusion. Flagged on q5/q4 as value-not-derivable; overall a solid reproduction with one genuine integrity concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

296 k
tokens (I/O) · 20 M incl. cache
36 min
runtime · 0 CPU-h
0.2 GB
peak RAM
1
HPC jobs
hummel
machine