Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Metapangenomics of wild and cultivated banana microbiome reveals a plethora of host-associated protective functions.

Environ Microbiome · 2023
L1 59/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
59/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 19% of all assessed papers rank 925 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the data-handling pipeline; partial overall. WHAT REPRODUCED: data provenance C1 (161.39 vs 161.5 Gbp, <0.1%) and C2 (14 metagenomes, exact) from ENA; C3 host-DNA removal (45.3 Gbp claim) independently reproduced to 43.63 Gbp -- within ~3.7% -- using the paper's EXACT aligner (BWA 0.7.17 BWA-MEM) and EXACT specified Musa reference genomes (M.acuminata GCF_000313855.2 lineage + M.balbisiana GCA_004837865.1) plus Trimmomatic 0.38; and T1, a deterministic code-artifact reproduction (P16) of the cited third-party tool lyijin/topGO_pipeline (88.5% of GO-enrichment values bit-identical to the repo's shipped reference, the rest explained by GO.db ontology version drift). DIFFERENT/NOT 1:1: none of the headline numbers appear fabricated, but several are not reproducible for legitimate reasons. NOT ATTEMPTED / BLOCKED: C4 metagenome co-assembly (metaSPAdes v3.13.0 genuinely launched on the 43.6 Gbp host-removed reads, but a paper-scale 44 Gbp co-assembly cannot complete under a 12h wall + 750GB single node with no big-partition access); C5 taxonomy (depends on a non-obtainable Dec-2020 NCBI nr snapshot); C6 pangenome (extreme compute, depends on C4); C7's specific 192-GO-term result (paper's curated 72-genus GO annotation + gene lists are unshipped). KEY DATA NOTE: the brief's accession PRJNA432894 is actually the M. balbisiana HOST GENOME used for host removal, not the microbiome reads -- the microbiome data is PRJNA837781 (confirmed). All grades are provisional and human-checkable (see AUDIT.md / agreement.json).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-25
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether Musa (banana) host genotype and tissue type (root vs. leaf) shape endosphere microbiome taxonomic and functional composition, predicting that root endospheres would be more diverse than leaf, wild diploid banana relatives would harbor more diverse endospheres than cultivated triploids, and that genotype-specific (vs. transient shared) microbes would be enriched for host-beneficial rather than pathogenic functions.

Core claims
  • Root and corm endosphere communities are significantly richer and compositionally distinct from leaf endosphere communities across Musa genotypes finding
  • Wild diploid M. balbisiana (BB genotype) endospheres harbor richer taxa and functions than cultivated triploid M. acuminata (AAA genotype) finding
  • Agrobacterium and Rhizobium dominate across all samples, while Chitinophagia and Actinomycetia are enriched in roots and Flavobacteria in leaves finding
  • Most microbial taxa and gene clusters are unshared (unique) between different Musa genotypes, indicating strong genotype-driven structuring of the endosphere finding
  • Some plant-protective gene clusters show phylosymbiosis signatures, suggesting long-standing or heritable host-microbiome associations in Musa mechanism
  • A culture-free density-gradient (Nycodenz) enrichment protocol effectively enriches endophytic bacterial DNA from banana root and leaf tissue for shotgun metagenomic sequencing method
  • The resulting metapangenomic dataset (taxa and gene clusters across 7 Musa genotypes) provides a baseline resource for future in planta bacterization or engineering of wild host endophytes resource
  • Gene ontology enrichment reveals core beneficial functions shared with other plant microbiomes as well as many previously unreported specialized prospective beneficial functions finding
Experimental setups
Assay System Perturbation Readout Platform
shotgun metagenomic sequencing root/corm and leaf endosphere tissue from 7 Musa genotypes (14 samples) none (comparison across host genotype and tissue) microbial community taxonomic and functional composition Illumina HiSeq 2500
culture-free microbiome enrichment (density gradient centrifugation) Musa root/corm and leaf tissue homogenates none enriched microbial cell pellet (bacteria-dominated) for DNA extraction Nycodenz density gradient
metagenome assembly and taxonomic binning assembled metagenomic scaffolds from Musa endosphere samples none taxonomic assignment of scaffolds/reads to species/strain level metaSPAdes; DIAMOND v2.0.9 vs NCBI nr
diversity analysis (OTU-based) Musa root and leaf endosphere microbial communities host genotype and tissue type as grouping variables species richness/diversity indices (Chao, ACE, Shannon, Simpson) and beta diversity (PCoA/NMDS, PERMANOVA) phyloseq and Vegan packages in R
metapangenome/ortholog clustering Prokka-annotated metagenome assemblies from each Musa genotype/tissue none shared vs. unique gene cluster presence/absence across genotypes and tissues Prokka v1.14.6; Roary v3.13.0 with PRANK
gene ontology (GO) enrichment analysis core and unique/shared gene clusters from banana microbiome metapangenome none enriched functional (GO) categories topGO v2.4.0; GO-Figure!
host plant phylogenetic reconstruction Musa, Ensete, and Musella marker gene (ycf1) sequences none host phylogenetic relationships/topology RAxML v4.0; MrBayes v2.2.4
phylosymbiosis analysis of microbial gene clusters microbial ortholog gene cluster phylogenies vs host Musa phylogeny none topological congruence between microbial gene trees and host tree FastTree 2.1.8
Key results
  • Root/corm endosphere communities are significantly richer and compositionally distinct from leaf communities
  • M. balbisiana (wild BB genotype) has richer taxa and functions than M. acuminata (cultivated AAA genotype)
  • Agrobacterium and Rhizobium were most abundant overall; Chitinophagia and Actinomycetia more abundant in roots, Flavobacteria more abundant in leaves
  • More than 2000 bacterial taxa were unique to each of M. acuminata and M. balbisiana genotypes >2000 taxa
  • About 20% of sequence reads did not match any taxon database ~20%
  • About 62% of gene clusters could not be annotated to function ~62%
  • Some gene clusters with plant-protective functions showed phylosymbiosis signatures
  • 24,325 species/strains and 559,108 predicted gene clusters were identified across 1.7 million metagenomic scaffolds 24,325 taxa; 559,108 gene clusters
Key statistics
  • count 24,325 (total species or strains distinguished across all samples)
  • count 1.7 million (metagenomic scaffolds assembled)
  • count 559,108 (predicted gene clusters)
  • other ~20% (sequence reads not matching any taxon database)
  • other ~62% (gene clusters unannotated to function)
  • count >2000 (bacterial taxa unique to each of M. acuminata and M. balbisiana genotypes)
  • count 161.5 Gbp (total raw sequence data generated across 14 Musa samples)
  • other 13.3% to 97.72% (range of percentage of sequenced nucleotides per sample not mapping to Musa host genome)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study compared endosphere microbiome diversity and metapangenome gene content across 7 Musa genotypes and 2 tissue types (root/corm vs. leaf), using one shotgun-metagenomics-derived sample per genotype-tissue combination. Community-level differences were assessed with PERMANOVA (Adonis, Bray-Curtis dissimilarity) and ordination (PCoA/NMDS) in the R vegan/phyloseq packages, richness indices (Chao, ACE, Shannon, Simpson) were computed on read-count-normalized OTU data, and individual taxon abundances were compared using ANOVA (across genotypes, with pairwise Welch's t-tests as post-hoc) and Welch's two-sample t-test (across tissues). Gene ontology functional enrichment on metapangenome gene clusters was performed with topGO using the weight01 Fisher's exact algorithm. The excerpted methods/results text reports outcomes largely in qualitative terms (e.g., 'significantly richer/different') without exact p-values, effect sizes, or dispersion measures visible in this portion of the text.

Replicationunclear Sample size14 samples total: 7 Musa genotypes/species x 2 tissues (root+corm, leaf), one sample per genotype-tissue combination as listed in Table 1; no explicit statement of biological replicate counts per group or a priori power/sample-size calculation Groupshost genotype/species (7 Musa genotypes) and tissue type (root+corm vs. leaf) Pairingunclear Randomization/blindingnot stated Dispersionunclear Multiplicity correctionFor GO enrichment, topGO's weight01 Fisher algorithm is described as accounting for GO-graph topology and returning 'multiple testing independent' p-values; no separately named correction (e.g., Bonferroni, Benjamini-Hochberg) is stated for the ANOVA/t-test taxon-abundance comparisons
Statistical tests used
Test Applied to n Assumptions
PERMANOVA (Adonis, vegan package) microbial community (OTU) composition differences by genotype and tissue location, based on Bray-Curtis dissimilarity not stated
ANOVA individual taxon abundance compared across host genotypes not stated
Welch's two-sample t-test (pairwise post-hoc following ANOVA, and independently for tissue comparisons) pairwise genotype comparisons of taxon abundance, and root/corm vs. leaf tissue comparisons not stated
topGO 'weight01.fisher' (topology-weighted Fisher's exact test) gene ontology enrichment of metapangenome gene clusters (overlapping and non-overlapping gene sets) not stated
Approaches that could also have been used
  • Individual taxon abundances were compared across genotypes (ANOVA with pairwise Welch's t-test post-hoc) and across tissues (Welch's t-test) on a per-taxon basis without a stated correction for the resulting large number of simultaneous tests.
    Could also: A false discovery rate procedure such as Benjamini-Hochberg applied across all taxon-level comparisons could also be used. — When many taxa are tested individually, an FDR adjustment helps control the expected proportion of false positives among the taxa called significant, which is a standard complement to per-taxon ANOVA/t-test screens in microbiome studies.
  • Community-level compositional differences were tested with PERMANOVA (Adonis) on Bray-Curtis dissimilarities.
    Could also: A complementary PERMDISP (betadisper) test of within-group multivariate dispersion could also be run alongside PERMANOVA. — PERMANOVA can detect both centroid (location) and dispersion differences between groups; pairing it with PERMDISP helps distinguish whether a significant PERMANOVA result reflects differing group centroids, differing within-group variability, or both.
  • GO enrichment was assessed using topGO's weight01 Fisher's exact algorithm, which the authors describe as producing multiple-testing-independent p-values based on GO graph topology.
    Could also: An alternative enrichment approach, such as gene set enrichment analysis (GSEA) or a classic hypergeometric/Fisher test with an explicit Benjamini-Hochberg-adjusted q-value per term, could also be applied. — Reporting an explicit FDR-adjusted q-value alongside a topology-aware test gives readers a directly comparable multiplicity-controlled statistic that is common across enrichment tools and studies.
  • Each of the 7 Musa genotypes was represented by one sampled specimen per tissue type (root/corm and leaf), as shown in Table 1, rather than multiple independent plants per genotype.
    Could also: Sampling multiple independent biological replicate plants per genotype (and tissue) is also a standard design in host-microbiome comparisons. — Additional biological replicates per genotype allow estimation of within-genotype (plant-to-plant) variability, which can be compared against between-genotype variability to strengthen inference that observed differences are genotype-driven rather than reflecting a single individual plant.
  • Diversity indices (Chao, ACE, Shannon, Simpson) were calculated from OTU data normalized to the median of sample read counts, with results described qualitatively (e.g., 'significantly richer').
    Could also: Presenting bootstrap-based or analytic confidence intervals around each diversity index estimate could also be reported. — Confidence intervals convey the precision of diversity estimates, which can be especially informative given the wide range of sequencing depth and percent-host-mapped values across samples noted in Table 1.
  • Statistical outcomes in the provided text are largely reported as significance statements (e.g., 'significantly different') without accompanying exact p-values.
    Could also: Reporting exact p-values (in addition to or instead of significance labels) is also standard practice. — Exact p-values let readers gauge the strength of evidence directly and support later meta-analyses or reanalyses that combine results across studies.
Software: R (phyloseq package) · R (vegan package, Adonis/PERMANOVA) · topGO 2.4.0 · QUAST 5.0.2 · RAxML 4.0 · MrBayes 2.2.4

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

C1
Reported
161.5 Gbp raw sequence data
Reproduced
161.39 Gbp (sum of ENA base_count over 14 runs)
within tolerance
C2
Reported
14 Musa specimens / metagenomes
Reproduced
14 SRA runs / 14 BioSamples observed
exact
C3
Reported
45.3 Gbp after plant-DNA removal
Reproduced
43.63 Gbp non-host (Trimmomatic 0.38 + BWA-MEM 0.7.17 vs M.acuminata GCF_000313855.2 + M.balbisiana GCA_004837865.1, all 14 metagenomes); within ~3.7%
within tolerance
C4
Reported
~2.2M scaffolds, 2.03 Gbp, N50 ~2191 bp
Reproduced
ATTEMPTED: metaSPAdes v3.13.0 launched (k=21,33,45,59,73,99) on 43.6 Gbp host-removed co-assembly input; cannot complete within 12h-wall/750GB-std limits (no big/2.2TB partition access)
partial
C5
Reported
24,325 species/strains
Reproduced
blocked: DIAMOND vs NCBI nr Dec-2020 snapshot not reproducibly obtainable (DB-version drift)
did not match
C6
Reported
559,108 gene clusters
Reproduced
not attempted: Prokka+Roary over ~1.7M genes, extreme compute + depends on resource-blocked C4 assembly
partial
C7
Reported
192 enriched GO terms (46 shared)
Reproduced
blocked: paper's curated 72-genus GO annotation + core-gene lists not shipped (repo or SRA)
did not match
T1
Reported
topGO_pipeline tool produces GO-enrichment tables = shipped reference
Reproduced
938/1060 (88.5%) GO terms numerically identical to shipped reference; residual 11.5% = GO.db ontology drift 2018->2026, not a code defect
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 59/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

What reproduced: the upstream half. C1 (161.39 vs 161.5 Gbp, <0.1%), C2 (14/14 runs and BioSamples), and C3 (43.63 vs 45.3 Gbp, ~3.7% low) all check out, the last using the authors' own BWA-MEM 0.7.17 and specified Musa references; T1 shows the cited third-party tool lyijin/topGO_pipeline is deterministic (938/1060 values bit-identical to its shipped reference, residual = GO.db 2018→2026 drift). What deviates and whose side: the only measurable deviation (C3's ~3.7%) is squarely our side — we dropped the paper's Pear v0.9.11 merge and had to invent Trimmomatic parameters the paper never states. What is blocked: all four headline results. C4 (2.2M scaffolds / 2.03 Gbp / N50 2191) and C6 (559,108 gene clusters) are our-side compute limits (12h wall, single 750 GB node); C5 (24,325 taxa) and C7 (192 GO terms, 46 shared) are authors'-side transparency gaps — a frozen, never-archived Dec-2020 nr snapshot and a curated 72-genus GO annotation that appears in neither commit f36e70e nor SRA. Severity: nothing contradicts the paper and nothing looks fabricated — every derivable number landed within tolerance — but the central functional conclusion was never actually tested, so this is a yellow on incompleteness, not on discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

65.6 k
tokens (I/O) · 3.1 M incl. cache
10 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.