The soil microbiome modulates the sorghum root metabolome and cellular traits with a concomitant reduction of Striga infection.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🔴A deviation arose in the data or preprocessing
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the RNA-seq DE pipeline: 1:1 in method and data, partial in exact numbers. The authors' own shipped scripts (DE-SQR2.R/-SQR3.R; edgeR + limma-voom) were re-run on the public GEO count matrix (GSE216351_20180321_Counts_SQR.csv; the RNA-seq lives in GEO GSE216351, NOT in the registry's amplicon accession PRJEB57848) and compared to the DEG lists in Data S2 (mmc3.xlsx). Sample structure matched the paper exactly (2wpi n=15: 3/4/4/4; 3wpi n=16: 4x4). Qualitative pattern reproduced exactly (2wpi many DEGs, Soil>>Striga>interaction; 3wpi very few). Exact DEG counts differ (Soil 2wpi 2516 vs 2563 = -1.8%; interaction 2wpi 589 vs 893 = -34%; 3wpi single-to-low-double digits), attributable to unpinned R/Bioconductor versions (repo has no README/DESCRIPTION/renv; ran R 4.3.3/edgeR 4.0.16/limma 3.58.1). Gene-level overlap is high and the smaller DEG set is a near-subset of the larger in every contrast (e.g. 2wpi Soil Jaccard 0.903; 2wpi interaction 570/589 reproduced genes are among the reported 893) -> threshold-boundary effect, NO fabrication signal; Data S2 ships full gene-level stats so the comparison is end-to-end auditable. NOT attempted (hard >20%): the GJAM microbiome modeling + taxa-rank scripts (4 of 6 repo scripts ship only post-processing of GJAM outputs; the model run, OTU tables, metabolome/trait inputs, amplicon->OTU step, and a curated isolate CSV are not provided and not derivable from PRJEB57848 alone). All wet-lab/phenotype results out of scope.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 62assessed: 2026-06-15 ⛓ 63e779a8b157
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan the soil microbiome suppress early-stage Striga hermonthica infection of sorghum, and if so, through what mechanisms involving host root chemistry and root cellular traits?
- ★ The natural soil microbiome reduces Striga attachment to sorghum roots at the post-germination stage by interfering with haustorium initiation rather than seed germination finding
- ★ The microbiome lowers levels of haustorium-inducing factors (HIFs) in the rhizosphere, likely via microbial degradation into conversion products mechanism
- ★ The soil microbiome induces root endodermal suberization and cortical aerenchyma formation independent of Striga infection finding
- ★ The microbiome alters transcription of suberin biosynthetic genes and aerenchyma-associated genes in sorghum roots finding
- ★ Pseudomonas strain VK46 reduces haustorium formation by degrading syringic acid mechanism
- ★ Arthrobacter strain VK49 increases sorghum endodermal suberization mechanism
- Gamma irradiation sterilizes the soil (lowers bacterial alpha diversity) without altering its physico-chemical properties, enabling natural-vs-sterilized comparison method
- The Clue Field soil serves as a discovery tool/resource for identifying Striga-suppressive bacterial taxa resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| 16S rRNA and ITS amplicon sequencing (microbiome profiling) | Clue Field bulk soil (natural vs gamma-irradiated) | gamma irradiation (sterilization) | bacterial and fungal alpha diversity / community composition | — |
| Striga infection / attachment assay | Sorghum bicolor cv. Shanqui Red (SQR) grown in natural vs sterilized soil | soil microbiome presence/absence; Striga hermonthica inoculation | number of Striga attachments per gram fresh root weight at 2 and 3 wpi | — |
| In vitro Striga germination assay | Striga hermonthica seeds + sorghum root exudates | root exudates from natural vs sterilized soil; GR24 positive control | germination percentage | — |
| In vitro haustorium formation assay | Striga hermonthica seeds + sorghum root exudates | root exudates from natural vs sterilized soil | percentage of seeds forming early haustoria | — |
| Targeted HIF metabolite profiling of root exudates | SQR root exudates (natural vs sterilized soil, ± Striga) | soil microbiome; Striga infection | relative abundance (peak area) of HIFs (syringic acid, vanillic acid, DMBQ, acetosyringone, vanillin) | — |
| Untargeted metabolite profiling of root exudates | SQR root exudates from 4-week-old plants (natural vs sterilized soil) | soil microbiome | abundance of features matching predicted HIF conversion products | — |
| Root anatomy / histology (fluorol yellow suberin staining, toluidine blue cross-sections) | Sorghum SQR crown and seminal roots | soil microbiome; Striga infection | endodermal suberin fluorescence intensity; aerenchyma proportion of cross-section area; cortex layers, metaxylem vessels, endodermal lignification | — |
| Root system architecture phenotyping | Sorghum SQR mature root systems (crown and seminal roots) | soil microbiome; Striga infection | root length, area, diameter, biomass | — |
- ▼ Significantly fewer Striga attachments on SQR roots in natural vs sterilized soil at 2 and 3 wpi
- – No difference in Striga germination between exudates from natural vs sterilized soil
- ▼ Haustorium formation strongly reduced with natural-soil exudates; >60% of seeds formed haustoria with sterilized-soil exudates vs few with natural-soil exudates >60% vs few
- ▼ Syringic acid and vanillic acid levels lower in natural-soil exudates both with and without Striga at 2 wpi; DMBQ and acetosyringone lower only in absence of Striga
- ▲ Of 82 features matching predicted HIF conversion products, 26 differed significantly; 73% accumulated higher in natural-soil exudates 73% of 26 compounds
- ▲ Greater endodermal suberization in crown roots of 5-week-old plants in natural soil, independent of Striga
- ▲ More aerenchyma in crown roots of 4- and 5-week-old plants in natural soil, independent of Striga
- ▲ Sorghum orthologs of maize aerenchyma-associated genes enriched among microbiome-regulated genes p=0.008
- other ~20% of sorghum yield lost annually in Africa (Striga infestation yield loss)
- count >6 million tons grain lost annually (annual cereal production losses from Striga)
- other >60% (Striga seeds forming haustoria with sterilized-soil exudates)
- count 26 of 82 features (HIF conversion product features differing significantly between soils)
- pvalue adjusted p < 0.05, log2 fold change > 1 or < −1 (threshold for differential untargeted metabolite features)
- other 73% (share of differential conversion products higher in natural soil)
- pvalue p = 0.008 (Fisher's exact test) (enrichment of aerenchyma-associated gene orthologs among microbiome-regulated genes)
- count 74 compounds predicted (potential HIF break-down products predicted by BioTransformer)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study employed a two-way ANOVA framework (soil type × Striga infection status) as its primary analytical structure for comparing root cellular anatomy traits, root system architecture metrics, and root exudate HIF metabolite abundances. Welch t-tests were used for pairwise group comparisons of microbiome diversity and in vitro bioassay outcomes, and Tukey HSD post hoc tests were applied when soil × Striga interaction effects were detected in two-way ANOVAs. For multi-feature untargeted metabolomics and transcriptomics, significance was assessed using adjusted p-values with an additional log2 fold-change threshold. Results were predominantly visualized as boxplots (25th–75th percentile, median-centered) with significance denoted by asterisk tiers or letter-based post hoc groupings rather than exact p-values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Welch t-test | Alpha diversity comparison of bacterial (Figure 1A) and fungal (Figure 1B) communities between natural and sterilized soil | n = 4 per soil type | not stated |
| Two-way ANOVA | Number of Striga attachments per gram fresh root weight at 2 and 3 weeks post-infection (Figure 1C) | n = 6 per group | not stated |
| Welch t-test | In vitro Striga seed germination (Figure 1D) and haustorium formation (Figure 1E) in response to root exudates from natural vs. sterilized soil | n = 6 exudates per soil; 3 technical replicates per exudate | not stated |
| Two-way ANOVA with Tukey HSD post hoc test | Relative abundance of individual HIFs (syringic acid, vanillic acid, DMBQ, acetosyringone, vanillin) in root exudates; Tukey applied where soil × Striga interaction was detected (Figures 2A–2E) | n = 6 per group | not stated |
| Two-way ANOVA | Root cellular anatomy (suberin, aerenchyma, cortex layers, metaxylem vessels, lignification) and root system architecture traits at 2 and 3 weeks post-infection (Figure 3A) | Varies per trait; listed in Data S1 | not stated |
| Fisher's exact test | Enrichment of sorghum orthologs of maize aerenchyma-associated genes among microbiome-regulated genes (Figure 3M) | — | na |
-
Welch t-tests were used to compare alpha diversity metrics between two soil conditions (n = 4 per group)↳ Could also: A non-parametric Mann-Whitney U (Wilcoxon rank-sum) test could also be applied — With n = 4 per group, normality cannot be reliably assessed; a rank-based test makes no distributional assumption and is often preferred for small samples or diversity indices that may not be normally distributed
-
Multiple separate two-way ANOVAs were conducted across a large panel of root anatomy, RSA, and metabolite traits, each tested independently↳ Could also: A multivariate approach such as MANOVA, PERMANOVA, or a linear mixed model with a global multiplicity correction across all traits could also be used — A single multivariate test controls the experiment-wide error rate across correlated traits; a global FDR correction (e.g., Benjamini-Hochberg applied across all trait p-values) would also clarify the study-wide false discovery rate
-
The multiplicity correction method for the adjusted p-values in metabolomics and gene expression analyses is not named↳ Could also: Explicitly stating the correction method (e.g., Benjamini-Hochberg FDR or Bonferroni) would also fully specify the analysis — Naming the method allows readers to understand the stringency applied and to reproduce or compare results; BH-FDR and Bonferroni differ substantially in power at the scales typical of untargeted metabolomics
-
Striga attachment counts (a count variable bounded at zero) were analyzed with two-way ANOVA↳ Could also: A generalized linear model with a Poisson or negative-binomial error distribution could also be applied to count data — Count outcomes often violate the normality and homoscedasticity assumptions of ANOVA, particularly when values cluster near zero; a GLM framework is designed for count data and accommodates overdispersion naturally
-
Results were summarized and displayed with boxplots showing the IQR and median↳ Could also: Reporting the mean ± SD (or a 95% CI around the mean) alongside or instead of the median/IQR would also convey central tendency and spread — When parametric tests (ANOVA, t-test) are used, reporting the mean and SD is internally consistent with those test assumptions; CIs additionally convey estimation uncertainty and are increasingly recommended by reporting guidelines (e.g., APA, Nature Portfolio)
-
Effect sizes (e.g., Cohen's d, eta-squared, fold-change for continuous traits) were not reported for the ANOVA or t-test comparisons↳ Could also: Reporting partial eta-squared (η²p) for ANOVA terms or Cohen's d for t-tests alongside p-values would also characterize the magnitude of observed differences — P-values depend on sample size and do not convey practical magnitude; effect sizes allow readers to judge biological relevance independently of n and are required by many journals for complete reporting
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Aerenchyma proportion of cross-sectional area is greater in crown roots of sorghum grown in natural soil versus sterilized soil, independent of Striga infection.imaging sorghum root up 2024×1papers★ This paper is the founder (earliest)
-
Endodermal suberization is greater in crown roots of sorghum grown in natural soil versus sterilized soil, independent of Striga infection.imaging sorghum root up 2024×1papers★ This paper is the founder (earliest)
-
73% of predicted HIF conversion products accumulate at higher levels in root exudates from sorghum grown in natural soil, consistent with microbial conversion of haustorium-inducing factors.metabolomics sorghum root up 2024×1papers★ This paper is the founder (earliest)
-
Haustorium-inducing factor (HIF) metabolites including syringic acid, vanillic acid, DMBQ, and acetosyringone are reduced in root exudates of sorghum grown in natural compared to sterilized soil.metabolomics sorghum root down 2024×1papers★ This paper is the founder (earliest)
-
Striga hermonthica attachment to sorghum roots is significantly reduced when plants are grown in natural versus gamma-irradiated sterilized soil, indicating a microbiome-mediated protective effect.other sorghum root down 2024×1papers★ This paper is the founder (earliest)
-
Striga hermonthica germination rate is not significantly different in response to root exudates from sorghum grown in natural versus sterilized soil.other striga hermonthica none 2024×1papers★ This paper is the founder (earliest)
-
Striga hermonthica haustorium formation is strongly reduced by root exudates from sorghum grown in natural soil versus sterilized soil, with >60% haustoria in sterilized-soil controls.other striga hermonthica down 2024×1papers★ This paper is the founder (earliest)
-
Sorghum orthologs of maize aerenchyma-associated genes are significantly enriched among microbiome-regulated genes in sorghum roots (p=0.008).RNA-seq sorghum root up 2024×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38537644
Paper: Kawa et al. 2024, Cell Reports 43(4):113971. "The soil microbiome modulates the sorghum root metabolome and cellular traits with a concomitant reduction of Striga infection." PMCID PMC11063626.
Code: https://github.com/DorotaKawa/Striga-suppressive-soil (public, main, last push 2022-10-10, no license, 6 R scripts, no README).
Repo inventory (6 R scripts, no data, no README)
| script | purpose | required input | input shipped? |
|---|---|---|---|
20210210-DE-SQR2.R |
RNA-seq DE, 2 weeks post-infection | 20180321_Counts_POP2.csv (raw count matrix) |
NO (repo); YES via GEO |
20210210-DE-SQR3.R |
RNA-seq DE, 3 weeks post-infection | same count matrix | NO; YES via GEO |
GJAM outputs analysis- Bacteria.R |
rank microbial taxa from GJAM residual corr. | ...Residual Correlation_M_complete.xlsx (= GJAM output) + curated isolate CSV |
NO; xlsx ≈ Data S3 |
GJAM outputs analysis- Fungi.R |
same, fungi | same | NO; xlsx ≈ Data S3 |
20220811-Post GJAM analysis -Bacteria.R |
post-processing | GJAM outputs (local Dropbox) | NO |
20220825 - Post GJAM analysis- Fungi.R |
post-processing | GJAM outputs (local Dropbox) | NO |
All scripts hard-code local Dropbox paths («path»).
Data availability (from paper "Data and code availability")
- Sorghum transcriptome → GEO GSE216351 (BioProject PRJNA893193).
Series supplementary =
GSE216351_20180321_Counts_SQR.csv.gz— this is the raw count matrix the DE scripts consume (same20180321_Countsdatestamp). - Microbiome amplicon reads → ENA PRJEB57848 (120 AMPLICON runs) — raw input to an OTU/GJAM pipeline whose code is NOT in the repo.
- Reported outputs in supplementary spreadsheets: Data S2 (mmc3.xlsx) = CPM values + DEG lists; Data S3 (mmc4) = GJAM residual correlations; Data S4 (mmc5) = taxa ranks.
In scope (pipeline-derived, attempted)
- C1/C2 — RNA-seq differential expression (edgeR + limma-voom). Fully
specified: the two DE scripts are self-contained given the count matrix, which
IS publicly obtainable from GEO. Reproduce the per-contrast DEG counts
(Soil
N, TreatmentSTRIGA, interactionNSTRIGA) at 2 weeks and 3 weeks at adj.P.Val < 0.05, and compare against the DEG lists in Data S2 (mmc3.xlsx). Pipeline: read counts → filter (rowSums(cpm>1)>=3) → TMM →model.matrix (~Soil*Treatment)→voomWithQualityWeights(normalize="quantile")→lmFit→eBayes→topTable(adjust="BH")→ count adj.P<0.05.
Out of scope (not attempted, with reason)
- GJAM microbiome modeling & rank scripts. The repo ships only POST-processing of GJAM model OUTPUTS. The GJAM model itself, its inputs (OTU tables, metabolome, trait matrices), the amplicon→OTU step, and the curated isolate CSV are NOT shipped and NOT derivable from the single public accession without re-deriving an unspecified multi-step pipeline (UPARSE → phyloseq → gjam v2.6.2). This is the hard >20%; skipped per 80/20. (The GJAM input xlsx ≈ Data S3 and the output ranks ≈ Data S4 exist as supplementary, so a partial rank reproduction is conceivable later, but the scripts' hard-coded paths/filenames and missing isolate CSV make it fiddly; not attempted now.)
- All wet-lab / phenotype results (Striga attachment counts, suberin, aerenchyma, metabolite/exudate measurements, isolate assays) — manual/experimental, out of scope by definition.
Known script caveat (faithfulness note)
Both DE scripts contain a dead, crashing line: contrasts.fit(vfit, contrasts=contr.matrix) where contr.matrix is never defined. The actually-saved
results come from the subsequent block (lmFit→eBayes→topTable on the design
coefficients), which does not use contr.matrix. We therefore run the
result-producing block verbatim and omit the dead line; this is faithful to the
outputs the script writes to disk.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Using the authors' own shipped DE scripts on the identical public GEO count matrix with exactly matching sample structure, the qualitative pattern and rank-order of DEGs reproduced 1:1, and every reported DEG set is recovered as a near-subset with high overlap (Jaccard 0.903/0.772/0.625 at 2wpi). Exact counts differ — most notably C3 interaction 893→589 (-34%) — but this is an our-side software limitation (unpinned R/Bioconductor; FDR-boundary genes), not an authors' defect, and there is no fabrication signal (Data S2 ships full gene-level stats). Overall a solid reproduction with explainable, version-attributable deviations: green on derivability and core claim, yellow overall and on severity.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.