Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Niche Modification by Sulfate-Reducing Bacteria Drives Microbial Community Assembly in Anoxic Marine Sediments.

mBio · 2023
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1) for all in-scope pipeline-derived results. The paper's own analysis repo (github.com/2015qyliang/InhibitingSRB_molybdate; NOT the brief's code_url github.com/hyattpd/Prodigal, which is only a Methods-cited tool) ships per-figure data + R scripts. Re-running them on «our HPC» (R 4.3.3, vegan 2.7.1->adonis2, microeco, GUniFrac, ape) regenerates the shipped derived values exactly: iCAMP relative importance (Fig 4, HoS+DL: C=96.72% M=96.06%, TEXTMATCH) and Cohen's d (TEXTMATCH); alpha diversity Observed range 341-1781 matching the paper exactly + ANOVA/Tukey NUMMATCH (1e-12); beta composition adonis/ANOSIM/MRPP byte-identical across all 12 pairs; betapart adonis R2 exact. Shipped OTU table integrity confirmed (3,337 OTUs x 180 samples; 9x20 design). Dataset profiled: SRP364228 = 208 runs observed (207 reported), deposited SINGLE-end though Methods describe PE300+FLASH. NOT attempted/blocked: the from-raw upstream 16S pipeline (SINGLE-end deposit + proprietary UPARSE v7 + missing run->group mapping), and out-of-scope wet-lab / MENAP web-server networks (Fig 3) / GTDB genome reconstruction (Fig 5/6). One honest note: shipped iCAMP data give HoS+DL96%, paper text rounds to ~90% (same dominance conclusion). All grades provisional; human must confirm.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 67
    assessed: 2026-06-19 ⛓ 2dd3bac944c1
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The combined effects of sulfate-reducing bacteria (SRB) strongly influence their localized environment through niche modification during anaerobic mineralization of organic matter, and this niche modification drives microbial community assembly and network structure in marine sediments.

Core claims
  • SRB modify niches (e.g., pH) to affect species coexistence during organic matter degradation finding
  • Molybdate specifically inhibits sulfate reduction/SRB activity in microcosms method
  • SRB and the order Marinilabiliales show strong coexistence via positive frequency-dependent selection finding
  • Inhibiting SRB reduces pH, which suppresses Marinilabiliales growth; adding HEPES buffer restores pH and Marinilabiliales relative abundance mechanism
  • Inhibiting SRB alters bacterial community composition, structure, molecular ecological networks, and assembly processes finding
  • Reduced Marinilabiliales abundance after SRB inhibition is caused by the inhibited SRB effect itself, not direct toxicity of the molybdate inhibitor mechanism
  • Marinilabiliales act as keystone taxa/module hubs contributing to preserved network modules and homogeneous ecological selection finding
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA gene high-throughput sequencing coastal marine sediment microcosms (n=20 replicates per stage) molybdate inhibition of SRB bacterial community composition, OTUs, alpha/beta diversity
molecular ecological network (MEN) construction marine sediment microcosm bacterial communities molybdate inhibition of SRB network topology, modularity, robustness, vulnerability, keystone nodes
physicochemical measurement (sulfate, sulfite, Fe, TOC, TIC, pH, phosphorus) marine sediment microcosms molybdate inhibition of SRB concentration changes over 30-day incubation
volatile fatty acid (VFA) quantification marine sediment microcosms molybdate inhibition of SRB acetate, propionate, butyrate and other VFA concentrations
molybdate tolerance/growth test isolated Marinilabiliales strains molybdate exposure growth/tolerance of isolates
LEfSe (linear discriminant analysis effect size) marine sediment microcosm bacterial community, family/order level taxa molybdate inhibition of SRB taxa with statistically significant and biologically consistent differential abundance
pH buffering rescue experiment SRB-inhibited marine sediment microcosms HEPES buffer addition pH and relative abundance of Marinilabiliaceae/Marinilabiliales
Key results
  • Sulfate concentration and multiple physicochemical parameters (Fe, TOC, TIC, pH, VFAs) differed significantly between control and SRB-inhibited groups from days 12-30, indicating sulfate reduction was fully blocked
  • Bacterial community structure significantly differed between control and SRB-inhibited groups across all time points (e.g., D00C_vs_D30C Adonis F=34.12, P=0.001) F=34.12, P=0.001
  • Families Synergistaceae, Peptostreptococcaceae, Dethiosulfatibacteraceae, Prolixibacteraceae, Marinilabiliaceae, and Marinifilaceae were simultaneously suppressed after inhibiting SRB
  • Molybdate tolerance test showed that inhibited SRB (not the molybdate inhibitor itself) triggered a decrease in relative abundance of Marinilabiliales
  • Inhibiting SRB reduced pH; adding HEPES buffer to SRB-inhibited microcosms restored pH and the relative abundance of Marinilabiliales-related bacteria
  • Inhibiting SRB significantly decreased network robustness and increased vulnerability at incubation days 5, 12, and 30
  • Marinilabiliales showed close association with SRB, contributed to preserved network modules, were keystone nodes, and contributed to homogeneous ecological selection
  • Inhibiting SRB increased node persistence and constancy of empirical networks compared to control
Key statistics
  • pvalue F=34.12, P=0.001 (Adonis test, D00C vs D30C bacterial community structure)
  • count 3,808,147 sequences from 180 samples (20,031-21,434 per sample) (16S rRNA gene sequencing depth)
  • count 3,337 OTUs (97% identity cutoff), ranging 341-1,781 per sample (OTU richness across samples)
  • count n=20 highly replicated microcosms per incubation stage (experimental design replication)
  • pvalue P<0.001 (community evenness lower in SRB-inhibited group on days 5 and 12)
  • other >40% of variance from nestedness, 16% from turnover, 25% from incubation days (beta diversity partitioning in SRB-inhibited group)
  • other 27 large modules (≥5 nodes) covering 75-88% of nodes (control) vs 26 modules covering 70-85% (SRB-inhibited) (molecular ecological network modularity comparison)
  • other proportion of novel OTUs ~8% (days 5-12) declining to ~2% (days 21-30) in control vs ~4% throughout in SRB-inhibited group (detection of previously unobserved OTUs over incubation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a highly replicated microcosm design (n=20 per stage per treatment, 180 total samples) comparing control vs. SRB-inhibited (molybdate-treated) sediment incubations across five timepoints. Physicochemical variables and diversity indices were compared using ANOVA with Tukey's HSD post-hoc tests, while bacterial community composition/structure differences were assessed with three nonparametric multivariate tests (Adonis/PERMANOVA, ANOSIM, MRPP) and LEfSe was used to identify differentially abundant taxa. Results are reported with exact p-values in some tables and significance-threshold stars (*, **, ***) elsewhere, with error bars shown as standard deviation.

Replicationbiological Sample sizeDescribed as a 'highly replicated' microcosm study with n=20 replicate microcosms per condition across five incubation stages, totaling 180 homogenized sediment samples; no formal power analysis mentioned GroupsControl vs. molybdate (SRB-inhibited) treatment microcosms sampled at 5 incubation days (0, 5, 12, 21, 30) Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionTukey's HSD (post-hoc correction following ANOVA)
Statistical tests used
Test Applied to n Assumptions
ANOVA with Tukey's HSD post-hoc test physicochemical factors and VFAs (Fig. 1), alpha diversity indices (Fig. 2A-C, Fig. S1), family-level abundance comparisons (Fig. S2A) between control and SRB-inhibited groups across incubation days n=20 replicate microcosms per stage/treatment as stated for the overall design not stated
Adonis (PERMANOVA) pairwise community structure comparisons across timepoints/treatments (Table 1, Table S1) not stated
ANOSIM same pairwise community structure comparisons (Table 1, Table S1) not stated
MRPP same pairwise community structure comparisons (Table 1, Table S1) not stated
LEfSe (LDA effect size) identifying taxa with significant and consistent differences between control and SRB-inhibited groups (Fig. S2B) not stated
Pearson and Spearman correlation relationships between network stability indices and complexity metrics (Fig. S3A-B, Fig. S4D); molecular ecological network (MEN) construction not stated
Approaches that could also have been used
  • Group differences at each incubation day were tested with ANOVA followed by Tukey's HSD across multiple physicochemical variables, VFAs, and diversity metrics.
    Could also: A mixed-effects or repeated-measures ANOVA (or a similar longitudinal model) treating incubation day as a within-microcosm factor — Since the same experimental system was sampled across five timepoints, a repeated-measures/longitudinal framework can account for time-based correlation structure and may increase power relative to treating each day as an independent ANOVA comparison.
  • Multiple pairwise community-level tests (Adonis, ANOSIM, MRPP) were run across many timepoint/treatment pairs in Table 1 without a stated multiple-testing correction for that family of tests.
    Could also: Applying a false discovery rate procedure (e.g., Benjamini-Hochberg) or Bonferroni correction across the full set of pairwise comparisons — When many pairwise tests are run within one table, an explicit correction step can help control the family-wise or false-discovery error rate across all reported comparisons.
  • Variability in Fig. 1 and Fig. 2 was displayed using standard deviation error bars.
    Could also: Reporting standard error of the mean (SEM) or 95% confidence intervals alongside or instead of SD — SEM/CI convey the precision of the estimated mean specifically, which some readers find more directly interpretable when comparing group means, whereas SD conveys spread of individual replicate values.
  • Significance for many figures is reported as threshold-based stars (*, **, ***) rather than exact p-values.
    Could also: Reporting exact p-values throughout, as already done in Table 1 — Exact p-values allow readers to gauge the precise strength of evidence and to perform their own multiple-testing corrections or meta-analyses, complementing the threshold-star convention.
  • Community composition differences were assessed with three complementary nonparametric methods (Adonis, ANOSIM, MRPP).
    Could also: A distance-based redundancy analysis (db-RDA) or a generalized linear model for compositional data (e.g., ALDEx2, ANCOM-BC) for taxon-level differential abundance — These approaches can directly model compositional (relative-abundance) data structure and covariates, offering an alternative lens to ordination-and-permutation-based dissimilarity testing for identifying which specific taxa drive community differences.
  • LEfSe (LDA effect size) was used to identify differentially abundant taxa between control and SRB-inhibited groups.
    Could also: DESeq2 or edgeR-style differential abundance testing with FDR correction — These count-based methods explicitly model library-size normalization and variance-mean relationships in sequencing data, which can be a useful complement to LEfSe's non-parametric rank-based approach.
Software: R package microeco R version 3.6.3

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36988509

Paper: Liang et al. 2023, mBio 14(2):e03535-22. "Niche Modification by Sulfate-Reducing Bacteria Drives Microbial Community Assembly in Anoxic Marine Sediments." DOI 10.1128/mbio.03535-22. PMCID PMC10128000.

Design: Intertidal sediment microcosms, control vs molybdate (SRB-inhibited), 5 timepoints (days 0,5,12,21,30), 20 replicates. 16S rRNA V3-V4 (338F/806R), Illumina MiSeq PE300. Day 0 has a single shared baseline group → 9 groups × 20 reps = 180 analyzed samples.

Code/data artifacts

  • Authors' own repo (CORRECTED): https://github.com/2015qyliang/InhibitingSRB_molybdate — per-figure folders, each shipping BOTH the input data (OTU table, taxonomy, tree, env tables) AND the R scripts that derive the reported numbers/figures. (The brief's code_url pointed to github.com/hyattpd/Prodigal — that is merely a tool cited in Methods for gene prediction, not the paper's analysis repo.)
  • Data: SRA SRP364228 (BioProject), 16S amplicon. Data-availability statement: "deposited ... under accession no. SRP364228 for 207 samples". Repo ships The list of 207 SRA runs under SRA SRP364228.txt.

Pipelines named in Methods

Step Tool (version)
PE merge FLASH 1.2.11
OTU clustering @97% + chimera UPARSE 7.0 (usearch)
Taxonomy RDP Classifier 2.11 + Silva 138, conf 0.7
Gene prediction (genomes) Prodigal 2.6.3
Functional annot KofamKOALA, KEGG Mapper
Community assembly iCAMP (phylo bin-based null model), IQ-TREE 1.6.12
Networks MENAP (RMT) web server + iDIRECT

IN SCOPE (pipeline-derived, attempted)

  1. OTU table integrity / dimensions — 3,337 OTUs × 180 samples (Results; UPARSE output shipped as Figure_2/02-OTUtable.txt).
  2. Upstream 16S pipeline from SRA (the genuine heavy reproduction): SRP364228 → FLASH → UPARSE/vsearch → high-quality reads (4,072,354), OTU count (3,337), per-sample reads (20,031–21,434), rarefaction (21,530).
  3. Alpha diversity (Observed/Simpson/Shannon + ANOVA/TukeyHSD), Fig 2A-C; per-sample Observed range 341–1,781.
  4. Beta diversity / betapart — turnover vs nestedness, PERMANOVA/ANOSIM/MRPP, NMDS (binary Jaccard) stress, Fig 2D-F.
  5. iCAMP community assembly — relative importance of HoS/DL/etc. ("HoS + DL ≈ 90%"), Cohen's d comparisons, Fig 4. Re-derived from the shipped bootstrap object PD.Boot.rds; the raw null-model run that produced it is not shipped as a script.

OUT OF SCOPE (not pipeline-reproducible here)

  • Wet-lab: physicochemistry (pH, sulfate, VFAs — Fig 1), 305 bacterial isolates, OD600 growth, molybdate-tolerance/metabolite-facilitation assays.
  • MENAP networks (Fig 3/S4): built on an external web server (ieg4.rccc.ou.edu/mena) — not locally re-runnable; only shipped NetInfo/module files can be inspected, not regenerated.
  • Genome/metabolic reconstruction (Fig 5/6, GTDB, Prodigal, KofamKOALA): depends on 3,596 external GTDB genomes + manual KEGG Mapper curation; the per-genome gene-prediction inputs are not shipped. Prodigal itself is runnable but the paper's specific genome set/curation is not reproducible from the repo.

Reproduction strategy

  • Tier A (fast, deterministic): re-run the shipped R scripts on the shipped inputs → verify they regenerate the shipped derived outputs that underlie the paper's reported numbers. Genuine reproduction of the derivation pipeline.
  • Tier B (heavy, bonus): re-run the upstream 16S pipeline from raw SRA to independently regenerate the OTU table / read counts. Exact OTU count is version-sensitive (UPARSE/usearch v7); report with honest tolerance.
Figures / tables: Fig 2Fig 2AFig 2DFig 4
C1_otu_dim
Reported
3,337 OTUs x 180 samples
Reproduced
shipped OTU table = 3,337 rows x 180 sample columns
exact
C2_design
Reported
9 groups x 20 reps = 180 samples
Reproduced
9 groups x 20 = 180 in samplesGroup
exact
C3_alpha_observed
Reported
Observed richness range 341-1781 (Fig 2A)
Reproduced
re-ran microeco cal_alphadiv: Observed min 341 max 1781 (n=180)
exact
C4_alpha_anova
Reported
alpha ANOVA/TukeyHSD Observed/Shannon/Simpson (Fig 2A-C)
Reproduced
NUMMATCH shipped, max abs diff ~4e-12
exact
C5_composition
Reported
adonis/ANOSIM/MRPP over 12 group pairs (Fig 2D-F)
Reproduced
all 12 pairs byte-identical (adonis.F, anosim.r, mrpp.delta)
exact
C6_betapart_adonis
Reported
betapart turnover/nestedness adonis R2 (Fig 2)
Reproduced
R2 exact for jac/tur/nes (C&M); only nestedness perm-p differs (0.078 vs 0.091)
exact
C7_icamp_relimp
Reported
HoS+DL dominate community assembly (~90%); Fig 4
Reproduced
TEXTMATCH shipped; HoS+DL Control=96.72%, Molybdate=96.06%
exact
C8_cohenD
Reported
iCAMP Cohen's d effect sizes (Fig 4)
Reproduced
TEXTMATCH shipped
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is a preliminary, work-in-progress reproduction of an author-friendly study: the authors shipped their own analysis repo (per-figure tables + R scripts) and raw reads are public on SRA. The directly checkable claims match — 180 samples (9x20) exact, 3,337 OTU rows present in the shipped table, and iCAMP shows HoS+DL dominance (C 96.7%, M 96.1%) confirming the paper's ~90% central claim. The open issues are on our/data-definition side, not the authors': the upstream FLASH+UPARSE and iCAMP re-runs are still PENDING, and there is an unresolved 180-analyzed-vs-207-SRA-run sample discrepancy. No fabrication concern; overall solid but incomplete, hence yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

87.4 k
tokens (I/O) · 4.7 M incl. cache
24 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.