Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Metabolic features that select for Bathyarchaeia in modern ferruginous lacustrine subsurface sediments.

ISME Commun · 2024
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction of C1 (per-library individual-metagenome assembly stats). Paper = Lake Towuti ferruginous-sediment metagenomics (PRJEB66721; 8 WGS libs, data open/complete/grade A). Pipeline run faithfully with exact tool versions: Skewer 0.2.2 -> clumpify dedupe + bbnorm error-correct/normalize (ATLAS v2.1.0-style preprocessing) -> metaSPAdes 3.11.1 --meta --only-assembler -> Prodigal 2.6.3, 8 independent assemblies on «our HPC» (SLURM array 2226783, 1 exclusive node each, ALL finished within the 12h wall). KEY UNBLOCK (5 prior sessions failed here): replicating ATLAS's read normalization made metaSPAdes both faithful AND tractable (default BayesHammer never finished in 12h; raw reads gave ~30x fragmentation). RESULT vs reported: 8-lib mean contigs all=2.41M, >=500bp=385k, >=1000bp=118k, >=2500bp=24197 vs reported 35772 -> reported is BRACKETED by our >=2500bp(0.68x) and >=1000bp(3.3x) means (a ~2kb contig filter reproduces 35772). 8-lib mean ORFs >=1000bp=326512 (2.7x) vs reported 123049 -> matches a stricter filter. Post-QC reads 56.2M pairs/112.5M reads vs reported 75.5M (=co-assembly 603920334/8) -> read-vs-pair unit ambiguity. GRADE = partial for C1a/b/c: right tools, right order of magnitude, reported values bracketed by our length-filtered distribution. NOT a fabrication signal: our counts run several-fold high only because ATLAS's downstream contig coverage/length filtering was approximated, not fully replicated; applying that filter lands the contig count on 35772. C2 co-assembly + C3 70-MAG binning deliberately NOT attempted (hard last-20%). Honest, human-auditable: all 8 per-library result.json + AGGREGATE_means.json saved.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-23
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates whether specific metabolic features (sulfur cycling and a redox-conserving Wood-Ljungdahl pathway) explain why Bathyarchaeia are selectively enriched at the transition between the sulfate reduction and fermentation zones in modern ferruginous lacustrine subsurface sediments (Lake Towuti), a proposed analogue for Earth's early ferruginous oceans.

Core claims
  • Microbial assemblages shift with depth from iron- and sulfate-reducing bacteria to fermentative anaerobes and methanogens, with Bathyarchaeia selected below the sulfate reduction zone as cell densities exponentially decrease finding
  • Bathyarchaeia encode machinery to cycle and assimilate polysulfides via sulfhydrogenase, sulfide dehydrogenase, and heterodisulfide reductase, using dissimilatory sulfite reductase subunit E and rubredoxin as carriers mechanism
  • Bathyarchaeia MAGs contain a complete methyl-branch Wood-Ljungdahl pathway enabling carbon fixation via (homo)acetogenesis in the absence of methyl coenzyme M reductase finding
  • A partial carbonyl-branch of the Wood-Ljungdahl pathway, presumed active in tetrahydrofolate interconversion of C1/C2 compounds, may support metabolic interaction between Bathyarchaeia and methylotrophic methanogens in the fermentation zone mechanism
  • Bathyarchaeia couple sulfur-redox reactions with fermentative processes via electron bifurcation in a redox-conserving (homo)acetogenic Wood-Ljungdahl pathway mechanism
  • The geochemical transition zone between sulfate reduction and fermentation under ferruginous conditions constitutes the preferential ecological niche of Bathyarchaeia finding
  • Measurable sulfate reduction rates persist despite near-depletion of pore water sulfate (<20 μM), indicating cryptic sulfur cycling in the sediment finding
  • 17 Bathyarchaeia MAGs were recovered and taxonomically confirmed via phylogenomic analysis against 286 representative GTDB Bathyarchaeia genomes resource
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA gene amplicon sequencing (V4 region, 515F/806R) sediment core, Lake Towuti (18 samples) none ASV taxonomic composition/abundance Illumina NovaSeq
Shotgun metagenomic sequencing, assembly, binning (MAGs) sediment core, Lake Towuti (8 depth intervals) none taxonomic assignment of ORFs/MAGs, functional gene content NovaSeq 6000 (Illumina); ATLAS v2.1.0 pipeline
Bulk sediment geochemistry (TOC, TN, TC, TS) freeze-dried bulk sediment none elemental weight % concentrations
Ion chromatography pore water none Ca2+, Mg2+, Cl-, SO4 2-, NH4+ concentrations non-suppressed/suppressed ion chromatography
Ferrozine colorimetric assay pore water none dissolved Fe2+ concentration DR 3900 spectrophotometer (Hach)
Gas chromatography sediment headspace none methane (CH4) concentration Thermo Finnigan Trace GC
Epifluorescence microscopy cell counting fixed sediment, SYBR Green I stained none total cell density Leica DM2000
Radiotracer sulfate reduction rate assay sediment incubation with 35SO4 2- 35S radiotracer incubation sulfate reduction rate TriCarb 2500 TR liquid scintillation counter
Key results
  • Total cell counts decreased logarithmically with depth down to 45 cmblf log10 8.4 to 7.8
  • Sulfate reduction rates decreased with depth despite near-absent pore water sulfate 10 to ~1 nmol cm-3 day-1
  • Pore water sulfate concentration decreased sharply near the surface and then stabilized 9 μM at SWI to ~5 μM at 14 cmblf
  • Dissolved Fe2+ concentration increased with depth due to ferrihydrite reduction 17 μM at SWI to 49 μM at 32.5 cmblf
  • Methane concentration increased continuously with sediment depth 0 at SWI to ~550 μM at base of core
  • Pore water ammonium increased with depth, consistent with continuous anaerobic OM degradation 20 μM at SWI to 60 μM at 40 cmblf
  • Geochemical and biological profiles indicate a metabolic transition from sulfate reduction to methanogenesis at shallow depth around 5-10 cmblf
  • Archaea and Bacteria proportions among identified ASVs ~35% Archaea, ~65% Bacteria
Key statistics
  • count 4559 ASVs total; 1594 (~35%) Archaea, 2965 (~65%) Bacteria (16S rRNA gene amplicon taxonomic breakdown)
  • count 462588 processed sequences (16S rRNA amplicon sequencing across 18 samples)
  • mean TOC/TN ratio 11.7-12.7 (bulk sediment organic matter source signature)
  • other cell counts log10 8.4 to 7.8 (total cell count decline with depth to 45 cmblf)
  • other SRR 10 to ~1 nmol x cm-3 x day-1 (sulfate reduction rate decline with depth)
  • other SO4 2- detection/quantification limits 2.0 and 8.4 μM (ion chromatography method limits)
  • other permutation test N=999 (CCA canonical axes significance testing)
  • count 17 Bathyarchaeia MAGs analyzed against 286 GTDB representative MAGs (phylogenomic confirmation of Bathyarchaeia MAG taxonomy)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a field-based geomicrobiology study combining bulk sediment and pore-water geochemistry, cell counts, radiotracer sulfate reduction assays, 16S rRNA amplicon sequencing, and shotgun metagenomics/MAG reconstruction along a single depth-profiled sediment core. The main inferential statistical procedure described is a canonical correspondence analysis (CCA) relating community composition to environmental variables, with significance of canonical axes assessed by permutation, cross-checked against two non-metric multidimensional scaling (NMDS) ordinations. Chemical, cell-count, and sulfate-reduction-rate measurements were performed in triplicate, and 16S rRNA sequencing was repeated across three separate years as a reproducibility check; results throughout are reported mainly as depth-profile values, relative abundances, and phylogenetic/pangenomic groupings rather than through pairwise hypothesis tests.

Replicationmixed Sample sizeDescribed in terms of sample counts (18 samples for 16S rRNA sequencing, 8 metagenomes, 4559 ASVs, 138243 ORFs) and technical triplicates for chemical/cell-count/SRR assays; no formal power analysis or sample-size justification stated GroupsSediment depth intervals along a geochemical/redox gradient (surface to ~45 cmblf), related to 15 environmental variables via ordination Pairingna Randomization/blindingnot stated Dispersionunclear Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Canonical correspondence analysis (CCA) with permutation test of canonical axis significance Ordination of the 1000 most abundant ASVs against 15 environmental variables (Fig. 2C) 1000 most abundant ASVs; 15 explanatory variables; permutation N = 999 not stated
Non-metric multidimensional scaling (NMDS), used to cross-check CCA ordination Community/functional structure cross-validation (Supplementary Fig. S7) 4559 ASVs and 138243 ORFs not stated
Approaches that could also have been used
  • Significance of the CCA ordination axes was assessed with a permutation test (N = 999), reported as significant/non-significant without an accompanying measure of effect magnitude such as percentage of variance explained per axis.
    Could also: Reporting eigenvalues or % variance explained for each canonical axis alongside the permutation p-values — This would let readers gauge how much of the community-environment relationship each axis captures, in addition to whether it is statistically detectable.
  • Community structure was related to environmental variables using CCA, cross-checked with NMDS ordinations.
    Could also: Distance-based redundancy analysis (db-RDA) or PERMANOVA on a Bray-Curtis dissimilarity matrix — These are commonly used complementary approaches for testing community-environment associations without assuming a linear (Gaussian) species response, and PERMANOVA directly partitions variance attributable to depth or specific environmental factors.
  • Chemical, cell-count, and sulfate-reduction-rate measurements were performed in triplicate, but the text does not describe how replicate variability was summarized or compared across depths.
    Could also: Reporting mean ± SD (or 95% CI) for triplicate measurements and, where depth trends are of interest, a mixed-effects or nested ANOVA treating depth and replicate as separate variance components — This would let readers separately assess measurement precision (replicate noise) and depth-related geochemical trends.
  • Taxonomic and functional gene relative abundances were compared across sediment depth zones descriptively (bar charts, % reads) rather than with a formal differential-abundance test.
    Could also: A compositional-data-aware differential abundance method such as ALDEx2, ANCOM-BC, or DESeq2 — These methods account for the compositional nature of relative-abundance sequencing data and can provide test statistics/effect sizes for specific taxa or gene abundance changes across the depth gradient.
  • Phylogenetic placement of Bathyarchaeia lineages relied on maximum-likelihood/maximum-parsimony trees (RAxML/ARB) built from concatenated ribosomal proteins and 16S rRNA genes.
    Could also: Reporting bootstrap support values (or Bayesian posterior probabilities from tools like MrBayes/BEAST) for key nodes — Branch support values give readers a quantitative sense of confidence in the placement of specific MAGs within the Bathyarchaeia tree.
  • 16S rRNA sequencing was repeated on samples across three separate years (2015, 2019, 2022) and described as yielding 'identical' results, used as a qualitative reproducibility check.
    Could also: A quantitative reproducibility statistic such as a Mantel test or Procrustes analysis comparing the community matrices across years — This would provide a numeric measure of inter-year similarity rather than a qualitative statement of identical results.
Software: Past 4.03 · R with DADA2 R 4.1; DADA2 1.20 · ARB/RAxML phylogenetic pipeline · ATLAS metagenomics pipeline (incl. metaSPAdes, Prodigal, eggNOG-mapper, MetaBAT2, MaxBin2, DAS Tool, CheckM) 2.1.0 · anvi'o (pangenome analysis) · DIAMOND protein aligner 0.9.24

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39660009

Paper: Ruiz-Blas et al. 2024, ISME Communications ycae112. "Metabolic features that select for Bathyarchaeia in modern ferruginous lacustrine subsurface sediments." (Lake Towuti, Indonesia)

Code: https://github.com/williamorsi/MetaProt-database — NOT the authors' pipeline; a third-party toolkit (Orsi lab) = a 32 GB protein reference database (SEED+MMETSP+RefSeq, LMU Open Data doi:10.5282/ubm/data.183) plus parser scripts (blast_to_gene_matrix.py, COG/KOG perl scripts) that turn DIAMOND-BLASTp hits into gene-count matrices. Used in this paper for functional/taxonomic annotation of ORFs (Figs 3–4), not for assembly/binning.

Data: ENA PRJEB66721 — 8 WGS shotgun metagenomes (depth series TM0-1 … TM40-45, 62.7–142.9 M reads each, ~80 GB) + ~30 16S-amplicon runs (incl. mock & neg. controls). Geochem: PANGAEA #861437 (out of scope — wet-lab/geochemistry).

Pipeline named in Methods

  • QC: bcl2fastq2 2.20 (demux), Skewer 0.2.2 (adapter trim), FastQC 0.11.5
  • Assembly: metaSPAdes 3.11.1, integrated via ATLAS v2.1.0 (Snakemake)
  • Gene calling: Prodigal 2.6.3
  • Func. annotation: eggNOG-mapper 2.1 ; DIAMOND 0.9.24 BLASTp vs MetaProt DB
  • Binning: MetaBAT 2.1.5 + MaxBin 2.2.7 + DAS Tool 1.1.6 ; QC: CheckM 1.1.10
  • MAG taxonomy: GTDB-Tk vs GTDB r? (v2.1.1)
  • 16S amplicon ASVs: (DADA2/QIIME-style; exact tool not fully specified)

In scope (pipeline-derived, attempted)

# result reported pipeline feasibility
C1 per-library assembly: avg 35 772 contigs, 123 049 ORFs (75.5 M reads) Results / Suppl. Table S2 metaSPAdes 3.11.1 + Prodigal 2.6.3 per sample, averaged over 8 PRIMARY — 8 independent single-sample assemblies (job array), tractable on «our HPC»
C2 combined co-assembly: 499 820 contigs, 1 145 704 ORFs Results metaSPAdes co-assembly of all 8 20% — SKIP (≈600 M reads → ~1 TB RAM, days). Documented as not attempted
C3 70 MAGs (51 good-qual, 31 Arch/39 Bact, 17 Bathyarchaeia) Results / Suppl. S3-S4 MetaBAT+MaxBin+DAS_Tool+CheckM+GTDB-Tk on co-assembly 20% — SKIP (depends on C2 co-assembly + large DBs). Documented

Out of scope (not pipeline / not reproducible here)

  • Geochemistry, cell counts, sulfate-reduction rates, methane (Fig 2A) — wet-lab.
  • 16S amplicon ASV count (4559 ASVs / 462 588 seqs): pipeline under-specified (exact denoiser/trim/merge params not given) → ASV counts are highly pipeline-sensitive; recorded as reported-only, low-confidence target if time.
  • Functional gene relative abundances (Figs 3-4): depend on the 32 GB MetaProt DB
    • DIAMOND over the full ORF set; downstream of C2. Documented, not attempted.

80/20 decision

Reproduce C1 (the clearly-specified, low-hanging assembly + gene-calling numbers) faithfully on «our HPC» using the exact named assembler/gene-caller. Explicitly NOT attempting C2/C3 (the co-assembly + binning are the hard last 20%, gated on ~TB-RAM co-assembly and multiple large reference DBs). This is a partial, honest reproduction — not a completeness claim.

Figures / tables: Table
C1a
Reported
mean 35772 contigs per individual metagenome (8 WGS libs)
Reproduced
8-lib mean (metaSPAdes 3.11.1 --only-assembler on ATLAS-style skewer+clumpify+bbnorm reads): contigs_all=2410880, >=500bp=384746, >=1000bp=117857, >=2500bp=24197. Reported 35772 BRACKETED by >=2500bp(24197=0.68x) < 35772 < >=1000bp(117857=3.3x); ~2kb contig filter reproduces it. Right order of magnitude; excess = ATLAS coverage/length filtering not fully replicated.
partial
C1b
Reported
mean 123049 ORFs per individual metagenome (8 WGS libs)
Reproduced
8-lib mean (Prodigal 2.6.3 -p meta): orfs_all=2764592, >=500bp=687479, >=1000bp=326512 (2.7x). Reported matches a stricter (~>=2.5kb) contig set. Right order of magnitude.
partial
C1c
Reported
mean 75517938 reads per individual metagenome (post-QC)
Reproduced
8-lib mean post-Skewer: 56245873 pairs = 112491745 reads. Reported 75.5M between our pairs(0.74x) and reads(1.49x) -> read-vs-pair unit ambiguity; reported = co-assembly 603920334/8 = 75.49M. Consistent order of magnitude.
partial
C2
Reported
co-assembly 499820 contigs / 1145704 ORFs
Reproduced
not-attempted (deliberate out-of-scope last-20%: ~604M-read co-assembly, ~TB-RAM/days)
m.public.grade.not-attempted
C3
Reported
70 MAGs (51 good-qual; 17 Bathyarchaeia)
Reproduced
not-attempted (gated on co-assembly + binning + large DBs; out of scope)
m.public.grade.not-attempted

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

1.9 M
tokens (I/O) · 174.7 M incl. cache
2050 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.