Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

GAL08, an Uncultivated Group of Acidobacteria, Is a Dominant Bacterial Clade in a Neutral Hot Spring.

Front Microbiol · 2022
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PROVISIONAL/IN-PROGRESS. Cited code = pyani (ANI of 23 GAL08 SAGs+MAGs); headline result (3 species; within >99.2% / between <88.1% ANI; 95-96% delineation; genome stats 3.17 Mb / 62.8% GC) depends on the 23 assembled genomes which are JGI IMG registered-access only (HTTP 403 anonymous; not mirrored at NCBI; no JGI account available) -> not runnable on the authors' exact data (data_restricted for C1-C3,C7-C8). The cited OPEN accession PRJNA779083 is 32 16S amplicon runs supporting a different pipeline (QIIME) for the abundance claims C4-C6; these are attemptable. Only 7/23 SAGs have open raw reads, allowing at best a partial deviating self-assembly+pyani corroboration. Heavy compute deferred: «our HPC»/VPN down at start (port 22 timeout) - waiting per protocol. This file will be refined once compute runs.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ a08ac43942a3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What is the abundance, distribution, and genomic/metabolic potential of the uncultivated GAL08 clade of Acidobacteria in a pH-neutral hot spring (Dewar Creek, British Columbia), and what is its phylogenetic placement and lifestyle?

Core claims
  • GAL08 is a dominant bacterial clade in a neutral hot spring, comprising up to 29.2% of the community by relative read abundance and up to 4.7 × 10^5 16S rRNA gene copies per gram sediment. finding
  • Absolute and relative abundances of GAL08 peaked at 65°C along three temperature gradients. finding
  • The three GAL08 bacteria represent a distinct subgroup of Acidobacteria, a candidate order (Ca. Frugalibacteriales) within the class Blastocatellia. finding
  • GAL08 genomes predict a heterotrophic metabolism with potential for aerobic respiration, incomplete denitrification, and fermentation. mechanism
  • GAL08 appears to have a microaerophilic lifestyle: qPCR counts declined rapidly under atmospheric oxygen but increased slightly at 1% (v/v) O2. finding
  • 25 SAGs and 7 MAGs were generated representing three separate species based on ANI, with average genome size 3.17 Mb and GC content 62.8%. resource
  • A GAL08-specific qPCR primer set was developed to detect the GAL08 16S rRNA gene. method
  • GAL08 was a consistently dominant clade across multiple sampling years and locations. finding
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA gene amplicon sequencing (Illumina MiSeq) Dewar Creek hot spring sediment, 22.5–79.8°C none relative abundance of GAL08 (OTUs/ASVs) in community Illumina MiSeq, MiSeq Reagent Kit v3 600 cycles; primers 341fw/785rv
qPCR of 16S rRNA genes (GAL08-specific) Dewar Creek and Hoodoo Creek hot spring sediment none absolute 16S rRNA gene copy number per gram sediment Rotor-Gene 6000, SYBR Green qPCR master mix; primers GAL08_842/GAL08_992
Single amplified genome (SAG) generation via single-cell sorting/WGA/sequencing Dewar Creek sediment supernatant, samples at 65 and 77°C none genome assembly, completeness, ANI, metabolic gene content FACS, SPAdes assembly, ProDeGe QC
Metagenome sequencing and MAG reconstruction Dewar Creek sediment at 45, 65, 66, 77°C none binned genomes (MAGs), completeness, contamination IMG/JGI (CheckM); IMG IDs 3300025105, 3300025094, 3300025775, 3300006767
Phylogenetic reconstruction (16S rRNA Bayesian tree; 56-marker concatenated ML tree) DChs_GAL08 sequences vs Acidobacteria/reference genomes none phylogenetic placement of GAL08 MrBayes v3.2.6; IQ-TREE; MAFFT; HMMER; Prodigal
Fluorescence in situ hybridization (FISH) Hoodoo Creek hot spring sediment, 61°C none in situ detection/visualization of GAL08 cells probes GAL08_185_cy3 and EUB338_Alexa488
Laboratory cultivation/enrichment with varied O2 Dewar Creek sediment (2017 samples) oxygen level (atmospheric vs 1% v/v O2) GAL08 abundance change by qPCR
Genome annotation and comparative analysis (ANI, COG, iron genes) GAL08 SAGs and MAGs none metabolic pathway prediction, core/unique COGs, species delineation IMG/JGI; pyani; FeGenie; CheckM
Key results
  • GAL08 relative and absolute abundance peaked at 65°C across three temperature transects
  • GAL08 comprised up to 29.2% of the microbial community by relative read abundance 29.2%
  • GAL08 reached up to 4.7 × 10^5 16S rRNA gene copies per gram sediment by qPCR 4.7 × 10^5 copies/g
  • SAGs and MAGs represented three separate species with average genome size 3.17 Mb and GC content 62.8% 3.17 Mb; 62.8% GC
  • GAL08 qPCR counts declined rapidly under atmospheric O2 but increased slightly at 1% O2
  • GAL08 placed as a distinct subgroup within class Blastocatellia (candidate order Ca. Frugalibacteriales)
  • GAL08 detected at Hoodoo Creek at 3.8 × 10^4 gene copies per gram sediment by qPCR 3.8 × 10^4 copies/g
  • DChs_GAL08 shows closest 16S rRNA gene identity (82–84%) to other phyla (Thermotogae, Aquificae, etc.) 82–84% identity
Key statistics
  • count 29.2% (maximum GAL08 relative read abundance in community)
  • count 4.7 × 10^5 16S rRNA gene copies per gram sediment (maximum GAL08 abundance by qPCR at Dewar Creek)
  • count 3.8 × 10^4 gene copies per gram sediment (GAL08 abundance at Hoodoo Creek (61°C))
  • mean 3.17 Mb (estimated average genome size of three GAL08 species)
  • mean 62.8% (average GC content of GAL08 genomes)
  • count 25 SAGs and 7 MAGs (genomes recovered identified as GAL08)
  • other completeness 97.2% to 14.2% (estimated completeness range of SAGs (CheckM))
  • other 82–84% (16S rRNA gene identity of GAL08 to closest reference phyla)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a descriptive environmental microbiology and comparative genomics study characterizing the uncultivated GAL08 clade in a neutral hot spring. Abundance was quantified using 16S rRNA amplicon sequencing (relative abundance) and qPCR (absolute gene copy numbers per gram sediment) across temperature gradients and multiple sampling years (2010–2017). Phylogenetic placement was established using Bayesian inference (MrBayes) and maximum-likelihood (IQ-TREE) methods on 16S rRNA and concatenated marker gene datasets. No formal inferential hypothesis tests (e.g., ANOVA, t-tests) were performed; all quantitative results are reported descriptively.

Replicationmixed Sample size32 samples for Illumina amplicon sequencing; qPCR samples from 3 separate field trips (2011, 2015, 2017); three parallel stream transects collected on one date (August 24, 2015); 25 SAGs from 2 sampling dates; 7 MAGs from 4 metagenomes; number of qPCR technical replicates not stated GroupsTemperature gradient (24.2–79.8°C); multiple sampling years (2010–2017); three stream transects from a single source; SAG/MAG species groupings (Species 1–3) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Bayesian MCMC phylogenetic inference (MrBayes, GTR+Gamma+I substitution model) 16S rRNA gene-based phylogenetic tree for Acidobacteria placement 153 sequences of ≥1,395 bp; 2×10^6 MCMC iterations, burn-in 1×10^5 stated
Maximum-likelihood phylogenetic inference with bootstrap support (IQ-TREE, 1,000 bootstraps) Concatenated 56-marker gene tree for genomic placement of GAL08 within Acidobacteria and broader bacterial phyla 25 SAGs/7 MAGs + 68 Acidobacteria genomes + 214 reference genomes; 6 Parcubacteria as outgroup not stated
OTU clustering at 97% similarity with BLAST taxonomic classification against Silva 119 (QIIME 1.9.1) Relative abundance of GAL08 in the broader microbial community from 16S rRNA amplicon libraries 32 samples collected 2012–2015 not stated
ASV denoising via DADA2 with Silva 138 classification (QIIME2 2021.4) Identification of all DChs_GAL08 amplicon sequence variants 32 samples collected 2012–2015 not stated
Average nucleotide identity (ANI) pairwise comparison (pyani) Species-level grouping of SAGs and MAGs using a standard 95% ANI threshold 25 SAGs and 7 MAGs not stated
SYBR Green qPCR quantification against a dilution series standard curve (Rotor-Gene 6000) Absolute abundance of DChs_GAL08 16S rRNA gene copies per gram sediment across temperature gradients and years Samples from field trips in 2011, 2015, and 2017; number of technical replicates not stated not stated
Approaches that could also have been used
  • qPCR absolute abundance was reported as single point values (e.g., up to 4.7×10^5 gene copies per gram sediment) without any dispersion measure
    Could also: Report mean ± SD or mean ± SEM across technical replicates, or provide 95% confidence intervals derived from the standard curve regression — Dispersion measures communicate measurement variability and allow readers to judge reproducibility; for qPCR data, reporting variability across technical or biological replicates is standard practice and aids interpretation of abundance differences across temperatures or years
  • The relationship between temperature and GAL08 abundance was described qualitatively (peaked at ~65°C along three transects)
    Could also: Fit a unimodal response model (e.g., Gaussian curve, second-degree polynomial regression, or a GAM) to the temperature–abundance data — A formal model would estimate the thermal optimum and niche width with confidence intervals, enabling quantitative comparison across sampling years and stream transects and placing the result in the context of other thermophile ecology studies
  • Two parallel amplicon analyses were run (OTU-based in QIIME 1.9.1; ASV-based in QIIME2/DADA2), apparently for different purposes
    Could also: Use ASVs (DADA2 or Deblur) as the single primary abundance unit throughout, replacing the 97% OTU clustering step — ASVs provide single-nucleotide resolution, are fully reproducible across studies, and avoid the arbitrary 97% threshold that can merge ecologically distinct populations; the paper already uses DADA2, so unifying on ASVs would harmonize both analyses
  • Relative community composition from amplicon data was reported descriptively without diversity metrics or statistical comparison across temperature bins or sampling years
    Could also: Calculate alpha-diversity indices (e.g., Shannon entropy, observed ASVs after rarefaction) and perform beta-diversity ordination (Bray-Curtis PCoA) with a permutation test (PERMANOVA/adonis) to test temperature or year as predictors — Ordination and PERMANOVA would quantify and statistically support the observed temperature-driven community structuring, making the pattern more directly comparable with other environmental gradient studies
  • Genome completeness and contamination were estimated using CheckM with a universal single-copy marker gene set
    Could also: Cross-validate with a taxon-specific CheckM lineage workflow or with BUSCO using a relevant lineage database — Universal marker sets may be under-represented or biased for deeply divergent, uncultivated lineages; a taxon-specific or complementary completeness tool would provide an independent quality estimate tailored to this novel Acidobacteria clade
  • Species-level delineation of SAGs and MAGs relied solely on pairwise ANI with a 95% threshold
    Could also: Supplement ANI with digital DNA–DNA hybridization (dDDH via the GGDC server) or tetranucleotide frequency-based clustering — Using multiple converging species delineation methods strengthens candidate species proposals; dDDH in particular is a widely accepted complement to ANI in formal prokaryotic taxonomic descriptions and provides an independent probability-based similarity estimate
Software: QIIME 1.9.1 · QIIME2 2021.4 · DADA2 · MrBayes 3.2.6 · MAFFT 1.3.7 and 7.221 · IQ-TREE · CheckM · Prodigal 2.6.3 · HMMER 3.1b2 · pyani · R (ape, ggtree) · Python (Venn diagram) · Microsoft Excel

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35087491

Paper: Ruhl et al. 2022, GAL08, an Uncultivated Group of Acidobacteria, Is a Dominant Bacterial Clade in a Neutral Hot Spring. Front Microbiol 12:787651. PMID 35087491 · PMCID PMC8787282 · DOI 10.3389/fmicb.2021.787651

Cited code: https://github.com/widdowquinn/pyani (third-party ANI tool — valid per P16) Cited data: SRA PRJNA779083 (amplicon 16S rRNA sequencing)

Datasets the paper relies on

dataset what access used for
PRJNA779083 (SRA) 32 Illumina 16S amplicon runs (341f/785r), 2012–2015 temperature transects OPEN abundance / QIIME pipeline (C4–C6)
IMG SAG/MAG genomes (Table 1 taxon IDs) 16 SAGs + 7 MAGs assembled genomes of DChs_GAL08 JGI IMG only — registered access (no anonymous download; taxon-detail pages HTTP 403; no JGI account in secret store) pyani ANI, genome stats, marker-gene tree (C1–C3, C7–C8)
GAL08 SAG raw reads (PRJNA469195/469184/469179, PRJNA364518/364596/364633/364655) 7 single-cell raw read sets ("hot springs metagenome", library OTHER) OPEN (SRA/ENA) only the raw reads for 7 of the 23 SAGs; the assembled genomes used by the paper are IMG-only; all 7 MAGs have no open product

Key mismatch: the cited code (pyani) operates on the assembled genomes, which are deposited in JGI IMG (registered access), while the cited open accession (PRJNA779083) is amplicon data for a different pipeline (QIIME). The assembled genomes are not in NCBI GenBank/WGS (assembly search for "GAL08" = 0 hits; data availability statement: "publicly available in the IMG database").

Reproducible results

In scope — OPEN data (attemptable)

  • C4 GAL08 relative abundance up to 29.2 %; 7.0–29.2 % across the 15 samples at 64.3–67.4 °C; community peak ~65 °C. (Results; Fig 1/2) — pipeline: QIIME2 on PRJNA779083.
  • C5 ASV1 = on average 95.9 % of reads identified as DChs_GAL08. (Results; Supp Fig 6) — QIIME2.
  • C6 At 60–85 °C GAL08 = on average 89.0 % of acidobacterial reads. (Results) — QIIME2.

In scope — but data RESTRICTED (cannot run on authors' genomes)

  • C1 (headline) pyani ANI of all SAGs+MAGs → 3 species-level clusters; species delineation 95–96 % ANI; within-cluster >99.2 %, between-cluster <88.1 %. (Fig 3A) — needs the 23 IMG genomes.
  • C2 estimated average genome size 3.17 Mb, GC 62.8 %. (Abstract/Results) — IMG genomes.
  • C3 16S of Species 1 vs 3 differ by 0.19 % (full-length). (Results) — IMG genomes.
  • C7 80.2 % average 16S divergence to cultivated Blastocatellia. (Results; Fig 5) — IMG genomes + refs.
  • C8 56-concatenated-marker-gene tree, 1000 bootstraps; 153-seq 16S tree. (Fig 4/5) — IMG genomes + refs (heavy).

Partial illustrative path (OPEN raw reads, deviation)

  • 7 GAL08 SAG raw-read sets are open → self-assemble (SPAdes) + run pyani to test whether the open subset reproduces the within/between-species ANI structure of C1. Clearly a deviation (self-assembly vs JGI assembly) and incomplete (7/23, no MAGs), so at best a partial corroboration of the headline.

Out of scope (wet-lab / manual / interpretive)

qPCR 16S copy numbers (4.7×10⁵/g), temperature measurements, FISH, metabolic reconstruction interpretation, candidate taxonomy naming.

Blockers

  • «our HPC»/VPN down at start («host».«infra».uni-hamburg.de:22 connection timed out). Waiting per protocol; not touching the VPN. All heavy compute deferred until tunnel returns.
  • IMG registered access for the 23 assembled genomes → headline ANI not runnable on the authors' exact data.
Figures / tables: Fig 3AFig 1Fig 6Fig 5Fig 4
C1
Reported
pyani ANI -> 3 species clusters; delineation 95-96% ANI
Reproduced
partial
C1b
Reported
within-species ANI >99.2%
Reproduced
partial
C1c
Reported
between-species ANI <88.1%
Reproduced
partial
C2
Reported
avg genome size 3.17 Mb
Reproduced
partial
C2b
Reported
GC 62.8%
Reproduced
partial
C4
Reported
GAL08 relative abundance up to 29.2%
Reproduced
partial
C5
Reported
ASV1 = 95.9% of DChs_GAL08 reads
Reproduced
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

83.1 k
tokens (I/O) · 3.9 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.