Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genomic Description of 'Candidatus Abyssubacteria,' a Novel Subsurface Lineage Within the Candidate Phylum Hydrogenedentes.

Front Microbiol · 2018
L1 91/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
91/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 82% of all assessed papers rank 197 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduced 1:1 (within version-driven tolerance). P16 approach: the paper's two MAGs are deposited as GenBank assemblies (SURF_5=GCA_003598085.1, SURF_17=GCA_003598055.1, BioProject PRJNA355136), so we re-derived genome statistics and re-ran CheckM lineage_wf/qa + Prodigal on the deposited FASTAs instead of re-running the non-deterministic Trimmomatic->IDBA-UD->CONCOCT->Anvi'o(manual) assembly. RESULT: all 14 in-scope claims reproduced. 8/8 genome-structure claims exact/within-tol with longest scaffolds bit-for-bit identical (208,377 & 111,957 bp) -> deposited assemblies ARE the paper's MAGs (no fabrication of genome stats). CheckM headline numbers reproduced within-tol: completeness 97.74% vs 96-97% (SURF_5) and 90.05% vs 91% (SURF_17) -- both <1pp off; contamination 3.30%/2.20% vs reported 4%/4% (CheckM completeness/contamination are version/marker-set dependent; paper used an earlier CheckM). Gene counts within 0.5-2.7% (4,359 vs 4,482; 4,086 vs 4,105). FIX vs prior requeued run: env build cached on shared «infra» prefix, all downloads via python-urllib (compute nodes lack curl), and numpy pinned <2 to defeat the CheckM np.float64()/ast.literal_eval crash. NOT ATTEMPTED (out of scope): re-assembly from raw reads, RAxML 16-ribosomal-protein phylogeny, ANI/AAI (non-deterministic or not single pinnable numbers).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 76
    assessed: 2026-06-16 ⛓ 44e37a6052f4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates the phylogenetic placement and metabolic potential of two metagenome-assembled genomes recovered from deep terrestrial subsurface fluids that could not be classified into any known bacterial taxon, testing whether they represent a novel lineage within the candidate phylum Hydrogenedentes with distinct adaptations to subsurface biogeochemical cycling.

Core claims
  • SURF_5 and SURF_17 are the first full genomes of a novel bacterial lineage, 'Candidatus Abyssubacteria,' within the candidate phylum Hydrogenedentes finding
  • Members of 'Ca. Abyssubacteria' are globally distributed in both marine and terrestrial subsurface environments finding
  • Metabolic reconstruction suggests versatile capability including nitrogen reduction, sulfite oxidation, sulfate reduction, and homoacetogenesis finding
  • Two high-quality MAGs were reconstructed using differential-coverage binning (CONCOCT) with manual curation in Anvi'o method
  • Phylogeny based on concatenated ribosomal proteins is more robust than single 16S rRNA gene phylogeny for classifying this lineage mechanism
  • Both genomes encode a complete reductive acetyl-CoA (Wood-Ljungdahl) carbon fixation pathway finding
  • Low AAI values (45-55%) between 'Ca. Abyssubacteria' and other Hydrogenedentes genomes suggest the SURF genomes represent a novel phylum or order finding
  • The genomes lack complete canonical secretion systems but encode a Sec-SRP system and motility/chemotaxis machinery finding
Experimental setups
Assay System Perturbation Readout Platform
whole-genome shotgun metagenomic sequencing deep subsurface fracture fluids (SURF, former Homestake gold mine, South Dakota) none paired-end genomic DNA sequence reads Illumina HiSeq 2500
de novo co-assembly and differential-coverage genome binning co-assembled metagenome from two fluid samples none MAG completeness and contamination IDBA-UD, CONCOCT, Anvi'o, CheckM
16S rRNA gene phylogenetic/BLAST analysis SURF_17 MAG (16S rRNA not recoverable from SURF_5) none sequence identity to isolates and environmental clones SILVA Incremental Aligner, NCBI BLAST
concatenated ribosomal protein phylogenomics SURF_5 and SURF_17 genomes plus ~18,000 public environmental genomes none maximum-likelihood phylogenetic placement Prodigal, HMMER, MUSCLE, TrimAl, FastTree, RAxML
pairwise average nucleotide/amino acid identity and POCP comparison SURF_5 and SURF_17 vs three closest Hydrogenedentes genomes and vs each other none ANI, AAI, and percentage of conserved proteins ChunLab ANI Calculator, CompareM
comparative metabolic pathway reconstruction/annotation SURF_5, SURF_17 genomes vs three related Hydrogenedentes genomes none presence/completeness of metabolic gene sets KEGG/KOALA (KEGGDecoder)
global 16S rRNA BLAST distribution survey NCBI nr database environmental and isolate sequences none geographic/environmental source of top matching sequences NCBI BLAST
Key results
  • SURF_5 and SURF_17 MAGs recovered at 96-97% and 91% completeness with 4% contamination each, qualifying as high-quality genomes
  • SURF_17 16S rRNA gene showed 80-83% identity to Deltaproteobacteria, Gammaproteobacteria, and Firmicutes with no taxonomic consensus 80-83%
  • Concatenated ribosomal protein phylogeny placed SURF_5 and SURF_17 within candidate phylum Hydrogenedentes, closest to three other Ca. Hydrogenedentes genomes
  • ANI values between Abyssubacteria and other Hydrogenedentes genomes were below 70%, too low to be reliable <70%
  • AAI values between Abyssubacteria and other Hydrogenedentes genomes fell in a relatively low range 45-55%
  • SURF_5 and SURF_17 showed AAI of ~65% and POCP of 67%, indicating same family/genus-level classification ~65% AAI; 67% POCP
  • 16S BLAST hits >90% identical to SURF_17 were exclusively from deep subsurface environments; the single hit >98% identical came from Zacatón sinkhole, Mexico >90%; >98%
  • Both genomes contain the complete reductive acetyl-CoA (Wood-Ljungdahl) pathway plus nitrate transport/reduction and sulfur metabolism genes
Key statistics
  • count 96-97% completeness, 4% contamination (SURF_5 MAG quality)
  • count 91% completeness, 4% contamination (SURF_17 MAG quality)
  • other 80-83% identity (SURF_17 16S rRNA BLAST identity to Deltaproteobacteria/Gammaproteobacteria/Firmicutes)
  • other <70% (ANI between Abyssubacteria and other Hydrogenedentes genomes)
  • other 45-55% (AAI between Abyssubacteria and other Hydrogenedentes genomes)
  • other ~65% (AAI between SURF_5 and SURF_17)
  • other 67% (POCP between SURF_5 and SURF_17)
  • count top 50 BLAST results, 84-98% identity (global distribution survey of SURF_17 16S rRNA relatives)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a descriptive comparative genomics study rather than a hypothesis-testing study; it reconstructs two metagenome-assembled genomes (SURF_5 and SURF_17) and characterizes a proposed novel lineage. The 'statistical' approach is centered on bioinformatic and phylogenetic estimation: maximum-likelihood phylogenetic inference from concatenated ribosomal proteins, pairwise genome-similarity metrics (ANI, AAI, POCP), completeness/contamination estimation from marker-gene suites, and percent-identity comparisons of 16S rRNA sequences. Results are reported as point estimates (percentages, identity values) and phylogenetic support values, with thresholds drawn from published taxonomic standards; no group-comparison significance tests are reported.

Replicationunclear Sample sizeTwo MAGs (SURF_5, SURF_17) co-assembled from two deep fracture-fluid samples collected at ~1.5 km below surface; no power/sample-size justification stated GroupsTwo SURF genomes vs three closest 'Ca. Hydrogenedentes' reference genomes Pairingna Randomization/blindingna Dispersionnone Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Maximum-likelihood phylogenetic inference (FastTree with gamma + LG; RAxML with GTR substitution model under gamma and invariable rate-heterogeneity models), support from 1000 replicates Figure 1, concatenated 16-ribosomal-protein bacterial phylogeny 16 concatenated ribosomal proteins; ~18,000 publicly available genomes screened then culled stated
Pairwise average nucleotide identity (ANI; one-way and two-way/reciprocal best hits) SURF_5/SURF_17 vs three closest 'Ca. Hydrogenedentes' genomes (Supplementary Data File 2) null stated (thresholds: >95% same species, <75% unreliable)
Pairwise average amino acid identity (AAI) via CompareM SURF genomes vs other 'Ca. Hydrogenedentes' genomes; SURF_5 vs SURF_17 null stated (45–55% range interpreted via published cutoffs)
Percentage of conserved proteins (POCP) via pairwise BLAST SURF_5 vs SURF_17 null stated (per Qin et al., 2014)
16S rRNA percent-identity comparison (SILVA aligner; NCBI BLAST against nr and ref_seq) SURF_17 16S phylogeny and global distribution (Supplementary Figure 1, Figure 4) top 50 BLAST results (Figure 4) stated (>90% family, >98% species cutoffs cited)
Genome completeness/contamination estimation from marker-gene suites Table 1 MAG statistics five marker gene suites stated (quality thresholds per Bowers et al., 2017)
Approaches that could also have been used
  • Branch support was estimated from 1000 replicates (reported as support values on the ML tree).
    Could also: Reporting the support method explicitly (e.g., RAxML rapid bootstrap, SH-aLRT, or FastTree's Shimodaira–Hasegawa local supports) and/or pairing ML with a Bayesian inference (e.g., MrBayes/PhyloBayes posterior probabilities) would also be a standard option. — Naming the support metric and cross-checking topology with a second inference framework would add interpretability and an independent line of evidence for the placement of the novel lineage.
  • The phylogeny was built from concatenated ribosomal proteins under a single ML model (GTR/LG with gamma).
    Could also: Model selection (e.g., ModelTest/ProtTest) and partitioned or mixture models across the concatenation could also be applied. — Per-partition model fitting can better accommodate rate variation among genes and is commonly used to support robustness of deep phylogenetic placements.
  • Taxonomic rank was inferred by comparing ANI/AAI/POCP point values against published threshold cutoffs.
    Could also: Reporting a range or distribution of pairwise values across multiple reference genomes, or noting uncertainty intervals around the metrics, could also accompany the point estimates. — Conveying the spread alongside thresholds helps communicate confidence, especially given the noted undersampling of the candidate phylum.
  • Genome completeness and contamination were estimated from marker-gene suites and reported as single percentages.
    Could also: Reporting estimates from more than one tool/lineage-specific marker set side by side (e.g., CheckM lineage-specific vs domain-level) is also a common practice. — Showing concordance across estimators conveys the robustness of the completeness/contamination figures for the MAGs.
  • Pathway/metabolic completeness was summarized as percentages of requisite genes present (Figure 3 heatmap).
    Could also: Pairing presence/absence calls with annotation-confidence indicators or bootstrap-style resampling of gene calls could also be reported. — Adding a measure of annotation confidence would help readers gauge how firmly each putative pathway is supported across the compared genomes.
  • Global distribution was characterized using the top 50 BLAST hits ranked by percent identity.
    Could also: A phylogenetic placement of environmental sequences (e.g., tree-based or probabilistic placement such as pplacer/EPA) could also complement raw percent-identity ranking. — Tree-based placement can corroborate identity-based groupings and reduce sensitivity to alignment length or local similarity when describing biogeographic distribution.
Software: Trimmomatic 0.36 · IDBA-UD 1.1.1 · Bowtie2 · BWA (BWA-SAMPE algorithm) · SAMtools 0.1.17 · CONCOCT · Anvi'o · CheckM · SILVA Incremental Aligner (SINA) · NCBI BLAST · Prodigal · HMMER (hmmsearch) · MUSCLE · TrimAl (-automated1) · FastTree · RAxML · ChunLab ANI Calculator / CompareM

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
46
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA355136 BioProject in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
SAMN08498999 BioSamples in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
SAMN08499011 BioSamples in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-30210471

Paper: Momper, Aronson, Amend (2018). Genomic Description of 'Candidatus Abyssubacteria', a Novel Subsurface Lineage Within the Candidate Phylum Hydrogenedentes. Front Microbiol 9:1993. PMCID PMC6121073.

Brief's code repo: github.com/Ecogenomics/CheckM · Data: SRA PRJNA355136.

What the paper reports (all from Table 1 + Results text)

Two metagenome-assembled genomes (MAGs), SURF_5 and SURF_17, from SURF (South Dakota) subsurface fluids, BioProject PRJNA355136.

Metric SURF_5 SURF_17 source
CheckM completeness 96–97 % 91 % Results + Table 1
CheckM contamination 4 % 4 % Results + Table 1
Genome size 5.1 Mbp 4.6 Mbp Table 1
GC content 54.7 % 54.8 % Table 1
# scaffolds 145 144 Table 1
longest scaffold 208,377 bp 111,957 bp Table 1
# genes 4,482 4,105 Table 1
16S rRNA present No Yes Table 1

Pipeline that produced them

Trimmomatic 0.36 (Q40, minlen36) → IDBA-UD 1.1.1 (min contig 10 kb) → Bowtie2 mapping → CONCOCT binning → Anvi'o manual refinement → CheckM for completeness/contamination + 16S detection. Phylogeny: 16 concatenated ribosomal proteins (Hug et al. 2016) + RAxML.

IN SCOPE (reproduce) — pipeline-derived, clearly specified, deterministic

The two MAGs are deposited as GenBank assemblies (the cheap, honest P16 path — validate the deposited genomes rather than re-run a non-deterministic assembler):

  • SURF_5 = GCA_003598085.1 (ASM…)
  • SURF_17 = GCA_003598055.1

On the deposited FASTAs we recompute:

  1. CheckM completeness & contamination (checkm lineage_wf) — the headline numbers and the exact tool named in the brief. Primary target.
  2. Genome size, GC %, # scaffolds, longest scaffold — deterministic from the FASTA; should match Table 1 ~exactly.
  3. # predicted genes (Prodigal, as CheckM/Prokka would call them) — compare to the reported gene count (caveat: annotation-pipeline dependent).

OUT OF SCOPE (descoped last-20%, not attempted) — why

  • Re-assembly from raw SRA reads (Trimmomatic→IDBA-UD→CONCOCT→Anvi'o): binning is stochastic and the Anvi'o refinement was manual/interactive → not 1:1 reproducible; re-deriving the exact MAGs is not honestly feasible. We instead validate the authors' deposited MAGs (equally valid under P16).
  • 16S rRNA presence: CheckM's ssu_finder would give it, but it is a yes/no flag dependent on the same deposited assembly; we report it if CheckM surfaces it, otherwise treat as low-priority.
  • Phylogenomic tree / ANI / AAI (RAxML, 16 ribosomal proteins): topology and AAI ranges are not single pinnable numbers and depend on a custom reference set → not attempted.

Compute plan

All on «our HPC» (SLURM, partition std, no --mem). conda env with checkm-genome (pulls pplacer/hmmer/prodigal); CheckM 2015 reference DB downloaded in-job to «infra». Download the two GCA assemblies in-job (compute node has internet). Small result tables pulled back to «host» dataset folder.

Figures / tables: Table
S5_size
Reported
5.1 Mbp
Reproduced
5.103 Mbp (5,103,455 bp)
exact
S5_gc
Reported
54.7 %
Reproduced
54.76 %
within tolerance
S5_scaf
Reported
145
Reproduced
145
exact
S5_long
Reported
208,377 bp
Reproduced
208,377 bp
exact
S17_size
Reported
4.6 Mbp
Reproduced
4.642 Mbp (4,642,259 bp)
exact
S17_gc
Reported
54.8 %
Reproduced
54.89 %
within tolerance
S17_scaf
Reported
144
Reproduced
144
exact
S17_long
Reported
111,957 bp
Reproduced
111,957 bp
exact
S5_comp
Reported
96-97 %
Reproduced
97.74 %
within tolerance
S5_cont
Reported
4 %
Reproduced
3.30 %
within tolerance
S17_comp
Reported
91 %
Reproduced
90.05 %
within tolerance
S17_cont
Reported
4 %
Reproduced
2.20 %
within tolerance
S5_genes
Reported
4,482
Reproduced
4,359
within tolerance
S17_genes
Reported
4,105
Reproduced
4,086
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 91/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0

All four deterministic genome-structure metrics (size, GC, #scaffolds, longest scaffold) reproduce exact or within-tolerance for both MAGs, with longest-scaffold lengths bit-for-bit identical to Table 1 — the deposited GenBank assemblies are demonstrably the paper's genomes, and there is no fabrication concern for the reported stats. The only deviations are rounding-level GC offsets (+0.06/+0.09 pp). The headline CheckM completeness/contamination and gene counts were not completed within the time-box due to our environment/build limitations (our side, not the authors'), so overall reproduction is solid but partial rather than fully clean.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

262 k
tokens (I/O) · 15.9 M incl. cache
76 min
runtime · 0.27 CPU-h
34.1 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine