Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

What defines a photosynthetic microbial mat in western Antarctica?

PLoS One · 2025
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (provisional, P16 third-party-tool-on-own-data). Ran SingleM 0.21.3 + fastp 1.3.4 (modern metapackage S6.5.0/GTDB_r232) on all 14 public Antarctic mat metagenomes (PRJEB14287 9 + PRJEB12762 5 = 14, ~109 Gbp, N concordant), with the paper's exact fastp QC. All three SingleM alpha-diversity claims reproduce within the paper's own reported uncertainty: C1 species OTUs 8964 vs 9414 (-4.8%, within-tol; gap explained by newer/larger reference), C2 Shannon 4.61+/-0.82 vs 4.98+/-0.83 (within 1 SD; SD matches), C3 Simpson 0.946+/-0.054 vs 0.94+/-0.05 (essentially exact). Diversity computed at genus level from the 14 ribosomal markers; alternative aggregation levels recorded for audit (paper's 4.98 brackets the OTU-table-genus method, adopted as primary). NOT attempted (out of scope, compute-prohibitive / different tools): Kaiju taxonomy, MegaHit 109-Gbp assembly + Prodigal/eggNOG, EukDetect/Metaxa2, PCA/NMDS/Mantel/STAMP. fastp QC pass-rate (~83% mean) corroborates the paper's ~80%.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-20 ⛓ a3c9d93e8736
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study investigates what taxonomically and functionally defines photosynthetic microbial mats across western Antarctica, testing whether mats from different regions (Maritime Antarctica, Antarctic Peninsula, McMurdo Dry Valleys) share common compositional and functional characteristics despite varying environmental conditions.

Core claims
  • Taxonomic composition of Antarctic microbial mat communities is characterized by similar bacterial groups across regions finding
  • Diatoms are the main taxonomic factor distinguishing rapidly warming Maritime Antarctica mats from Peninsula and Dry Valleys mats finding
  • Bacteria are the predominant component (>90%) of all microbial mats, followed by Eukarya, Archaea, and Viruses finding
  • All mats, despite varied environmental characteristics across sites, show nitrogen limitation and share functional patterns finding
  • This is the first study to analyze western Antarctica microbial mats at a continental scale using shotgun metagenomic sequencing coupled with physicochemical characterization method
  • Certain microeukaryotes identified may play essential roles in the functioning of Antarctic microbial mats finding
Experimental setups
Assay System Perturbation Readout Platform
shotgun metagenomic sequencing 14 microbial mats from meltwater streams, western Antarctica (Maritime, Peninsula, Dry Valleys) none taxonomic composition and functional gene abundance (GPM) Illumina HiSeq2x150 (Nextera DNA Flex library prep)
taxonomic classification (Kaiju) quality-filtered metagenomic reads from 14 microbial mats none taxonomic profile at genus/phylum level against NCBI nr database Kaiju v1.9.2
alpha diversity profiling (marker genes) 14 microbial mat metagenomes none OTU tables / Shannon-Wiener and Simpson diversity indices from 14 ribosomal proteins SingleM; vegan R package
eukaryote detection from metagenomic reads 14 microbial mat metagenomes none presence/identification of eukaryotic taxa Eukdetect; Metaxa2
metagenome assembly and functional annotation 14 microbial mat metagenomes none gene clusters, orthology assignments, functional/metabolic capacity MegaHit v1.2.9; Augustus v3.5.0; Prodigal v2.6.3; eggNOG Mapper v2 with DIAMOND
rRNA gene phylogenetics microbial mat metagenomes (focus on Adineta vaga) none predicted rRNA sequences, phylogenetic placement Barrnap; ACT Silva; cd-hit-est; MAFFT; BLASTn; RAxML
water nutrient analysis (NH4+, NO3-, NO2-, SRP, SRSi) overflowing water from 14 mat sampling sites none dissolved nutrient concentrations (µM), DIN, DIN:SRP Skalar San Plus continuous-flow autoanalyzer
elemental analysis (C, N, N:P) of mat biomass 14 microbial mat biomass samples none % carbon, nitrogen, phosphorus content PerkinElmer 2400 Elemental Analyzer; Valderrama high-temperature persulfate oxidation method
Key results
  • Bacteria were the predominant component of all microbial mats >90%
  • Eukarya represented the second most abundant domain >3%
  • Archaea were a minor component of the mats <1%
  • Viruses were the least abundant component detected <0.1%
  • Bacteroidota and Pseudomonadota were the dominant bacterial phyla across mats, with Cyanobacteriota also prominent Bacteroidota 35%, Pseudomonadota 29%, Cyanobacteriota 19%
  • Diatoms (Bacillariophyta) distinguished Maritime Antarctica mats from Peninsula and Dry Valleys mats average 2% abundance
  • All sampled mats exhibited nitrogen limitation and shared functional patterns despite differing environmental characteristics
Key statistics
  • mean >90% (Bacterial relative abundance across microbial mats)
  • mean >3% (Eukarya relative abundance across microbial mats)
  • mean <1% (Archaea relative abundance across microbial mats)
  • mean <0.1% (Virus relative abundance across microbial mats)
  • mean Bacteroidota 35%, Pseudomonadota 29%, Cyanobacteriota 19%, Verrucomicrobiota 3%, Bacillariophyta 2%, Planctomycetota 2%, Acidobacteriota 2%, Actinomycetota 2%, Bacillota 1%, Chloroflexota 1% (Average phylum-level abundance composing Antarctic microbial mats)
  • count mean 7.82 Gb per metagenome; 109 Gbps total sequenced DNA (Metagenomic sequencing depth across 14 mat samples)
  • count n=14 microbial mats (6 MA, 3 AP, 5 DV), 5 subsamples each (Sampling design across three Antarctic regions)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is an observational, cross-sectional metagenomic survey comparing 14 Antarctic microbial mats across three regions (Maritime Antarctica, Antarctic Peninsula, Dry Valleys). Community composition and environmental data were explored mainly with multivariate/ordination methods (PCA on water nutrients, NMDS on Bray-Curtis dissimilarities of taxonomic counts, diversity indices), while a focused pairwise comparison (Fildes vs. Garwood) used Welch's t-tests on relative abundances, and associations among community composition, geography, and environment were assessed with Mantel tests and Spearman correlations. Results were reported primarily as ordination plots, barplots, and correlation/association statistics (rho thresholds, FDR-adjusted p-values) rather than as replicate-level means with dispersion measures.

Replicationmixed Sample sizeNo formal power/sample-size calculation was described; n reflects the field sampling design (14 microbial mats across 3 regions, each mat sampled as 5 pooled subsamples for DNA; water sampled in triplicate; elemental analysis run with 5 replicates per sample) GroupsFildes Peninsula (Maritime Antarctica) vs. Garwood Valley (Dry Valleys) for direct statistical testing; three regions (Maritime, Peninsula, Dry Valleys) compared descriptively/via ordination Pairingunpaired Randomization/blindingnot stated Dispersionunclear Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
Principal component analysis (PCA), via princomp on standardized data Water nutrient concentrations across 14 sampling sites 14 sampling sites not stated
Shannon-Wiener and Simpson diversity indices (vegan package) Genus-level diversity from 14 ribosomal marker proteins/OTU tables 14 metagenomes not stated
Non-metric multidimensional scaling (NMDS) on Bray-Curtis dissimilarity (TMM-normalized counts) Genus-level taxonomic composition, separately for eukaryotes and prokaryotes 14 metagenomes not stated (stress <0.2 used as a model-fit criterion)
Welch's t-test (via STAMP platform) Relative abundance of prokaryotes, eukaryotes, and microbial functions (read and gene level), Fildes Peninsula (MA) vs. Garwood Valley (DV) not explicitly stated (comparison restricted to these two sites based on sample size/physicochemical similarity) not stated
Mantel test Comparison of mat community composition, environmental variables, and geographic position 14 mats not stated
Spearman correlation Genus composition/abundance vs. environmental variables 14 mats not stated; results filtered at FDR-adjusted p < 0.01 and |rho| ≥ 0.75
Approaches that could also have been used
  • Two-group comparisons of relative abundance (Fildes vs. Garwood) were made with Welch's t-test for many taxa/functions.
    Could also: A non-parametric approach such as the Mann-Whitney U test, or a permutation-based test (e.g., ALDEx2/ANCOM-style compositional tests), could also be used. — Relative-abundance/compositional metagenomic data are often non-normal and compositionally constrained, so rank-based or compositional-aware methods are a commonly used alternative to a t-test on percentage data.
  • Multiple taxa and functional categories were each tested individually with Welch's t-tests between the two sites.
    Could also: A family-wise or FDR-based multiple-testing correction (as was already applied to the Spearman correlations) could also be extended to the t-test comparisons. — When many features are tested in parallel, applying a correction across that whole family of tests is a standard way to control the overall false-positive rate, complementing the correction already used elsewhere in the study.
  • Community composition differences among the 14 mats were visualized with NMDS on Bray-Curtis distances.
    Could also: A formal significance test for group differences, such as PERMANOVA (adonis) or ANOSIM on the same distance matrix, could also be reported alongside the ordination. — NMDS provides a visual summary of dissimilarity, while PERMANOVA/ANOSIM give a hypothesis test with a p-value for whether groups (e.g., regions) differ significantly in composition, which can complement the ordination plot.
  • Associations between community composition, environment, and geography were assessed with Mantel tests.
    Could also: A partial Mantel test or a distance-based redundancy analysis (db-RDA) could also be used. — Partial Mantel tests or db-RDA allow the effect of geographic distance to be separated from environmental variables, which can help disentangle overlapping spatial and environmental influences on community composition.
  • Elemental analysis (C, N) was run with five replicates per sample, and nutrient/biomass values are reported in Table 1 as point estimates.
    Could also: Reporting a measure of spread (SD or 95% CI) alongside each mean could also be included. — Showing dispersion for replicate analytical measurements conveys measurement precision and lets readers judge the reliability of site-to-site differences.
  • Diversity was summarized using Shannon-Wiener and Simpson indices from ribosomal marker genes.
    Could also: Complementary richness estimators (e.g., Chao1) or rarefaction/accumulation curves could also be used. — Rarefaction-based approaches help confirm that differences in diversity metrics are not simply driven by differences in sequencing depth across metagenomes.
Software: R (princomp, vegan, factoextra, ggplot2) R 4.4.1 · STAMP platform · GIMP / Inkscape (figure editing, not statistical)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40043057

Paper: Mercado-Juárez et al. 2025, "What defines a photosynthetic microbial mat in western Antarctica?" PLoS One. DOI 10.1371/journal.pone.0315919.

Data (public, confirmed on ENA): 14 shotgun metagenomes of Antarctic microbial mats, total ~109 Gbp — exactly the paper's "14 mats / total of 109 Gbps".

  • PRJEB14287 — 9 runs, "Microbial mats from Maritime peninsula" (Fildes / MA): ERR1456907–ERR1456915.
  • PRJEB12762 — 5 runs, "Dry-Valleys" (Garwood / DV): ERR1303297–ERR1303301.

In scope — SingleM-derived alpha diversity (the clearly-specified, low-hanging output)

SingleM (github.com/wwood/singlem) is the paper's stated tool for alpha-diversity profiling: "singlem pipeline was used to generate OTUs tables for each metagenome", "alpha diversity via relative abundances of single marker genes". Applying this established third-party tool to the paper's own public data is a valid, equal-weight reproduction (Brief rule P16). Targets:

id reported value location
C1 9,414 species-level OTUs (14 ribosomal-protein markers) Results, diversity
C2 Shannon = 4.98 ± 0.83 Results, diversity
C3 Simpson = 0.94 ± 0.05 Results, diversity

Pipeline: singlem pipe on each of the 14 metagenomes → per-sample OTU tables → aggregate distinct species-level OTUs (C1); compute Shannon & Simpson per sample from OTU relative abundances → mean ± sd (C2, C3).

Known reproducibility caveats (documented up front, honest 1:1):

  • The paper does not pin a SingleM version or metapackage. SingleM OTU counts are highly version/metapackage-sensitive (the modern default metapackage uses ~59 single-copy markers; the paper restricted to 14 ribosomal proteins). So C1 is expected to be same-order, not exact; Shannon/Simpson (C2/C3) are more robust.
  • No QC/host-removal parameters that change marker recovery are pinned beyond "fastp v1.9.2, trim first 10 bp". We run SingleM on the reads as-is (SingleM is designed to tolerate adapters); this is the main controlled deviation.

Out of scope (compute-prohibitive or different tool — 80/20 skip, not attempted)

  • Kaiju taxonomic composition (Bacteria 93.45%, Bacteroidota 34.61%, etc.): needs NCBI nr 2021-02 (~hundreds of GB) — prohibitive; different tool from SingleM.
  • MegaHit assembly (1,118,337 contigs >1 kb) + Prodigal/Augustus gene prediction + mmseqs2 clustering + eggNOG functional annotation: assembling 109 Gbp is many node-days — out of the 80/20 budget.
  • EukDetect / Metaxa2 eukaryote detection; PCA/NMDS/Mantel/STAMP downstream stats: derived from the above heavy steps, not attempted.

Rationale: SingleM diversity is the single clearly-specified, runnable-from-reads result; the rest are heavy or depend on a 109-Gbp de-novo assembly.

C1
Reported
9414 species-level OTUs (SingleM, 14 ribosomal-protein markers)
Reproduced
8964 distinct species-level OTU sequences (14 markers; 23987 with all 59 markers)
within tolerance
C2
Reported
Shannon 4.98 +/- 0.83
Reproduced
Shannon 4.61 +/- 0.82 (genus level, OTU-table marker abundances)
within tolerance
C3
Reported
Simpson 0.94 +/- 0.05
Reproduced
Simpson 0.946 +/- 0.054 (genus level)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

All three SingleM-derived alpha-diversity claims reproduce within the paper's own reported uncertainty on the identical public data (Simpson essentially exact, Shannon within 1 SD, richness within ~5%), with no fabrication signal. The only deviations sit on the input/methodology side: the paper pins neither the SingleM metapackage version nor the diversity aggregation level, so we adopted a modern reference (driving the −4.8% richness gap) and a self-chosen OTU-table-genus aggregation. These are technical/expected and our-method effects rather than authors' defects, hence a solid-but-not-1:1 yellow overall.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

428.8 k
tokens (I/O) · 30.6 M incl. cache
178 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.