Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Metatranscriptomics-based investigation of bacterial community dynamics across a dissolved organic matter gradient in southern Lake Michigan.

Appl Environ Microbiol · 2026
L1 81/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
81/100
Reproducibility score
0.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 59% of all assessed papers rank 468 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for a PARTIAL 1:1. The paper's only listed code is the generic R vegan package, used for a single non-significant PERMANOVA whose input (a 5,316 feature x 11 sample community matrix) was never deposited and for which no numeric statistic is printed -- so the literal vegan result is not reproducible (C5), as is the 5,316-feature count (C6), because re-deriving it needs JGI's proprietary IMG/M annotation of the raw SRA (the hard, non-faithful 80%, deliberately not attempted; no «our HPC» job warranted). Instead I audited the paper's headline numbers against its OWN deposited derived data (supplementary Tables S1/S3/S4 + the deposited MAG fasta repo) and reproduced three EXACTLY: the central claim of 130 differentially expressed gene families nearshore-vs-offshore (Table S4 = exactly 130 MaAsLin2 features, all q<0.05); 11 libraries = 3 nearshore + 8 offshore (Table S1); and 7 MAG populations (Table S3, independently corroborated by 7 matching fasta in github.com/Aditchaudhary/Lake-Michigan-MAGs). One within-tol flag: deposited per-library read counts (43.53-58.17 M) sit just below the paper's stated 43.8-58.6 M range (<1%, likely a counting/rounding convention). No fabrication signal on the checkable claims; the unverifiable ones (C5/C6) simply lack released data. NOT attempted: full assembly/IMG-M annotation/normalization pipeline, vegan PERMANOVA, PCoA (ecodist), taxonomy/DOC/transporter/stress-marker content (wet-lab/external).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 81
    assessed: 2026-06-16 ⛓ aed1e4ca9c34
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How do environmental controls—particularly the quality and gradient of dissolved organic matter (DOM)—shape bacterioplankton community function and substrate-acquisition strategies across a nearshore-to-offshore transect in southern Lake Michigan?

Core claims
  • DOM composition changes significantly across the nearshore-to-offshore transect, with more terrestrially derived and high-molecular-weight DOM nearshore, despite only minor reductions in DOC and similar inorganic N and P levels. finding
  • Differences in DOM quality across the transect are associated with differential expression of gene families between nearshore and offshore bacterioplankton. finding
  • Genes for acquiring DOM, N, and P substrates (peptidases, proteases, and transporters for amino acids, nucleobases, sugars, urea, and inorganic phosphate) are over-represented in offshore bacterioplankton. finding
  • Offshore bacterial communities are more substrate-limited (particularly carbon) than nearshore and invest more energy in acquiring DOM substrates. mechanism
  • Focused analysis of transporter gene expression for C, N, and P substrates shows higher expression of DOM transporter genes offshore versus nearshore. finding
  • Metatranscriptomics can be applied to assess bacterioplankton metabolism in a large freshwater lake in the context of rich environmental DOM characterization data. method
  • Coastal-to-offshore spatial gradients in large lakes provide a study system to investigate bacteria-water chemistry relationships with limited confounding abiotic effects. finding
Experimental setups
Assay System Perturbation Readout Platform
metatranscriptomics (mRNA-based gene expression) free-living bacterioplankton (0.2-µm filtered, 1.6-µm prefiltered) from surface waters of southern Lake Michigan nearshore-to-offshore transect none (spatial environmental gradient; nearshore vs offshore, spring vs summer 2017-2018) transcript abundance of gene families Illumina NovaSeq S4, paired-end 150 bp; RiboCop rRNA Depletion kit + CORALL Total RNA-Seq Library Prep kit
metagenomics (community gDNA sequencing) bacterioplankton community DNA from summer 2017 Lake Michigan samples none genomic sequence content Illumina NextSeq, paired-end 150 bp
dissolved organic carbon (DOC) measurement 0.2-µm filtered Lake Michigan surface water none DOC concentration (µM) high-temperature combustion method
nutrient analysis 0.2-µm filtered Lake Michigan surface water none orthophosphate (PO4 3-/SRP) and nitrate+nitrite (NOx) concentrations autoanalyzer AQ300, SEAL Analytical
CDOM UV-vis absorption spectroscopy Lake Michigan surface water filtrate (spring 2018 + some summer 2018) none absorption coefficient a254, spectral slope S275-295, slope ratio Sr (proxies for CDOM concentration and molecular weight)
fluorescence excitation-emission matrix (EEM) / FDOM characterization Lake Michigan surface water filtrate none biological index (BIX), humification index (HIX), fluorescence peaks/components spectrofluorometer
Key results
  • Higher presence of terrestrially derived and high-molecular-weight DOM in nearshore versus offshore
  • Minor reduction in DOC levels from nearshore to offshore
  • Inorganic N and P measurements similar across the transect
  • DOM-, N-, and P-acquisition genes (peptidases, proteases, transporters for amino acids, nucleobases, sugars, urea, inorganic phosphate) over-represented in offshore bacterioplankton
  • Higher expression of DOM transporter genes for C, N, P substrates offshore versus nearshore
  • Filtered feature table retained 5,316 unique gene families after thresholding 5,316 gene families
  • Metatranscriptome libraries yielded 43.8-58.6 million paired-end reads per library 43.8-58.6 million reads
Key statistics
  • count 5,316 unique gene families retained (features with ≥10 transcript counts in ≥30% (4/11) of samples)
  • count 43.8-58.6 million paired-end reads per library (metatranscriptome sequencing yield)
  • count 11 samples total (3 nearshore NRS, 8 offshore OFS1/OFS2/OFS3) (sampling design across transect 2017-2018)
  • other DOC range 93-191 µM (e.g., NRS 191±1.2; OFS3 93±1.4) (DOC concentrations nearshore vs offshore, Table 1)
  • other NRS ~3.5 km from shore; OFS ~10-50 km from shore (distance of sampling sites from shore)
  • pvalue P < 0.05 (Welch's T test for DOC/DOM parameters significantly different between nearshore and offshore)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This observational metatranscriptomic study compared bacterioplankton gene expression between nearshore (n=3) and offshore (n=8) sites along a southern Lake Michigan transect in 2017–2018 across two seasons. Water chemistry parameters (DOC, BIX, HIX, CDOM indices) were compared between locations using Welch's t-tests. Gene family transcript abundances were size-factor-normalized via DESeq2, ordinated with Bray-Curtis dissimilarity/PCoA, and tested for differential expression using MaAsLin2 mixed-effects models with location and season as fixed effects and sampling station and year as random effects (text truncated before full model specification was provided).

Replicationbiological Sample size3 nearshore (NRS) and 8 offshore (OFS1/OFS2/OFS3) surface-water samples collected across summer 2017, spring 2018, and summer 2018; no formal power calculation stated GroupsNearshore vs. offshore; spring vs. summer season as secondary factor Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionMaAsLin2 applies Benjamini-Hochberg FDR correction by default; whether this was used or reported is not confirmed in the provided text. No correction stated for the Welch's t-tests on water chemistry parameters.
Statistical tests used
Test Applied to n Assumptions
Welch's t-test (two-sample, unequal variance) Comparison of DOC concentration, BIX, and HIX between nearshore and offshore (Fig. 1B–D) 3 nearshore vs. up to 8 offshore; CDOM/FDOM subset restricted to spring 2018 and partial summer 2018 samples — exact n per test not stated not stated
Bray-Curtis dissimilarity + principal coordinate analysis (PCoA) Ordination of DESeq2-normalized gene family profiles across all 11 metatranscriptome samples 11 samples total na
MaAsLin2 linear mixed-effects model Detection of gene families differentially associated with location (nearshore/offshore) and season (spring/summer); sampling station and year included as random effects (model specification truncated in provided text) 11 samples; feature table filtered to genes with ≥10 counts in ≥4 of 11 samples (5,316 gene families retained) not stated
DESeq2 size-factor normalization Library-size normalization of gene family transcript count matrix prior to all downstream analyses 11 metatranscriptome libraries na
Approaches that could also have been used
  • Welch's t-tests were used to compare individual DOM/water chemistry parameters between nearshore and offshore with P < 0.05 as the threshold, with multiple parameters tested
    Could also: Apply a Benjamini-Hochberg FDR correction or Bonferroni correction across the family of water chemistry t-tests, or use a single MANOVA to test all DOM indices jointly — When multiple parameters are tested simultaneously, a family-wise or FDR correction reduces the probability of at least one false positive across the test family; MANOVA additionally accounts for correlations among the DOM indices
  • With 3 nearshore and up to 8 offshore samples, Welch's t-tests assume approximate normality in small groups
    Could also: Use a Mann-Whitney U (Wilcoxon rank-sum) test as a non-parametric alternative — Non-parametric tests make no distributional assumption and are often preferred when group sizes are small (n=3 in one group), where normality is difficult to assess
  • Bray-Curtis PCoA was used to visualize community-level gene expression structure
    Could also: Use non-metric multidimensional scaling (NMDS) on the same Bray-Curtis matrix, optionally with a PERMANOVA (adonis2 in vegan) to formally test location/season effects on community composition — NMDS relaxes the linearity assumption of PCoA and often provides better stress-minimized low-dimensional representation; PERMANOVA provides an omnibus statistical test of group separation in multivariate space to complement the ordination plot
  • DESeq2 size-factor normalization was applied to the gene family count matrix before MaAsLin2 modeling
    Could also: Use trimmed mean of M-values (TMM) normalization (edgeR) or centered log-ratio (CLR) transformation, then apply a linear mixed model or permutation-based approach — TMM and CLR are widely used normalization strategies for compositional count data; CLR in particular is compositionally appropriate and is sometimes preferred when downstream analyses assume log-linearity
  • DOC and other continuous measurements were reported as mean ± SD
    Could also: Report 95% confidence intervals alongside or instead of SD — CIs directly communicate uncertainty about the group mean and facilitate comparison across studies; with small n (e.g., n=3 at NRS), CIs are often more informative than SD for conveying precision of the estimate
  • Significance for water chemistry differences was reported as a binary P < 0.05 threshold
    Could also: Report exact p-values and an effect size metric (e.g., Cohen's d or rank-biserial correlation) for each comparison — Exact p-values allow readers to apply alternative thresholds; effect sizes convey the magnitude of the difference independently of sample size, which is particularly informative when n is small
Software: R 4.3.2 · DESeq2 2.0 · vegan 2.6-10 · ecodist 2.1.3 · ggplot2 3.5.2 · MaAsLin2 1.16.0 · Trimmomatic 0.33 · Megahit 1.1.1-2 · IMG/M (IMGAP) 5.1.17 · SortMeRNA · blastn · ArcGIS

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Table
C1
Reported
130 gene families with significant differential expression between nearshore and offshore (MaAsLin2 mixed-effects, FDR q<0.05)
Reproduced
130 features in Table S4, all Location/offshore contrast, all q<0.05 (max qval 0.0496); coef split 80 higher-offshore / 50 lower
exact
C2
Reported
11 metatranscriptome libraries = 3 nearshore + 8 offshore
Reproduced
Table S1: 11 rows = 3 NRS + 8 OFS (OFS1/2/3)
exact
C3
Reported
seven MAG-based populations tracked
Reproduced
Table S3 lists 7 MAGs (bin009/010/035/079/181/004/040); GitHub Aditchaudhary/Lake-Michigan-MAGs holds 7 matching fasta with same bin IDs and taxa
exact
C4
Reported
43.8-58.6 million paired-end reads per library
Reproduced
Table S1 sequence-cluster counts: min 43.53 M, max 58.17 M (both ~0.3-0.4 M / <1% below the stated bounds)
within tolerance
C5
Reported
PERMANOVA P>0.05 (vegan, Bray-Curtis) for sample type/season/year
Reproduced
not reproducible
partial
C6
Reported
filtered feature table retained 5,316 unique gene families
Reproduced
not reproducible
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 81/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

The paper's three checkable headline numbers — 130 differentially expressed gene families (C1, exact from Table S4, max q=0.0496), 11 libraries = 3 nearshore + 8 offshore (C2), and 7 MAG populations (C3, independently corroborated by 7 matching GitHub fasta) — reproduce 1:1 from the authors' own deposited derived data, with no fabrication signal. The only factual deviation is C4: deposited read-pair counts (43.53–58.17 M) sit <1% below the stated 43.8–58.6 M range, a likely counting/rounding convention. The two non-reproducible claims (C5 PERMANOVA, C6 5,316 features) fail on data availability — the count/community matrices were never deposited and the JGI IMG/M annotation step is proprietary — not on any authors' defect or contradiction. Net: a solid partial reproduction; the central conclusion holds while the heavy upstream community-matrix results remain unverifiable.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

94.6 k
tokens (I/O) · 4.7 M incl. cache
20 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.