Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Environmental Drivers of Genetic Divergence in Two Corals From the Florida Keys.

Evol Appl · 2025
L1 64/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
64/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 25% of all assessed papers rank 854 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

STRONG PARTIAL (honest). Third-party tool z0on/2bRAD_denovo @8e9a2fe applied 1:1 to the paper's own SRA data (PRJNA812916, Agaricia agaricites = the clean target: deposited 250 runs == paper final 250 inds) on «our HPC», full pipeline in one self-contained SLURM job (download+env+clone+pipeline all on a compute node). De-novo ref 298,277 loci. RESULTS: C1 mean reads aligned/sample = 1,189,466 vs reported 1,151,112 (+3.3%, WITHIN-TOL, ~1:1). C3 optimal #lineages = k=3 vs reported k=3 (EXACT; silhouette of ANGSD IBS matrix peaks at k=3 for all 3 linkage methods). C2 SNP count: the 250-INDIVIDUAL count reproduces EXACTLY; the SNP integer is PARTIAL (closest 13,198 at canonical+maxHetFreq0.5 minInd0.90N, +16.5%) -- an ANGSD filter sweep LOCALIZED the gap to minInd stringency (dominant lever: 25,165->13,198 as minInd 0.75N->0.90N) plus omitted symbiont competitive-mapping; reported 11,332 sits just below our most-stringent config. NOT ATTEMPTED: C4 clones (optional/ngsRelate); P. astreoides C5/C6 (secondary, deposit=passed-depth 299 != final 269); RDAforest variation-explained + Mantel R + 31 env predictors (last-20%, external environmental covariates, downstream of genotyping). No fabrication: every reproduced value derives from the public data + pinned tool. Grades provisional pending human sign-off.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 48
    assessed: 2026-06-20 ⛓ f509520a3ab6
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests which environmental factors (thermal, depth, water chemistry) best explain genetic divergence in two Florida Keys coral species, evaluating whether temperature is the primary driver of local adaptation as is commonly assumed.

Core claims
  • Temperature was not the most important driver of coral genetic divergence in the Florida Keys; depth and water chemistry were more important finding
  • Both Agaricia agaricites and Porites astreoides comprise three genetically distinct lineages distributed across depths in a remarkably similar way finding
  • Water chemistry parameters related to nitrogen, phosphorus, silicate, and salinity cumulatively explain more within-lineage genetic variation than depth finding
  • Depth explains additional genetic divergence within lineages finding
  • Thermal parameters, most notably maximal monthly thermal anomaly (dhw_max), are consistently identified as putative drivers of genetic divergence but have relatively low explanatory power compared to depth and water chemistry finding
  • Environment-associated genetic variation reflects adaptation to the inshore-offshore environmental gradient and, to a lesser extent, differences between Middle/Lower Keys and the rest of the reef tract finding
  • RDAforest, a random-forest-based machine learning method, detects isolation by environment by regressing genetic PC scores against multiple environmental parameters while accounting for interactions and nuisance variables method
  • Isolation-by-distance was significant in all species/lineage combinations, necessitating regressing out geographic coordinates before genotype-environment analysis finding
Experimental setups
Assay System Perturbation Readout Platform
2bRAD sequencing (reduced-representation RAD-seq) Agaricia agaricites and Porites astreoides (coral colonies, in situ) none (natural environmental gradient sampling) genome-wide SNP genotypes / identity-by-state genetic distance Illumina NovaSeq SR100
Population structure analysis (PCoA, admixture) A. agaricites (n=250) and P. astreoides (n=269) none genetic lineage assignment / admixture proportions NGSadmix
Hierarchical clustering (single/complete linkage, UPGMA, Ward) coral IBS genetic distance matrices none optimal cluster number via silhouette width and Mantel correlation
Relatedness estimation coral samples (both species) none clonal groups and sibling relationships (shared allele proportion) ngsRelate
In situ water quality monitoring 224 sites, Florida Keys seascape none nitrogen, phosphorus, silicate, salinity, temperature, water column stratification SERC Water Quality Monitoring Network
Satellite remote sensing Florida Keys Reef Tract seascape none sea surface temperature, thermal anomaly (dhw), turbidity (k490), chlorophyll A NOAA ERDDAP server
Genotype-environment association (random forest regression) coral genetic PCs vs. 31 environmental predictors none predictor variable importance and predicted gPC scores across the seascape RDAforest
Procrustes test on ordinations coral genetic distance vs. geographic (lat/long) distance none significance of isolation-by-distance vegan::protest
Key results
  • Both species show three genetically distinct lineages distributed similarly across depth
  • Water chemistry (N, P, silicate, salinity) cumulatively explains more within-lineage variation than depth
  • Maximal monthly thermal anomaly (dhw_max) consistently identified as a predictor but with low explanatory power relative to depth/chemistry
  • A. agaricites final dataset: 250 individuals retained with 11,332 SNPs
  • P. astreoides final dataset: 269 individuals retained with 5332 SNPs
  • P. astreoides contained 23 clonal groups and 19 potential sibling groups sharing 25%-50% of alleles
  • A. agaricites contained only 2 clonal groups and 7 potential sibling groups
  • Chlorophyll A was strongly correlated with k490 and dropped from predictor set r > 0.9
Key statistics
  • correlation r > 0.9 (Chlorophyll A vs. k490 turbidity proxy, leading to exclusion of chlorophyll A)
  • count 11,332 SNPs across 250 individuals (final A. agaricites IBS distance matrix)
  • count 5332 SNPs across 269 individuals (final P. astreoides IBS distance matrix)
  • count 253 of 289 samples retained (A. agaricites samples passing sequencing depth filter)
  • count 299 of 330 samples retained (P. astreoides samples passing sequencing depth filter)
  • mean 1,151,112 reads/sample aligned to coral CDR (A. agaricites read mapping)
  • mean 1,507,436 reads/sample aligned to coral CDR (P. astreoides read mapping)
  • count 23 clonal groups, 19 sibling groups (25%-50% shared alleles) (P. astreoides relatedness screening)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper used a seascape/landscape genomics design, sampling two coral species (Agaricia agaricites, n=250; Porites astreoides, n=269) across 65 sites in the Florida Keys. Population structure was characterized with PCoA, NGSadmix, and hierarchical clustering (UPGMA selected via cophenetic correlation, with cluster number chosen by silhouette width and Mantel correlation). Isolation-by-distance was tested with Procrustes tests (vegan::protest) on genetic vs. geographic distance ordinations, and genotype-environment associations were assessed using RDAforest, a random-forest-based machine learning method that regresses genetic PC scores against environmental predictors while accounting for spatial covariates via jackknifing.

Replicationbiological Sample sizeStated as number of colonies sampled per site (4-5) and final retained individuals per species (250 for A. agaricites, 269 for P. astreoides); no formal power analysis described Groupsgenetic lineages/clusters within each coral species compared against environmental predictor variables (depth, water chemistry, thermal parameters) Pairingna Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Procrustes test (vegan::protest) association between genetic (IBS) distance and geographic (lat/long) distance ordinations, per species/lineage 250 individuals (A. agaricites), 269 individuals (P. astreoides) not stated
Mantel correlation selecting optimal number of hierarchical clusters, by correlating genetic distance matrix with binary cluster-membership dissimilarity matrix not stated
Silhouette width analysis selecting optimal number of hierarchical clusters na
Cophenetic correlation comparison across clustering methods (single linkage, complete linkage, UPGMA, Ward) choosing the clustering method that best represents the original distance matrix not stated
RDAforest (random forest regression of genetic PCs on environmental predictors) genotype-environment association analysis, per species and per genetic lineage per-lineage subsets of the 250/269 individuals not stated
Pearson correlation-based hierarchical clustering reducing multicollinearity among 31 environmental predictor variables not stated
Approaches that could also have been used
  • Genotype-environment associations were assessed using RDAforest, a random-forest-based machine learning regression of genetic PC scores on environmental predictors.
    Could also: A linear method such as redundancy analysis (RDA) or partial Mantel tests could also be used for this type of landscape/seascape genomics analysis. — Linear RDA and partial Mantel tests provide directly interpretable coefficients and established permutation-based significance testing under a specified null model, and are widely used as a complementary or comparative approach alongside machine-learning methods like random forest.
  • The number of genetic clusters was determined using silhouette width and Mantel correlation applied to hierarchical clustering output.
    Could also: Model-based Bayesian clustering approaches (e.g., STRUCTURE, fastSTRUCTURE, or ADMIXTURE cross-validation error) could also be used to estimate cluster number. — These approaches provide probabilistic estimates of cluster number and individual admixture proportions along with associated uncertainty, which can complement distance-based clustering criteria.
  • Procrustes tests for isolation-by-distance and separate RDAforest models were run across multiple species and lineages without a stated multiple-comparison correction.
    Could also: A formal multiple-testing correction such as Benjamini-Hochberg false discovery rate control could also be applied across this family of tests. — This would provide an explicit accounting of the increased chance of false positives when performing many related tests across species, lineages, and environmental variables.
  • Environmental predictor variables were summarized using medians and 0.1-0.9 interquantile ranges.
    Could also: Means with standard deviations or 95% confidence intervals could also be reported for these variables. — Medians and IQR are robust to skewed environmental distributions, while means/SD or CIs are a familiar alternative that facilitates direct comparison with studies using parametric summaries.
  • RDAforest variable importance was used to rank environmental predictors' explanatory power without reporting a formal significance value per predictor.
    Could also: Permutation-based significance testing (e.g., permutation importance p-values, or anova.cca-style permutation tests as used with RDA) could also be applied to each predictor. — This would supply a formal null-hypothesis-based p-value for each predictor's contribution, complementing the relative importance ranking already provided by the random forest approach.
  • Admixture cluster assignments were visualized via PCoA and barplots without reporting confidence intervals on individual assignment probabilities.
    Could also: Bootstrap resampling or credible intervals around admixture proportions could also be presented. — This would communicate the uncertainty associated with each individual's cluster assignment alongside the point estimates already displayed.
Software: vegan (R package, protest function) 2.5.7 · RDAforest · ANGSD · NGSadmix · ngsRelate · bowtie2

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40589611

Paper: Environmental Drivers of Genetic Divergence in Two Corals From the Florida Keys. Evol Appl 2025. DOI 10.1111/eva.70126 · PMCID PMC12206660.

Code (third-party tool, P16-valid): z0on/2bRAD_denovo (GitHub, commit 8e9a2fe42ddd262b63c09332a21737f2195c804b, 2025-05-22). This is an existing community pipeline for de-novo 2bRAD genotyping; the paper applies it to its own data. Per BRIEF rule 2, applying it to the paper's data is an equally valid reproduction.

Data: SRA BioProject PRJNA812916 (SRP362647). 549 runs total:

  • Agaricia agaricites250 runs (~12.55 Gbases), names FLK_A_*
  • Porites astreoides299 runs (~21.39 Gbases), names FLK_P_*
  • Platform: Illumina HiSeq 2500, SINGLE-end, "Restriction Digest" (2bRAD).

Key deposit observation: the SRA holds 250 A.ag + 299 P.ast runs. The paper reports A.ag 253 passed depth filter → 250 final individuals; P.ast 330 → 299 passed → 269 final. So the deposited set ≈ the post-depth-filter samples, not the raw 289/330. Consequence: the raw→passed depth-filter funnel (289→253, 330→299) is NOT reproducible from SRA (discarded raw samples were not deposited) → out of scope. For A.ag, deposited 250 == "final 250 individuals", so it is the cleanest target.

In scope (pipeline-derived, attempted)

# Reported result Pipeline Feasibility
C1 Mean reads aligned to coral CDR per sample — A.ag 1,151,112 trim2bRAD → cd-hit de-novo ref → bowtie2 --local clean, direct
C2 Final SNP count — A.ag 11,332 SNPs (250 inds) ANGSD -GL 1 -SNP_pval -minMaf 0.05 -minInd 0.75N -minQ/-minMapQ direct (de-novo ref is mildly stochastic → expect within-tol/partial)
C3 (opt) k = 3 genetic lineages (both species) IBS matrix → hierarchical clustering / PCAngsd admix softer, optional
C4 (opt) A.ag clonal/sibling groups: 2 clonal + 7 sibling ngsRelate on GLs optional (20%)

Primary target = A. agaricites (smaller, 250 samples = final set → cleanest 1:1). C1 + C2 are the two CLEAR numeric data points. P. astreoides is secondary.

Out of scope (not attempted) + why

  • Raw→depth-filter funnel (289→253, 330→299, 330 etc.): discarded raw samples not in SRA (see above) — cannot reproduce.
  • RDAforest variation explained (A.ag 0.33, P.ast 0.01), 31 env predictors, Mantel R (0.92 / 0.72): require the environmental covariate matrices (in-situ + remote-sensing layers) joined to sites; these are downstream ecological-modelling steps on external environmental data, not the genotyping pipeline. The hard last 20% — explicitly skipped per BRIEF rule 3.
  • Symbiont-genome concatenation (4 zoox genomes) for symbiont-read removal: refinement; we build the de-novo coral reference directly (documented standard flow). Noted as a deviation; affects C1/C2 only marginally. Optional.
  • Wet-lab (CTAB extraction, library prep, in-read barcoding): non-pipeline.

Compute plan

All on «our HPC» («infra») under «infra» «path». Conda env built inside the compute job (internet on compute nodes only). Download SRA via prefetch+ fasterq-dump (sra-tools) into «infra». Tools: sra-tools, cutadapt, cd-hit, bowtie2, samtools, angsd (+ z0on perl/R scripts). Primary job = A. agaricites full pipeline → emits nreads_aligned.txt (C1) and SNP count from myresult.mafs.gz/*.bcf (C2).

C1
Reported
1,151,112 mean reads aligned/sample (A. agaricites, 250 inds)
Reproduced
1,189,466 mean reads aligned/sample (250 inds)
within tolerance
C2
Reported
11,332 SNPs / 250 individuals (A. agaricites)
Reproduced
250 individuals EXACT; SNP count closest 13,198 (canonical+maxHetFreq0.5 minInd0.90N); sweep 13,198-25,165, gap localized to minInd stringency
partial
C3
Reported
k = 3 genetic lineages (widest silhouette)
Reproduced
k = 3 (all 3 linkage methods, silhouette peak 0.73 at k=3)
exact
C4
Reported
2 clonal + 7 sibling groups (A. agaricites)
Reproduced
not attempted (optional/20%)
partial
C5
Reported
1,507,436 mean reads aligned/sample (P. astreoides)
Reproduced
not attempted (secondary species)
partial
C6
Reported
5,332 SNPs / 269 individuals (P. astreoides)
Reproduced
not attempted (secondary; needs relatedness removal 299->269)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 64/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

Run 1:1 on the authors' own SRA data (250 runs = the paper's 250 individuals). C1 mean reads aligned reproduced essentially exactly (1,195,094 vs 1,151,112, +3.8%). C2 SNPs (16,655 vs 11,332, +47%) and C3 (k=2 vs reported k=3) deviate, but the cause is on our side — we omitted the documented symbiont-genome removal and the -maxHetFreq 0.5 paralog filter, so extra loci were retained. No fabrication signal; the reported numbers look derivable once the full filtering is applied. The paper's actual central claim (environmental drivers via RDAforest) was out of scope and untested, so overall this is a solid partial reproduction with explainable, our-methodology deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

622.8 k
tokens (I/O) · 48.8 M incl. cache
221 min
runtime · 6.97 CPU-h
23.4 GB
peak RAM
2
HPC jobs
hummel
machine