Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Early life-stage thermal resilience is determined by climate-linked regulatory variation.

Proc Natl Acad Sci U S A · 2026
L1 64/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
64/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 25% of all assessed papers rank 854 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL (substantial) reproduction of the primary pipeline result, plus a flagged secondary. PRIMARY = FET introgression mapping, reproduced from RAW pool-seq (8 pools, PRJNA1285949) via a fully independent pipeline on «our HPC»: bwa-mem (dmel-r6.12) -> sambamba markdup -> bcftools mpileup/call (1,343,424 biallelic SNPs; 1,238,207 after all-8>=10x) -> per-SNP Fisher exact test each F16 cross vs VT8 parent using the authors' own contingency table c(cross_AD,cross_DP,VT8_AD,VT8_DP) -> Bonferroni. RESULT: the chrX 15.5 Mb locus reproduces EXACTLY (densest consensus peak 15.5-16.0 Mb); the chr2R locus reproduces as a broad high-signal region with the paper's named top SNP 2R:20,551,633 recovered at the SAME delta_p=0.41; the X top SNP X:15,607,604 is recovered (sig 5/6 crosses); the genome-wide sig-SNP count at a 3/6-cross consensus (455) is the same order as the reported 391. Exact count + exact 2R peak position differ - expected because the authors' PoolSNP/DEST SNP set + FET-input tables are NOT deposited (had to rebuild with bcftools; valid third-party stand-in P16) and the cross-consensus rule is under-specified. SECONDARY = DGRP embryonic ANOVA on the authors' shipped CORRECTION data: test magnitude reproduces strikingly (F11.6 vs reported 11.22; P~0.002 vs 0.0014) but the effect DIRECTION is reversed (temperate higher) and group-Ns are swapped vs the main text -> FLAGGED for human review against the published correction. NOT attempted: Gowinda GO enrichment, pi diversity, SLiM FST-window outliers (compare-only, lower priority), RNA-seq regulatory-variation + wet-lab phenotyping (out of scope). All grades provisional; AUDIT.md + claims.tsv + agreement.json let a human re-derive 1:1.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 64
    assessed: 2026-06-20 ⛓ d8a71d3850e1
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Enhanced embryonic heat tolerance in tropical (Saint Kitts) versus temperate (Vermont) Drosophila melanogaster is determined by identifiable, climate-linked regulatory genetic variants that are under natural selection in wild populations.

Core claims
  • Embryonic heat tolerance in D. melanogaster is determined by climate-linked regulatory genetic variation at loci on chromosomes 2R and X finding
  • 16 generations of thermal selection via advanced introgression-backcrossing consistently targeted two loci on 2R and X across six replicate introgressions finding
  • SNP 2R:20,551,633 lies in a putative regulatory region near SP70, shows clinal and seasonal allele frequency patterns tied to precipitation/temperature variability, and correlates with heat tolerance finding
  • SNP X:15,607,604 lies in a putative regulatory region near sog and shows seasonal and clinal signals paralleling the 2R locus finding
  • Tropical alleles at the 2R and X loci increase embryonic heat tolerance and interact epistatically, validated in an independent DGRP panel finding
  • An advanced introgression-backcrossing design combined with Pool-Seq and simulation-based FST/FET scans can map the genomic basis of embryonic heat tolerance method
  • The 2R regulatory SNP falls within a transcription factor binding hotspot containing motifs for Medea and dorsal, genes involved in dorso-ventral patterning mechanism
  • The tropical genotype (2R^C X^T) shows a minor decline in adult thermal performance (CTmax) despite higher embryonic heat survival finding
Experimental setups
Assay System Perturbation Readout Platform
Heat-shock selection and backcrossing (advanced introgression) D. melanogaster embryos, VT x SK cross, 6 replicate F16 lines heat shock (80% mortality, 1-h-old embryos) alternated with backcrossing to VT background over 16 generations embryonic heat tolerance/survival
Pool-Seq whole-genome resequencing VT and SK parental pools and F16 introgression pools none genome-wide allele frequencies, FST, PCA of genomic background Pool-Seq
Drift-only population genetic simulations in silico simulated introgression pools none FST outlier windows relative to neutral expectation
SNP-wise Fisher's Exact Test (FET) synthetic pool aggregating all F16 replicates none SNP-level significance of association with heat tolerance
Colocalization with clinal/seasonal genomic datasets (DEST 2.0) wild-collected D. melanogaster populations, >520 populations worldwide, plus Virginia/Odesa/Yesiloz time series none clinal and seasonal allele frequency patterns
Embryonic heat-shock survival assay 64 Drosophila Genetic Reference Panel (DGRP) lines heat shock, 45 min at 35°C, 1-h-old embryos proportion surviving by 2R(C/A) x X(T/A) genotype
Adult thermal performance assay (CTmax; reanalysis of Lecheta et al. data) DGRP line adults (male and female) acute thermal ramp critical thermal maximum (CTmax)
Heat-shock RNA-seq SK, VT, and F16 embryos 34°C heat shock vs 25°C control normalized transcript read counts for SP70 and sog RNA-seq
Key results
  • Introgressed F16 populations show elevated embryonic heat tolerance after 16 generations of selection
  • Two loci on 2R and X are consistently the top hits across all six introgression replicates, diverging from the otherwise VT-like genomic background
  • 2R:20,551,633 allele frequency shifted far beyond the neutral introgression expectation, dropping from ~0.60 (A allele, VT) to 0.18 in F16 pools 41.7% shift (Δp=0.41 vs expected 0.00021)
  • 2R:20,551,633 shows clinal frequency increase with latitude in North America and a decrease in South America in DEST 2.0 r=0.540 (N. America); r=-0.9816 (S. America)
  • 2R:20,551,633 undergoes seasonal selection tied to temperature variance 0-45 d prior to collection in Virginia β=-0.089
  • X:15,607,604 allele frequency shifted toward the tropical (T) allele during introgression Δp=0.33
  • DGRP lines homozygous for both tropical alleles (2R^C X^T) show significantly higher embryonic heat-shock survival than all other genotype combinations F(1,55)=11.22
  • Fully tropical genotype (2R^C X^T) DGRP lines show a minor decline in adult CTmax relative to other genotype combinations
Key statistics
  • pvalue P=0.013 (Welch's t-test, elevated embryonic heat tolerance in introgressed populations after 16 generations)
  • pvalue P=0.0003 (Wald test, LT50 difference between embryo and adult heat resistance (prior study))
  • count 391 SNPs (genome-wide SNPs with significant FET P-values (Bonferroni corrected, threshold 0.01))
  • fold_change Δp=0.41 (expected Δp=0.00021) (allele frequency shift at 2R:20,551,633 during introgression)
  • correlation r=0.540, P=0.0012 (North America); r=-0.9816, P=8.70×10^-5 (South America) (clinal allele frequency pattern of 2R:20,551,633 in DEST 2.0 dataset)
  • pvalue β_seasonal=-0.089, P_seasonal=0.0053 (beats 96% of permutations) (generalized linear model of seasonal selection at 2R SNP using Virginia genomic time series)
  • pvalue F(1,55)=11.22, P=0.0014 (ANOVA test of 2R(C/A)-by-X(T/A) genotype interaction effect on DGRP embryonic survival)
  • pvalue P=0.017 (F16 vs VT); P=0.54 (F16 vs SK) (Wald test, genetic background x temperature interaction on SP70 expression)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combines a multi-generation experimental evolution (introgression/backcrossing) design with population genomics, quantitative genetics in a reference panel, and RNA-seq to map loci underlying embryonic heat tolerance in Drosophila melanogaster. Group comparisons used t-tests, Fisher's exact tests, ANOVA, and Wald tests, while genome-wide scans were benchmarked against neutral simulations and corrected with a Bonferroni threshold; locus-level seasonal/clinal associations were assessed with correlations and a permutation-based generalized linear model. Results are reported primarily as point estimates with exact P-values, effect estimates (allele-frequency shifts, beta coefficients, F-statistics), and, in one case, a 95% binomial confidence interval.

Replicationbiological Sample sizeSix replicate introgression lines carried through 16 generations; 64 DGRP lines with genotype subgroups (n = 23, 13, 8, 22); a priori power/sample-size justification not described GroupsVT (temperate) vs. SK (tropical) parental lines vs. F16 introgressed pools; DGRP genotype classes at 2R and X loci; control vs. heat-shock RNA-seq conditions Pairingunclear Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionBonferroni correction for the genome-wide Fisher's exact test scan; empirical permutation-based null distributions for locus-specific seasonal GLM tests
Statistical tests used
Test Applied to n Assumptions
Wald test LT50 comparison of adult vs. embryo heat-stress survival (Fig. 1A inset, citing prior study) not stated
Welch's t-test embryonic heat tolerance of F16 introgressed populations vs. parental lines (SI Appendix, Fig. S1) not stated
Fisher's exact test (SNP-wise, Bonferroni corrected) genome-wide scan for allele-frequency differentiation in the synthetic F16 pool vs. parentals (Fig. 2 B and C) not stated
Pearson correlation (r) clinal allele frequency vs. latitude (Fig. 3D) and seasonal allele frequency vs. temperature variance (Fig. 4 B–D) not stated
Generalized linear model with permutation-based significance (100 permutations) locus-specific seasonal selection test at top 2R and X SNPs (Fig. 4A; SI Appendix, Fig. S6A) 100 permutations not stated
Two-way ANOVA (2R×X genotype interaction) embryonic survival across DGRP genotype combinations (Fig. 5B), F1,55 = 11.22 64 DGRP lines (subgroups n = 23, 13, 8, 22) not stated
Pairwise group comparison (test not specified) embryonic survival comparisons between individual genotype classes (tropical vs. temperate; single-locus effects, Fig. 5 B–D) subsets of the 64 DGRP lines not stated
Wald test (RNA-seq differential expression model) genotype × temperature interaction on SP70 and sog transcript levels (Fig. 5 I and K) not stated
Approaches that could also have been used
  • Genome-wide Fisher's exact test P-values were corrected using a Bonferroni threshold.
    Could also: A false discovery rate procedure such as Benjamini-Hochberg — FDR control is often used in genome-wide SNP scans because it is less conservative than Bonferroni while still bounding the expected proportion of false positives, which can be useful when many correlated tests are performed across the genome.
  • Significance of the seasonal GLM beta coefficients was assessed by comparing the observed estimate to a distribution of 100 permutations.
    Could also: A mixed-effects or autoregressive time-series model with parametric standard errors — Explicitly modeling temporal autocorrelation in allele-frequency trajectories can complement a permutation-based approach and may provide additional information such as confidence intervals on the effect size.
  • Embryonic survival proportions across DGRP genotype combinations were analyzed with ANOVA and pairwise comparisons.
    Could also: A generalized linear model with a binomial error structure (logistic regression) — Because survival is a proportion/count outcome, a binomial GLM directly models the underlying data-generating process and avoids relying on the normality assumptions that ANOVA requires for continuous outcomes.
  • Clinal relationships between allele frequency and latitude/temperature variance were summarized with Pearson correlation coefficients.
    Could also: A nonparametric Spearman correlation or a spatial regression accounting for autocorrelation among sampling sites — These alternatives can be preferred when sample sizes are small (as in the South American cline) or when spatial non-independence among sampled populations is a concern.
  • Differential expression of SP70 and sog was tested per-gene with a Wald test for the genotype × temperature interaction.
    Could also: A transcriptome-wide multiple-testing correction (e.g., FDR) applied across all tested genes — If additional genes beyond the two focal candidates were tested in the same RNA-seq dataset, an FDR-adjusted framework would jointly control error rates across the full set of comparisons.
  • The 41.7% allele-frequency shift observed during introgression was compared to an analytically expected value under the introgression design.
    Could also: A simulation-based null distribution (as was already used for the F ST scan) applied similarly to this specific allele-frequency shift — Using the same simulation-based benchmarking approach across all key statistics would provide a directly comparable measure of how unusual the observed shift is relative to drift alone.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

fet_loci_2R_X
Reported
two introgression loci: chr2R (~19.5 Mb) and chrX (~15.5 Mb)
Reproduced
chrX densest consensus peak 15.5-16.0 Mb (exact match); chr2R broad signal 17-20.5 Mb with named top SNP recovered (densest sub-window 17.0-17.5 Mb)
partial
fet_top_2R
Reported
top SNP 2R:20,551,633, delta_p=0.41
Reproduced
2R:20,551,633 genome-wide significant (minp=6.45e-13); |meanF16_af-VT8_af|=0.414
within tolerance
fet_top_X
Reported
top SNP X:15,607,604, delta_p=0.33
Reproduced
X:15,607,604 significant in 5/6 crosses (minp=2.31e-15); delta_p=0.393
within tolerance
fet_sig_count
Reported
391 sig SNPs genome-wide (FET P<=0.01 Bonferroni)
Reproduced
455 SNPs sig in >=3/6 crosses (Bonf0.01) - same order of magnitude via independent caller
partial
dgrp_embryo_tropVtemp
Reported
ANOVA F(1,55)=11.22 P=0.0014 / P=0.0096, tropical HIGHER hatching
Reproduced
F(1,29)=11.60 P=0.00195 on shipped corrected data, but TEMPERATE higher (0.461 vs 0.259) and group-Ns swapped (23 vs 8)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 64/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

The primary finding — two-locus introgression mapping on chr2R and chrX — reproduces well from raw pool-seq via a fully independent pipeline: the chrX 15.5 Mb locus is exact, both named top SNPs (2R:20,551,633 delta_p≈0.41; X:15,607,604) are recovered as genome-wide significant, and the sig count (455) is the same order as 391. The count and exact chr2R peak differ because the authors' PoolSNP/DEST SNP set and FET-input tables were not deposited and the cross-consensus rule is under-specified — these are our-method/data-availability deviations, not authors' defects. The secondary DGRP embryonic ANOVA reproduces in magnitude but reverses direction (temperate higher) with swapped group-Ns on the authors' own shipped correction data — an authors-side main-text/correction inconsistency that is flagged for human reconciliation. Overall: a solid partial reproduction with explainable genomic deviations plus one genuine direction-flip flag, hence yellow throughout.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

588.7 k
tokens (I/O) · 61.4 M incl. cache
214 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.