Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

QTLs for heat-induced stomatal anatomy underpin gas exchange variation in field-grown wheat.

Front Plant Sci · 2026
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduction of Chaplin et al. 2026 (Front Plant Sci), a field-wheat heat-stress study mapping QTLs for stomatal anatomy. Described well enough to reproduce without author contact; the clearly-specified, low-hanging pipeline outputs reproduce essentially 1:1. INDEPENDENT re-derivation (C1-C9) from the public per-leaf trait data (Zenodo grdc2023/2024.rds; 3,200 + 1,200 leaf x surface records) reproduces the paper's reported phenotypic percentage-changes, trait means and trait correlations to 3-4 significant figures: gs decline 28.9% (adax 33.9% / abax 14.8%), SD increase +6.4/+6.7% (S1) and 35.4->40.2 (+13.5%) / 46.9->50.4 (+7.3%) (S2), SA decline -8.2% (S1) and -4.7/-6.2% (S2), gsmax +7.1% (S2), and correlations SDgs R2=0.1383, SDSA R2=0.321/0.312, gsmax~gs R2=0.213 - all exact. CONSISTENCY check (C10-C12): the QTL/heritability driver genome-to-phenome-grdc.R depends on the COMMERCIAL licence-keyed asreml (ASReml-R) package, which cannot be installed on «our HPC» without a paid licence, so a fresh re-run was not possible (env_unresolvable); instead the deposit's own grdc_gwas_results.rds was loaded and shown to reproduce Table 3 (all 56 Cullis heritabilities to 3 dp), Table 4 (60 QTLs in 2023 = 25 abaxial + 35 adaxial; 62 across environments) and the chr1A gsmax LOD 8.8 EXACTLY - a deposit-vs-paper internal-consistency check, weaker than C1-C9 and flagged as such. NOT attempted (the hard ~20%): (1) the fresh asreml+wgaim QTL/heritability re-run (commercial licence); (2) the YOLOv8-M image-model metrics in Table 1 (Precision 0.9568, Recall 0.9602, mAP50 0.9441, mAP50-95 0.6775) - the annotated microscopy image dataset / train-val-test split is NOT in the Zenodo deposit (phenotype+genotype only) and not in the GitHub FieldDinoMicroscopy repo, which ships only a trained .pt from a different run (326 epochs vs the paper's 119), so the metrics cannot be recomputed (data_unavailable); (3) the exact LMM ANOVA F/p-values - Physiology Figures.R reads 20240423_grdc.csv (with the Rep/Leaf columns the lme4 model needs) which is not deposited, so the descriptive effects were reproduced instead. No fabrication concern: every checkable value is independently derivable from the public CC-BY deposit, the descriptive statistics reproduce exactly with no author code, and the shipped GWAS object is internally consistent with the published tables. Technical note for the next agent: «our HPC» COMPUTE nodes have no outbound internet (first «job» failed on Zenodo downloads + conda build); data and the conda R env (reused: .conda_envs/repro-wgcna-39754813, R 4.5.3) were staged from front1 (python3/wget) onto «infra», then the compute «job» only read local files. All grades provisional; human reviewer decides via AUDIT.md.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.19706766

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-14 ⛓ 401e488ba881
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Delayed sowing exposes wheat anthesis to higher temperatures (~1-3°C), which the authors hypothesised would reduce operating stomatal conductance, increase stomatal density and reduce stomatal size, with substantial genotypic variation enabling identification of QTLs for heat tolerance; the study tests how stomatal anatomy and physiology integrate to determine wheat heat tolerance under field conditions.

Core claims
  • Heat stress (delayed sowing) uncouples stomatal anatomy from physiological performance: delayed sowing impaired stomatal function despite similar theoretical anatomical capacity (g_smax). finding
  • The adaxial leaf surface consistently shows higher g_s, stomatal density and g_smax than the abaxial surface, making it the dominant surface for gas exchange and stomatal anatomical variation. finding
  • 62 putative QTLs were identified across environments for stomatal traits, with recurring/co-localised loci on 6B and pleiotropic QTLs on 1A and 2A; the majority (36) were for the adaxial surface. resource
  • No QTLs were detected for stomatal physiological traits, indicating limited potential for indirect selection relative to more stable anatomical traits. finding
  • Delayed sowing induced plastic anatomical shifts toward smaller, denser stomata, particularly on the adaxial surface. finding
  • 21 QTLs were consistent with chromosomal regions previously reported for stomatal anatomical traits in wheat, particularly on chromosome 7A. finding
  • Significant genotypic variation and moderate heritability were observed for stomatal anatomical traits. finding
  • A high-throughput field phenotyping pipeline (LI-600 porometer, handheld digital microscope/FieldDino app, YOLOv8-based deep learning image analysis) enables breeding-scale stomatal measurement. method
Experimental setups
Assay System Perturbation Readout Platform
Stomatal conductance (porometry) field-grown wheat (Triticum aestivum) flag leaves, 200 genotypes S1 / 50 genotypes S2 heat stress via delayed sowing (TOS2 vs timely TOS1) stomatal conductance g_s of adaxial and abaxial surfaces LI-COR LI-600 porometer/fluorometer
Stomatal anatomy imaging / deep-learning image analysis field-grown wheat flag leaves, adaxial and abaxial surfaces timely vs delayed sowing stomatal density, guard cell width/length, stomatal area, stomatal size, g_smax Dino-Lite handheld USB microscope (200× S1; 400× S2), YOLOv8-M deep learning model, Roboflow annotation
Grain yield and yield component analysis field-grown wheat plots, Narrabri NSW timely vs delayed sowing grain yield per hectare, thousand kernel weight, screenings, grain protein, test weight, moisture optical seed counter (Contador), near-infrared spectroscopy (FOSS)
Derived gas-exchange efficiency calculation field-grown wheat flag leaves timely vs delayed sowing stomatal conductance operating efficiency g_se (g_sop/g_smax)
QTL mapping / linkage analysis wheat genotype panel (CIMMYT/ICARDA/University of Sydney germplasm) across environments and seasons two times of sowing across two seasons QTLs for stomatal anatomical and physiological traits
Key results
  • Early (timely) sowing supported higher g_s and g_se while delayed sowing impaired stomatal function despite similar g_smax
  • Adaxial surface exhibited higher g_s, stomatal density and g_smax than abaxial surface
  • 62 putative QTLs detected across environments for stomatal traits 62 QTLs
  • 36 of the QTLs were detected for the adaxial surface 36 QTLs
  • 21 QTLs consistent with previously reported wheat stomatal anatomical loci, particularly chromosome 7A 21 QTLs
  • No QTLs detected for stomatal physiological traits 0 QTLs
  • Delayed sowing produced smaller, denser stomata, especially adaxially
  • Delayed sowing (TOS2) raised mean daily max temperature at anthesis relative to TOS1 (S1: 24.2°C vs 27.1°C; S2: 22.0°C vs 26.7°C) ~1-3°C
Key statistics
  • pvalue <0.001 (TOS effect on stomatal conductance, guard cell width/length, stomatal area, stomatal density, g_se, yield (2023 LMM ANOVA))
  • pvalue 0.490 (TOS effect on g_smax (non-significant) in 2023 LMM ANOVA)
  • pvalue <0.001 (Surface effect on stomatal conductance, guard cell width, stomatal density, g_smax, g_se)
  • other Precision 0.9568, Recall 0.9602, mAP50 0.9441, mAP50-95 0.6775 (YOLOv8-M deep learning model performance metrics for stomatal detection)
  • count 200 genotypes (S1), 50 genotypes (S2) (genotype panel sizes across two seasons of field trials)
  • other 6%–10% yield reduction per 1°C (cited literature on temperature effect on wheat yield)
  • count n=4 per genotype per TOS (S1); n=6 per genotype per TOS (S2) (leaf sampling per genotype per time of sowing)
  • count 121 images initially annotated (images annotated via instance segmentation in Roboflow to train the model)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a randomised complete block design with two time-of-sowing (TOS) treatments across two consecutive field seasons (200 genotypes in season 1, 50 in season 2) to examine stomatal anatomical and physiological traits under contrasting temperature regimes. Linear mixed models with factorial ANOVA were applied to partition variance attributable to TOS, genotype (variety), leaf surface, and all two-way and three-way interactions. QTL analysis was conducted across multiple environments to identify 62 putative genomic loci for stomatal traits, and moderate heritability was reported for anatomical traits. Stomatal anatomy was quantified from field microscopy images using a YOLOv8-M deep-learning pipeline.

Replicationbiological Sample sizeS1: 200 genotypes, 2 TOS treatments, 2 field replicates, 4 leaves per genotype per TOS; S2: 50 genotypes, 2 TOS treatments, 2 field replicates, 6 leaves per genotype per TOS; no formal power calculation or sample-size justification stated GroupsTwo TOS treatments (timely vs. delayed/heat-stressed sowing) across 200 or 50 wheat genotypes; adaxial vs. abaxial leaf surface Pairingpaired Randomization/blindingstated Dispersionunclear Exact p-valuesyes Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Linear mixed model ANOVA (full factorial with TOS × variety × surface interactions) Main effects and interactions of TOS, variety, and leaf surface on stomatal conductance, guard cell width, guard cell length, stomatal area, stomatal density, g_smax, g_se, and yield (Table 2) S1: 200 genotypes × 2 TOS × 2 field replicates, n=4 leaves per genotype per TOS; S2: 50 genotypes × 2 TOS × 2 replicates, n=6 leaves per genotype per TOS not stated
QTL mapping (specific algorithm — e.g. composite interval mapping, GWAS — not described in provided text) Detection of 62 putative QTLs across environments for stomatal anatomical and physiological traits and yield 200 genotypes (S1); 50 genotypes (S2) not stated
Heritability estimation (method not specified in provided text) Stomatal anatomical traits described as showing 'moderate heritability'; specific H² or h² values not given in provided text not stated
YOLOv8-M instance segmentation deep-learning model (precision, recall, mAP evaluation) Automated quantification of stomatal anatomical traits from field microscopy images (Table 1) 121 images used for initial annotation na
Approaches that could also have been used
  • Eight traits were each tested with ANOVA terms from the same factorial mixed model, generating a family of p-values without explicit multiplicity correction
    Could also: Applying a Benjamini-Hochberg FDR correction across the family of trait-level tests, or a multivariate mixed model (MANOVA) treating all traits jointly, would also be standard approaches — When several correlated traits are tested in the same experiment, a joint or corrected approach makes the expected false-discovery rate explicit; this is complementary to per-trait p-values and aids interpretation of which effects are robust across the full trait set
  • QTL detection across 62 loci and multiple environments used an unspecified mapping method; genome-wide significance thresholds are not described in the provided text
    Could also: Permutation-based genome-wide LOD thresholds (e.g., 1,000 permutations at α = 0.05) or mixed-model GWAS (e.g., BLINK, FarmCPU, or GAPIT) accounting for population structure and kinship would also be standard options for a diverse panel — For germplasm panels with population structure (CIMMYT/ICARDA diversity), GWAS with kinship correction reduces spurious associations; permutation thresholds calibrated to the specific marker density and population size make QTL detection criteria reproducible and directly comparable across studies
  • Heritability of stomatal anatomical traits was described qualitatively as 'moderate' without specifying the estimation method or reporting a numerical estimate with uncertainty
    Could also: Broad-sense heritability (H²) or narrow-sense heritability (h²) estimated via REML from the mixed model variance components, reported with bootstrap or likelihood-based confidence intervals, is a standard approach in multi-environment field trials — A numerical heritability estimate with confidence interval and a clearly stated method (H² vs. h², REML vs. ANOVA-based) supports cross-study comparisons, informs expected selection response, and distinguishes genetic from residual and genotype-by-environment variance
  • Each TOS treatment had two field replicates per season, limiting residual degrees of freedom for variance component estimation
    Could also: An augmented or alpha-lattice design with three or more replicates, or inclusion of replicated check varieties across incomplete blocks, would also support more precise spatial error correction and variance partitioning — Two replicates constrain the precision of residual variance estimates and heritability calculations; augmented designs are widely used in large wheat field trials to balance genotype coverage with replication, and spatial models (e.g., P-spline or AR1 × AR1 error structure) further improve efficiency on heterogeneous field soils
  • Dispersion around genotype or treatment means is not reported in the provided text (no SD, SEM, or CI values are given for trait means)
    Could also: Reporting best linear unbiased predictions (BLUPs) or least-squares means extracted from the fitted mixed model, paired with their standard errors or 95% confidence intervals, is standard practice for multi-environment trials — BLUPs from the mixed model correctly propagate the variance-covariance structure and shrink noisy estimates; confidence intervals allow readers to judge the precision of individual genotype rankings and to assess whether observed differences are practically meaningful beyond statistical significance
  • The same 50 genotypes were evaluated in season 2 as a subset of season 1, but the provided text does not describe a formal joint cross-season analysis (e.g., a combined multi-environment model)
    Could also: A single combined linear mixed model including season as a factor (with genotype-by-season interaction as a random effect) or a factor analytic model for multi-environment trials would also be a standard approach to estimate genotype × environment interaction and stability — Joint analysis across seasons with an explicit genotype-by-environment (G×E) term quantifies the stability of QTL effects across thermal environments and distinguishes QTLs with consistent versus environment-specific effects, which is directly relevant to selection across years
Software: YOLOv8-M (Ultralytics, via Chaplin et al. 2025a pipeline) · Roboflow (image annotation platform) · Python / PyQt5 (FieldDino image capture application) · LI-COR LI-600 porometer/fluorometer (hardware; stomatal conductance measurement)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-42231893

Paper: Chaplin E, Tanaka E, Merchant A, Sznajder B, Trethowan R, Salter W. (2026) QTLs for heat-induced stomatal anatomy underpin gas exchange variation in field-grown wheat. Front Plant Sci. DOI 10.3389/fpls.2026.1769384.

Code: github.com/williamtsalter/FieldDinoMicroscopy (FieldDino image-capture app

  • YOLOv8 stomatal-inference model/script). Data + analysis code: Zenodo 10.5281/zenodo.19706766 (CC-BY-4.0) — ships per-season trait .rds (grdc2023.rds, grdc2024.rds), genotype cross (data_cross_imputed.rds), the already-computed GWAS results object (grdc_gwas_results.rds), a literature-QTL table, the raw field-data .xlsx, and 4 R scripts (pc-selection.R, genome-to-phenome-grdc.R, Physiology Figures.R, Genome Phenome Figures & Tables.R).

Computational pipelines in this paper

# Pipeline Software Produces In scope?
P1 Stomatal image analysis FieldDino + YOLOv8-M (Ultralytics) + OpenCV ellipse fit SD, GCL, GCW, SA, gsmax per leaf; Table 1 model metrics Partial / NO — see below
P2 Phenotypic statistics R lme4 / lmerTest / emmeans / MASS trait means, TOS %-changes, correlations, ANOVA p-values YES (descriptive part)
P3 Heritability asreml + heritable::H2 (Cullis 2006) Table 3 broad-sense H² Verify-only (see below)
P4 QTL / GWAS asreml + wgaim (genome-to-phenome-grdc.R) Table 4 / S4 / S5, 60–62 putative QTLs, LOD Verify-only (see below)

In scope — attempted (open-source, independent re-derivation)

  • P2 descriptive: From grdc2023.rds/grdc2024.rds (per-leaf trait records: length,width,area,density,gs,gsmax,gse × year,tos,surface,name+design) we recompute trait means by year/surface/TOS, the reported TOS1→TOS2 % changes (SD, SA, gsmax, gs), and the reported trait correlations (R²) via lm. These need no author code and no license.

In scope — VERIFY-ONLY (shipped results object, provisional)

  • P3/P4: A fresh asreml+wgaim re-run is blocked: asreml (ASReml-R, VSN International) is a commercial, license-keyed package — unobtainable on «our HPC» without a paid license (drop_reason: env_unresolvable for the fresh re-run). Instead we load the shipped grdc_gwas_results.rds (the authors' own list(gwas=…, H2=…)) and check it reproduces the paper's Table 3 heritabilities and Table 4 / text QTL counts + the chr1A gsmax LOD 8.8. This is an internal-consistency check of the deposit against the publication, not an independent re-computation — graded and flagged as such.

Out of scope / NOT attempted (the hard ~20%)

  • P1 / Table 1 YOLOv8-M metrics (Precision 0.9568, Recall 0.9602, mAP50 0.9441, mAP50-95 0.6775): the annotated microscopy image dataset / train-val-test split is NOT in the Zenodo deposit (phenotype+genotype only) and not in the GitHub repo (which ships only a trained .pt from a different run — 326 epochs vs the paper's 119). Without the images the model metrics cannot be recomputed → effectively data_unavailable for this sub-result. Not attempted.
  • Fresh P3/P4 re-runasreml license (env_unresolvable), see above.
  • Exact P2 ANOVA p-values from the LMMPhysiology Figures.R reads 20240423_grdc.csv (with the Rep/Leaf columns the lme4 model needs); that CSV is not in the deposit, so the exact mixed-model F/p cannot be reproduced 1:1. We report the descriptive effects (means, %-changes, correlations) instead.
  • All wet-lab / field instrumentation (LI-600 porometry, NIR grain quality, soil moisture) — measured, not pipeline-derived.

Heavy-compute compliance

All compute on «our HPC» (SLURM job, «infra» workdir); Zenodo data downloaded inside the job onto «infra»; only small result TSVs pulled back to «host». No data on «host».

Figures / tables: TableFig 1AFigsFig 1BFig 5B
C1
Reported
S1 gs decline TOS1->TOS2 28.9% (adax 33.9%, abax 14.8%)
Reproduced
-28.9% / -33.89% / -14.77%
exact
C2
Reported
S1 SD increase +6.4% abax / +6.7% adax
Reproduced
+6.37% / +6.70%
exact
C3
Reported
S1 SA decline -8.2%
Reproduced
-8.21%
exact
C4
Reported
S1 GCL decline -5.0%
Reproduced
-5.04%
exact
C5
Reported
S1 correlations SD~gs R2=0.1383, SD~SA R2=0.321, gsmax~gs R2=0.213
Reproduced
0.13830, 0.32104, 0.21279 (n=3200)
exact
C6
Reported
S2 SD abax 35.4->40.2 (+13.5%), adax 46.9->50.4 (+7.3%)
Reproduced
35.38->40.16 (+13.52%), 46.94->50.36 (+7.29%)
exact
C7
Reported
S2 SA decline abax -4.7%, adax -6.2%
Reproduced
-4.70%, -6.16%
exact
C8
Reported
S2 gsmax abax 0.97->1.04 (+7.1%), adax stable 1.32
Reproduced
0.972->1.042 (+7.14%), 1.321->1.318
exact
C9
Reported
S2 SD~SA correlation R2=0.312
Reproduced
R2=0.31229 (n=1200)
exact
C10
Reported
Table 3 broad-sense heritability (Cullis), 56 values
Reproduced
all 56 reproduce to 3 dp from shipped grdc_gwas_results.rds$H2 (verify-only)
exact
C11
Reported
60 QTLs in 2023 (25 abaxial, 35 adaxial); 62 across environments
Reproduced
25+35=60 (2023); 62 total (verify-only, shipped gwas object)
exact
C12
Reported
gsmax QTL chr 1A, LOD = 8.8
Reproduced
chr1A gsmax 1:575536483 LOD=8.8015 (verify-only)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

167.7 k
tokens (I/O) · 10.6 M incl. cache
18 min
runtime · 0 CPU-h
0.2 GB
peak RAM
1
HPC jobs
hummel
machine