QTLs for heat-induced stomatal anatomy underpin gas exchange variation in field-grown wheat.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduction of Chaplin et al. 2026 (Front Plant Sci), a field-wheat heat-stress study mapping QTLs for stomatal anatomy. Described well enough to reproduce without author contact; the clearly-specified, low-hanging pipeline outputs reproduce essentially 1:1. INDEPENDENT re-derivation (C1-C9) from the public per-leaf trait data (Zenodo grdc2023/2024.rds; 3,200 + 1,200 leaf x surface records) reproduces the paper's reported phenotypic percentage-changes, trait means and trait correlations to 3-4 significant figures: gs decline 28.9% (adax 33.9% / abax 14.8%), SD increase +6.4/+6.7% (S1) and 35.4->40.2 (+13.5%) / 46.9->50.4 (+7.3%) (S2), SA decline -8.2% (S1) and -4.7/-6.2% (S2), gsmax +7.1% (S2), and correlations SDgs R2=0.1383, SDSA R2=0.321/0.312, gsmax~gs R2=0.213 - all exact. CONSISTENCY check (C10-C12): the QTL/heritability driver genome-to-phenome-grdc.R depends on the COMMERCIAL licence-keyed asreml (ASReml-R) package, which cannot be installed on «our HPC» without a paid licence, so a fresh re-run was not possible (env_unresolvable); instead the deposit's own grdc_gwas_results.rds was loaded and shown to reproduce Table 3 (all 56 Cullis heritabilities to 3 dp), Table 4 (60 QTLs in 2023 = 25 abaxial + 35 adaxial; 62 across environments) and the chr1A gsmax LOD 8.8 EXACTLY - a deposit-vs-paper internal-consistency check, weaker than C1-C9 and flagged as such. NOT attempted (the hard ~20%): (1) the fresh asreml+wgaim QTL/heritability re-run (commercial licence); (2) the YOLOv8-M image-model metrics in Table 1 (Precision 0.9568, Recall 0.9602, mAP50 0.9441, mAP50-95 0.6775) - the annotated microscopy image dataset / train-val-test split is NOT in the Zenodo deposit (phenotype+genotype only) and not in the GitHub FieldDinoMicroscopy repo, which ships only a trained .pt from a different run (326 epochs vs the paper's 119), so the metrics cannot be recomputed (data_unavailable); (3) the exact LMM ANOVA F/p-values - Physiology Figures.R reads 20240423_grdc.csv (with the Rep/Leaf columns the lme4 model needs) which is not deposited, so the descriptive effects were reproduced instead. No fabrication concern: every checkable value is independently derivable from the public CC-BY deposit, the descriptive statistics reproduce exactly with no author code, and the shipped GWAS object is internally consistent with the published tables. Technical note for the next agent: «our HPC» COMPUTE nodes have no outbound internet (first «job» failed on Zenodo downloads + conda build); data and the conda R env (reused: .conda_envs/repro-wgcna-39754813, R 4.5.3) were staged from front1 (python3/wget) onto «infra», then the compute «job» only read local files. All grades provisional; human reviewer decides via AUDIT.md.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-14 ⛓ 401e488ba881
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDelayed sowing exposes wheat anthesis to higher temperatures (~1-3°C), which the authors hypothesised would reduce operating stomatal conductance, increase stomatal density and reduce stomatal size, with substantial genotypic variation enabling identification of QTLs for heat tolerance; the study tests how stomatal anatomy and physiology integrate to determine wheat heat tolerance under field conditions.
- ★ Heat stress (delayed sowing) uncouples stomatal anatomy from physiological performance: delayed sowing impaired stomatal function despite similar theoretical anatomical capacity (g_smax). finding
- ★ The adaxial leaf surface consistently shows higher g_s, stomatal density and g_smax than the abaxial surface, making it the dominant surface for gas exchange and stomatal anatomical variation. finding
- ★ 62 putative QTLs were identified across environments for stomatal traits, with recurring/co-localised loci on 6B and pleiotropic QTLs on 1A and 2A; the majority (36) were for the adaxial surface. resource
- ★ No QTLs were detected for stomatal physiological traits, indicating limited potential for indirect selection relative to more stable anatomical traits. finding
- ★ Delayed sowing induced plastic anatomical shifts toward smaller, denser stomata, particularly on the adaxial surface. finding
- 21 QTLs were consistent with chromosomal regions previously reported for stomatal anatomical traits in wheat, particularly on chromosome 7A. finding
- ★ Significant genotypic variation and moderate heritability were observed for stomatal anatomical traits. finding
- A high-throughput field phenotyping pipeline (LI-600 porometer, handheld digital microscope/FieldDino app, YOLOv8-based deep learning image analysis) enables breeding-scale stomatal measurement. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Stomatal conductance (porometry) | field-grown wheat (Triticum aestivum) flag leaves, 200 genotypes S1 / 50 genotypes S2 | heat stress via delayed sowing (TOS2 vs timely TOS1) | stomatal conductance g_s of adaxial and abaxial surfaces | LI-COR LI-600 porometer/fluorometer |
| Stomatal anatomy imaging / deep-learning image analysis | field-grown wheat flag leaves, adaxial and abaxial surfaces | timely vs delayed sowing | stomatal density, guard cell width/length, stomatal area, stomatal size, g_smax | Dino-Lite handheld USB microscope (200× S1; 400× S2), YOLOv8-M deep learning model, Roboflow annotation |
| Grain yield and yield component analysis | field-grown wheat plots, Narrabri NSW | timely vs delayed sowing | grain yield per hectare, thousand kernel weight, screenings, grain protein, test weight, moisture | optical seed counter (Contador), near-infrared spectroscopy (FOSS) |
| Derived gas-exchange efficiency calculation | field-grown wheat flag leaves | timely vs delayed sowing | stomatal conductance operating efficiency g_se (g_sop/g_smax) | — |
| QTL mapping / linkage analysis | wheat genotype panel (CIMMYT/ICARDA/University of Sydney germplasm) across environments and seasons | two times of sowing across two seasons | QTLs for stomatal anatomical and physiological traits | — |
- ▼ Early (timely) sowing supported higher g_s and g_se while delayed sowing impaired stomatal function despite similar g_smax
- ▲ Adaxial surface exhibited higher g_s, stomatal density and g_smax than abaxial surface
- – 62 putative QTLs detected across environments for stomatal traits 62 QTLs
- – 36 of the QTLs were detected for the adaxial surface 36 QTLs
- – 21 QTLs consistent with previously reported wheat stomatal anatomical loci, particularly chromosome 7A 21 QTLs
- – No QTLs detected for stomatal physiological traits 0 QTLs
- – Delayed sowing produced smaller, denser stomata, especially adaxially
- ▲ Delayed sowing (TOS2) raised mean daily max temperature at anthesis relative to TOS1 (S1: 24.2°C vs 27.1°C; S2: 22.0°C vs 26.7°C) ~1-3°C
- pvalue <0.001 (TOS effect on stomatal conductance, guard cell width/length, stomatal area, stomatal density, g_se, yield (2023 LMM ANOVA))
- pvalue 0.490 (TOS effect on g_smax (non-significant) in 2023 LMM ANOVA)
- pvalue <0.001 (Surface effect on stomatal conductance, guard cell width, stomatal density, g_smax, g_se)
- other Precision 0.9568, Recall 0.9602, mAP50 0.9441, mAP50-95 0.6775 (YOLOv8-M deep learning model performance metrics for stomatal detection)
- count 200 genotypes (S1), 50 genotypes (S2) (genotype panel sizes across two seasons of field trials)
- other 6%–10% yield reduction per 1°C (cited literature on temperature effect on wheat yield)
- count n=4 per genotype per TOS (S1); n=6 per genotype per TOS (S2) (leaf sampling per genotype per time of sowing)
- count 121 images initially annotated (images annotated via instance segmentation in Roboflow to train the model)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used a randomised complete block design with two time-of-sowing (TOS) treatments across two consecutive field seasons (200 genotypes in season 1, 50 in season 2) to examine stomatal anatomical and physiological traits under contrasting temperature regimes. Linear mixed models with factorial ANOVA were applied to partition variance attributable to TOS, genotype (variety), leaf surface, and all two-way and three-way interactions. QTL analysis was conducted across multiple environments to identify 62 putative genomic loci for stomatal traits, and moderate heritability was reported for anatomical traits. Stomatal anatomy was quantified from field microscopy images using a YOLOv8-M deep-learning pipeline.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Linear mixed model ANOVA (full factorial with TOS × variety × surface interactions) | Main effects and interactions of TOS, variety, and leaf surface on stomatal conductance, guard cell width, guard cell length, stomatal area, stomatal density, g_smax, g_se, and yield (Table 2) | S1: 200 genotypes × 2 TOS × 2 field replicates, n=4 leaves per genotype per TOS; S2: 50 genotypes × 2 TOS × 2 replicates, n=6 leaves per genotype per TOS | not stated |
| QTL mapping (specific algorithm — e.g. composite interval mapping, GWAS — not described in provided text) | Detection of 62 putative QTLs across environments for stomatal anatomical and physiological traits and yield | 200 genotypes (S1); 50 genotypes (S2) | not stated |
| Heritability estimation (method not specified in provided text) | Stomatal anatomical traits described as showing 'moderate heritability'; specific H² or h² values not given in provided text | — | not stated |
| YOLOv8-M instance segmentation deep-learning model (precision, recall, mAP evaluation) | Automated quantification of stomatal anatomical traits from field microscopy images (Table 1) | 121 images used for initial annotation | na |
-
Eight traits were each tested with ANOVA terms from the same factorial mixed model, generating a family of p-values without explicit multiplicity correction↳ Could also: Applying a Benjamini-Hochberg FDR correction across the family of trait-level tests, or a multivariate mixed model (MANOVA) treating all traits jointly, would also be standard approaches — When several correlated traits are tested in the same experiment, a joint or corrected approach makes the expected false-discovery rate explicit; this is complementary to per-trait p-values and aids interpretation of which effects are robust across the full trait set
-
QTL detection across 62 loci and multiple environments used an unspecified mapping method; genome-wide significance thresholds are not described in the provided text↳ Could also: Permutation-based genome-wide LOD thresholds (e.g., 1,000 permutations at α = 0.05) or mixed-model GWAS (e.g., BLINK, FarmCPU, or GAPIT) accounting for population structure and kinship would also be standard options for a diverse panel — For germplasm panels with population structure (CIMMYT/ICARDA diversity), GWAS with kinship correction reduces spurious associations; permutation thresholds calibrated to the specific marker density and population size make QTL detection criteria reproducible and directly comparable across studies
-
Heritability of stomatal anatomical traits was described qualitatively as 'moderate' without specifying the estimation method or reporting a numerical estimate with uncertainty↳ Could also: Broad-sense heritability (H²) or narrow-sense heritability (h²) estimated via REML from the mixed model variance components, reported with bootstrap or likelihood-based confidence intervals, is a standard approach in multi-environment field trials — A numerical heritability estimate with confidence interval and a clearly stated method (H² vs. h², REML vs. ANOVA-based) supports cross-study comparisons, informs expected selection response, and distinguishes genetic from residual and genotype-by-environment variance
-
Each TOS treatment had two field replicates per season, limiting residual degrees of freedom for variance component estimation↳ Could also: An augmented or alpha-lattice design with three or more replicates, or inclusion of replicated check varieties across incomplete blocks, would also support more precise spatial error correction and variance partitioning — Two replicates constrain the precision of residual variance estimates and heritability calculations; augmented designs are widely used in large wheat field trials to balance genotype coverage with replication, and spatial models (e.g., P-spline or AR1 × AR1 error structure) further improve efficiency on heterogeneous field soils
-
Dispersion around genotype or treatment means is not reported in the provided text (no SD, SEM, or CI values are given for trait means)↳ Could also: Reporting best linear unbiased predictions (BLUPs) or least-squares means extracted from the fitted mixed model, paired with their standard errors or 95% confidence intervals, is standard practice for multi-environment trials — BLUPs from the mixed model correctly propagate the variance-covariance structure and shrink noisy estimates; confidence intervals allow readers to judge the precision of individual genotype rankings and to assess whether observed differences are practically meaningful beyond statistical significance
-
The same 50 genotypes were evaluated in season 2 as a subset of season 1, but the provided text does not describe a formal joint cross-season analysis (e.g., a combined multi-environment model)↳ Could also: A single combined linear mixed model including season as a factor (with genotype-by-season interaction as a random effect) or a factor analytic model for multi-environment trials would also be a standard approach to estimate genotype × environment interaction and stability — Joint analysis across seasons with an explicit genotype-by-environment (G×E) term quantifies the stability of QTL effects across thermal environments and distinguishes QTLs with consistent versus environment-specific effects, which is directly relevant to selection across years
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-42231893
Paper: Chaplin E, Tanaka E, Merchant A, Sznajder B, Trethowan R, Salter W. (2026) QTLs for heat-induced stomatal anatomy underpin gas exchange variation in field-grown wheat. Front Plant Sci. DOI 10.3389/fpls.2026.1769384.
Code: github.com/williamtsalter/FieldDinoMicroscopy (FieldDino image-capture app
- YOLOv8 stomatal-inference model/script). Data + analysis code: Zenodo
10.5281/zenodo.19706766 (CC-BY-4.0) — ships per-season trait
.rds(grdc2023.rds,grdc2024.rds), genotype cross (data_cross_imputed.rds), the already-computed GWAS results object (grdc_gwas_results.rds), a literature-QTL table, the raw field-data.xlsx, and 4 R scripts (pc-selection.R,genome-to-phenome-grdc.R,Physiology Figures.R,Genome Phenome Figures & Tables.R).
Computational pipelines in this paper
| # | Pipeline | Software | Produces | In scope? |
|---|---|---|---|---|
| P1 | Stomatal image analysis | FieldDino + YOLOv8-M (Ultralytics) + OpenCV ellipse fit | SD, GCL, GCW, SA, gsmax per leaf; Table 1 model metrics | Partial / NO — see below |
| P2 | Phenotypic statistics | R lme4 / lmerTest / emmeans / MASS | trait means, TOS %-changes, correlations, ANOVA p-values | YES (descriptive part) |
| P3 | Heritability | asreml + heritable::H2 (Cullis 2006) |
Table 3 broad-sense H² | Verify-only (see below) |
| P4 | QTL / GWAS | asreml + wgaim (genome-to-phenome-grdc.R) |
Table 4 / S4 / S5, 60–62 putative QTLs, LOD | Verify-only (see below) |
In scope — attempted (open-source, independent re-derivation)
- P2 descriptive: From
grdc2023.rds/grdc2024.rds(per-leaf trait records:length,width,area,density,gs,gsmax,gse×year,tos,surface,name+design) we recompute trait means by year/surface/TOS, the reported TOS1→TOS2 % changes (SD, SA, gsmax, gs), and the reported trait correlations (R²) vialm. These need no author code and no license.
In scope — VERIFY-ONLY (shipped results object, provisional)
- P3/P4: A fresh
asreml+wgaimre-run is blocked:asreml(ASReml-R, VSN International) is a commercial, license-keyed package — unobtainable on «our HPC» without a paid license (drop_reason: env_unresolvablefor the fresh re-run). Instead we load the shippedgrdc_gwas_results.rds(the authors' ownlist(gwas=…, H2=…)) and check it reproduces the paper's Table 3 heritabilities and Table 4 / text QTL counts + the chr1A gsmax LOD 8.8. This is an internal-consistency check of the deposit against the publication, not an independent re-computation — graded and flagged as such.
Out of scope / NOT attempted (the hard ~20%)
- P1 / Table 1 YOLOv8-M metrics (Precision 0.9568, Recall 0.9602, mAP50 0.9441,
mAP50-95 0.6775): the annotated microscopy image dataset / train-val-test split
is NOT in the Zenodo deposit (phenotype+genotype only) and not in the GitHub
repo (which ships only a trained
.ptfrom a different run — 326 epochs vs the paper's 119). Without the images the model metrics cannot be recomputed → effectivelydata_unavailablefor this sub-result. Not attempted. - Fresh P3/P4 re-run —
asremllicense (env_unresolvable), see above. - Exact P2 ANOVA p-values from the LMM —
Physiology Figures.Rreads20240423_grdc.csv(with theRep/Leafcolumns the lme4 model needs); that CSV is not in the deposit, so the exact mixed-model F/p cannot be reproduced 1:1. We report the descriptive effects (means, %-changes, correlations) instead. - All wet-lab / field instrumentation (LI-600 porometry, NIR grain quality, soil moisture) — measured, not pipeline-derived.
Heavy-compute compliance
All compute on «our HPC» (SLURM job, «infra» workdir); Zenodo data downloaded inside the job onto «infra»; only small result TSVs pulled back to «host». No data on «host».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.