Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Ancient variation of the AvrPm17 gene in powdery mildew limits the effectiveness of the introgressed rye Pm17 resistance g

Proc Natl Acad Sci U S A · 2022
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1 on the shipped computational results). The AvrPm17 repo ships per-figure derived data + R scripts; applying the authors' own code to their own data on «our HPC» (SLURM «job», R 4.3.3) reproduced all 9 attempted pipeline-derived statistics. The headline statistic C1 (paired Wilcoxon varA vs varB) reproduces to the EXACT reported digits: 4.6566e-10 vs reported 4.657e-10. C2-C8 Wilcoxon tests all reproduce as significant with the same direction. C9 copy-number distribution matches the claim (majority 2 copies, 3 isolates >3) via a control-gene normalization reconstruction (the normalized intermediate av_covA_new.txt is not shipped). C11 QTL reproduces the peak chromosome+position exactly (chr9 ~339-340) and the chr1 secondary LOD within ~5%, but the absolute chr9 LOD is higher (24.9 vs 16.0) because the authors ship only the precomputed scan, not the scanone parameters. NOT attempted (out of scope, not shipped as runnable pipeline): raw-read mapping/SNP calling of the 151 SRA isolates, PacBio ISR7 assembly, E003 phylogeny + divergence dating, IntFOLD structure, and C12 (separate transgenic cross, wet-lab phenotypes). No reported value was found to be non-derivable from the shipped data -> no fabrication concern. Dataset profiling flagged PRJNA625429 N discrepancy: 151 isolates reported vs 114 WGS runs in the bioproject (rest from prior projects).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-21 ⛓ 230cd22b06a6
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests why the rye-introgressed wheat powdery mildew resistance gene Pm17 broke down rapidly in the field, by seeking to identify its corresponding fungal avirulence gene AvrPm17 and determine whether pre-existing virulence variation in the mildew population explains the fast resistance breakdown.

Core claims
  • AvrPm17 is encoded by a paralogous, tandemly duplicated effector gene pair located in a pericentromeric, mildew sublineage-specific effector cluster (family E003) showing signs of recurring gene conversion. finding
  • Ancient virulent AvrPm17 haplovariants were already present as standing genetic variation in wheat powdery mildew populations before the Pm17 introgression, explaining the rapid resistance breakdown. finding
  • Transient coexpression of AvrPm17 candidates with Pm17 in N. benthamiana induces a hypersensitive response, functionally validating BgTH12-04537/BgTH12-04538 (and Bgt-51729/Bgt-51731) as AvrPm17. method
  • A second, previously undetected resistance gene was co-introgressed with Pm17 on the rye 1AL.1RS translocation, masked by suppressed recombination in the introgressed segment. finding
  • AVRPM17 is predicted to adopt a ribonuclease-like fold structurally related to other characterized wheat mildew Avr effectors (e.g., AvrPm3 family), despite low primary sequence similarity. mechanism
  • Amino acid polymorphisms between virulent and avirulent AvrPm17 variants affect protein abundance in a heterologous system, correlating with differential strength of Pm17 recognition. mechanism
Experimental setups
Assay System Perturbation Readout Platform
QTL mapping (biparental cross, genetic linkage mapping) B.g. tritici Bgt_96224 x B.g. triticale THUN-12 F1 progeny (n=55), tested on transgenic Pm17 wheat lines none (natural avirulence/virulence segregation) avirulence/virulence phenotype (LOD score) mapped to chromosome 1 locus
Comparative genome assembly and sequence alignment B.g. tritici Bgt_96224 and B.g. triticale THUN-12 chromosome-scale genome assemblies none gene content, structural variation (deletion), SNPs within QTL interval
RNA-sequencing Parental mildew isolates Bgt_96224 and THUN-12, infection time course infection stage (2 dpi, haustorial establishment) gene expression level and differential expression (logFC) of AvrPm17 candidates
Agrobacterium-mediated transient coexpression / hypersensitive response (HR) assay Nicotiana benthamiana coexpression of Pm17-HA with AvrPm17 candidate effectors (with/without epitope tags) HR cell death, quantified via Fusion FX imager Fusion FX imager system
Western blot N. benthamiana leaves expressing FLAG-tagged AVRPM17 variants and HA-tagged PM17 AVRPM17 variant (THUN12 vs 96224), FLAG tagging protein abundance/detection
Disease phenotyping / infection assay Transgenic wheat lines Pm17#34 and Pm17#181 infection with isolate Bgt_96224, THUN-12, or Bgt_96224 x THUN-12 progeny mildew leaf coverage (disease severity)
In silico protein structure modeling AVRPM17 protein sequence (computational) none predicted protein fold (ribonuclease-like fold) with confidence P-value IntFOLD5.0
Haplovariant mining / population sequence survey Wheat mildew and related mildew sublineages (haplotype diversity survey) none AvrPm17 allelic/haplotype diversity and virulence status across isolates
Key results
  • Single significant QTL for Pm17 avirulence mapped to pericentromeric chromosome 1 LOD 9.2 (Pm17#34), LOD 7.0 (Pm17#181)
  • 50-kb deletion in avirulent THUN-12 genome relative to virulent Bgt_96224 within the QTL interval, leaving only the paralogous effector pair as candidates 61.8 kb (THUN-12) vs 114.3 kb (Bgt_96224) interval; 50-kb deletion
  • AvrPm17 candidate genes highly expressed early in infection and not differentially expressed between virulent and avirulent isolates logFC < 1.5
  • Coexpression of AvrPm17_THUN12 or AvrPm17_96224 with Pm17-HA induces hypersensitive response in N. benthamiana, confirming AvrPm17 identity n=18 leaves, 3 independent experiments
  • AvrPm17_96224 variant induces significantly weaker HR than AvrPm17_THUN12, corresponding to reduced disease resistance on Pm17 wheat P = 4.657e-10 (paired Wilcoxon rank-sum test)
  • AVRPM17_THUN12 shows higher protein abundance than weakly-recognized AVRPM17_96224 in Western blot
  • Ancient virulent AvrPm17 haplovariants identified as standing genetic variation in wheat mildew predating Pm17 introgression
  • AVRPM17 predicted to adopt ribonuclease-fold structure P = 1.145E-4
Key statistics
  • other LOD 9.2 (QTL mapping significance for Pm17#34 avirulence locus)
  • other LOD 7.0 (QTL mapping significance for Pm17#181 avirulence locus)
  • pvalue P = 4.657e-10 (Paired Wilcoxon rank-sum test comparing HR induced by AvrPm17_THUN12 vs AvrPm17_96224)
  • pvalue P = 1.145E-4 (IntFOLD5.0 prediction confidence for AVRPM17 ribonuclease-fold)
  • count n = 18 leaves (HR induced by Pm17 + AvrPm17 coexpression across three independent experiments)
  • other 61.8 kb vs 114.3 kb (Genetic confidence interval physical size in THUN-12 vs Bgt_96224 assemblies)
  • fold_change logFC < 1.5 (Differential expression of AvrPm17 candidates between Bgt_96224 and THUN-12)
  • count 50 kb deletion (Deletion in THUN-12 genome relative to Bgt_96224 within AvrPm17 interval)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used forward-genetics QTL mapping (single-interval LOD score analysis with a permutation-derived significance threshold) in a biparental F1 mildew progeny population to localize the AvrPm17 avirulence locus, then functionally validated candidate effectors using Agrobacterium-mediated transient co-expression assays in Nicotiana benthamiana scored for hypersensitive response (HR). A quantitative HR comparison between two effector haplotypes was assessed with a paired Wilcoxon rank-sum test and an exact p-value. Results were presented mainly as genetic/physical maps, QTL plots, and individual leaf-level data points from independent experiments rather than as pooled summary tables with dispersion measures.

Replicationbiological Sample sizen = 18 leaves in three independent experiments for HR presence/absence (Fig. 2 B–C); at least n = 8 leaves per experiment across three independent experiments for quantitative HR comparison (Fig. 2D); 55 progeny used for QTL mapping GroupsAvrPm17_THUN12 vs AvrPm17_96224 effector variants co-expressed with Pm17 in N. benthamiana; mildew progeny genotypes scored on Pm17 transgenic wheat lines for QTL mapping Pairingpaired Randomization/blindingnot stated Dispersionunclear Exact p-valuesyes Multiplicity correction1,000 permutations used to set a genome-wide LOD significance threshold for the QTL scan; no explicit multiplicity correction stated for the single paired Wilcoxon comparison
Statistical tests used
Test Applied to n Assumptions
Single-interval QTL mapping (LOD score analysis) Mapping the AvrPm17 avirulence locus using 55 F1 progeny of the Bgt_96224 × THUN-12 cross, scored on two independent Pm17 transgenic wheat lines (Fig. 1 A–C) 55 progeny not stated
Permutation test (1,000 permutations) for LOD significance threshold Establishing the genome-wide significance line for the QTL scan (Fig. 1A) 1,000 permutations not stated
Paired Wilcoxon rank-sum test Comparing HR intensity (Fusion FX imager quantification) induced by AvrPm17_THUN12 versus AvrPm17_96224 when co-expressed with Pm17 in N. benthamiana (Fig. 2D) at least n = 8 leaves per experiment across three independent experiments not stated
Approaches that could also have been used
  • The QTL significance threshold was established empirically via 1,000 permutations of the LOD score analysis.
    Could also: A parametric genome-wide threshold (e.g., Bonferroni correction based on the number of independent markers/linkage groups, or an extreme-value-theory-based threshold) could also be used to set the significance cutoff. — A parametric approach can be faster to compute and is sometimes preferred when permutation is computationally costly, though permutation-based thresholds are already considered a standard, robust method for QTL mapping.
  • The difference in HR intensity between the two AvrPm17 haplotypes was assessed with a paired Wilcoxon rank-sum (nonparametric) test.
    Could also: If the quantitative HR measurements approximate a normal distribution, a paired t-test could also be used. — A paired t-test would additionally yield a mean difference with a confidence interval, which can convey effect magnitude alongside the significance level.
  • Data from three independent experiments were pooled for the paired Wilcoxon test, with individual leaves shown color-coded by experiment.
    Could also: A mixed-effects (hierarchical) model treating experiment as a random effect and leaf as the observational unit could also be used. — This would let experiment-to-experiment variability be explicitly modeled rather than pooled, which can be informative when replicate batches differ in baseline HR response.
  • Two separate QTL scans (one per Pm17 transgenic line, Pm17#34 and Pm17#181) were each evaluated against their own permutation-derived threshold.
    Could also: A joint or multi-trait QTL model, or a Bonferroni-style adjustment across the two scored phenotypes, could also be applied. — Considering the two transgenic-line phenotypes jointly (or correcting across them) can guard against inflated false-positive rates when the same genetic interval is tested against multiple related phenotypes.
  • HR quantification results (Fig. 2D) are shown as individual data points without an explicitly stated summary dispersion measure (e.g., SD, SEM, or CI) in the visible text.
    Could also: Reporting a summary statistic such as the median with an interquartile range, or a bootstrap-based confidence interval for the median difference, could also accompany the individual-point display. — This would give readers a compact numerical sense of central tendency and spread in addition to the raw distribution of points, which can aid comparison across figures or studies.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35857869 (Müller et al. 2022, PNAS, AvrPm17 / Pm17)

  • Paper: Ancient variation of the AvrPm17 gene in powdery mildew limits the effectiveness of the introgressed rye Pm17 resistance gene. PMID 35857869 · PMC9335242 · DOI 10.1073/pnas.2108808119
  • Code: https://github.com/MarionCMueller/AvrPm17 (commit 1692dd2992bc0e2ae693b40a52c03627664f866e, 2022-07-07)
  • Data: SRA PRJNA625429 (151 resequenced isolates), PRJNA783175 (PacBio ISR7), ENA PRJEB41382 (ISR7 assembly), GenBank OM258717–OM258731 (15 AvrPm17 haplovariants).

Repo structure (what the authors actually shipped)

The repository is organized per figure: each folder ships the small derived input data (plain-text tables, alignments, coverage tables, genetic maps) plus an R script that produces the figure / statistic. The repo does NOT ship the upstream raw-read mapping / variant-calling / assembly pipeline. So the authors' own reproducible computational artifacts are the figure-level R analyses.

IN SCOPE (pipeline-derived, reproducible from shipped data + R)

id result script shipped data paper location
C1 Wilcoxon paired test varA vs varB HR intensity (avr vs vir) Figure2D/Boxplots_Figure2D_varAvarB.R AllvarBWilkocon.txt Fig 2D
C2 Wilcoxon varA vs varC Figure4C/Boxplots_varC.R AllvarCWilkocon.txt Fig 4C
C3 Wilcoxon varA vs varD Figure4C/Boxplots_varD.R AllvarDWilkocon.txt Fig 4C
C4 Wilcoxon varA vs varE Figure4C/Boxplots_varE.R AllDatavarEWilcoson.txt Fig 4C
C5 Wilcoxon varA vs varF Figure4C/Boxplots_varF.R AllvarF_wilkocon... Fig 4C
C6 Wilcoxon varA vs varG Figure4C/Boxplots_varG.R AllvarGWilkocon.txt Fig 4C
C7 Wilcoxon varA vs varBgd Figure4C/Boxplots_Bgd.R AllBgdWilkocon.txt Fig 4C
C8 Wilcoxon varB vs varC FigureS22/Boxplots_varBvarC.R AllvarBvarCWilkocon.txt Fig S22
C9 Copy-number distribution of AvrPm17 across isolates (majority 2 copies; few >3) FigureS20/Rscript_histogram.R av_cov.zip Fig S20, Dataset S3
C10 Protoplast HR assay Wilcoxon FigureS22/Rscript_protoplast_Figure.R Protoplast_data_all.txt Fig S22
C11 QTL LOD on chr1 / Amigo cross (scanone) Figure5/Script_*.R, Figure1 GeneticMap*.zip, scanonePm17cg.csv Fig 1A-C, Fig 5
C12 Nucleotide-identity sliding window of duplicated copies FigureS16/Visualization_nucleotide_identity.R muscle-*.clw/.aln Fig S16, Fig 3D
C13 E003 effector-family phylogeny input FigureS11/Flat_E003_final Flat_E003_final Fig 3A, Fig S11

OUT OF SCOPE (not shipped as a runnable pipeline; not attempted)

  • Raw-read mapping of 151 SRA isolates, SNP calling, haplotype calling from reads.
  • PacBio assembly of ISR7; chromosome-scale assemblies.
  • E003 phylogenetic tree inference (only the flat input file is shipped, not the tree-build command / model); divergence dating (~250 ky).
  • IntFOLD5.0 structural model (web server, not in repo).
  • All wet-lab phenotyping (HR assays, transgenics) — the numeric HR-intensity tables are shipped, so the statistical test on them is in scope, but generation of the phenotype data is wet-lab and out of scope.

Primary reproduction target

Recompute the Wilcoxon paired-test p-values (C1–C8, C10), the copy-number distribution (C9), and the QTL LOD (C11) from the shipped data and check they match the paper's reported statistics / qualitative claims. This is the "apply the shipped code to the shipped data" path (P16: equally valid).

Figures / tables: Fig 2DFig 4CFig S22Fig S20Fig 5DFig 1A
C1
Reported
paired Wilcoxon P=4.657e-10 (varA vs varB HR intensity, Fig 2D)
Reproduced
P=4.6566e-10 (V=528, n=32)
exact
C2
Reported
paired Wilcoxon varA vs varC, *** P<0.001 (Fig 4C)
Reproduced
P=2.61e-08 (V=459, n=30)
within tolerance
C3
Reported
paired Wilcoxon varA vs varD, *** (Fig 4C)
Reproduced
P=1.86e-09 (V=465, n=30)
within tolerance
C4
Reported
paired Wilcoxon varA vs varE, significant (Fig 4C)
Reproduced
P=1.86e-08 (V=432, n=29)
within tolerance
C5
Reported
paired Wilcoxon varA vs varF, significant (Fig 4C)
Reproduced
P=4.71e-07 (V=447, n=30)
within tolerance
C6
Reported
paired Wilcoxon varA vs varG, significant (Fig 4C)
Reproduced
P=1.86e-09 (V=465, n=30)
within tolerance
C7
Reported
paired Wilcoxon varA vs varBgd, significant (Fig 4C)
Reproduced
P=3.73e-09 (V=435, n=29)
within tolerance
C8
Reported
paired Wilcoxon varB vs varC, significant (Fig S22)
Reproduced
P=3.62e-05 (V=7, n=20)
within tolerance
C9
Reported
AvrPm17 copy number: majority 2 copies, ~3 isolates >3 (Fig S20)
Reproduced
majority=2 (147/166), n>3=3, median=2 (control-gene reconstruction)
within tolerance
C10
Reported
protoplast HR boxplot (Fig S22)
Reproduced
no statistical test in shipped script (boxplot only)
partial
C11
Reported
Amigo QTL scanone: chr9 LOD 15.97 (pos 340) + chr1 LOD 4.24 (shipped scan)
Reproduced
chr9 LOD 24.86 (pos 339) + chr1 LOD 4.46; peak chr+pos match, absolute chr9 LOD differs (method=em; authors' scanone params not shipped)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4

This is a clean 1:1 reproduction on the authors' own shipped per-figure data and R code: the headline statistic C1 reproduces digit-for-digit (P=4.6566e-10 vs 4.657e-10) and all C2–C8 paired Wilcoxon tests remain highly significant in the same direction, with copy number matching (majority 2; 3 isolates >3). The only deviations are on our/method side: C9 needed a control-gene reconstruction (normalized intermediate not shipped) and C11's absolute chr9 LOD differs (24.86 vs 15.97) because the authors shipped only the precomputed scan, not the scanone parameters — yet the QTL peak chromosome+position reproduce exactly. No reported value was non-derivable from the shipped data, so there is no fabrication concern; partial raw SRA coverage (114/151) is a data-availability limitation, not an authors' analytical defect.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

203.4 k
tokens (I/O) · 13.6 M incl. cache
46 min
runtime · 0.02 CPU-h
1.3 GB
peak RAM
1
HPC jobs
hummel
machine