Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Utility of Triti-Map for bulk-segregated mapping of causal genes and regulatory elements in Triticeae.

Plant Commun · 2022
L1 95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • Any deviation was negligible
What did not (or only partly)
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (core). Triti-Map = authors' own Snakemake BSA pipeline (P16-valid), run end-to-end on wheat ChIP-seq PRJNA725543 (6 runs = 2 bulks x 3 histone marks, all present, grade A) reused as datatype=dna, against IWGSC RefSeq v2.1, pinned commit baa0157 with config defaults. INTERVAL-MAPPING MODULE fully executed on «our HPC» («job», COMPLETED exit 0). C1 mapping interval reproduced WITHIN-TOL: reproduced chr7A:724,111,912-730,215,718 vs reported chr7A:724,111,912-730,119,678 -- START IDENTICAL to the base pair, END +96kb (1.6%), same #1 deltaSNP-index window. C2 EXACT: both reported candidate genes TraesCS7A02G551900 (5 nonsyn) & TraesCS7A02G555200 (4 nonsyn) recovered inside the interval carrying significant nonsynonymous SNPs (confirmed by deterministic codon translation, independent of paper ANNOVAR). C2-loc EXACT. NOT attempted: the assembly module (C3/C4) -- a separate compute-heavy stretch demonstration (ABySS k=90 + EBI web-API domain annotation); honestly recorded as not-attempted, not a blocker. No fabrication indicators: reported C1 start is bit-identical to pipeline output. The paper's pipeline-derived headline result reproduces faithfully.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 60
    assessed: 2026-06-19 ⛓ f05c8b1e41a2
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-26
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Bulk-segregated sequencing (DNA-, RNA-, or ChIP-seq) combined with an optimized computational pipeline that integrates multi-omics data across Triticeae species can efficiently locate causal genes and regulatory elements in Triticeae, including genes absent from the reference genome, despite the large genome size and frequent introgressions that hinder mapping in this group.

Core claims
  • Triti-Map is a computational package suite plus web interface specifically optimized for bulk-segregated gene mapping in Triticeae, accepting DNA-seq, RNA-seq/ChIP-seq, and traditional QTL data as input resource
  • The assembly module of Triti-Map can detect candidate genes/sequences that are absent from the reference genome via de novo assembly and comparison to colinear Triticeae regions and public databases method
  • Using bulk-segregated ChIP-seq with Triti-Map, the authors identified Pm60 (and its extended promoter region) as the candidate gene underlying powdery mildew resistance, a gene not present in the wheat reference genome finding
  • Pm60 originated in the common ancestor of diploid wheat but was lost from the A subgenome progenitor before tetraploidization, while B and D subgenome homologs diverged and no longer confer resistance finding
  • Two reference-genome-annotated genes (TraesCS7A02G551900 and TraesCS7A02G555200) carrying nonsynonymous mutations in the candidate interval did not fully segregate with the resistance phenotype finding
  • Triti-Map retrieves colinear regions across six Triticeae genomes (Hordeum vulgare, Aegilops tauschii, Triticum urartu, T. dicoccoides, T. turgidum, T. aestivum) to help identify genes missing from the reference assembly method
  • Software/parameter optimizations (GIGGLE for interval comparison, chromosome-split parallel GATK HaplotypeCaller, BWA-mem2 alignment) substantially reduce analysis time for Triticeae's large genomes method
  • The TraesCS7A02G551900 promoter region shows DNase I hypersensitivity and H3K36me3/H3K4me3/H3K9ac enrichment, indicating high accessibility to transcription factors such as ABF1 finding
Experimental setups
Assay System Perturbation Readout Platform
bulk-segregated ChIP-seq (H3K4me3, H3K27me3, H3K36me3) wheat F2 population (Xuezao x 3D249 cross), resistant vs susceptible bulks none (natural segregating disease-resistance trait) histone mark enrichment used for BSA-based interval mapping of powdery mildew resistance locus
interval/SNP mapping (variant calling and ΔSNP-index analysis) wheat bulked segregant pools none candidate genomic interval and SNP annotations associated with resistance trait GATK HaplotypeCaller, BWA-mem2, GIGGLE
de novo assembly of bulk-specific sequencing reads wheat resistant-pool ChIP-seq reads none assembled scaffolds absent from reference genome; protein domain content (NB-ARC, LRR) EMBL-EBI hmmscan and BLAST APIs
collinearity and phylogenetic analysis Hordeum vulgare, Aegilops tauschii, Triticum urartu, T. dicoccoides, T. turgidum, T. aestivum genomes none presence/absence and evolutionary relationships of Pm60 homologs across subgenomes and species
Key results
  • A 6-Mb candidate region on chromosome 7A (724,111,912–730,119,678) was identified as highly associated with powdery mildew resistance 6-Mb interval
  • Nonsynonymous mutations found in TraesCS7A02G551900 and TraesCS7A02G555200, but experimental validation showed these genes did not fully segregate with the resistance trait
  • 1,704 of 10,429 resistant-pool-specific scaffolds partially mapped to Triticeae regions colinear with the candidate interval 1,704/10,429
  • Two assembled sequences (one NB-ARC domain, one LRR domain) were >99.7% identical to Pm60, together covering 69% of its total length >99.7% identity; 69% of gene length
  • Pm60 sequence was extended by 235 bp at the 5' end and 203 bp at the 3' end using assembly data, adding 10% of the gene's total length 235 bp + 203 bp (~10% of length)
  • Colinear regions on chromosomes 7B and 7D each contain an R gene highly homologous to Pm60, but no highly homologous gene was found in tetraploid or hexaploid A subgenomes
  • hmmscan functional annotation detected nine bulk-specific sequences with R-gene-associated protein domains (four NB-ARC, five LRR) 9 sequences (4 NB-ARC, 5 LRR)
Key statistics
  • count 6-Mb region, chr7A:724,111,912–730,119,678 (candidate interval mapped for powdery mildew resistance)
  • count 10,429 resistant-pool-specific scaffolds; 1,704 mapped to colinear regions (de novo assembly module output)
  • other >99.7% sequence identity (assembled sequences vs. Pm60 via EMBL-EBI BLAST API)
  • other 69% of total Pm60 length covered (combined length of two matching assembled sequences)
  • count 235-bp 5' extension and 203-bp 3' extension (10% of total gene length) (Pm60 sequence extension using assembly data)
  • count nine sequences with R-gene domains (4 NB-ARC, 5 LRR) (hmmscan functional annotation of assembled scaffolds)
  • other 16 Gb (genome size of common wheat (background statistic))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods/resource paper describing the Triti-Map computational pipeline for bulk-segregant gene mapping in Triticeae, illustrated with a single case study using bulk-segregated ChIP-seq from two F2-derived phenotype pools (powdery-mildew resistant vs. susceptible). Candidate genomic intervals were identified using the ΔSNP-index method typical of bulked segregant analysis (BSA), and candidate genes/regulatory elements were then narrowed down via SNP annotation, cross-species collinearity, and sequence-similarity/homology comparisons (e.g., BLAST/hmmscan identity to Pm60) rather than classical inferential hypothesis testing. No formal statistical tests (e.g., t-tests, ANOVA), p-values, confidence intervals, or replicate-based dispersion measures are reported in the text.

Replicationunclear Sample sizeTwo bulked-segregant pools (resistant and susceptible) were collected from F2 progeny of a Xuezao × 3D249 cross; no numerical pool sizes, replicate counts, or power calculation are stated in the provided text. GroupsPowdery-mildew-resistant bulk vs. susceptible bulk (ChIP-seq for H3K4me3, H3K27me3, H3K36me3) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
ΔSNP-index method (BSA-based interval detection) Interval mapping module output; identification of the 6-Mb candidate region on chromosome 7A (Figure 4A) not stated
Approaches that could also have been used
  • Candidate intervals were identified using the ΔSNP-index method for BSA.
    Could also: Other established BSA statistics, such as the G-statistic (e.g., implemented in QTLseqr) or permutation-based simulated confidence intervals for SNP-index, could also be applied. — These alternatives attach an explicit statistical confidence threshold or significance boundary to the interval calls, which can complement a ΔSNP-index scan that is primarily descriptive.
  • The case study used one bulk pool per phenotype class without stated biological replicate pools.
    Could also: Constructing multiple independent replicate bulks per phenotype and analyzing them with a variance-aware BSA method could also be used. — Replicate pools allow estimation of pool-to-pool sampling variance, which can support formal significance thresholds rather than a single-pool interval estimate.
  • No genome-wide multiple-comparison correction is described for the SNP/interval scan.
    Could also: A genome-wide significance threshold derived from permutation testing or a Bonferroni-type adjustment across scanned windows could also be reported. — This would provide an explicit false-positive control across the many windows/SNPs examined in a genome-wide scan.
  • Candidate gene identity (e.g., the match to Pm60) was reported using percent sequence identity from BLAST/hmmscan.
    Could also: Reporting standard alignment statistics such as E-values or bit scores alongside percent identity could also be done. — E-values/bit scores give a probabilistic measure of match significance that complements percent identity, particularly for shorter or partial sequence matches.
  • The pipeline also accepts traditional QTL data as an input type but the case study itself does not use QTL-style significance testing.
    Could also: For QTL-type inputs, LOD score profiling with permutation-derived significance thresholds could also be used to declare significant loci. — LOD/permutation thresholds are a standard way to assign statistical confidence to QTL peaks, complementing the interval-mapping approach described for BSA data.
Software: Snakemake · GIGGLE · GATK HaplotypeCaller · BWA-mem2 · EMBL-EBI RESTful APIs (hmmscan, BLAST) · Conda

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35605195 (Triti-Map)

Paper: Zhao et al. 2022, Plant Commun 3:100304. "Utility of Triti-Map for bulk-segregated mapping of causal genes and regulatory elements in Triticeae." Code: https://github.com/fei0810/Triti-Map (MIT, pinned commit baa0157be702e1b74f518786f5b19974f3679441, 2022-03-05). Also on bioconda (tritimap) and Docker (fei0810/tritimap:v0.9.7). Data: SRA PRJNA725543 — 6 ChIP-seq runs (SRR14336994–SRR14336999).

Triti-Map is the authors' own Snakemake pipeline (P16: even so, it is a tool applied to the paper's own data → fully reproducible). Two modules:

  1. Interval Mapping — BSA via ΔSNP-index on variants called from the bulks.
  2. Assembly — de-novo assembly of bulk-specific scaffolds to find genes/ alleles absent from the reference.

The case study (what the deposited data supports)

Wheat (Triticum aestivum) F3 lines from Xuezao (powdery-mildew susceptible) × 3D249 (WEW introgression line, resistant). Two bulks (Res / Sus), each sequenced as ChIP-seq for 3 histone marks (H3K27me3, H3K4me3, H3K36me3). The pipeline treats these ChIP-seq reads as datatype: dna and calls SNPs from them for the BSA — that is the paper's methodological novelty (epigenomic reads reused as genomic markers for mapping).

IN SCOPE (pipeline-derived → attempt to reproduce)

id result pipeline paper location
C1 Mapping interval chr7A:724,111,912–730,119,678 (~6 Mb) for powdery-mildew resistance Interval Mapping (fastp→bwa-mem2→GATK HaplotypeCaller→QTLseqr ΔSNP-index) Results / Fig. 4
C2 Two candidate genes with nonsynonymous SNPs: TraesCS7A02G551900, TraesCS7A02G555200 Interval Mapping + ANNOVAR SNP annotation Results
C2-loc Both candidate genes lie within the C1 interval reference annotation (independently checkable) derived
C3 Assembly module: 10,429 resistant-pool-specific scaffolds, 1,704 partially mapped, nine R-gene-domain sequences, Pm60 homolog at >99.7% identity Assembly (ABySS k=90 → minimap2/bwa → HMMER/BLAST domain annotation) Results
C4 Pm60 assembly extension: 235-bp 5′ + 203-bp 3′ (~10% of Pm60 length) Assembly Results

Pipeline reference genome/annotation: IWGSC RefSeq v1.0 (Chinese Spring), Ensembl Plants annotation (iwgsc_high_conf, gene-ID scheme TraesCS7A02G…), chromosomes renamed with a chr prefix (region.csv uses chr7A).

OUT OF SCOPE (wet-lab / manual / external — not attempted)

  • Disease phenotyping / resistance scoring of the F3 lines (wet-lab).
  • Experimental validation that TraesCS7A02G551900/555200 do not fully segregate with resistance (wet-lab; this is why the paper pivots to assembly).
  • ChIP-seq antibody experiments / library construction (wet-lab).
  • Triti-Map web annotation platform (external service).

Reproduction strategy

Run the actual tritimap tool on the deposited data, IWGSC RefSeq v1.0 ref, default config (datatype: dna, merge_lib: merge, QTLseqr pop_struc: RIL, bulksize: 30, winsize: 1e6, filter_percentage: 0.75, fisher_p: 1e-4, min_length: 1e6, gatk min_SNP_DP: 10). Heavy compute on «our HPC»/«infra». Quick-minimum target = C1 + C2 (interval-mapping module). Stretch = C3/C4 (assembly module; ABySS needs ~300–400 G RAM per config).

Honest feasibility note

This is a large reproduction: ~135 GB ChIP-seq, the ~14 Gb hexaploid wheat genome, genome-wide GATK (paper: 1–2 days on 30 threads), and an assembly step needing very high RAM. C2-loc is already confirmed from public annotation. C1/C2 require the full interval-mapping run; C3/C4 are compute-heavy stretch goals.

Figures / tables: Fig.4
C1
Reported
chr7A:724,111,912-730,119,678 (~6 Mb) interval
Reproduced
chr7A:724,111,912-730,215,718 (6,103,806 bp); top-ranked QTL window, peakDeltaSNP=0.954
within tolerance
C2
Reported
candidate genes TraesCS7A02G551900 & TraesCS7A02G555200 carry nonsynonymous SNPs
Reproduced
both recovered in interval; 551900=5 nonsyn (R201G,F215Y,V225G,E257V,N495D), 555200=4 nonsyn (H86R,H70R,I64T,L54P); all filtered SNPs deltaSNP 0.9-1.0
exact
C2-loc
Reported
TraesCS7A02G551900 & TraesCS7A02G555200 lie within chr7A:724,111,912-730,119,678
Reproduced
551900=chr7A:725224414-725228601; 555200=chr7A:727220813-727221872 (both inside reproduced interval)
exact
C3
Reported
10429 scaffolds / 1704 mapped / 9 R-domain seqs / Pm60 >99.7% id
Reproduced
NOT_ATTEMPTED (assembly module, stretch)
m.public.grade.not-attempted
C4
Reported
Pm60 extension 235bp 5' + 203bp 3'
Reproduced
NOT_ATTEMPTED (assembly module, stretch)
m.public.grade.not-attempted

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

114.1 k
tokens (I/O) · 5.8 M incl. cache
20 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.