Utility of Triti-Map for bulk-segregated mapping of causal genes and regulatory elements in Triticeae.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓Any deviation was negligible
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (core). Triti-Map = authors' own Snakemake BSA pipeline (P16-valid), run end-to-end on wheat ChIP-seq PRJNA725543 (6 runs = 2 bulks x 3 histone marks, all present, grade A) reused as datatype=dna, against IWGSC RefSeq v2.1, pinned commit baa0157 with config defaults. INTERVAL-MAPPING MODULE fully executed on «our HPC» («job», COMPLETED exit 0). C1 mapping interval reproduced WITHIN-TOL: reproduced chr7A:724,111,912-730,215,718 vs reported chr7A:724,111,912-730,119,678 -- START IDENTICAL to the base pair, END +96kb (1.6%), same #1 deltaSNP-index window. C2 EXACT: both reported candidate genes TraesCS7A02G551900 (5 nonsyn) & TraesCS7A02G555200 (4 nonsyn) recovered inside the interval carrying significant nonsynonymous SNPs (confirmed by deterministic codon translation, independent of paper ANNOVAR). C2-loc EXACT. NOT attempted: the assembly module (C3/C4) -- a separate compute-heavy stretch demonstration (ABySS k=90 + EBI web-API domain annotation); honestly recorded as not-attempted, not a blocker. No fabrication indicators: reported C1 start is bit-identical to pipeline output. The paper's pipeline-derived headline result reproduces faithfully.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 60assessed: 2026-06-19 ⛓ f05c8b1e41a2
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-26
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetBulk-segregated sequencing (DNA-, RNA-, or ChIP-seq) combined with an optimized computational pipeline that integrates multi-omics data across Triticeae species can efficiently locate causal genes and regulatory elements in Triticeae, including genes absent from the reference genome, despite the large genome size and frequent introgressions that hinder mapping in this group.
- ★ Triti-Map is a computational package suite plus web interface specifically optimized for bulk-segregated gene mapping in Triticeae, accepting DNA-seq, RNA-seq/ChIP-seq, and traditional QTL data as input resource
- ★ The assembly module of Triti-Map can detect candidate genes/sequences that are absent from the reference genome via de novo assembly and comparison to colinear Triticeae regions and public databases method
- ★ Using bulk-segregated ChIP-seq with Triti-Map, the authors identified Pm60 (and its extended promoter region) as the candidate gene underlying powdery mildew resistance, a gene not present in the wheat reference genome finding
- ★ Pm60 originated in the common ancestor of diploid wheat but was lost from the A subgenome progenitor before tetraploidization, while B and D subgenome homologs diverged and no longer confer resistance finding
- ★ Two reference-genome-annotated genes (TraesCS7A02G551900 and TraesCS7A02G555200) carrying nonsynonymous mutations in the candidate interval did not fully segregate with the resistance phenotype finding
- ★ Triti-Map retrieves colinear regions across six Triticeae genomes (Hordeum vulgare, Aegilops tauschii, Triticum urartu, T. dicoccoides, T. turgidum, T. aestivum) to help identify genes missing from the reference assembly method
- Software/parameter optimizations (GIGGLE for interval comparison, chromosome-split parallel GATK HaplotypeCaller, BWA-mem2 alignment) substantially reduce analysis time for Triticeae's large genomes method
- The TraesCS7A02G551900 promoter region shows DNase I hypersensitivity and H3K36me3/H3K4me3/H3K9ac enrichment, indicating high accessibility to transcription factors such as ABF1 finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk-segregated ChIP-seq (H3K4me3, H3K27me3, H3K36me3) | wheat F2 population (Xuezao x 3D249 cross), resistant vs susceptible bulks | none (natural segregating disease-resistance trait) | histone mark enrichment used for BSA-based interval mapping of powdery mildew resistance locus | — |
| interval/SNP mapping (variant calling and ΔSNP-index analysis) | wheat bulked segregant pools | none | candidate genomic interval and SNP annotations associated with resistance trait | GATK HaplotypeCaller, BWA-mem2, GIGGLE |
| de novo assembly of bulk-specific sequencing reads | wheat resistant-pool ChIP-seq reads | none | assembled scaffolds absent from reference genome; protein domain content (NB-ARC, LRR) | EMBL-EBI hmmscan and BLAST APIs |
| collinearity and phylogenetic analysis | Hordeum vulgare, Aegilops tauschii, Triticum urartu, T. dicoccoides, T. turgidum, T. aestivum genomes | none | presence/absence and evolutionary relationships of Pm60 homologs across subgenomes and species | — |
- – A 6-Mb candidate region on chromosome 7A (724,111,912–730,119,678) was identified as highly associated with powdery mildew resistance 6-Mb interval
- – Nonsynonymous mutations found in TraesCS7A02G551900 and TraesCS7A02G555200, but experimental validation showed these genes did not fully segregate with the resistance trait
- – 1,704 of 10,429 resistant-pool-specific scaffolds partially mapped to Triticeae regions colinear with the candidate interval 1,704/10,429
- – Two assembled sequences (one NB-ARC domain, one LRR domain) were >99.7% identical to Pm60, together covering 69% of its total length >99.7% identity; 69% of gene length
- – Pm60 sequence was extended by 235 bp at the 5' end and 203 bp at the 3' end using assembly data, adding 10% of the gene's total length 235 bp + 203 bp (~10% of length)
- – Colinear regions on chromosomes 7B and 7D each contain an R gene highly homologous to Pm60, but no highly homologous gene was found in tetraploid or hexaploid A subgenomes
- – hmmscan functional annotation detected nine bulk-specific sequences with R-gene-associated protein domains (four NB-ARC, five LRR) 9 sequences (4 NB-ARC, 5 LRR)
- count 6-Mb region, chr7A:724,111,912–730,119,678 (candidate interval mapped for powdery mildew resistance)
- count 10,429 resistant-pool-specific scaffolds; 1,704 mapped to colinear regions (de novo assembly module output)
- other >99.7% sequence identity (assembled sequences vs. Pm60 via EMBL-EBI BLAST API)
- other 69% of total Pm60 length covered (combined length of two matching assembled sequences)
- count 235-bp 5' extension and 203-bp 3' extension (10% of total gene length) (Pm60 sequence extension using assembly data)
- count nine sequences with R-gene domains (4 NB-ARC, 5 LRR) (hmmscan functional annotation of assembled scaffolds)
- other 16 Gb (genome size of common wheat (background statistic))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a methods/resource paper describing the Triti-Map computational pipeline for bulk-segregant gene mapping in Triticeae, illustrated with a single case study using bulk-segregated ChIP-seq from two F2-derived phenotype pools (powdery-mildew resistant vs. susceptible). Candidate genomic intervals were identified using the ΔSNP-index method typical of bulked segregant analysis (BSA), and candidate genes/regulatory elements were then narrowed down via SNP annotation, cross-species collinearity, and sequence-similarity/homology comparisons (e.g., BLAST/hmmscan identity to Pm60) rather than classical inferential hypothesis testing. No formal statistical tests (e.g., t-tests, ANOVA), p-values, confidence intervals, or replicate-based dispersion measures are reported in the text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| ΔSNP-index method (BSA-based interval detection) | Interval mapping module output; identification of the 6-Mb candidate region on chromosome 7A (Figure 4A) | — | not stated |
-
Candidate intervals were identified using the ΔSNP-index method for BSA.↳ Could also: Other established BSA statistics, such as the G-statistic (e.g., implemented in QTLseqr) or permutation-based simulated confidence intervals for SNP-index, could also be applied. — These alternatives attach an explicit statistical confidence threshold or significance boundary to the interval calls, which can complement a ΔSNP-index scan that is primarily descriptive.
-
The case study used one bulk pool per phenotype class without stated biological replicate pools.↳ Could also: Constructing multiple independent replicate bulks per phenotype and analyzing them with a variance-aware BSA method could also be used. — Replicate pools allow estimation of pool-to-pool sampling variance, which can support formal significance thresholds rather than a single-pool interval estimate.
-
No genome-wide multiple-comparison correction is described for the SNP/interval scan.↳ Could also: A genome-wide significance threshold derived from permutation testing or a Bonferroni-type adjustment across scanned windows could also be reported. — This would provide an explicit false-positive control across the many windows/SNPs examined in a genome-wide scan.
-
Candidate gene identity (e.g., the match to Pm60) was reported using percent sequence identity from BLAST/hmmscan.↳ Could also: Reporting standard alignment statistics such as E-values or bit scores alongside percent identity could also be done. — E-values/bit scores give a probabilistic measure of match significance that complements percent identity, particularly for shorter or partial sequence matches.
-
The pipeline also accepts traditional QTL data as an input type but the case study itself does not use QTL-style significance testing.↳ Could also: For QTL-type inputs, LOD score profiling with permutation-derived significance thresholds could also be used to declare significant loci. — LOD/permutation thresholds are a standard way to assign statistical confidence to QTL peaks, complementing the interval-mapping approach described for BSA data.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35605195 (Triti-Map)
Paper: Zhao et al. 2022, Plant Commun 3:100304. "Utility of Triti-Map for
bulk-segregated mapping of causal genes and regulatory elements in Triticeae."
Code: https://github.com/fei0810/Triti-Map (MIT, pinned commit
baa0157be702e1b74f518786f5b19974f3679441, 2022-03-05). Also on bioconda
(tritimap) and Docker (fei0810/tritimap:v0.9.7).
Data: SRA PRJNA725543 — 6 ChIP-seq runs (SRR14336994–SRR14336999).
Triti-Map is the authors' own Snakemake pipeline (P16: even so, it is a tool applied to the paper's own data → fully reproducible). Two modules:
- Interval Mapping — BSA via ΔSNP-index on variants called from the bulks.
- Assembly — de-novo assembly of bulk-specific scaffolds to find genes/ alleles absent from the reference.
The case study (what the deposited data supports)
Wheat (Triticum aestivum) F3 lines from Xuezao (powdery-mildew susceptible)
× 3D249 (WEW introgression line, resistant). Two bulks (Res / Sus), each
sequenced as ChIP-seq for 3 histone marks (H3K27me3, H3K4me3, H3K36me3). The
pipeline treats these ChIP-seq reads as datatype: dna and calls SNPs from them
for the BSA — that is the paper's methodological novelty (epigenomic reads reused
as genomic markers for mapping).
IN SCOPE (pipeline-derived → attempt to reproduce)
| id | result | pipeline | paper location |
|---|---|---|---|
| C1 | Mapping interval chr7A:724,111,912–730,119,678 (~6 Mb) for powdery-mildew resistance | Interval Mapping (fastp→bwa-mem2→GATK HaplotypeCaller→QTLseqr ΔSNP-index) | Results / Fig. 4 |
| C2 | Two candidate genes with nonsynonymous SNPs: TraesCS7A02G551900, TraesCS7A02G555200 | Interval Mapping + ANNOVAR SNP annotation | Results |
| C2-loc | Both candidate genes lie within the C1 interval | reference annotation (independently checkable) | derived |
| C3 | Assembly module: 10,429 resistant-pool-specific scaffolds, 1,704 partially mapped, nine R-gene-domain sequences, Pm60 homolog at >99.7% identity | Assembly (ABySS k=90 → minimap2/bwa → HMMER/BLAST domain annotation) | Results |
| C4 | Pm60 assembly extension: 235-bp 5′ + 203-bp 3′ (~10% of Pm60 length) | Assembly | Results |
Pipeline reference genome/annotation: IWGSC RefSeq v1.0 (Chinese Spring),
Ensembl Plants annotation (iwgsc_high_conf, gene-ID scheme TraesCS7A02G…),
chromosomes renamed with a chr prefix (region.csv uses chr7A).
OUT OF SCOPE (wet-lab / manual / external — not attempted)
- Disease phenotyping / resistance scoring of the F3 lines (wet-lab).
- Experimental validation that TraesCS7A02G551900/555200 do not fully segregate with resistance (wet-lab; this is why the paper pivots to assembly).
- ChIP-seq antibody experiments / library construction (wet-lab).
- Triti-Map web annotation platform (external service).
Reproduction strategy
Run the actual tritimap tool on the deposited data, IWGSC RefSeq v1.0 ref,
default config (datatype: dna, merge_lib: merge, QTLseqr pop_struc: RIL,
bulksize: 30, winsize: 1e6, filter_percentage: 0.75, fisher_p: 1e-4,
min_length: 1e6, gatk min_SNP_DP: 10). Heavy compute on «our HPC»/«infra».
Quick-minimum target = C1 + C2 (interval-mapping module). Stretch = C3/C4
(assembly module; ABySS needs ~300–400 G RAM per config).
Honest feasibility note
This is a large reproduction: ~135 GB ChIP-seq, the ~14 Gb hexaploid wheat genome, genome-wide GATK (paper: 1–2 days on 30 threads), and an assembly step needing very high RAM. C2-loc is already confirmed from public annotation. C1/C2 require the full interval-mapping run; C3/C4 are compute-heavy stretch goals.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.