A newly identified glycosyltransferase AsRCOM provides resistance to purple curl leaf disease in agave.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.
▸Reproduction agent’s raw note
DROP (data_unavailable, borders data_restricted). The declared code (Trinity v2.4.0, github.com/trinityrnaseq/trinityrnaseq) is fully available, but the pipeline INPUT -- raw RNA-seq reads of BioProject PRJNA852945 -- was never released to the public. NCBI BioProject page states 'not public in BioProject: 852945'; NCBI esearch db=sra/bioproject returns Count 0 / PhraseNotFound; the ENA project shell exists (agave sisalana Transcriptome, taxon 442491) but every run/experiment/analysis filereport returns 0 rows and there is no secondary study (SRP) accession and nothing on fastq_ftp. The paper's 'Availability of data and materials' points raw reads only to a PRIVATE reviewer-preview link (dataview.ncbi.nlm.nih.gov/...?reviewer=<token>), a peer-review metadata SPA that is not a programmatic download path (prefetch/fasterq-dump cannot use a reviewer token). GenBank BankIt 'Submission IDs' are tracking numbers, not retrievable accessions (and would at most hold a few assembled RCOM mRNAs, not the raw reads or count matrix). No shipped intermediate (clean FASTQ, Trinity FASTA, count/RPKM matrix) exists to bypass the missing reads; Suppl. Table S1 is DEG-only RPKM, not raw counts. Therefore ALL in-scope pipeline claims (Table 1 QC stats C1-C5; Fig 1C DEG counts C6-C9) are unreproducible -- blocked upstream of compute, so no «our HPC»/SLURM job was spent (a fetch job would deterministically fail). NOT attempted: Trinity assembly, QC (FastQC/Trimmomatic), DEG calling, WGCNA -- all require the unavailable reads. INTEGRITY NOTE: results rest on data the community cannot access (reviewer-only link, BioProject non-public ~3y post-publication); reported numbers are not independently verifiable -- flagged, not asserted wrong.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-15 ⛓ 8eaec905bc7b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether RNA-seq and WGCNA can reveal the molecular mechanism underlying resistance of the Agave sisalana hybrid 11648 resistant mutant (A.H11648R) to purple curl leaf disease and identify the pivotal disease-resistance gene(s) responsible.
- ★ The glycosyltransferase gene AsRCOM is the most critical disease-resistance gene against agave purple curl leaf disease, and its overexpression significantly enhances resistance. finding
- ★ Hypersensitive response (HR) accompanied by systemic acquired resistance is the primary pathway increasing resistance to purple curl leaf disease in agave. mechanism
- ★ Overexpression of AsRCOM enhances resistance with abundant reactive oxygen species (ROS) accumulation. finding
- ★ WGCNA identified eight candidate resistance genes (4'OMT2, ACLY, NCS1, GTE10, SMO2, FLS2, SQE1, RCOM) potentially mediating resistance. finding
- ★ DEGs were mainly enriched in alpha-linolenic acid metabolism, starch and sucrose metabolism, phenylpropanoid biosynthesis, flavonoid biosynthesis, HR and systemic acquired resistance pathways. finding
- Integrative RNA-seq combined with WGCNA is an effective strategy to excavate disease-resistance genes in agave. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq | Agave sisalana hybrid 11648 disease-susceptible (A.H11648) and resistant mutant (A.H11648R) leaf blade tips | Dysmicoccus neobrevipes / phytoplasma infection (purple curl leaf disease) at 0, 60, 90 days | gene expression (FPKM), differentially expressed genes | — |
| WGCNA (weighted gene co-expression network analysis) | A.H11648 and A.H11648R RNA-seq expression data | none | co-expression modules / candidate hub resistance genes | — |
| qPCR validation | A.H11648 and A.H11648R agave | purple curl leaf disease infection | expression of candidate genes (AsRCOM) | — |
| transgenic overexpression | agave (A.H11648) | AsRCOM overexpression | disease resistance phenotype and ROS accumulation | — |
| phenotypic disease assessment | A.H11648 and A.H11648R field plants | D. neobrevipes infection over time | disease symptoms (blade tip yellowing/blackening) | — |
- ▲ Overexpression of AsRCOM significantly enhanced resistance to purple curl leaf disease with abundant ROS accumulation.
- – A.H11648 showed yellow blade tips at 60 d and blackened/dead blade tips at 90 d, while A.H11648R showed no disease symptoms throughout infection.
- ▲ DEG numbers increased with infection time: 547 in CK-1-vs-T-1, 4137 in CK-2-vs-T-2, 22712 in CK-3-vs-T-3. 22712 DEGs at 90 d
- – HR-related genes (LOX, PR4, NPR1) showed significant differential expression between A.H11648 and A.H11648R.
- – Plant-pathogen interaction pathway enriched; Ca2+ signal and ROS burst trigger HR and cell wall reinforcement.
- – 210 DEGs were shared across all three pairwise comparison groups. 210 DEGs
- ▼ Severe purple curl leaf disease can reduce leaf production by more than 30%. >30%
- count 547 DEGs (95 up, 451 down) (CK-1-vs-T-1 at 0 d)
- count 4137 DEGs (1782 up, 2355 down) (CK-2-vs-T-2 at 60 d)
- count 22712 DEGs (14786 up, 7926 down) (CK-3-vs-T-3 at 90 d)
- count 18 RNA-seq libraries (6 samples × 3 replicates) (total sequencing libraries)
- other 43,825,614 to 55,003,390 raw reads; 41,657,756 to 52,397,778 clean reads per library (sequencing read counts)
- other 79.22 to 81.64% (coverage of mapped reads)
- other Q30 above 95.57%; GC 48.85–51.1% (RNA-seq quality metrics)
- count 210 DEGs common to all three pairwise groups (shared DEGs)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used RNA-seq to compare leaf transcriptomes of disease-resistant (A.H11648R) and disease-susceptible (A.H11648) agave across three time points (0d, 60d, 90d post-infection), with three biological replicates per group (18 libraries total). DEGs were identified using an adjusted p-value threshold of <0.05, and KEGG/GO pathway enrichment was assessed with FDR-corrected Q-values (≤0.05). WGCNA was used to identify co-expression modules and candidate resistance genes, which were further validated by qPCR and functional overexpression assays.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression testing (specific algorithm not named; adjusted p-value reported) | Pairwise comparisons CK-1-vs-T-1, CK-2-vs-T-2, CK-3-vs-T-3 for DEG identification | n=3 biological replicates per group | not stated |
| KEGG pathway enrichment (specific test not named; FDR-corrected Q-value reported) | Top 20 enriched pathways per pairwise group and for common DEGs across all three groups | — | not stated |
| GO enrichment analysis (specific test not named; FDR-corrected Q-value reported) | Common DEGs across all three pairwise groups | — | not stated |
| Principal components analysis (PCA) | All 18 samples based on FPKM values, for sample-level clustering and quality assessment | 18 samples | na |
| Weighted gene co-expression network analysis (WGCNA) | Identification of co-expression modules and candidate disease-resistance genes | 18 samples | not stated |
| qPCR validation (specific statistical test not stated in excerpt) | Verification of candidate gene expression levels | — | not stated |
-
The specific differential expression algorithm was not named; only the adjusted p-value threshold (<0.05) was reported↳ Could also: Naming the tool (e.g., DESeq2 with negative binomial Wald test, edgeR GLM, or limma-voom) and its version is a standard practice in RNA-seq reporting — Identifying the tool and version allows readers to reproduce the analysis and understand the underlying statistical model, including how dispersion is estimated — an important factor with only n=3 replicates per group
-
Pairwise between-group comparisons were conducted independently at each time point (CK-1-vs-T-1, CK-2-vs-T-2, CK-3-vs-T-3)↳ Could also: A linear model with genotype, time point, and genotype-by-time interaction terms (e.g., via DESeq2 likelihood-ratio test or edgeR GLM) could be fit to the full dataset simultaneously — A joint model formally tests whether the genotype effect changes across time points, pools information across groups, and avoids the need for post-hoc intersection of independently derived DEG lists
-
Gene expression was quantified as FPKM (Fragments Per Kilobase of exon per Million mapped reads)↳ Could also: TPM (Transcripts Per Million) is an alternative normalization metric widely used in more recent transcriptomics studies — TPM normalizes for transcript length before library size, making values more directly comparable across samples with different sequencing depths, and is now the metric more commonly recommended for cross-sample comparisons
-
GO and KEGG enrichment was assessed using a binary DEG list (genes passing the adjusted p-value threshold) with RichFactor and Q-value↳ Could also: Gene set enrichment analysis (GSEA) or ranked-list methods such as fgsea use the full ranked gene list rather than a binary cutoff — Ranked-list enrichment avoids sensitivity to the chosen DEG threshold and can detect coordinated but subtle shifts across a pathway that fall below the individual-gene significance cutoff
-
No measures of dispersion (SD, SEM, or CI) were reported for expression values or qPCR results↳ Could also: Reporting SD or 95% CI alongside means for qPCR and overexpression assay outcomes is standard practice — With n=3 biological replicates, dispersion metrics allow readers to assess within-group variability relative to observed effect magnitudes, which is particularly informative for interpreting small-n experiments
-
Candidate genes identified by WGCNA and qPCR were validated by overexpression in agave↳ Could also: Loss-of-function approaches (RNAi, CRISPR knockout) or validation across additional independent plant accessions or environments could also be used — Complementing gain-of-function evidence with loss-of-function data, or replication in additional genetic backgrounds, would further characterize the role of AsRCOM in the resistance pathway
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37936069
Paper: Lu Z et al. 2023, A newly identified glycosyltransferase AsRCOM provides resistance to purple curl leaf disease in agave. BMC Genomics. DOI 10.1186/s12864-023-09700-y · PMCID PMC10629022.
Declared code: Trinity (github.com/trinityrnaseq/trinityrnaseq), v2.4.0. A third-party de-novo transcriptome assembler applied to the paper's own data (P16: equally valid to reproduce).
Declared data: SRA BioProject PRJNA852945 (18 RNA-seq libraries, Agave sisalana, 6 conditions × 3 reps).
Reported pipeline-derived results (candidate claims)
| id | result | reported | location | pipeline |
|---|---|---|---|---|
| C1 | Raw reads / library | 43,825,614 – 55,003,390 | Table 1 | FastQC on SRA reads |
| C2 | Clean reads / library | 41,657,756 – 52,397,778 | Table 1 | Trimmomatic v0.36.5 |
| C3 | Q30 | >95.57% all libs | Table 1 | FastQC v0.11.2 |
| C4 | GC content | 48.85–51.1% | Table 1 | FastQC |
| C5 | Mapped reads | 79.22–81.64% | Table 1 | align to Trinity assembly |
| C6 | DEGs CK-1 vs T-1 | 547 (95 up / 451 down) | Fig 1C | Trinity→quant→DE |
| C7 | DEGs CK-2 vs T-2 | 4,137 (1,782 / 2,355) | Fig 1C | Trinity→quant→DE |
| C8 | DEGs CK-3 vs T-3 | 22,712 (14,786 / 7,926) | Fig 1C | Trinity→quant→DE |
| C9 | Common DEGs | 210 | Fig 1C | overlap |
In scope vs out of scope
- All of C1–C9 are pipeline-derived and would normally be in scope. Every one of them requires the raw RNA-seq reads (PRJNA852945) as the pipeline input. There is no shipped intermediate (no count matrix, no assembly FASTA, no clean-read FASTQs) that would let any claim be reproduced without the raw reads.
- The reported Trinity assembly statistics (n transcripts, unigenes, N50, total length) are not stated in the paper at all → no_expected_result for the assembly step itself even if data were available.
- Wet-lab results (qRT-PCR, transgenic line phenotypes, Suppl. Fig 1, primer table S4) are out of scope (manual/experimental).
Blocker (see AUDIT.md / agreement.json)
The pipeline input data is not publicly obtainable → the whole in-scope set cannot be run. This is a data-availability DROP, documented in detail in AUDIT.md.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean data-availability failure, not a computational discrepancy. The declared code (Trinity v2.4.0) is fully available, but the raw reads of PRJNA852945 were never released - the paper's statement offers only a private peer-review token and a still-non-public BioProject ~3y post-publication - so none of C1-C9 (Table 1 QC stats; Fig 1C DEG counts 547/4137/22712) could be put against any reproduced output. The blocker sits on the authors'/data-availability side, so q1/q2 are red and q4 red; I keep q5/q7 at yellow rather than red because there is no evidence the numbers are wrong, only that they are unverifiable. q8 is red overall because complete unreproducibility plus a non-functional data-availability statement is a critical outcome a human should flag.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.