Exploring candidate genes for pericarp russet pigmentation of sand pear (Pyrus pyrifolia) via RNA-Seq data in two genotypes contrasting for pericarp color.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the pipeline-derived core of this paper using a reference-genome-based substitute pipeline (SeqPrep [paper's own cited tool] -> minimap2 -ax splice [substitute for unmaintained TopHat] -> samtools -> featureCounts [substitute for Cufflinks] against the paper's actual cited reference genome Pbr_v1.0/GCF_000315295.1, not the paper's own de novo unigene assembly). Both RNA-Seq libraries (S1 non-russet, S2 russet; true accessions SRR1582086/SRR1582087, NOT the SRR925365 named in this room's Data field, which does not exist in SRA) downloaded, trimmed, and aligned cleanly: 82.73%/82.01% mapping rates for S1/S2 vs the paper's reported 77.4%/85.1% (TopHat) -- within-tolerance given the aligner substitution. All 15 candidate genes (GALR TSA accessions) discussed in the paper's Results/Discussion were retrieved and uniquely localized to Pbr_v1.0 gene models at 99.7-100% identity. Of these, 13/15 show CPM-normalized expression-direction concordance with the paper's stated up/down calls; 1/15 (GALR01022575, HOTHEAD-like) is a clear mismatch (paper: down in russet; reproduced: up); 1/15 (GALR01011224, peroxidase) could not be cleanly graded because the paper's own text states its direction inconsistently in two places. NOT attempted: the paper's genome-wide de novo assembly + unigene-based statistical DEG pipeline (29,100 unigenes; 206 DEGs at FDR<=0.001 & |log2FC|>1; 137 after length filtering) -- this room used a reference-genome BLAST approach targeted at the 15 explicitly-named candidate genes instead, which is a methodologically different (and in this case, feasible-within-budget) but not identical reproduction of the paper's own DEG-calling pipeline. Two data-provenance errors inherited from a prior room checkpoint were found and corrected: (1) the reference genome was previously mislabeled 'Pbr_v1.0' while actually being PPY_r1.0 (Pyrus pyrifolia, the study species, not the paper's cited P. bretschneideri reference); (2) a prior checkpoint falsely claimed 15/15 candidate gene FASTA files were already fetched and saved, when none existed on disk. This status is 'partial' (not 'reproduced') because the genome-wide DEG-calling pipeline was not attempted.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-31
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper asks which and how many genes underlie russet versus green pericarp pigmentation in sand pear (Pyrus pyrifolia), hypothesizing that RNA-seq-based bulked segregant analysis of contrasting F1 pericarp pools can identify candidate genes for cuticle and cork (suberin/cutin/wax/lignin) layer formation.
- ★ RNA-seq-based bulked segregant analysis of russet- vs green-pericarp F1 pools identified 29,100 unigenes, 206 of which were significantly differentially expressed (|log2 fold change| > 1). finding
- ★ Candidate genes for suberin, cutin and wax biosynthesis showed repressed expression in russet pericarps. finding
- ★ Genes encoding putative cinnamoyl-CoA reductase (CCR), cinnamyl alcohol dehydrogenase (CAD) and peroxidase (POD) of the lignin biosynthesis pathway are candidates for sand pear russet pericarp pigmentation. mechanism
- ★ GO analysis detected 123 unigenes in terms related to 'cellular_component' and 'biological_process', suggesting developmental and growth differentiation between the two pericarp types. finding
- ★ GO categories for 'lipid metabolic processes', 'transport', 'response to stress' and 'oxidation-reduction process' were enriched among genes with divergent expression between the two libraries. finding
- ★ qRT-PCR of nine differentially expressed genes gave results consistent with the Illumina RNA-sequencing data, validating the expression profiling. method
- The assembled sand pear pericarp transcriptome (23,002 assemblies, accession GALR00000000) is a new annotated resource for pear pericarp research. resource
- This is the first report to use a global (transcriptome-wide) approach to decipher the molecular biology of fruit cork and suberin. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk RNA-seq (101 bp paired-end, bulked segregant analysis of two pooled libraries) | Pericarp (periderm) of 80-day-after-pollination fruits from Pyrus pyrifolia 'Qingxiang' × 'Cuiguan' F1 offspring; pools of 10 russet (S2) and 10 green (S1) individuals | none (natural genotype contrast: russet vs green pericarp segregants) | Transcript abundance as FPKM; differentially expressed unigenes (FDR ≤ 0.001, |log2 fold change| > 1) | Illumina HiSeq2000, single lane; TruSeq RNA Sample Preparation Kit (Illumina); Plant RNA Purification Reagent (Invitrogen); SuperScript II reverse transcriptase (Life Technologies); TBS380 Picogreen quantification |
| qRT-PCR | S1 (green) and S2 (russet) sand pear pericarp pools | none (russet vs green comparison) | Relative transcript levels of nine selected genes, normalized to EF1-a (AY338250) by the 2^-ΔΔCT method; three biological replicates | ABI 7500 Real-Time PCR System (Applied Biosystems); TransStart Top Green qPCR SuperMix (TransGen); primers designed with Primer 6.0 |
| Light microscopy (histology, 200× magnification) | 'Cuiguan' young fruit pericarp (green) | none | Pericarp tissue structure: cuticle layer, epidermal cell layer, cork cambium, parenchyma cell layer, stone cells | — |
| Read processing, genome mapping and transcript assembly (bioinformatics) | Illumina reads from S1 and S2 pericarp libraries mapped to sand pear reference genome Pbr_v1.0 | none | Mapping rate, ORF prediction, transfrag/transcript assembly, FPKM quantification | SeqPrep, Condetri, TopHat, Cufflinks/Cuffmerge, Trinity, R/cummeRbund |
| Functional annotation and GO/KEGG enrichment analysis | Assembled sand pear pericarp unigenes/transcripts (with and without ORFs) | none | BlastP/BlastX hits against NT, NR, String and KEGG-Gene databases (E-value 1.0×10^-5); GO category assignment; enriched GO terms (q-value cutoff 0.05) | blast2go, WEGO, GOseq (R package), Kobas |
- – 206 of 29,100 unigenes were significantly differentially expressed between russet- and green-pericarp pools log2 fold change > 1; FDR ≤ 0.001
- ▼ Candidate genes for suberin, cutin and wax biosynthesis were repressed in russet pericarps
- – 123 unigenes fell into GO terms related to 'cellular_component' and 'biological_process'
- – 77.4% (S1, green) and 85.1% (S2, russet) of filtered reads mapped to the sand pear reference genome Pbr_v1.0 77.4% and 85.1%
- – 22,046 non-redundant mapping coordinates had a predicted ORF while 4,101 had none; predicted ORFs ranged 150–15,285 bp (mean 1,236 bp) mean ORF 1236 bp
- – 19,862 unigenes were assigned at least one GO functional category, distributed over 28 Biological Process, 9 Cellular Component and 14 Molecular Function subsets
- – qRT-PCR expression of nine differentially expressed genes agreed with the Illumina RNA-seq results
- – 58,524,312 ESTs of 101 bp were obtained and assembled into 29,100 unigenes; ~24 million filtered high-quality reads (average 96 bp) per sample 58,524,312 ESTs; ~24 million reads/sample
- count 29,100 (total unigenes identified from the pericarp transcriptome)
- count 206 (unigenes significantly differentially expressed between russet and green pericarp pools)
- fold_change log2 fold change > 1 (threshold for significant differential expression (with FDR ≤ 0.001))
- pvalue FDR ≤ 0.001 (false discovery rate cutoff for differential expression)
- count 58,524,312 (total ESTs (101 bp each) generated by Illumina sequencing)
- other 77.4% and 85.1% (filtered reads of S1 (green) and S2 (russet) mapped to reference genome Pbr_v1.0)
- count 19,862 (unigenes with at least one GO functional category)
- pvalue E-value 1.0×10^-5; GO q-value cutoff 0.05 (BLAST significance threshold and GO enrichment cutoff)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study compared gene expression between two RNA-seq libraries built from pooled pericarp tissue (10 green-pericarp offspring pooled as S1; 10 russet-pericarp offspring pooled as S2), an RNA-seq-based bulked segregant analysis design. Differential expression was called using a Cufflinks/Cuffmerge-based workflow with an FDR ≤0.001 and |log2 fold change|>1 threshold, GO term enrichment was assessed with GOseq using a Z-score/T-statistic-derived p-value with multiple-testing correction (q<0.05), and nine candidate genes were validated by qRT-PCR using the 2^-ΔΔCT method averaged over three biological replicates. Results were reported primarily as fold-change values and FDR/q-value cutoffs rather than as tables of per-gene test statistics or confidence intervals.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Cufflinks/Cuffmerge-based differential expression calling (FDR and log2 fold-change threshold) | Comparison of digital gene expression between the green-pericarp (S1) and russet-pericarp (S2) pools | Two pooled RNA-seq libraries, each pooled from 10 offspring individuals (one pool per pericarp type) | not stated |
| GOseq enrichment test (Z-score derived from mean T-statistic, converted to p-values, multiple-test corrected) | GO term enrichment among the 206 differentially expressed unigenes relative to the 29,100 total unigenes | 206 differentially expressed unigenes out of 29,100 total unigenes | not stated |
| qRT-PCR relative quantification, 2^-ΔΔCT method | Validation of 9 selected differentially expressed genes comparing S1 and S2 pericarp pools | Average of three biological replicates | not stated |
-
RNA-seq differential expression was assessed from a single pooled library per pericarp type (bulked segregant analysis), without additional independently sequenced biological replicates per group.↳ Could also: Sequencing multiple independent biological replicates per group and analyzing counts with a replicate-aware tool such as DESeq2, edgeR, or limma-voom — Replicate-based tools estimate within-group biological variance directly and provide a per-gene test statistic and p-value, which can complement the pooled-sample fold-change/FDR approach used here.
-
Differentially expressed genes were defined using an FDR ≤0.001 threshold combined with a |log2 fold change|>1 cutoff.↳ Could also: Reporting the exact adjusted p-value (q-value) and fold-change confidence interval for each gene alongside the cutoff — Exact values and interval estimates let readers gauge how far individual genes fall from the threshold and the precision of the fold-change estimate, complementing a binary significant/non-significant cutoff.
-
GO term enrichment significance was corrected for multiple testing using a q-value cutoff of 0.05, without naming the specific correction algorithm.↳ Could also: Explicitly specifying and reporting a Benjamini-Hochberg FDR correction (as used in tools like topGO or clusterProfiler) — Naming the exact multiple-testing procedure makes the enrichment analysis fully reproducible and lets readers compare results against other studies using the same standard method.
-
qRT-PCR validation of 9 candidate genes was summarized as an average of three biological replicates using the 2^-ΔΔCT method, without a stated inferential test or dispersion measure.↳ Could also: Reporting SD or SEM alongside the mean, and applying a t-test or non-parametric Mann-Whitney U test to compare ΔΔCT values between pools — Adding a dispersion measure and an inferential test would let readers assess the variability among replicates and the statistical confidence behind the RNA-seq/qPCR concordance claim.
-
The study design pools tissue from 10 individuals per pericarp type into a single sequencing library per condition (BSA), rather than sequencing individuals separately.↳ Could also: An individually-barcoded, multi-replicate RNA-seq design — Individual-level sequencing allows estimation of inter-individual variability within each phenotype class, which pooled BSA designs are not able to capture directly.
-
Effect sizes for expression differences were reported as log2 fold change without accompanying confidence intervals.↳ Could also: Reporting a confidence interval (e.g., from a shrinkage estimator such as DESeq2's apeglm/ashr) around each fold-change estimate — An interval estimate conveys the precision of the fold-change value in addition to its point estimate and significance threshold.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The pipeline-derived core that could be checked largely holds: both real libraries (SRR1582086/SRR1582087) trimmed and aligned cleanly at 82.73%/82.01% against the paper's reported 77.4%/85.1%, a deviation fully explicable by substituting minimap2 -ax splice for the unmaintained TopHat, and all 15 GALR TSA candidate accessions localize uniquely to Pbr_v1.0 gene models at 99.7-100% identity with 13/15 showing concordant expression direction. Two defects sit on the authors'/deposit side: the accession recorded for this study (SRR925365) does not exist in SRA, and the paper's own text gives contradictory up/down statements for GALR01011224. The largest unverified block is the genome-wide branch — 29,100 unigenes, 206 DEGs at FDR<=0.001, 137 after filtering, R2=0.88 and R2=0.85 — which is not derivable because the de novo unigene assembly was never deposited and the single-replicate-per-condition design provides no variance for the paper's own FDR threshold; this was not attempted, so it is unchecked rather than contradicted. Net: a solid partial reproduction with explainable deviations, limited (not refuted) support for the central candidate-gene claim, and no fabrication signal.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.