Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

From bud formation to flowering: transcriptomic state defines the cherry developmental phases of sweet cherry bud dormancy.

BMC Genomics · 2019
L1 88/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
88/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 74% of all assessed papers rank 276 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the DOWNSTREAM pipeline; reproduced 1:1 on the strongest result and partially on the DEG count. The hierarchical clustering of Garnet DEGs into 10 clusters is EXACT: re-running the described method (1-Pearson distance on Garnet TPM, complete-linkage hclust, cutree k=10) on the authors' deposited 6683-DEG list (Additional file 2 Table S2) reproduces all ten cluster sizes (1548,989,924,884,739,648,612,156,113,70) with Adjusted Rand Index = 1.000 (every gene in the same cluster). Reference genome (P. persica v2.0), gene count (26873) and Garnet sample count (31) confirmed exactly from the deposited GSE130426 count tables. The DESeq2 DEG count (6683 'dormant vs non-dormant') is only PARTIALLY reproduced: the paper does not specify which dates/stages form the binary dormant/non-dormant grouping nor the CV-filter scale, and DESeq2 has drifted from the 2018 version (~1.18-1.20) to 1.42.0. Sweeping 3 filters x 8 plausible groupings gives DEG counts of 5969-7696 (bracketing 6683); the best matches the authors' exact list at 73% gene overlap (Jaccard 0.54), 7166 DEGs. 100% of the authors' 6683 DEGs survive my filter, so the gap is the DESeq2 statistical call, not filtering. NOT attempted: the upstream alignment (Trimmomatic/TopHat/Picard) and Table S6 mapping statistics, which would require re-running a deprecated aligner on 82 SRA FASTQ libraries + the peach genome (deposited counts were used instead); and the Cristobalina/Regina cultivars, wet-lab phenology, RT-qPCR, GO/TF enrichment, and ML flowering models (out of pipeline-DEG/clustering scope). No evidence of fabrication: the one numeric discrepancy is cluster 1 = 1549 (paper text) vs 1548 (deposited Table S2), a benign +1.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 88
    assessed: 2026-06-16 ⛓ c92c3340fcd4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study asks whether fine-resolution, genome-wide transcriptomic changes throughout sweet cherry (Prunus avium L.) flower bud development define distinct dormancy stages (organogenesis, paradormancy, endodormancy, ecodormancy), whether these transcriptional signatures are conserved across cultivars with contrasted flowering dates, and whether a small set of marker genes can predict dormancy stage.

Core claims
  • Flower buds in organogenesis, paradormancy, endodormancy and ecodormancy stages are each defined by expression of genes in specific pathways, and the transcriptional state accurately captures the dormancy state. finding
  • These stage-specific transcriptional changes are conserved between sweet cherry cultivars with contrasted dormancy release dates. finding
  • DORMANCY ASSOCIATED MADS-box (DAM), floral identity and organogenesis genes are up-regulated during pre-dormancy stages, while endodormancy is characterized by cold response, ABA and oxidation-reduction pathways. mechanism
  • Endodormancy is separable into two distinct transcriptional periods (Oct/Nov vs Dec), indicating dormancy is a chain of biological events rather than an on/off mechanism. finding
  • A model based on the transcriptional profiles of just seven genes can accurately predict the main bud dormancy stages. method
  • After dormancy release, genes for cell activity, division, differentiation, transport, cell wall biogenesis and oxidation-reduction are activated during ecodormancy. finding
  • Specific transcription factors (e.g. ABF2, ABI5, MADS-box AP3/AG, ERF/DREB) have over-represented targets and target promoter motifs in stage-specific gene clusters. mechanism
  • 81 transcriptomes spanning bud development provide a fine-resolution time-course resource for dormancy in three cultivars. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (transcriptomics) sweet cherry (Prunus avium L.) flower buds, cultivars 'Cristobalina', 'Garnet', 'Regina' none (seasonal time-course sampling, 11 dates July–April) gene expression (TPM, differential expression between dormant/non-dormant stages)
forcing assay / phenological observation sweet cherry flower buds, three cultivars forcing conditions bud break percentage (50% at BBCH stage 53 = dormancy release)
differential expression analysis 'Garnet' RNA-seq transcriptomes none DEGs between dormant and non-dormant bud stages DESeq2 (adjusted p-value threshold 0.05)
hierarchical clustering / PCA 'Garnet' DEGs none expression clusters and sample separation by stage
GO enrichment analysis 'Garnet' DEG clusters none enriched biological process GO terms per cluster topGO (classic Fisher algorithm)
transcription factor target / motif enrichment peach (Prunus persica) reference regulation applied to cherry gene clusters none TFs and promoter motifs with over-represented targets per cluster PlantTFDB 4.0; FIMO; hypergeometric tests (FDR)
predictive modelling sweet cherry bud transcriptomes none prediction of main bud dormancy stage from gene expression seven-gene transcriptional model
Key results
  • 6683 genes differentially expressed between dormant and non-dormant bud stages in 'Garnet' 6683 DEGs
  • PCA of DEGs cleanly separates bud stages (organogenesis and paradormancy projecting together); PC1 represents dormancy strength PC1 = 41.63% variance
  • PC2 distinguishes phases before and after dormancy release PC2 = 20.24% variance
  • DEGs grouped into ten clusters with distinct expression peaks across organogenesis/paradormancy, endodormancy, and ecodormancy 10 clusters
  • PavDAM1, PavDAM3, PavDAM6 highly expressed during paradormancy/early endodormancy; PavDAM4 peaks at end of endodormancy 4 of 6 DAM genes differentially expressed
  • ABA pathway genes PavABF2, PavATHB7, PavCYP707A2 and stress gene PavHVA22 highly expressed during endodormancy
  • Floral identity genes PavAGL20 and PavFD up-regulated before dormancy; PavAG and PavAP3 peak during ecodormancy
  • PavGH17 (1,3-β-glucanases) and PavPDCB3 repressed during dormancy
Key statistics
  • count 6683 differentially expressed genes (DEGs between dormant and non-dormant 'Garnet' bud stages (DESeq2, adj. p<0.05))
  • count 81 transcriptomes (total RNA-seq samples across three cultivars and 11 dates)
  • other 41.63% (variance explained by PC1 (dormancy strength))
  • other 20.24% (variance explained by PC2 (before vs after dormancy release))
  • pvalue adj. p = 7.5E-04 (***) (PavABF2 target over-representation in cluster 8)
  • pvalue adj. p = 2.8E-05 (***) (PavAP3 / PavAGL15 target motif enrichment in clusters)
  • count 7 genes (genes used in transcriptional model to predict dormancy stages)
  • count 11 sampling dates (harvest dates spanning bud stages July to April)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This RNA-seq time-course study profiled sweet cherry flower buds across developmental and dormancy stages and used a largely descriptive/exploratory analytical pipeline. Differentially expressed genes were identified with DESeq2 using an adjusted p-value threshold of 0.05, samples were summarized by principal component analysis, genes were grouped by hierarchical clustering, and functional interpretation relied on GO enrichment (topGO, Fisher) plus hypergeometric tests for transcription-factor target enrichment with FDR correction. A predictive model based on seven genes was then developed to classify dormancy stages.

Replicationbiological Sample sizeThree trees sampled per date over 11 dates and three cultivars, yielding 81 transcriptomes; no formal power/sample-size calculation described GroupsBud developmental/dormancy stages (organogenesis, paradormancy, endodormancy, dormancy release, ecodormancy) across cultivars Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg / false discovery rate (DESeq2 adjusted p-value; FDR for TF-target hypergeometric tests)
Statistical tests used
Test Applied to n Assumptions
DESeq2 differential expression (Wald test, adjusted p-value < 0.05) DEGs between dormant and non-dormant bud stages in cultivar 'Garnet' (6683 genes) three trees (biological replicates) per sampling date; 81 transcriptomes total across cultivars not stated
GO term enrichment using a classic Fisher algorithm (topGO) biological-process GO enrichment for each of the 10 gene clusters not stated
Hypergeometric test for over-representation of transcription-factor target genes Table 1, enrichment of TF targets within clusters not stated
Motif occurrence detection (FIMO) for enriched target promoter motifs Table 2, over-represented target motifs in clusters na
Approaches that could also have been used
  • Differential expression was identified with DESeq2 using an adjusted p-value threshold of 0.05.
    Could also: Edge-case-robust pipelines such as edgeR or limma-voom could also have been used, and a combined adjusted-p-value plus log fold-change threshold could define DEGs. — Adding an effect-size (fold-change) criterion alongside significance, or cross-checking with an alternative pipeline, can emphasize biologically larger changes and characterize robustness of the gene set.
  • Genes were grouped into ten clusters using hierarchical clustering of expression profiles.
    Could also: Model-based or soft-clustering approaches (e.g., k-means with a chosen-k criterion, or fuzzy c-means as in Mfuzz) could also have been applied. — Soft clustering allows genes to have partial membership across temporal profiles, which can be informative for continuous time-course transcriptomic data.
  • GO enrichment used a classic Fisher algorithm within topGO.
    Could also: The topGO 'elim' or 'weight' algorithms, or other tools (e.g., GSEA), could also have been used. — GO-structure-aware algorithms account for term dependency, and GSEA-style ranked analyses avoid reliance on a hard DEG cutoff, offering complementary views of pathway signal.
  • The time course was analyzed primarily by stage-based comparisons, clustering, and PCA.
    Could also: Explicit time-series / spline-based differential expression frameworks (e.g., maSigPro, ImpulseDE2) could also have been used. — Time-aware models directly leverage the ordering of the 11 sampling dates and can identify genes with specific temporal trajectories.
  • Cluster expression patterns and key genes were presented using TPM values and z-scores.
    Could also: Adding dispersion summaries (SD, IQR, or 95% CI) across the biological replicates could also accompany the displayed values. — Showing spread across replicates conveys variability and is often preferred, particularly with a small number of replicate trees per date.
Software: DESeq2 · topGO · FIMO (MEME suite) · PlantTFDB 4.0

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
78
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31830909 (Vimont et al. 2019, BMC Genomics)

"From bud formation to flowering: transcriptomic state defines the cherry developmental phases of sweet cherry bud dormancy." DOI 10.1186/s12864-019-6348-z.

The pipeline (from Methods)

RNA-seq of flower buds, 3 cultivars (Cristobalina/Garnet/Regina), 11 dates each (82 GSM in GSE130426). Per the paper: FastQC -> Trimmomatic (trim) -> TopHat (map to Prunus persica peach genome v2.0; cherry has no reference) -> Picard (remove optical duplicates; this is the only "code" link in the brief, a generic third-party tool) -> per-gene raw counts + TPM -> pre-filter -> DESeq2 DEGs (padj<0.05 BH) -> hierarchical clustering of Garnet DEGs into 10 clusters (1-Pearson distance on TPM).

In scope (pipeline-derived; attempted)

GEO ships per-sample count tables (raw counts + TPM, 26873 P. persica v2.0 genes) in GSE130426_RAW.tar — i.e. the output of the upstream alignment. The deposited counts let us reproduce the downstream statistical pipeline directly:

  1. DESeq2 DEG calling — "dormant vs non-dormant" for Garnet, padj<0.05 -> 6683 DEGs.
  2. Hierarchical clustering of those DEGs into 10 clusters (sizes 1548..70).

Ground truth for both is deposited in Additional file 2 (Table S2 = the exact 6683 DEGs with cluster IDs; Table S1 = sample->stage; Table S6 = mapping stats), which makes this auditable at the gene level (overlap, ARI) rather than just by counts.

Out of scope / not attempted

  • Upstream alignment (Trimmomatic/TopHat/Picard) and Table S6 mapping statistics: would require downloading 82 SRA FASTQ libraries + the peach genome and re-running TopHat (a deprecated aligner) + Picard. The processed counts are deposited and were used instead. Mapping-rate reproduction left as not-attempted.
  • Wet-lab phenology (bud-break %, dormancy-release dating), RT-qPCR validation, the DorPatterns Shiny app, GO/TF-target enrichment, and the machine-learning flowering-prediction models (Additional file 3) — manual/external, not pipeline DEG/clustering.
  • Cristobalina / Regina cultivars: the headline DEG+clustering result is Garnet-only; focused there.

Notes on the "code" link

The brief's code URL is github.com/broadinstitute/picard — a generic duplicate-marking tool, not an authors' analysis repo. The authors stated analysis scripts would be on GitHub "upon acceptance"; no public repo was found (searched GitHub users/repos + DorPatterns). Per P16 we reproduce by re-running the described pipeline (DESeq2 + hierarchical clustering) on the paper's own data — equally valid.

Figures / tables: Fig3Table2 Table
n_clusters
Reported
10
Reproduced
10
exact
cluster_sizes
Reported
1549,989,924,884,739,648,612,156,113,70
Reproduced
1548,989,924,884,739,648,612,156,113,70
exact
clustering_ARI_vs_TableS2
Reported
Table S2 gene->cluster partition
Reproduced
ARI=1.000 (identical partition)
exact
garnet_DEGs
Reported
6683
Reproduced
7166 (best 73% / 4878 of 6683 overlap, Jaccard 0.54; range 5969-7696 over designs)
partial
reference_genome
Reported
Prunus persica v2.0
Reproduced
Prupe.*_v2.0.a1 gene IDs confirmed
exact
n_genes_quantified
Reported
26873
Reproduced
26873
exact
n_garnet_samples
Reported
31
Reproduced
31
exact
mapping_statistics_TableS6
Reported
per-sample mapped reads
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 88/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

Strong, partial reproduction. The central result — hierarchical clustering of Garnet DEGs into ten developmental-phase clusters — reproduces 1:1 from the deposited Table S2 set, matching every cluster size and reaching ARI = 1.000, so the core conclusion holds and there is no fabrication concern (the only blemish is a benign +1 text typo, 1549 vs deposited 1548). The single material deviation is the DESeq2 DEG count (7166 vs reported 6683, 73% overlap), which sits on our side: the paper underspecifies the binary dormant/non-dormant grouping and DESeq2 has drifted from ~v1.18 to 1.42.0, while all 6683 authors' DEGs remain real and testable. Upstream alignment and Table S6 mapping stats were not attempted. Net: solid with explainable, non-authors-defect deviations → overall yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

223.4 k
tokens (I/O) · 21.7 M incl. cache
52 min
runtime · 0.11 CPU-h
2.7 GB
peak RAM
4 (1 failed)
HPC jobs
hummel
machine