De novo transcriptomic analysis of leaf and fruit tissue of Cornus officinalis using Illumina platform.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🔴Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🔴The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
MISMATCH (described well enough, ran 1:1 on the data that IS public, headline numbers do NOT reproduce). Paper: Hou et al. 2018, de novo transcriptome of Cornus officinalis leaf+fruit (PLoS ONE, PMC5815590). Linked 'code' = SeqPrep, a third-party paired-end read merger; per P16 I reproduced by running the read-preprocessing chain on the paper's own data (ENA SRP115440: SRR5936587 leaf, SRR5936588 fruit). On «our HPC» («job», seqkit v2.13.0 / sickle 1.33 / fastqc v0.12.1) I downloaded the FASTQ (md5 == ENA, so genuine), counted reads/bases/GC/Q20, and ran sickle q20 l20. RESULT: the public deposit is SINGLE-END ~12M reads/run (1.81 Gbp), but the paper reports ~58-61M raw reads / 8.6-9.0 Gbp per sample -- a ~4.8-5.1x shortfall (deposit ~21% of reported). GC% matches the paper to within 0.05 (46.57 vs 46.54; 46.03 vs 45.98) and md5 matches ENA, confirming the deposit is authentic and the right samples -- so the reported raw-read table is NOT derivable from the shipped data (subsampled deposit and/or count inflation). The SeqPrep citation (a PE merger) is also internally inconsistent with a single-end deposit. Reproduced fine: GC% (within-tol), qualitative cleaning claim (<1% reads removed, within-tol). Flagged as possible-fabrication / data-integrity for human audit (see AUDIT.md). NOT ATTEMPTED (the hard ~20%, by design): Trinity assembly Table 1 (56,392 unigenes / N50 1445 / 70,329 transcripts), NR/GO/KOG/KEGG annotation, and DEG counts -- de-novo assembly is non-deterministic, Trinity version unspecified, the input volume does not match so Table 1 cannot reproduce regardless, and there are no biological replicates for the DE analysis.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 45assessed: 2026-06-16 ⛓ cd629a35f202
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBecause the genetics and molecular biology of Cornus officinalis are poorly understood, this study performs the first de novo transcriptome sequencing of its leaf and fruit tissue to obtain genomic information and explore the molecular mechanisms of secondary metabolite (terpene/iridoid glucoside) biosynthesis.
- ★ This is the first de novo transcriptomic analysis of Cornus officinalis, providing fundamental gene and biosynthetic pathway information. resource
- ★ Pooled leaf and fruit reads assembled into 56,392 unigenes with an average length of 856 bp, of which 41,146 matched the NCBI NR protein database. resource
- ★ 4,585 significant differentially expressed genes were identified between fruit and leaf (1,392 up-regulated, 3,193 down-regulated in fruit vs leaf). finding
- ★ 581 transcription factors spanning 50 transcription factor gene families were identified in the transcriptome. finding
- ★ Most DEGs and transcription factors were related to terpene biosynthesis and secondary metabolic regulation. mechanism
- 18,435 unigenes were mapped to 371 KEGG pathways, with biosynthesis of secondary metabolites a top represented pathway. finding
- De novo assembly with Trinity and downstream annotation against NR/GO/KOG/KEGG is an effective approach for a non-model medicinal plant lacking a reference genome. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (de novo transcriptome sequencing) | Cornus officinalis mature fruit (coGS) and leaf (coYP) tissue | none (tissue comparison: leaf vs fruit) | clean reads, assembled unigenes/transcripts, transcript abundance | Illumina HiSeq 4000 SBS Kit (300 cycles); TruSeq RNA sample preparation kit; library quantified by TBS 380 Mini-Fluorometer |
| RNA extraction and quality assessment | frozen C. officinalis leaf and fruit tissue (~80 mg) | none | RNA quality (260/280 ratio) | TRIzol Reagent (Invitrogen); NanoDrop 2000 spectrophotometer |
| de novo transcriptome assembly | high-quality leaf and fruit reads of C. officinalis | none | unigenes and transcripts (count, length, N50, GC) | Trinity (assembly); SeqPrep and Sickle (read trimming) |
| functional annotation (BlastX sequence homology) | C. officinalis assembled unigenes | none | unigene matches/annotation in NR, Swissprot, GO, KOG, KEGG databases | NCBI BlastX (E-value <1e-5); Blast2GO; KOBAS |
| expression quantification | C. officinalis fruit (coGS) and leaf (coYP) libraries | none | read alignment counts and FPKM expression levels | RSEM package (fragment lengths 200-300 bp) |
| differential expression analysis | C. officinalis fruit vs leaf | none | differentially expressed genes (up/down-regulated) | R Bioconductor edgeR (FDR<0.05, |log2FC|>=1) |
| transcription factor prediction (BlastP) | C. officinalis leaf and fruit transcriptome unigenes | none | predicted transcription factors and gene families | PlantTFDB 3.0; BlastP (E-value <1e-8) |
- – 60,971,652 clean reads obtained for fruit (coGS) and 57,954,134 clean reads for leaf (coYP) 60,971,652 and 57,954,134 clean reads
- – Assembly yielded 56,392 unigenes (avg 856 bp, N50 1445 bp) and 70,329 transcripts 56,392 unigenes; N50 1445 bp
- – 41,146 unigenes matched the NR database; top species hit Vitis vinifera 7,722 unigenes (18.77%) matched Vitis vinifera
- – 24,336 unigenes assigned GO terms across biological process (83.26%), cellular components (53.58%), molecular function (83.93%) 24,336 unigenes
- – 10,808 unigenes assigned to 25 KOG functional categories 10,808 unigenes
- – 18,435 unigenes mapped to 371 KEGG pathways 18,435 unigenes; 371 pathways
- – 4,585 significant DEGs in fruit vs leaf, 1,392 up-regulated and 3,193 down-regulated 1,392 up / 3,193 down
- – 581 transcription factors in 50 gene families identified 581 TFs; 50 families
- count 56,392 unigenes (48,264,743 bp; average 856 bp; N50 1445 bp) (de novo assembly result)
- count 41,146 unigenes matched NR database (NR annotation)
- count 4,585 significant DEGs (1,392 up, 3,193 down) (edgeR DEGs, FDR<0.05, log2FC>=1, p<0.05)
- count 26,136 total DEGs without significance threshold (11,735 up, 14,401/14,402 down) (DEGs ignoring significance)
- count 18,435 unigenes mapped to 371 KEGG pathways (KEGG pathway mapping)
- other Q20 98.08% / Q30 97.94% (fruit); Q20 93.97% / Q30 93.58% (leaf); GC 45.98% (fruit), 46.54% (leaf) (clean read quality metrics)
- count 581 transcription factors in 50 TF gene families (PlantTFDB prediction)
- other 82.05% of fruit clean reads and 82.93% of leaf clean reads aligned to unigene database (RSEM read alignment rate)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study sequenced a single Illumina HiSeq 4000 cDNA library per tissue type (leaf and fruit of Cornus officinalis), performed de novo assembly with Trinity, and quantified expression as FPKM using RSEM. Differential expression between the two tissues was assessed with edgeR using thresholds of FDR < 0.05 and |log2FC| ≥ 1. GO and KEGG pathway enrichment of the resulting DEG set was conducted with GOatools and KOBAS at p < 0.05, and results were reported as DEG counts, FPKM distributions, and pathway/term assignments rather than traditional summary statistics.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| edgeR negative binomial exact test / GLM | Differential expression between leaf (coYP) and fruit (coGS) tissue across 56,392 unigenes | 1 library per tissue; no biological replicates stated | not stated |
| Fisher exact / hypergeometric enrichment test (GOatools) | GO term enrichment among significant DEGs | — | not stated |
| Fisher exact / hypergeometric enrichment test (KOBAS) | KEGG pathway enrichment among significant DEGs | — | not stated |
| BlastP homology search (E-value < 1e-8) | Transcription factor identification against Plant TFDB 3.0 | — | na |
-
Differential expression was estimated from a single library per tissue with no stated biological replicates↳ Could also: Include biological replicates (≥ 3 independent samples per tissue) before applying edgeR or DESeq2 — Biological replicates enable empirical estimation of within-group variance; edgeR and DESeq2 both rely on replicate-based dispersion estimates to set appropriate significance thresholds and control false discovery rates — without replicates, dispersion must be fixed or estimated from the data under assumptions that may not reflect true biological variability
-
edgeR was used as the sole DEG-calling tool↳ Could also: DESeq2 (Love et al. 2014) could also be applied to the same count matrix — DESeq2 uses shrinkage estimation for log2 fold changes and a different variance-stabilization strategy; comparing results from both tools is a common cross-validation approach in RNA-seq workflows and can highlight findings that are robust across methods
-
Expression levels were quantified and reported as FPKM↳ Could also: TPM (transcripts per million) is an alternative abundance unit — TPM normalizes within-sample before between-sample, causing values to sum to the same constant across all samples and facilitating more direct cross-sample comparisons; it has become the more widely recommended unit in recent transcriptomics literature
-
GO and KEGG enrichment analyses used a raw p-value threshold of 0.05 with no stated multiple-testing correction↳ Could also: Apply Benjamini-Hochberg FDR correction across the full family of GO terms or KEGG pathways tested — Enrichment analyses simultaneously evaluate hundreds to thousands of terms; FDR correction on this family controls the expected proportion of false discoveries among reported enriched terms, which raw p-value thresholds do not
-
Transcriptome assembly quality was characterized by N50, N90, and size statistics↳ Could also: BUSCO (Benchmarking Universal Single-Copy Orthologs) assessment could also be applied — BUSCO provides a gene-content-based completeness estimate by searching for conserved single-copy orthologs expected in the lineage, complementing contiguity statistics like N50 with an informative measure of biological completeness
-
Transcript abundance was estimated with RSEM mapped against the de novo assembly↳ Could also: Pseudoalignment tools such as kallisto or Salmon could also quantify transcript abundances directly from reads — Pseudoalignment approaches are computationally faster and propagate quantification uncertainty to downstream analyses; they have become widely used alternatives for expression quantification when a reference or de novo assembly is available
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
4585 DEGs identified between C. officinalis fruit and leaf; 1392 up-regulated in fruit and 3193 down-regulated in fruit relative to leafRNA-seq cornus officinalis fruit-leaf mixed 2018×1papers★ This paper is the founder (earliest)
-
581 transcription factors spanning 50 gene families identified in C. officinalis leaf and fruit transcriptomeRNA-seq cornus officinalis fruit-leaf 2018×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
0 downstream papers · 1 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-29451882
Paper: Hou et al. (2018) De novo transcriptomic analysis of leaf and fruit tissue of Cornus officinalis using Illumina platform. PLoS ONE 13(2):e0192610. PMCID PMC5815590 · DOI 10.1371/journal.pone.0192610.
Linked "code": https://github.com/jstjohn/SeqPrep — a third-party tool that merges overlapping paired-end Illumina reads and trims adapters. Per brief P16, applying this third-party tool to the paper's own data is a valid reproduction. The paper also names Sickle (quality trimming) and Trinity (de novo assembly).
Data: SRA study SRP115440 (BioProject PRJNA398165), 2 runs:
SRR5936587= coYP = leaf RNA-seqSRR5936588= coGS = fruit RNA-seq
Pipeline-derived results in the paper
| # | Result | Pipeline / tool | In scope? |
|---|---|---|---|
| R1 | Raw read counts per sample (leaf 57,954,134; fruit 60,971,652) and bases | sequencing → SRA | YES (cheap, hard number) |
| R2 | GC content of reads (leaf 46.54%, fruit 45.98%) | FastQC-class QC | YES (FastQC) |
| R3 | Q20 % (leaf 93.97%, fruit 98.08%) | FastQC-class QC | YES (FastQC) |
| R4 | Clean reads after SeqPrep+Sickle (Q<20 / Q<10 / drop N / len<20) | SeqPrep + Sickle (linked tool) | YES (run linked tool) |
| R5 | Trinity assembly: 56,392 unigenes, 70,329 transcripts, N50 1445/1536, total 48.26/66.77 Mbp, GC 43.74/43.34%, max 124,880, min 201 (Table 1) | Trinity | PARTIAL / optional 20% — non-deterministic, Trinity version unspecified, and input data does not match (see below); not chased for 1:1 |
| R6 | Annotation counts (NR 41,146; GO 24,336; KOG 10,808; KEGG 18,435 / 371 pathways) | BLAST/diamond + KAAS etc. | OUT (downstream of R5; depends on an assembly we cannot match) |
| R7 | DEGs leaf vs fruit (4,585 sig; up 1,392 / down 3,193; 26,136 total) | RSEM/edgeR-class | OUT (no replicates → DE is not properly defined; downstream of R5) |
Out of scope / not attempted
- R5–R7 are downstream of a de novo assembly that is (a) non-deterministic, (b) built with an unspecified Trinity version/params, and (c) cannot use the reported input because the deposited data is ~1/4.8 the reported volume (see the critical discrepancy below). Reproducing 56,392 unigenes exactly is not feasible and is the explicit "hard 20%" the brief says to skip. We document why.
CRITICAL discrepancy found at screening (control-plane, no compute)
The public deposited data does not match the reported raw reads:
| Sample | Paper raw reads | Paper bases | SRA/ENA deposited reads | SRA/ENA bases | Layout |
|---|---|---|---|---|---|
| leaf (SRR5936587/coYP) | 57,954,134 | 8,579,879,557 bp | 11,976,344 | 1,808,427,944 bp | SINGLE |
| fruit (SRR5936588/coGS) | 60,971,652 | 9,044,801,270 bp | 11,976,272 | 1,808,417,072 bp | SINGLE |
→ Deposited volume is ~4.8× smaller than reported (≈21% of the reported reads/bases). Additionally the deposit is single-end, yet the named preprocessing tool SeqPrep merges paired-end reads — inconsistent with a single-end deposit. These are recorded as possible-fabrication / data-integrity flags for the human auditor (source: ENA filereport + NCBI SRA runinfo, below).
Plan
Run real compute on «our HPC» against the data that IS public (the only honest option): count reads/bases (seqkit), measure GC% and Q20% (FastQC), and run the linked tool chain (SeqPrep/Sickle) with the paper's parameters to report a clean read count. Compare 1:1 to R1–R4. R5–R7 documented as not-reproduced-by-design.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Sample identity is solid — GC% reproduces to within 0.05 (46.57 vs 46.54; 46.03 vs 45.98) and md5 matches ENA — yet the paper's headline raw-data table reports ~58-61M reads / 8.6-9.0 Gbp per sample against an authentic public deposit of only ~12M single-end reads / 1.81 Gbp (a 4.8-5.1x shortfall). This is squarely an authors'/data-availability problem, not our method: the reported values are not derivable from the shared data, and citing SeqPrep (a paired-end merger) against a single-end deposit adds an internal inconsistency. Because the assembly (Table 1) rests on ~5x data that is not public, the central transcriptome resource cannot be reproduced — flagged as a data-integrity / possible-fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.