Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genome-wide identification of conserved and novel microRNAs in one bud and two tender leaves of tea plant (Camellia sinensis) by small RNA sequencing, microarra

BMC Plant Biol · 2017
L1 91/100 PQI 98
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
91/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 82% of all assessed papers rank 197 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH + reproduced. The paper's pipeline-derived headline numbers are reproduced 1:1 from the shipped supplementary data: 175 conserved miRNAs, 83 novel miRNAs, 716 miRNA-target rows (=579 distinct targets), and the 116-conserved + 71-novel split of miRNAs-with-targets all match exactly; the TargetFinder score ceiling (max=4.0) is consistent with the named tool's default cutoff. The named third-party tool TargetFinder (commit 848b2dd) was independently installed and its scoring validated deterministically (perfect site -> score 0). An end-to-end re-run of TargetFinder on the paper's 258 miRNAs (default cutoff 4) against an obtainable annotated C. sinensis RefSeq CDS (GCF_055761805.1) is a PARTIAL match (different DB): a similar fraction of miRNAs find targets (201/258 vs 187/258) but ~5x more pairs (3799 vs 716), expected because the authors' actual transcriptome assembly (target IDs CL####.Contig##) was NEVER deposited - only raw reads SRR1979118/PRJNA281432 are public - so an exact 1:1 of the 716 list is impossible. NOT ATTEMPTED (declared 20%/out-of-scope): de-novo re-assembly of SRR1979118 (would not yield a 1:1 comparison anyway), miRNA discovery from raw reads (mapping tool/params under-specified), microarray, qRT-PCR/5'RLM-RACE, GO/KEGG enrichment. No evidence of fabrication: all reported pipeline numbers are exactly derivable from the shipped data.

💻 Code ↗ 🗄 Data: GSE68267

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 91
    assessed: 2026-06-14 ⛓ 1e24355dbb18
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Due to the lack of a reference genome, miRNAs in tea plant (Camellia sinensis) are poorly characterized; this study aims to identify and validate conserved and novel miRNAs in one bud and two tender leaves using small RNA sequencing combined with genome survey scaffold sequences, and to predict and validate their target genes.

Core claims
  • 175 conserved and 83 novel miRNAs were identified mainly in one bud and two tender leaves of tea plant via small RNA sequencing combined with genome survey data finding
  • Combining genome survey scaffold sequences with small RNA sequencing enables prediction of pre-miRNA secondary structures and miRNA discovery in a non-model plant lacking a reference genome method
  • 716 potential target genes were predicted for 187 miRNAs (116 conserved, 71 novel), enriched for transcription factors, stress response, and phenylpropanoid biosynthesis enzymes finding
  • A negative correlation was observed between expression of 3 of 4 conserved miRNAs and their target transcription factors, supporting post-transcriptional regulation mechanism
  • miRNA-guided cleavage of target mRNAs (ARF17, NAC100, WER, MYB12) was experimentally verified by 5'RLM-RACE finding
  • This work enriches the Theaceae miRNA database and provides a resource of potential miRNA regulators of secondary metabolism in C. sinensis resource
  • The tea genome size was estimated at 3.22 Gb with 27.2x coverage from genome survey sequencing finding
Experimental setups
Assay System Perturbation Readout Platform
small RNA sequencing (high-throughput sequencing) Camellia sinensis, one bud and two tender leaves none small RNA read counts / miRNA identification and abundance
whole genome shotgun genome survey sequencing Camellia sinensis genomic DNA (180 bp, 500 bp, 800 bp libraries) none k-mer/genome size estimate, contig/scaffold assembly
miRNA microarray hybridization Camellia sinensis, one bud and two tender leaves (small-MW RNA pool) none miRNA detection/expression signal (258 probes)
stem-loop qRT-PCR Camellia sinensis leaves/tissues (different leaf positions and tissues), three biological replicates none relative miRNA expression (2^△CT / 2^-△△CT), U6 snRNA control
qRT-PCR of target genes Camellia sinensis seven tissues (bud, 1st-3rd leaf, stem, root, flower) none target transcription factor expression, GAPDH control
5'RLM-RACE Camellia sinensis none mRNA cleavage site mapping for miRNA targets
catechin content measurement Camellia sinensis leaf tissues from different shoot positions, three biological samples none catechin content
target prediction and GO/KEGG analysis C. sinensis transcriptome sequence data none predicted target genes, functional/pathway classification Target Finder program
Key results
  • 175 conserved miRNAs in 39 families and 83 novel miRNAs identified 175 conserved; 83 novel
  • 93 conserved and 18 novel miRNAs (111 total) validated by microarray hybridization 111 miRNAs
  • csn-miR396 was the most highly expressed conserved family 59,922 reads (48% of conserved reads)
  • 716 potential target genes predicted for 187 miRNAs (116 conserved, 71 novel) 716 targets
  • 3 of 4 conserved miRNAs (miR160a-5p, miR164a, miR858a) negatively correlated with targets ARF17, NAC100, MYB12; miR828/WER partially positively correlated 3 of 4
  • Cleavage sites verified between 11th-12th base (ARF17, WER, MYB12) and 10th-11th base (NAC100) from 5' pairing
  • Catechin content gradually decreased from 1st to 5th leaf; csn-miRn23 negatively correlated with catechin
  • 24-nt small RNAs predominant in size distribution 24-nt: 43.0% total, 58.29% unique; 21-nt: 11.16% total, 5.76% unique
Key statistics
  • count 6,211,111 raw reads; 3,455,797 unique small RNA reads (55.64%) (small RNA library sequencing output)
  • count 175 conserved miRNAs (39 families); 83 novel miRNAs (identified miRNAs)
  • count 716 potential target genes for 187 miRNAs (target prediction)
  • other folding free energy −4.7 to −138.3 kcal/mol (average −57.08); MFEI 0.4 to 2.8 (novel miRNA precursor secondary structures)
  • other genome size 3.22 Gb; 27.2x coverage; 115.7 Gbp data; N50 668 bp (tea genome survey assembly)
  • count 59,922 reads (48%); miR166 25,026 (20.05%); miR159 11,904 (9.54%) (expression of top conserved miRNA families)
  • other enzyme activity 33.12%; response to stress 19.06%; phenylpropanoid biosynthesis 16% of targets (GO/KEGG functional annotation of targets)
  • pvalue P ≤ 0.05 (KEGG pathway enrichment (16 pathways))

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a descriptive small-RNA sequencing and discovery study that profiled miRNAs in tea plant from a single small RNA library by high-throughput sequencing, with confirmation by microarray hybridization, stem-loop qRT-PCR (2^△CT and 2^-△△CT methods), 5'RLM-RACE, and bioinformatic target prediction with GO/KEGG enrichment. Quantitative comparisons are largely reported as means ± SD of three biological replicates, with significance markers and Duncan's Multiple Range Test (DMRT) letters used for the catechin and miRNA expression panels. KEGG pathway enrichment used a P ≤ 0.05 threshold, and miRNA–target relationships were described qualitatively as correlations in expression patterns.

Replicationbiological Sample sizethree biological replicates (n = 3) stated for qRT-PCR and catechin assays; no power/sample-size calculation described Groupsleaf positions (1st-5th leaf) and tissues (bud, leaves, stem, root, flower) Pairingunclear Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated for the per-comparison significance tests; DMRT inherently provides multiple-comparison grouping, and KEGG used a P ≤ 0.05 / q-value display
Statistical tests used
Test Applied to n Assumptions
unspecified significance test for group comparison (denoted by * p<0.05, ** p<0.001) catechin content across leaf positions (Fig. 4) n = 3 biological samples not stated
Duncan's Multiple Range Test (DMRT), letter-based grouping relative expression of four novel csn-miRNAs across leaf positions (Fig. 5) three biological replicates not stated
KEGG pathway enrichment (significance threshold P ≤ 0.05; q-values shown) enrichment of predicted target genes (Fig. 7) na
correlation of expression patterns (described qualitatively as positive/negative correlation) miRNA vs target transcription factor expression across tissues (Fig. 10) and miRNA vs catechin content (Fig. 5) na
Approaches that could also have been used
  • Group differences in catechin content were marked with significance symbols (* p<0.05, ** p<0.001) without the specific test being named.
    Could also: Explicitly naming the test (e.g., one-way ANOVA followed by a post-hoc test, or a non-parametric Kruskal–Wallis with Dunn's test) and reporting it in the methods. — Stating the exact test and its assumptions helps readers match the analysis to the data type and reproduce the comparison.
  • Expression results were summarized as mean ± SD of three biological replicates.
    Could also: Also reporting a 95% confidence interval or showing individual data points alongside the mean. — With small n, displaying individual replicate values or a CI conveys the precision of the estimate and the spread directly, which many journals now encourage.
  • Multiple miRNA/tissue comparisons were each evaluated at a per-test threshold (p<0.05), and DMRT was used for the expression panels.
    Could also: Applying a family-wise (e.g., Tukey HSD, Bonferroni) or false-discovery-rate (Benjamini–Hochberg) correction across the full set of comparisons. — A correction across the family of comparisons controls the overall error rate when many groups or miRNAs are tested together.
  • miRNA–target and miRNA–catechin relationships were described qualitatively as positive/negative correlations across tissues.
    Could also: Computing a correlation coefficient (Pearson or Spearman) with its p-value, or a regression estimate. — A quantitative correlation statistic with an interval gives a reproducible measure of the strength and direction of the association beyond a verbal description.
  • miRNA identification and expression were validated using a single small RNA library plus microarray and qRT-PCR.
    Could also: Incorporating replicate sequencing libraries with a count-based differential-expression framework (e.g., DESeq2 or edgeR) for the high-throughput data. — Replicated libraries analyzed with a dedicated model would allow formal statistical inference on expression differences from the sequencing data itself, complementing the qRT-PCR validation.
  • qRT-PCR quantification used the 2^△CT and 2^-△△CT methods with a single reference gene (U6 / GAPDH).
    Could also: Normalizing to multiple validated reference genes (geometric mean, e.g., geNorm/NormFinder selection). — Using several stably-expressed reference genes can improve normalization robustness across tissues that differ physiologically.
Software: mfold (precursor secondary structure prediction) · Target Finder (target prediction) · miRBase (Release 21) reference database 21 · Repbase / Rfam (read filtering databases)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
95
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE68267 GEO in Acknowledgments (http://purl.org/orb/Acknowledgments)
no other assessed paper uses this yet
SRR1979118 ENA in Acknowledgments (http://purl.org/orb/Acknowledgments)
no other assessed paper uses this yet

Downstream reach in the literature

2 downstream papers · 2 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

SRR1979118 ENA reused by 3 papers in the literature
Most-cited downstream papers:
GSE68267 GEO reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-29157210

Paper: Jeyaraj A et al. (2017) Genome-wide identification of conserved and novel microRNAs in one bud and two tender leaves of tea plant (Camellia sinensis) by small RNA sequencing, microarray-based hybridization and genome survey scaffold sequences. BMC Plant Biol 17:212. PMID 29157210 / PMC5697157 / DOI 10.1186/s12870-017-1169-1.

Named code: TargetFinder (https://github.com/carringtonlab/TargetFinder) — a third-party plant miRNA target-prediction tool. Per brief rule P16, applying this existing tool to the paper's own data is a fully valid reproduction.

Data: GEO GSE68267 (small-RNA-seq reads); NCBI SRA SRR1979118 (C. sinensis transcriptome used as the TargetFinder target database). Supplementary tables: Additional file 1 (conserved miRNAs), file 3 (novel miRNAs), file 5 (predicted target transcripts).

Reported pipeline results (candidate claims)

# Reported value Location Pipeline In scope?
C1 175 conserved miRNAs in 39 families Results / Add. file 1 sRNA→miRBase r21 mapping (0–3 mm) partial — mapping tool not named
C2 83 novel miRNAs Results / Add. file 3 mfold v3.6 + MFEI≥0.85 criteria out (custom, under-specified)
C3 716 putative target genes (conserved+novel miRNAs) Results / Add. file 5 TargetFinder vs SRR1979118 transcriptome IN — primary
C4 111 miRNAs detected by microarray Results / Add. file 4 microarray hybridisation out (wet-lab)
C5 qRT-PCR / 5'RLM-RACE validation Results / Add. file 8 wet-lab out

Primary reproduction target

C3 — TargetFinder target prediction (the named tool). TargetFinder is deterministic given (miRNA sequences, target FASTA, score cutoff). Strategy:

  1. Extract the paper's mature miRNA sequences (Add. files 1+3) as TargetFinder query.
  2. Build the target database from SRR1979118 (the transcriptome the paper used).
  3. Run TargetFinder (default score cutoff 4) → count predicted targets, compare to 716.
  4. Determinism check: re-run TargetFinder on the paper's reported miRNA→target pairs (Add. file 5) and confirm the tool reproduces the reported alignments/ scores — an anti-fabrication check that does not depend on re-assembly noise.

The hard ~20% (declared, not chased)

  • Exact transcriptome re-assembly of SRR1979118 (Trinity version/params unstated) is non-deterministic, so the exact count 716 is not expected to match 1:1.
  • miRNA identification (C1/C2): the read-mapping/novel-prediction tools and exact parameters are under-specified; reproduced only as far as the supplementary sequences allow.

Out of scope (not attempted)

Microarray (C4), qRT-PCR/RACE (C5), GO/KEGG enrichment of targets (downstream of C3).

C1
Reported
175 conserved miRNAs
Reproduced
175 records with mature sequence (Additional file 1)
exact
C2
Reported
83 novel miRNAs
Reproduced
83 records with mature sequence (Additional file 3)
exact
C3
Reported
716 putative target genes
Reproduced
716 target rows in Additional file 5 (716 pairs; 579 distinct transcripts)
exact
C4
Reported
targets for 116 conserved + 71 novel miRNAs
Reproduced
116 conserved + 71 novel miRNAs have >=1 target
exact
C5
Reported
TargetFinder default cutoff (unstated) = 4
Reproduced
all 716 scores in [0,4], max = 4.0
within tolerance
C6
Reported
TargetFinder documented position-weighted scoring
Reproduced
independent install reproduces score 0 for a perfect target site, correct alignment
exact
C7
Reported
187/258 miRNAs with targets; 716 pairs / 579 targets (authors' assembly)
Reproduced
201/258 miRNAs with targets; 3799 pairs / 3023 targets (RefSeq CDS GCF_055761805.1, 89613 tx)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 91/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Largely a documentary substantiation: the headline counts (175 conserved + 83 novel miRNAs; 716 target rows = 579 distinct transcripts; 116+71 miRNAs-with-targets) match the shipped supplementary tables exactly and are internally consistent, the named TargetFinder tool reproduces its documented scoring deterministically, and the score ceiling (=4) matches its default cutoff. The actual generative pipeline is not independently reproducible because the authors' transcriptome assembly was never deposited (only raw reads public) — re-running TargetFinder on a public CDS yields a ~5x-different pair count (3799 vs 716), as expected for a different reference. No fabrication signal; the limitation is on the authors' data-deposit side, not a discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

202.2 k
tokens (I/O) · 12.3 M incl. cache
27 min
runtime · 1.14 CPU-h
9.4 GB
peak RAM
4
HPC jobs
hummel
machine