Genome-wide identification of conserved and novel microRNAs in one bud and two tender leaves of tea plant (Camellia sinensis) by small RNA sequencing, microarra
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🔴Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH + reproduced. The paper's pipeline-derived headline numbers are reproduced 1:1 from the shipped supplementary data: 175 conserved miRNAs, 83 novel miRNAs, 716 miRNA-target rows (=579 distinct targets), and the 116-conserved + 71-novel split of miRNAs-with-targets all match exactly; the TargetFinder score ceiling (max=4.0) is consistent with the named tool's default cutoff. The named third-party tool TargetFinder (commit 848b2dd) was independently installed and its scoring validated deterministically (perfect site -> score 0). An end-to-end re-run of TargetFinder on the paper's 258 miRNAs (default cutoff 4) against an obtainable annotated C. sinensis RefSeq CDS (GCF_055761805.1) is a PARTIAL match (different DB): a similar fraction of miRNAs find targets (201/258 vs 187/258) but ~5x more pairs (3799 vs 716), expected because the authors' actual transcriptome assembly (target IDs CL####.Contig##) was NEVER deposited - only raw reads SRR1979118/PRJNA281432 are public - so an exact 1:1 of the 716 list is impossible. NOT ATTEMPTED (declared 20%/out-of-scope): de-novo re-assembly of SRR1979118 (would not yield a 1:1 comparison anyway), miRNA discovery from raw reads (mapping tool/params under-specified), microarray, qRT-PCR/5'RLM-RACE, GO/KEGG enrichment. No evidence of fabrication: all reported pipeline numbers are exactly derivable from the shipped data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 91assessed: 2026-06-14 ⛓ 1e24355dbb18
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDue to the lack of a reference genome, miRNAs in tea plant (Camellia sinensis) are poorly characterized; this study aims to identify and validate conserved and novel miRNAs in one bud and two tender leaves using small RNA sequencing combined with genome survey scaffold sequences, and to predict and validate their target genes.
- ★ 175 conserved and 83 novel miRNAs were identified mainly in one bud and two tender leaves of tea plant via small RNA sequencing combined with genome survey data finding
- ★ Combining genome survey scaffold sequences with small RNA sequencing enables prediction of pre-miRNA secondary structures and miRNA discovery in a non-model plant lacking a reference genome method
- ★ 716 potential target genes were predicted for 187 miRNAs (116 conserved, 71 novel), enriched for transcription factors, stress response, and phenylpropanoid biosynthesis enzymes finding
- ★ A negative correlation was observed between expression of 3 of 4 conserved miRNAs and their target transcription factors, supporting post-transcriptional regulation mechanism
- ★ miRNA-guided cleavage of target mRNAs (ARF17, NAC100, WER, MYB12) was experimentally verified by 5'RLM-RACE finding
- ★ This work enriches the Theaceae miRNA database and provides a resource of potential miRNA regulators of secondary metabolism in C. sinensis resource
- The tea genome size was estimated at 3.22 Gb with 27.2x coverage from genome survey sequencing finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| small RNA sequencing (high-throughput sequencing) | Camellia sinensis, one bud and two tender leaves | none | small RNA read counts / miRNA identification and abundance | — |
| whole genome shotgun genome survey sequencing | Camellia sinensis genomic DNA (180 bp, 500 bp, 800 bp libraries) | none | k-mer/genome size estimate, contig/scaffold assembly | — |
| miRNA microarray hybridization | Camellia sinensis, one bud and two tender leaves (small-MW RNA pool) | none | miRNA detection/expression signal (258 probes) | — |
| stem-loop qRT-PCR | Camellia sinensis leaves/tissues (different leaf positions and tissues), three biological replicates | none | relative miRNA expression (2^△CT / 2^-△△CT), U6 snRNA control | — |
| qRT-PCR of target genes | Camellia sinensis seven tissues (bud, 1st-3rd leaf, stem, root, flower) | none | target transcription factor expression, GAPDH control | — |
| 5'RLM-RACE | Camellia sinensis | none | mRNA cleavage site mapping for miRNA targets | — |
| catechin content measurement | Camellia sinensis leaf tissues from different shoot positions, three biological samples | none | catechin content | — |
| target prediction and GO/KEGG analysis | C. sinensis transcriptome sequence data | none | predicted target genes, functional/pathway classification | Target Finder program |
- – 175 conserved miRNAs in 39 families and 83 novel miRNAs identified 175 conserved; 83 novel
- – 93 conserved and 18 novel miRNAs (111 total) validated by microarray hybridization 111 miRNAs
- ▲ csn-miR396 was the most highly expressed conserved family 59,922 reads (48% of conserved reads)
- – 716 potential target genes predicted for 187 miRNAs (116 conserved, 71 novel) 716 targets
- – 3 of 4 conserved miRNAs (miR160a-5p, miR164a, miR858a) negatively correlated with targets ARF17, NAC100, MYB12; miR828/WER partially positively correlated 3 of 4
- – Cleavage sites verified between 11th-12th base (ARF17, WER, MYB12) and 10th-11th base (NAC100) from 5' pairing
- ▼ Catechin content gradually decreased from 1st to 5th leaf; csn-miRn23 negatively correlated with catechin
- – 24-nt small RNAs predominant in size distribution 24-nt: 43.0% total, 58.29% unique; 21-nt: 11.16% total, 5.76% unique
- count 6,211,111 raw reads; 3,455,797 unique small RNA reads (55.64%) (small RNA library sequencing output)
- count 175 conserved miRNAs (39 families); 83 novel miRNAs (identified miRNAs)
- count 716 potential target genes for 187 miRNAs (target prediction)
- other folding free energy −4.7 to −138.3 kcal/mol (average −57.08); MFEI 0.4 to 2.8 (novel miRNA precursor secondary structures)
- other genome size 3.22 Gb; 27.2x coverage; 115.7 Gbp data; N50 668 bp (tea genome survey assembly)
- count 59,922 reads (48%); miR166 25,026 (20.05%); miR159 11,904 (9.54%) (expression of top conserved miRNA families)
- other enzyme activity 33.12%; response to stress 19.06%; phenylpropanoid biosynthesis 16% of targets (GO/KEGG functional annotation of targets)
- pvalue P ≤ 0.05 (KEGG pathway enrichment (16 pathways))
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a descriptive small-RNA sequencing and discovery study that profiled miRNAs in tea plant from a single small RNA library by high-throughput sequencing, with confirmation by microarray hybridization, stem-loop qRT-PCR (2^△CT and 2^-△△CT methods), 5'RLM-RACE, and bioinformatic target prediction with GO/KEGG enrichment. Quantitative comparisons are largely reported as means ± SD of three biological replicates, with significance markers and Duncan's Multiple Range Test (DMRT) letters used for the catechin and miRNA expression panels. KEGG pathway enrichment used a P ≤ 0.05 threshold, and miRNA–target relationships were described qualitatively as correlations in expression patterns.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| unspecified significance test for group comparison (denoted by * p<0.05, ** p<0.001) | catechin content across leaf positions (Fig. 4) | n = 3 biological samples | not stated |
| Duncan's Multiple Range Test (DMRT), letter-based grouping | relative expression of four novel csn-miRNAs across leaf positions (Fig. 5) | three biological replicates | not stated |
| KEGG pathway enrichment (significance threshold P ≤ 0.05; q-values shown) | enrichment of predicted target genes (Fig. 7) | — | na |
| correlation of expression patterns (described qualitatively as positive/negative correlation) | miRNA vs target transcription factor expression across tissues (Fig. 10) and miRNA vs catechin content (Fig. 5) | — | na |
-
Group differences in catechin content were marked with significance symbols (* p<0.05, ** p<0.001) without the specific test being named.↳ Could also: Explicitly naming the test (e.g., one-way ANOVA followed by a post-hoc test, or a non-parametric Kruskal–Wallis with Dunn's test) and reporting it in the methods. — Stating the exact test and its assumptions helps readers match the analysis to the data type and reproduce the comparison.
-
Expression results were summarized as mean ± SD of three biological replicates.↳ Could also: Also reporting a 95% confidence interval or showing individual data points alongside the mean. — With small n, displaying individual replicate values or a CI conveys the precision of the estimate and the spread directly, which many journals now encourage.
-
Multiple miRNA/tissue comparisons were each evaluated at a per-test threshold (p<0.05), and DMRT was used for the expression panels.↳ Could also: Applying a family-wise (e.g., Tukey HSD, Bonferroni) or false-discovery-rate (Benjamini–Hochberg) correction across the full set of comparisons. — A correction across the family of comparisons controls the overall error rate when many groups or miRNAs are tested together.
-
miRNA–target and miRNA–catechin relationships were described qualitatively as positive/negative correlations across tissues.↳ Could also: Computing a correlation coefficient (Pearson or Spearman) with its p-value, or a regression estimate. — A quantitative correlation statistic with an interval gives a reproducible measure of the strength and direction of the association beyond a verbal description.
-
miRNA identification and expression were validated using a single small RNA library plus microarray and qRT-PCR.↳ Could also: Incorporating replicate sequencing libraries with a count-based differential-expression framework (e.g., DESeq2 or edgeR) for the high-throughput data. — Replicated libraries analyzed with a dedicated model would allow formal statistical inference on expression differences from the sequencing data itself, complementing the qRT-PCR validation.
-
qRT-PCR quantification used the 2^△CT and 2^-△△CT methods with a single reference gene (U6 / GAPDH).↳ Could also: Normalizing to multiple validated reference genes (geometric mean, e.g., geNorm/NormFinder selection). — Using several stably-expressed reference genes can improve normalization robustness across tissues that differ physiologically.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
catechin content gradually decreases from 1st to 5th leaf and is negatively correlated with csn-miRn23metabolomics camellia sinensis leaf down 2017×1papers★ This paper is the founder (earliest)
-
111 miRNAs (93 conserved, 18 novel) validated by microarray hybridizationmicroarray camellia sinensis leaf 2017×1papers★ This paper is the founder (earliest)
-
miRNA-guided cleavage sites on targets ARF17, WER, MYB12, NAC100 verified by 5'RLM-RACEother camellia sinensis leaf 2017×1papers★ This paper is the founder (earliest)
-
conserved miRNAs (miR160a-5p, miR164a, miR858a) negatively correlate with targets ARF17, NAC100, MYB12, while miR828/WER is partially positively correlatedqPCR camellia sinensis leaf mixed 2017×1papers★ This paper is the founder (earliest)
-
24-nt small RNAs are the predominant size class in the tea plant small RNA populationRNA-seq camellia sinensis leaf 2017×1papers★ This paper is the founder (earliest)
-
csn-miR396 is the most highly expressed conserved miRNA family in tea plant bud and tender leavesRNA-seq camellia sinensis leaf up 2017×1papers★ This paper is the founder (earliest)
-
175 conserved miRNAs (39 families) and 83 novel miRNAs identified in tea plant bud and tender leaves by small RNA sequencingRNA-seq camellia sinensis leaf 2017×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
2 downstream papers · 2 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Genome-wide identification of microRNAs responsive t... 2017 · 39 cites
- Identification of Regulatory Networks of MicroRNAs a... 2019 · 37 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-29157210
Paper: Jeyaraj A et al. (2017) Genome-wide identification of conserved and novel microRNAs in one bud and two tender leaves of tea plant (Camellia sinensis) by small RNA sequencing, microarray-based hybridization and genome survey scaffold sequences. BMC Plant Biol 17:212. PMID 29157210 / PMC5697157 / DOI 10.1186/s12870-017-1169-1.
Named code: TargetFinder (https://github.com/carringtonlab/TargetFinder) — a third-party plant miRNA target-prediction tool. Per brief rule P16, applying this existing tool to the paper's own data is a fully valid reproduction.
Data: GEO GSE68267 (small-RNA-seq reads); NCBI SRA SRR1979118 (C. sinensis transcriptome used as the TargetFinder target database). Supplementary tables: Additional file 1 (conserved miRNAs), file 3 (novel miRNAs), file 5 (predicted target transcripts).
Reported pipeline results (candidate claims)
| # | Reported value | Location | Pipeline | In scope? |
|---|---|---|---|---|
| C1 | 175 conserved miRNAs in 39 families | Results / Add. file 1 | sRNA→miRBase r21 mapping (0–3 mm) | partial — mapping tool not named |
| C2 | 83 novel miRNAs | Results / Add. file 3 | mfold v3.6 + MFEI≥0.85 criteria | out (custom, under-specified) |
| C3 | 716 putative target genes (conserved+novel miRNAs) | Results / Add. file 5 | TargetFinder vs SRR1979118 transcriptome | IN — primary |
| C4 | 111 miRNAs detected by microarray | Results / Add. file 4 | microarray hybridisation | out (wet-lab) |
| C5 | qRT-PCR / 5'RLM-RACE validation | Results / Add. file 8 | wet-lab | out |
Primary reproduction target
C3 — TargetFinder target prediction (the named tool). TargetFinder is deterministic given (miRNA sequences, target FASTA, score cutoff). Strategy:
- Extract the paper's mature miRNA sequences (Add. files 1+3) as TargetFinder query.
- Build the target database from SRR1979118 (the transcriptome the paper used).
- Run TargetFinder (default score cutoff 4) → count predicted targets, compare to 716.
- Determinism check: re-run TargetFinder on the paper's reported miRNA→target pairs (Add. file 5) and confirm the tool reproduces the reported alignments/ scores — an anti-fabrication check that does not depend on re-assembly noise.
The hard ~20% (declared, not chased)
- Exact transcriptome re-assembly of SRR1979118 (Trinity version/params unstated) is non-deterministic, so the exact count 716 is not expected to match 1:1.
- miRNA identification (C1/C2): the read-mapping/novel-prediction tools and exact parameters are under-specified; reproduced only as far as the supplementary sequences allow.
Out of scope (not attempted)
Microarray (C4), qRT-PCR/RACE (C5), GO/KEGG enrichment of targets (downstream of C3).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Largely a documentary substantiation: the headline counts (175 conserved + 83 novel miRNAs; 716 target rows = 579 distinct transcripts; 116+71 miRNAs-with-targets) match the shipped supplementary tables exactly and are internally consistent, the named TargetFinder tool reproduces its documented scoring deterministically, and the score ceiling (=4) matches its default cutoff. The actual generative pipeline is not independently reproducible because the authors' transcriptome assembly was never deposited (only raw reads public) — re-running TargetFinder on a public CDS yields a ~5x-different pair count (3799 vs 716), as expected for a different reference. No fabrication signal; the limitation is on the authors' data-deposit side, not a discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.