Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Enhanced microRNA accumulation and gene silencing efficiency through optimized precursor base pairing.

Plant J · 2026
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. This is a mostly wet-lab paper; the one clear pipeline-derived numeric result (Fig 5c processing accuracy = 90% of mature amiR-NbDXS from both shc and shc-A18G precursors) was reproduced END-TO-END from the paper's own raw SRA (SRR24210313 shc; SRR35183675 shc-A18G) using the documented pipeline (FASTX fastx_collapser + acarbonell/map_sRNA_reads exact-match mapping + the Methods' +/-4nt accuracy metric) on «our HPC»/«infra». Independent results: 89.89% and 89.68%, both rounding to the reported 90%; and the independently re-mapped mature-read counts (31410, 113922) match the authors' deposited Data S3 table EXACTLY -- strong no-fabrication evidence. Precursor refs were reconstructed from Data S3 and differ at exactly position 18 (A->G), validating the A18G claim. NOT attempted (80/20 + out of scope): P-SAMS amiRNA design for AtELF3 (Data S2; MySQL+BLAST web-backend) and all wet-lab quantities (Northern blots, RT-qPCR, chlorophyll, transgenic phenotypes). Brief corrections: code repo is carringtonlab/p-sams (brief's 'psams' 404s); new data accession is PRJNA1312446.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 87
    assessed: 2026-06-16 ⛓ 16b0e32c909e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Does base pairing at or near DCL1 cleavage sites within the minimal shc artificial miRNA (amiRNA) precursor affect amiRNA biogenesis and silencing efficiency, and can introducing a base pair upstream of the first cleavage site enhance amiRNA accumulation?

Core claims
  • Introducing a G–C pair immediately upstream of the mature amiRNA (A18G substitution) markedly enhances amiRNA accumulation and gene silencing efficiency in shc precursors. finding
  • Eliminating the mismatch at the DCL1 first cleavage site of AtMIR390a (A18G) increases miR390a and amiRNA accumulation and silencing of NbSu and NbDXS. finding
  • Base pairing at the DCL1 second cleavage site (positions 40/51) has no significant effect on amiRNA accumulation. finding
  • Base pairing at internal basal-stem positions 12/79 and 15/76 generally increases amiRNA accumulation. finding
  • A18G-modified shc precursors are accurately processed and predominantly release the intended authentic amiRNAs, as confirmed by deep sequencing. finding
  • A single structural modification in the amiRNA precursor provides an optimized, highly specific RNAi tool suited for functional genomics and crop engineering. resource
  • Silencing sensor systems in N. benthamiana targeting NbSu and NbDXS enable functional screening of amiRNA precursor variants via bleaching phenotypes. method
Experimental setups
Assay System Perturbation Readout Platform
sRNA/Northern blot Nicotiana benthamiana leaves (transient agroinfiltration) AtMIR390a A18G mutation vs wild-type; amiRNA precursor variants miR390a / amiRNA (amiR-NbSu, amiR-NbDXS) accumulation
Quantitative RT-PCR (RT-qPCR) Nicotiana benthamiana leaves (transient agroinfiltration) amiRNAs expressed from wild-type vs modified precursors NbSu and NbDXS target mRNA levels normalized to PP2A
Chlorophyll a quantification / phenotype imaging Nicotiana benthamiana leaves (transient agroinfiltration) amiR-NbSu / amiR-NbDXS vs GUS control relative chlorophyll a content / bleaching phenotype
sRNA/Northern blot of shc precursor variants Nicotiana benthamiana leaves (transient agroinfiltration) all base-pair combinations at positions 18/73, 40/51, 12/79, 15/76 and combined mutants amiR-NbSu and amiR-NbDXS accumulation
Phenotypic analysis in transgenic plants Arabidopsis thaliana transgenic lines amiRNAs against endogenous genes from A18G-modified vs wild-type shc precursors visible/quantifiable silencing phenotype
High-throughput / deep sequencing of small RNAs A18G-modified shc precursors (Arabidopsis/N. benthamiana) A18G modification precursor processing accuracy and amiRNA identity
Key results
  • miR390a accumulation increased from mismatch-corrected AtMIR390a-A18G precursor vs wild-type 26.5% increase
  • amiR-NbSu and amiR-NbDXS accumulation higher from AtMIR390a-A18G precursor 19% (NbSu) and 133% (NbDXS) increase
  • NbSu and NbDXS mRNA reduced more strongly with A18G-modified AtMIR390a precursor to 13.3% (NbSu) and 18.9% (NbDXS) vs 40.7%/44.9% for wild-type
  • shc A18G/C73 variant increased amiR-NbSu accumulation vs wild-type 42.3% increase
  • shc A18G, C73U, A18G/C73G variants increased amiR-NbDXS accumulation 137%, 126%, 85% higher respectively
  • shc A18G/C73 precursor reduced target mRNA levels NbSu to 7.7% and NbDXS to 16.2% (vs 12.1%/25.7% wild-type)
  • Base pairing at DCL1 second cleavage site (40/51) had no significant effect on amiRNA accumulation
  • Internal positions 12/79 and 15/76 base pairing increased amiRNA accumulation (e.g. A12G/C79U for NbSu; A12G/C79G for NbDXS) 54%/53% (NbSu); 102%/129% (NbDXS)
Key statistics
  • fold_change 26.5% increase (miR390a accumulation, AtMIR390a-A18G vs wild-type)
  • fold_change 133% increase (amiR-NbDXS accumulation from AtMIR390a-A18G precursor)
  • mean 13.3% and 18.9% (NbSu/NbDXS mRNA remaining with A18G AtMIR390a precursor vs amiR-GUS control)
  • fold_change 42.3% increase (amiR-NbSu from shc A18G/C73 variant)
  • fold_change 137%, 126%, 85% increases (amiR-NbDXS from shc A18G, C73U, A18G/C73G variants)
  • mean 7.7% and 16.2% (NbSu/NbDXS mRNA remaining with shc A18G/C73 precursor)
  • fold_change 54% and 53% increases (amiR-NbSu from shc A12G or C79U variants (P<0.05))
  • pvalue P < 0.05 (Student's t-test threshold for significance throughout)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used transient expression in Nicotiana benthamiana (n = 3 biological replicates) to test how base pairing modifications in amiRNA precursors affect miRNA accumulation (Northern/sRNA blot densitometry) and target mRNA levels (RT-qPCR). Multiple independent pairwise Student's t-tests were used to compare each precursor variant against a wild-type control at a P < 0.05 threshold. Results were reported as means with either SD or SEM and as percentage changes relative to a reference control; the provided text is truncated before results from Arabidopsis transgenic lines and deep-sequencing analyses are reached.

Replicationbiological Sample sizen = 3 biological replicates (plants) stated throughout; two leaves per plant agroinfiltrated per construct GroupsamiRNA precursor variants with single or combined mutations at basal stem positions (18/73, 12/79, 15/76, 40/51) vs wild-type shc or AtMIR390a precursor controls, across two amiRNA target sequences (amiR-NbSu, amiR-NbDXS) Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
pairwise Student's t-test miR390a accumulation from sRNA Northern blot densitometry (Figure 1b) n = 3 biological replicates not stated
pairwise Student's t-test relative chlorophyll a content across agroinfiltrated leaf sectors (Figure 1e) n not explicitly stated for this figure not stated
pairwise Student's t-test amiRNA accumulation from sRNA blot densitometry, multiple precursor variants vs wild-type (Figures 1f, 2c, 3b, 3c, 3d) n = 3 biological replicates not stated
pairwise Student's t-test RT-qPCR target mRNA levels (NbSu, NbDXS) for precursor variants vs control (Figures 1g, 2d) n = 3 biological replicates not stated
Approaches that could also have been used
  • Multiple pairwise Student's t-tests were used to compare each of many precursor variants independently against a single wild-type control within the same experiment
    Could also: One-way ANOVA followed by Dunnett's test (optimized for many-vs-one-control comparisons) or Tukey's HSD could also have been applied — An ANOVA framework with a post-hoc correction explicitly accounts for the inflation of Type I error that accumulates across many simultaneous pairwise comparisons, and Dunnett's test is specifically designed for the many-variants-vs-one-reference structure used here
  • Dispersion was reported as SD for blot-densitometry data and as SEM for RT-qPCR data within the same paper
    Could also: Consistent use of SD across all figures, or 95% confidence intervals throughout, could also have been applied — A uniform dispersion measure facilitates cross-figure comparison; 95% CIs additionally convey precision of the mean estimate, which is often considered informative for small-n data and is increasingly recommended by journals
  • Statistical significance was reported only as a binary threshold (P < 0.05) with no exact p-values given
    Could also: Reporting exact p-values (e.g., P = 0.018) alongside the threshold decision could also have been done — Exact p-values let readers assess the strength of evidence continuously rather than dichotomously, and they are required for inclusion in systematic reviews and meta-analyses
  • Parametric Student's t-tests were applied to blot densitometry data from n = 3 replicates per group
    Could also: A non-parametric alternative such as the Mann-Whitney U (Wilcoxon rank-sum) test could also have been used — With only three observations per group, the normality assumption of the t-test cannot be empirically verified; non-parametric rank-based tests make no distributional assumption and are a common alternative in small-n molecular biology experiments
  • Effect magnitudes were expressed as percentage changes relative to a control (e.g., '42.3% increase in amiRNA abundance')
    Could also: A standardized effect size metric such as Cohen's d could also have been reported alongside the percentage change — Standardized effect sizes are scale-independent and allow comparison of effect magnitudes across different assay types (blot vs. qPCR) and across studies, and they support prospective power calculations for follow-up experiments
  • amiRNA accumulation was quantified by Northern/sRNA blot densitometry as the primary continuous readout
    Could also: Small RNA sequencing read counts with a count-based statistical model (e.g., DESeq2 or edgeR negative binomial framework) could also have served as the primary quantitative comparison — Count-based models are well suited to the overdispersion characteristic of sequencing data and provide variance-stabilized estimates; the paper does describe deep-sequencing validation later, and integrating that as the primary quantitative endpoint is a standard alternative for sRNA studies
Software: not stated in provided text

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

246 715 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
plasmid_199560 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
plasmid_227963 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
plasmid_246716 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
plasmid_51778 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
PRJNA1312446 BioProject in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
PRJNA957136 BioProject in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
S69414 ENA in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41505763

Enhanced microRNA accumulation and gene silencing efficiency through optimized precursor base pairing. Llorens-Gámez et al., Plant J 2026. DOI 10.1111/tpj.70665.

This is a predominantly wet-lab paper (sensor systems in N. benthamiana, transgenic Arabidopsis, Northern blots, RT-qPCR, chlorophyll/phenotyping). Those results are OUT OF SCOPE (manual/experimental, not pipeline-derived).

Pipeline-derived results (IN SCOPE)

# result pipeline repro target
C1 Processing accuracy = 90% of amiR-NbDXS from BOTH shc and shc-A18G precursors (Fig 5c) sRNA-seq → FASTX collapse → acarbonell/map_sRNA_reads (exact match to precursor +strand) → proportion of 19-24nt(+) reads within ±4nt of the amiRNA 5' end that are the exact 21-nt mature re-run pipeline on raw SRA (SRR24210313 shc, SRR35183675 shc-A18G)
C2 21-nt size class dominates; reads confined to amiRNA/amiRNA* region; sharp 5' end at position 0 (Fig 5a,b) same pipeline size distribution + top read = mature at pos 19
C3 amiR-NbDXS mature read counts/RPM (Data S3 deposited mapping) same exact read counts vs Data S3
C4 P-SAMS optimal amiRNA designs for AtELF3 (Data S2); amiR-AtELF3 = TTCGCCTTGACCTGATCCCTT (Optimal Result 2) carringtonlab/p-sams (P-SAMS, Perl) on Araport11/TAIR10.1 secondary; cross-check the chosen guide

Out of scope (not attempted)

miRNA accumulation fold-changes from Northern blots, RT-qPCR target levels, chlorophyll/silencing phenotypes, Arabidopsis transgenic phenotyping — all wet-lab.

Notes / corrections

  • Brief's code URL github.com/carringtonlab/psams is 404; correct repo is carringtonlab/p-sams. The actual sRNA-mapping code is acarbonell/map_sRNA_reads (a self-contained exact-substring mapper; full source in its README).
  • Brief lists data PRJNA957136 (= Cisneros 2023 reference shc data); the new data for THIS paper is PRJNA1312446. Both used (one run from each).
  • Primary focus (80/20): C1 processing accuracy 90% — the single clearest pipeline-derived number, reproduced end-to-end from raw SRA.
Figures / tables: Fig 5cFig 5a
C1a
Reported
90% processing accuracy (amiR-NbDXS from shc precursor, Fig 5c)
Reproduced
89.89% (rounds to 90%)
within tolerance
C1b
Reported
90% processing accuracy (amiR-NbDXS from shc-A18G precursor, Fig 5c)
Reproduced
89.68% (rounds to 90%)
within tolerance
C2
Reported
Dominant 21-nt class; top read = authentic mature TAAACCGCGGGTTCCTAACAG at 5' pos 19 (Fig 5a,b)
Reproduced
21-nt dominant; top + read = mature @ pos19 in both samples
exact
C3a
Reported
31410 mature amiR-NbDXS reads in 35S-shc-NbDXS (Data S3)
Reproduced
31410
exact
C3b
Reported
113922 mature amiR-NbDXS reads in 35S-shc-A18G-NbDXS (Data S3)
Reproduced
113922
exact
C4
Reported
P-SAMS amiR-AtELF3 = TTCGCCTTGACCTGATCCCTT (Data S2)
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Strong, essentially 1:1 reproduction. The pipeline-derived result (Fig 5c, 90% processing accuracy of mature amiR-NbDXS from the shc and shc-A18G precursors) was re-derived end-to-end from the authors' own raw SRA and gave 89.89%/89.68% (rounds to 90%), while independently re-mapped mature-read counts matched the deposited Data S3 table exactly (31,410 and 113,922). The only deviation is sub-0.5% rounding — on our/technical side, negligible — and the exact count match is strong evidence against fabrication. The single un-attempted claim (C4, P-SAMS amiR-AtELF3 design via a MySQL+BLAST web backend) is secondary/optional and out of scope, so it does not weaken the core conclusion.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

170.4 k
tokens (I/O) · 13.5 M incl. cache
28 min
runtime · 0.04 CPU-h
1.3 GB
peak RAM
1
HPC jobs
hummel
machine