CIRCprimerXL: Convenient and High-Throughput PCR Primer Design for Circular RNA Quantification.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
CIRCprimerXL (authors' own Nextflow DSL1 tool, P16) reproduced on «our HPC» via the paper-era Singularity container oncornalab/primerxl_circ:v0.27 (NUPACK 4.0.0.23), commit c365a01, all compute on SLURM compute nodes. DETERMINISTIC EXAMPLE: C1 (3/3=100%, 0 fail), C2 (20 candidate pairs) and C3 (the 3 exact selected primer pairs, all fields) reproduce BIT-EXACT vs committed expected_output (confirmed on a clean compute node). FINDING (not fabrication): the shipped example known_exons_GRCh38_small.bed lacks the exons that expected_output references, so the example does NOT run out-of-the-box (errors); supplying the author-intended exons makes every output byte-identical -> isolates it as a stale shipped-fixture bug (identical at HEAD v25.1). HEADLINE 88.4% (Fig 3) is NOT bit-reproducible by construction: the 2000/15000-circRNA SW480 subsets are undeposited and the find_circ/CIRCexplorer2 detection + random RNG are unreported. We built the full GRCh38 reference (Ensembl 104 genome+exons, 101 cDNA, dbSnp153Common) and ran a same-distribution re-estimate on 1998 REAL human circRNAs (circBase hg19->hg38, seed 42) with default parameters: 79.7% (1592/1998) get a passing primer pair, with a failure-mode breakdown (primer3 GC/Tm/hairpin; off-target specificity 10.5% of pairs; amplicon secondary structure 0.3%) matching the paper. 79.7% vs 88.4% (~9pp lower) is a same-ballpark plausibility confirmation, expected to undershoot because circBase has no SW480 and is a more heterogeneous/older catalog than the authors' high-confidence set. NOT attempted: wet-lab PCR validation (Fig 4, bench) and Table 3 timing (hardware-specific). Verdict: described well enough to reproduce the deterministic core exactly and to confirm the headline rate in the right range; the central quantitative figure is not independently bit-checkable due to a data-deposition gap (no fabrication evidence).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 88assessed: 2026-06-22 ⛓ 036a665152a4
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors aim to show that CIRCprimerXL, a new high-throughput primer design pipeline, can overcome the limitations of existing circRNA primer design tools (lack of scalability, species restriction, and failure to account for assay specificity, secondary structure, and SNPs) and produce validated, back-splice-junction-specific RT-qPCR primers.
- ★ CIRCprimerXL is a high-throughput, user-friendly circRNA RT-qPCR primer design pipeline available as both a web tool and a standalone Nextflow/Docker pipeline resource
- ★ The pipeline flags common SNPs and secondary structures in the template prior to primer design, and filters candidate primer pairs based on predicted off-target specificity and amplicon secondary structure method
- ★ CIRCprimerXL achieves a primer design success rate of 88.4% (almost 90%) using default settings finding
- ★ The pipeline is scalable and can design primers for tens of thousands of circRNAs within a couple of hours on HPC infrastructure finding
- ★ Twenty empirically validated circRNA primer pairs designed by CIRCprimerXL all showed PCR efficiency between 90 and 110% finding
- ★ All 20 tested circRNAs remained stable (unchanged Cq) after RNase R treatment, while 4 linear control genes were degraded (increased Cq), confirming circRNA-specific amplification finding
- ★ CIRCprimerXL supports primer design for human, mouse, rat, zebrafish, Xenopus tropicalis, and C. elegans via the web tool, and for any species via the Nextflow pipeline with a user-supplied reference genome resource
- CIRCprimerXL is freely available on GitHub under an MIT license, restricted to non-commercial academic use due to dependency licensing resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RT-qPCR primer efficiency test | synthetic DNA template positive control (6-point 10-fold dilution series) | none (dilution series) | PCR efficiency (E-value, Cq) | — |
| RT-qPCR with RNase R treatment | SW480 colon carcinoma cells, total RNA (2 treatment replicates, 2 qPCR replicates) | RNase R treatment vs buffer control | Cq value stability of circRNAs vs linear control genes | — |
| circRNA detection from RNA sequencing | SW480 colon carcinoma cells, deeply sequenced total RNA (SRA: SRS11316475) | none | circRNA back-splice junction identification | find_circ, CIRCexplorer2 |
| in silico primer design success rate test | 2000 circRNAs (subset of SW480-derived circRNA dataset) | none | primer design success rate and failure reasons | Primer3, NUPACK, bowtie (CIRCprimerXL pipeline) |
| in silico scalability/run time test | 15,000 circRNAs | none | pipeline run time | HPC (40 GB memory, 16 CPUs) and standard laptop |
- – Primer pairs could be designed for 88.4% of 2000 circRNAs using default settings 88.4%
- – All 20 validated primer pairs showed PCR efficiency between 90 and 110% E-value 1.9-2.1
- – 20 circRNAs remained stable (similar Cq) after RNase R treatment while 4 linear control genes showed increased Cq (degradation)
- – Primer design for 15,000 circRNAs completed on HPC 25 h 58 min 52 s
- ▼ Primer design for 200 circRNAs was substantially faster on HPC than on a standard laptop 9 min 37 s (HPC) vs 1 h 2 min 58 s (laptop)
- other 88.4% (primer design success rate for 2000 circRNAs with default settings)
- count 2000 circRNAs (subset used for success rate evaluation)
- count 15,000 circRNAs (subset used for scalability/run time evaluation)
- other E-value 1.9-2.1 (90-110% efficiency) (PCR efficiency of 20 validated circRNA primer pairs)
- count 20 circRNAs (empirically validated by RT-qPCR)
- count 4 linear control genes (degraded upon RNase R treatment, contrasting with stable circRNAs)
- other 25 h 58 min 52 s (run time for designing primers for 15,000 circRNAs on HPC)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper describes a bioinformatics software tool (CIRCprimerXL) for circRNA-specific PCR primer design and reports its performance largely through descriptive metrics: a primer design success rate (percentage) across a subset of circRNAs, single-run timing benchmarks, and empirical wet-lab validation of 20 primer pairs via PCR efficiency (E-values from a dilution series) and RNase R stability (Cq value comparisons across a small number of replicates). No formal inferential statistical hypothesis tests, p-values, or multiplicity corrections are reported; results are presented as percentages, ranges, and qualitative comparisons.
-
The primer design success rate (88.4%) for the 2,000-circRNA subset is reported as a single point estimate.↳ Could also: A binomial or Wilson confidence interval around the proportion could also be reported — This would convey the precision of the estimate, which can be useful when comparing success rates across different parameter settings or circRNA subsets.
-
PCR efficiency across the 20 validated primer pairs was summarized as a range (E-values 1.9-2.1, i.e., 90-110% efficiency).↳ Could also: Reporting the mean and SD (or SEM) of efficiency values, or displaying the full distribution (e.g., a dot plot or boxplot), could also be used — This would show central tendency and variability across primer pairs in addition to the extremes captured by a range.
-
RNase R stability of the 20 circRNAs (two treatment replicates, two qPCR replicates) was assessed qualitatively by comparing Cq values between treated and control samples.↳ Could also: A paired comparison such as a paired t-test or Wilcoxon signed-rank test on ΔCq (treated vs. control) could also be applied — This would provide a formal inferential summary (e.g., effect size and p-value) of the stability difference, complementing the qualitative Cq comparison, particularly useful given the small replicate numbers.
-
Run times for different input sizes (2 to 15,000 circRNAs) are presented as single-run examples on each computing platform.↳ Could also: Repeating each run multiple times and reporting mean ± SD (or a range) of run time could also be used — This would characterize run-to-run variability in the scalability benchmark rather than relying on one observation per condition.
-
CircRNA subsets for success-rate and scalability testing were described as 'random subsets' drawn from a larger detected set.↳ Could also: Explicitly documenting the randomization procedure (e.g., seed, sampling method) could also be reported — This would support reproducibility of the reported subsampling and any downstream comparisons based on it.
-
No multiple-comparison correction is discussed, consistent with no formal multi-group hypothesis testing being performed in this study.↳ Could also: If future comparisons across many species, parameter settings, or circRNAs were tested statistically, a false discovery rate method (e.g., Benjamini-Hochberg) could also be applied — This would control the family-wise error rate in any expanded analysis involving numerous simultaneous comparisons.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 36304334 (CIRCprimerXL)
Paper: Marchal et al. (2022) "CIRCprimerXL: Convenient and High-Throughput PCR Primer Design for Circular RNA Quantification." Front. Bioinform. 2:834655. DOI 10.3389/fbinf.2022.834655 · PMCID PMC9580850.
Code: https://github.com/OncoRNALab/CIRCprimerXL (Nextflow DSL2 pipeline;
P16 = authors' own tool). Container oncornalab/primerxl_circ (paper era: v0.11,
NUPACK 4.0.0.23 / primer3 2.5.0 / bowtie 1.3.0 / fastahack 1.0; current main:
v25.1, NUPACK 4.0.1.12 / bowtie 1.3.1).
Validation data: SRA SRS11316475 = SW480 colon-carcinoma total RNA-seq
(BioSample SAMN24046636; deep RNA-seq run SRR17235468, 295M paired reads, ~86 Gbp).
circRNAs detected with find_circ + CIRCexplorer2 (versions/params NOT reported).
Pipeline (what CIRCprimerXL does, per result)
BED of circRNA back-splice junctions (BSJ) → fastahack pulls BSJ-flanking sequence →
SNP masking → primer3 (settings in assets/primer3plus_settings.txt) proposes primer
pairs spanning the BSJ → bowtie specificity filter → NUPACK secondary-structure filter
on primer + amplicon → 10_filter.py selects one passing pair per circRNA →
filtered_primers.txt + summary_run.txt.
In scope (pipeline-derived — attempted)
| # | Reported result | Location | Pipeline | Repro plan |
|---|---|---|---|---|
| C1 | Example run: 3 circRNAs, 100% get a passing primer pair, 20 primer pairs generated, 0 fail (spec/SNP/sec-str) | repo example/expected_output/ (committed 2022-01-31) |
CIRCprimerXL -profile example (self-contained subset indexes) |
DETERMINISTIC. Run example profile via Singularity on «our HPC», diff vs shipped expected_output. 80% floor. |
| C2 | Exact primer pairs/Tm/GC/amplicon for circ0/circ1/circ2 | example/expected_output/filtered_primers.txt |
same | bit/value-level compare of the 3 selected pairs |
| C3 | Success rate 88.4% of circRNAs get a primer pair (random subset of 2000 SW480 circRNAs) | Fig 3, §3.2 | full pipeline on 2000-circRNA BED | HARD / statistical. Exact 2000 subset NOT deposited; detection method undocumented. Attempt: detect SW480 circRNAs from SRS11316475, random 2000, run pipeline, check rate ≈ 88.4% (within-tol, not exact). |
| C4 | Runtime: 2000 circRNAs ≈ 2 h 47 m; 15000 ≈ 25 h 59 m (16 CPU) | Table 3, §3.3 | full pipeline timing | WEAK — hardware/cluster-specific; record as context only, not a 1:1 numeric match. |
Out of scope (not attempted — not pipeline-derived)
- 20-circRNA wet-lab PCR validation (efficiency 90–110%, RNase R stability) — Fig 4, §3.4–3.5. Bench RT-qPCR, not computational.
- Web-tool (circprimerxl.cmgg.be) UX claims.
Blockers / caveats (honesty)
- 2000- and 15000-circRNA subsets are not deposited ("further inquiries to corresponding author"); no supplementary BED. → C3 cannot be reproduced exactly; only a same-distribution re-estimate of the success rate.
- find_circ/CIRCexplorer2 versions + the random-subset RNG are unspecified → C3 input is not bit-reproducible by construction.
- NUPACK is license-gated; only obtainable inside the authors' container → must run via the published Docker/Singularity image, not a hand-built conda env.
- «our HPC» has no Docker; will use Singularity/Apptainer to run
docker://oncornalab/....
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The deterministic bundled example reproduces bit-exactly under the paper-era container (C1 3/3=100%, C2 20 pairs, C3 all three selected primer pairs byte-identical) once the author-intended exon bed is supplied. The only out-of-box discrepancy is a shipped-fixture inconsistency on the authors' side — the committed known_exons_GRCh38_small.bed lacks the chr1:16606-18061 exons that the committed expected_output references — so the deviation is an input/preprocessing artefact, not an algorithm failure, and is not fabrication. The headline 88.4% (Fig3) is not bit-reproducible because the 2000-circRNA subset, detection-tool versions and RNG were never deposited/documented. Overall: solid reproduction with explainable, authors'-side deviations.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.