Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comparative analysis of circular RNAs between soybean cytoplasmic male-sterile line NJCMS1A and its maintainer NJCMS1B by high-throughput sequencing.

BMC Genomics · 2018
50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (1:1 in shape, ~10-20% off in magnitude), independently re-run and confirmed deterministic (identical to an earlier completed run). The harvested code link kadenerlab/cORF_pipeline is a confirmed text-mining false positive (Drosophila circ-ORF tool); per brief P16 we reproduced with the paper's ACTUAL tool, find_circ (Memczak 2013), on the paper's own data (SRP160000, 6 RNase-R PE soybean libraries) via a self-contained «our HPC» SLURM job (bowtie2 2.3.5.1 + samtools 1.9 + find_circ HEAD, genome Glycine_max_v2.0 = GCF_000004515.4; DESeq2 1.34 / R 4.1; bedtools 2.30). All 6 reported numbers land same order of magnitude with the correct qualitative pattern: C1 total 2524/2867 (88%), NJCMS1A 1547/1643 (94%), NJCMS1B 1377/1722 (80%); per-sample circRNA A1=529 A2=645 A3=740 B1=802 B2=682 B3=162; C2 DE circRNAs 1220/1009 (121%) at |log2FC|>=2; C3 225up/995down vs 360up/649down (same down>up direction in A vs B); C4 host genes 949(contained)/545. No claim is exact, none is a hard mismatch -> honest grade PARTIAL. The systematic offset is explained by ONE documented deviation: the paper aligned with TopHat 2.0.9 (spliced) before find_circ, whereas we used plain bowtie2 -> a different unmapped read set feeds back-splice detection (TopHat 2.0.9 = python2 + samtools 0.1.x install-hell, not forced). Secondary: the B3 replicate genuinely yields few circRNAs (162; 93% genome alignment but low back-spliced-read content), depressing the NJCMS1B union and reversing the paper's B>A ordering. Fabrication check: every reported number IS derivable from the shipped data via the described tool; differences are tool-version drift, not unsupported values. Dataset SRP160000 profiled in the same pass (grade B: open, complete, 6/6 libraries align cleanly at 91-93%; line/replicate assignment only inferable from free-text SRA titles). NOT attempted (out of scope): wet-lab qRT-PCR validation, GO/KEGG enrichment, miRNA-sponge prediction. Run notes: required several iterations - transient ENA download blip (hardened with resume+gzip-check), a «infra» per-user inode/space quota (mitigated via shared pkg-cache + freeing big BAMs after each sample), and find_circ flakiness under parallelism (fixed by running samples sequentially; also fixed an f-string-under-python2 bug in helper scripts).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-20 ⛓ 39d174287794
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study investigates whether circular RNAs (circRNAs) differentially expressed between the soybean cytoplasmic male-sterile line NJCMS1A and its maintainer line NJCMS1B participate in the regulation of flower and pollen development, potentially contributing to the molecular mechanism of cytoplasmic male sterility (CMS) in soybean.

Core claims
  • 2867 circRNAs were identified in soybean flower buds via high-throughput sequencing with RNase R enrichment, of which 1009 were differentially expressed between NJCMS1A and NJCMS1B finding
  • Parental genes of differentially expressed circRNAs are mainly involved in metabolic process, biological regulation, and reproductive process finding
  • 83 miRNAs were predicted from differentially expressed circRNAs, some related to pollen development and male fertility, with target mRNAs enriched in signal transduction and programmed cell death finding
  • 165 soybean circRNAs contain at least one IRES element and an ORF, indicating potential to encode polypeptides or proteins finding
  • CircRNAs identified by RNA-seq were experimentally validated by RNase R resistance, divergent-primer PCR, Sanger sequencing, and qRT-PCR method
  • CircRNAs show tissue-specific expression patterns, being detected in leaves and flower buds but not roots or stems for two tested circRNAs finding
  • CircRNAs may regulate expression of their parental genes mechanism
  • CircRNAs act as miRNA sponges, preventing miRNAs from inhibiting target mRNAs mechanism
Experimental setups
Assay System Perturbation Readout Platform
high-throughput RNA sequencing (circRNA-seq) with RNase R enrichment flower buds of soybean CMS line NJCMS1A and maintainer NJCMS1B (3 biological replicates each) cytoplasmic genotype (CMS vs. maintainer) circRNA identification and differential expression (TPM, DESeq2) Illumina HiSeq 2500, 125bp paired-end
quantitative real-time PCR (qRT-PCR) flower buds of NJCMS1A and NJCMS1B cytoplasmic genotype (CMS vs. maintainer) relative circRNA expression (2^-ddCt) validating RNA-seq results for 12 circRNAs Bio-Rad CFX96
quantitative real-time PCR (qRT-PCR) roots, stems, leaves, flower buds of NJCMS1A and NJCMS1B tissue type tissue-specific expression of gma-circRNA0002 and gma-circRNA2483 Bio-Rad CFX96
PCR with divergent/convergent primers and Sanger sequencing total RNA and RNase R-treated RNA from flower buds RNase R treatment (presence/absence) confirmation of circular back-splice junctions
bioinformatic prediction (psRobot_tar, Cytoscape network) in silico analysis of differentially expressed circRNA sequences none predicted circRNA-miRNA binding interactions psRobot, Cytoscape
GO term enrichment (WEGO, agriGO) and KEGG pathway enrichment (KOBAS) parental genes of DE circRNAs and predicted miRNA target mRNAs none functional category and pathway enrichment WEGO, agriGO, KOBAS
IRES and ORF prediction in silico analysis of circRNA sequences none protein-coding potential (IRES element presence, ORF spanning junction, conserved domains) iresite.org BLAST, cORF_pipeline, NCBI CDD
Key results
  • Total circRNAs identified 2867 (1722 from NJCMS1B, 1643 from NJCMS1A)
  • Differentially expressed circRNAs between NJCMS1A and NJCMS1B 1009 total; 360 up-regulated, 649 down-regulated in NJCMS1A vs. NJCMS1B
  • qRT-PCR validation consistency with RNA-seq for 12 randomly selected circRNAs 10/12 consistent (83.3% coincidence rate)
  • miRNAs predicted from differentially expressed circRNAs 83 miRNAs
  • circRNAs with IRES element and ORF (protein-coding potential) 165 circRNAs
  • Tissue-specific expression: gma-circRNA0002 and gma-circRNA2483 detected only in leaves and flower buds, not roots/stems
  • Classification of circRNAs by genomic origin 452 (15.8%) exonic, 821 (28.6%) intronic, 293 (10.2%) exon-intron, 1301 (45.4%) intergenic
  • Homology comparison with prior soybean circRNA study (Zhao et al.) only 78 circRNAs homologous (e-value<1e-5, identity>85%)
Key statistics
  • count 2867 circRNAs identified (total circRNAs from NJCMS1A and NJCMS1B flower bud libraries)
  • count 1009 differentially expressed circRNAs (NJCMS1A vs NJCMS1B, |log2(fold-change)| >= 2)
  • fold_change |log2(fold-change)| >= 2 cutoff (threshold for defining differential expression)
  • count 360 up-regulated, 649 down-regulated (direction of differential expression in NJCMS1A vs NJCMS1B)
  • other 83.3% coincidence rate (10/12) (qRT-PCR vs RNA-seq validation consistency)
  • count 83 miRNAs predicted (predicted from differentially expressed circRNAs)
  • count 165 circRNAs with IRES+ORF (protein-coding potential prediction)
  • pvalue P < 0.05 significance threshold; P < 0.01 (**) for tissue-specific qRT-PCR (Student's t-test comparing NJCMS1A vs NJCMS1B expression)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study profiled circRNA expression in soybean CMS line NJCMS1A versus maintainer NJCMS1B using strand-specific high-throughput sequencing with three biological replicates per line and RNase R enrichment. CircRNA expression levels were normalized by TPM, and differential expression was assessed with DESeq2 using a fold-change threshold of |log2FC| ≥ 2. Validation of selected circRNAs was performed by qRT-PCR (2^-ΔΔCt method) with Student's t-test at P < 0.05. Downstream analyses included bioinformatic prediction of miRNA binding sites, GO/KEGG enrichment of parental and target genes, and IRES/ORF scanning for coding potential.

Replicationbiological Sample sizeThree biological replicates per line stated for both RNA-seq and qRT-PCR; no formal power analysis mentioned GroupsNJCMS1A (CMS line) vs NJCMS1B (maintainer/near-isogenic line) Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
DESeq2 (negative binomial Wald test) Differential expression of circRNAs between NJCMS1A and NJCMS1B across the full set of 2867 identified circRNAs 3 biological replicates per group not stated
Student's t-test (two-sample) qRT-PCR validation of 12 randomly selected differentially expressed circRNAs between NJCMS1A and NJCMS1B 3 biological replicates per group not stated
Student's t-test (two-sample) Tissue-specific expression comparisons of gma-circRNA0002 and gma-circRNA2483 across roots, stems, leaves, and flower buds between NJCMS1A and NJCMS1B 3 biological replicates per group not stated
Approaches that could also have been used
  • Differential expression was defined by |log2FC| ≥ 2 alone; no adjusted p-value threshold is explicitly stated for the DESeq2 output
    Could also: Apply a combined criterion of |log2FC| ≥ 1 (or ≥ 2) AND a BH-adjusted p-value threshold (e.g., padj < 0.05), which is the standard DESeq2 workflow — A fold-change filter alone does not control the false discovery rate among the thousands of circRNAs tested; a statistical threshold complements the biological-relevance filter and is routine in RNA-seq differential expression reporting
  • CircRNA expression was normalized to TPM before being supplied to DESeq2
    Could also: Supply raw integer read counts directly to DESeq2 and use its internal median-of-ratios size-factor normalization — DESeq2's statistical model is parameterized for raw counts; it applies its own normalization as part of the estimation procedure, and the tool's documentation recommends against pre-normalized values as input, since the distributional assumptions differ
  • Multiple Student's t-tests were performed across 12 circRNAs and multiple tissue-by-line comparisons without a stated correction
    Could also: Apply a Benjamini-Hochberg FDR correction or Bonferroni correction across the family of simultaneous comparisons — Running many independent t-tests inflates the experiment-wide type I error rate; a correction procedure would quantify and limit the expected proportion of false-positive findings in the validation panel
  • Dispersion around qRT-PCR means was summarized with SEM at n = 3 per group
    Could also: Report SD or 95% confidence intervals alongside or instead of SEM — SD describes the observed variability in the data directly and does not depend on n in the same way; 95% CIs convey estimation uncertainty and support direct inferential interpretation, both of which are often preferred for small-n experiments
  • With n = 3 biological replicates per group, a parametric Student's t-test was used for qRT-PCR comparisons
    Could also: Use a non-parametric alternative such as the Mann-Whitney U (Wilcoxon rank-sum) test — Normality of the 2^-ΔΔCt distribution cannot be reliably assessed with n = 3; non-parametric tests make no distributional assumption, though statistical power at n = 3 is inherently limited regardless of the test chosen
  • GO and KEGG enrichment analyses were performed on parental gene sets and predicted miRNA target sets
    Could also: Apply BH-FDR correction to enrichment p-values and report adjusted p-values alongside raw p-values — Enrichment analyses test many GO terms or pathways simultaneously; correcting for multiple comparisons is standard practice to limit false-positive enrichment findings, and adjusted p-values are expected by most journals
Software: DESeq2 1.6.3 · TopHat 2.0.9 · Bowtie 2.0.6 · find_circ · psRobot · Cytoscape

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-30208848

Paper: Chen et al. 2018, Comparative analysis of circular RNAs between soybean cytoplasmic male-sterile line NJCMS1A and its maintainer NJCMS1B by high-throughput sequencing. BMC Genomics. PMID 30208848 / PMC6134632 / DOI 10.1186/s12864-018-5054-6.

Code-link note (P16)

The harvested code link https://github.com/kadenerlab/cORF_pipeline is a text-mining false positive — that repo is the Kadener-lab circular-ORF translation pipeline (Drosophila), unrelated to this soybean comparative-circRNA study. The paper ships no authors' own repository. Per brief rule P16, we reproduce by running the third-party tool the paper actually used on the paper's own data. Methods name that tool explicitly: find_circ (Memczak et al. 2013), with Bowtie v2.0.6 + TopHat v2.0.9, soybean genome v2.0 (Phytozome), and DESeq2 v1.6.3 for differential expression.

Data

  • SRA SRP160000 (BioProject PRJNA489367), 8 runs. The 6 paired-end, RNase-R-enriched runs are the circRNA libraries (3×NJCMS1A, 3×NJCMS1B):
    • NJCMS1A: SRR7800885 (1A1), SRR7800886 (1A2), SRR7800887 (1A3)
    • NJCMS1B: SRR7800888 (1B1), SRR7800889 (1B2), SRR7800890 (1B3)
    • (A/B mapping from SRA experiment TITLE "sample 1A1…1B3" + alias "NJCMS1A1_…/NJCMS1B3_…clean.fq.gz"; DESIGN_DESCRIPTION="RNase R enrichment".)
    • SRR7800744 / SRR7800745 are single-end runs (other assay) → not used.
  • Reference: GCF_000004515.4_Glycine_max_v2.0 (Wm82.a2.v1 assembly = the same genome sequence as Phytozome "Gmax v2.0"; NCBI mirror, no login needed).

IN SCOPE (pipeline-derived, attempted)

id reported result paper location pipeline
C1 2867 circRNAs total (1722 NJCMS1B + 1643 NJCMS1A) Results / Abstract find_circ on 6 libs, union per line
C2 1009 differentially expressed circRNAs (|log2FC|≥2) Results DESeq2 v1.6.3 on circRNA TPM counts
C3 360 up / 649 down in NJCMS1A vs NJCMS1B Results DESeq2 direction
C4 545 parental genes from the 1009 DE circRNAs Results overlap circ junctions with gene annotation

80/20 priority: C1 (total + per-line counts) is the clearest, most directly pipeline-derived number → primary target. C2/C3 (DESeq2 DE) and C4 (parent genes) are downstream and depend on count-matrix construction details the paper under-specifies → attempted only if C1 lands cleanly; otherwise documented as not-attempted.

OUT OF SCOPE (not pipeline / not attempted)

  • qRT-PCR validation of selected circRNAs (wet-lab).
  • GO / KEGG enrichment of parent genes (depends on C4; downstream, optional).
  • miRNA-sponge / ceRNA target prediction discussion.
  • Any RNA extraction / library-prep / sequencing wet-lab steps.

Methods VERIFIED from PMC6134632 full text (2026-06-20)

  • circRNA filter = "back-spliced reads with at least two supporting reads" → our minimal (CIRCULAR + n_reads≥2) is the PRIMARY criterion to compare against; strict (adds UNAMBIGUOUS_BP/ANCHOR_UNIQUE) is reported as a sensitivity.
  • Alignment: paper used TopHat v2.0.9 (built on Bowtie2 v2.0.6) then find_circ for anchor extension / GU-AG breakpoint. We use the canonical find_circ front-end (bowtie2 unspliced → unmapped2anchors → bowtie2 → find_circ). Deviation noted: TopHat is a spliced aligner; its unmapped set differs from plain bowtie2's, which can change counts. TopHat 2.0.9 (2013) needs python2 + samtools 0.1.x + old bowtie2 = install-hell; not forced (brief: don't force). Documented as the main caveat.
  • Per-line counting: paper reports per-line totals (1722 B / 1643 A) from per-sample detection; intended comparison = union of the 3 replicates per line (our stage-1 aggregation). We ALSO try a pooled-per-line call as a sensitivity if union undershoots.
  • Reference: "soybean genome version 2.0 in Phytozome" = GCF_000004515.4 (Wm82.a2.v1). ✓

Known reproduction cavea

C1
Reported
2867 circRNAs total
Reproduced
2524 (88%)
partial
C1a
Reported
1722 circRNAs in NJCMS1B
Reproduced
1377 (80%)
partial
C1b
Reported
1643 circRNAs in NJCMS1A
Reproduced
1547 (94%)
partial
C2
Reported
1009 DE circRNAs (|log2FC|>=2)
Reproduced
1220 (121%)
partial
C3
Reported
360 up / 649 down (NJCMS1A vs NJCMS1B)
Reproduced
225 up / 995 down
partial
C4
Reported
545 parental genes
Reproduced
949 (contained)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

942.5 k
tokens (I/O) · 88.1 M incl. cache
689 min
runtime · 103.83 CPU-h
16.2 GB
peak RAM
6 (3 failed)
HPC jobs
hummel
machine