DeepRNA-Reg: a deep-learning based approach for comparative analysis of CLIP experiments.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the SOFTWARE 1:1, and partially the results. DeepRNA-Reg is the authors' own MIT tool (one script + pretrained DeepRNAreg.h5). EXACT 1:1 on the concrete checkable numbers: the shipped model is precisely the reported architecture (3 LSTM 150/150/50 + Dense1, 312,051 params, MAE loss, Adam) and the GSE273503 processed deposits are this tool's own output tables (44,517 / 40,820 predictions); the recovered WT-vs-KO count 44,517 equals the paper's reported Fig.2 F-test denominator df 44,517, an internal-consistency cross-check. The pipeline also runs end-to-end and emits the documented enriched_cond{1,2}.csv (needs >=2 BED loci; multiprocess.cpu_count() must be capped on HPC nodes). NOT ATTEMPTED (the hard ~20%): re-deriving the figure-level dCLIP comparisons (Fig 1-8), TargetScan-overlap counts and icSHAPE/viennaRNA panels — these need SRA re-alignment with unspecified CLIP params plus running dCLIP and external annotation/assay data not pinned in the repo; several headline phrases ('>80% more','~twice as many') are qualitative, not exact numbers. No fabrication signal. All heavy compute on «our HPC»; only small results on «host»; raw/intermediate data kept on «infra».
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-14 ⛓ 495843fbe2ec
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a deep-learning approach (DeepRNA-Reg) outperform the current best method (dCLIP) for comparative/differential analysis of paired HITS-CLIP experiments, yielding more sensitive, precise, and biologically valid predictions of differential RBP (Ago2/miRNA) binding?
- ★ DeepRNA-Reg, a recurrent neural network-based algorithm, predicts differentially enriched sites in paired HITS-CLIP datasets and outperforms dCLIP 1.7. method
- ★ DeepRNA-Reg identifies a significantly larger set of differential target sites containing miRNA seed binding sequences (>80% more canonical seeds for miR-23/24/27 in Th2) than dCLIP, indicating greater sensitivity. finding
- ★ DeepRNA-Reg predictions localize smaller genomic regions with less size variance and greater centring of seed motifs, increasing positional precision relative to dCLIP. finding
- ★ DeepRNA-Reg's differential binding enrichment (DBE) score shows greater concordance with miRNA-induced mRNA expression changes and with biologically relevant sequence features than dCLIP's DBE. finding
- ★ DeepRNA-Reg confirmed known miR-24/miR-27 targets (Ikzf1, Gata3, Aff4, Gpr174, Cnot6, Clcn3) and identified CD28 as a novel direct target of miR-24/miR-27 in Th2 cells. mechanism
- DeepRNA-Reg predictions show greater translatability across distinct biological milieux (Th2 miR-23,24,27 and Th17 miR-29 paradigms). finding
- DeepRNA-Reg provides an orthologous deep-learning tool addressing the scarcity of computational methods for comparative HITS-CLIP analysis. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| AGO2 HITS-CLIP (AHC) | In vitro differentiated mouse WT vs miR-23,24,27 KO (Mirc11/Mirc22 deletion) Th2 cells | miRNA cluster gene knock-out | Ago2 RNA binding occupancy / differential enrichment at 3'UTR sites | — |
| AGO2 HITS-CLIP (AHC) | In vitro differentiated mouse WT vs miR-29ab1 KO (Mirc33 conditional mutant) Th17 cells | miRNA gene knock-out | Ago2 RNA binding occupancy / differential enrichment at 3'UTR sites | — |
| Gene expression profiling / RNA-sequencing | Dgcr8 Δ/Δ Tbx21 -/- (miRNA-deficient) Th2 cells transfected with miR-23a, miR-24, miR-27a, or control mimic | oligonucleotide miRNA mimic transfection | mRNA fold change in expression upon miRNA mimic vs control | — |
| Flow cytometry | WT and miR-23,24,27 KO mouse CD4+ T (Th2) cells | miRNA cluster gene knock-out | CD28 cell surface protein expression (mean fluorescent intensity, MFI) | — |
| Computational pathway analysis (Ingenuity Pathway Analysis / TargetScan) | DeepRNA-Reg high-confidence prediction set (Th2) | none | Intersection with IPA IL-4 upstream regulators and TargetScan 7.2 predicted target sites | Ingenuity Pathway Analysis; TargetScan 7.2 |
- ▲ DeepRNA-Reg Th2 prediction set contained >80% more canonical (6-8 nt) miR-23/24/27 seed binding sequences than dCLIP >80% more
- ▲ In Th2, DeepRNA-Reg overlapped 306 of 341 dCLIP-captured TargetScan sites and added 293 additional TargetScan predictions +293 sites (306/341 overlap)
- ▲ In Th17, DeepRNA-Reg overlapped 65 of 78 dCLIP-captured TargetScan miR-29 targets and added 119 additional predictions +119 sites (65/78 overlap)
- ▼ DeepRNA-Reg prediction region sizes showed significantly less variance than dCLIP in both paradigms miR-23,24,27: F=2.41; miR-29: F=2.56
- ▲ DeepRNA-Reg offers almost twice as many predictions at top 10% (DBE percentile 90-100) confidence level with similar miRNA-induced expression reduction as dCLIP ~2-fold more predictions
- ▲ More restrictive DBE percentile subsets of DeepRNA-Reg showed strong trend of greater seed motif enrichment (trend less clear for dCLIP)
- ▲ miR-23,24,27 KO Th2 cells express more surface CD28 than WT, consistent with direct miR-24/27 targeting
- ▲ DeepRNA-Reg detected differential Ago2 binding at conserved canonical miR-27 and miR-24 binding motifs in Cd28 3'UTR
- other F(15219,44517) = 2.41, p < 2.2e-16 (size distribution variance, miR-23,24,27 Th2 paradigm) (DeepRNA-Reg vs dCLIP prediction region size variance, F-test)
- other F(11446,47168) = 2.56, p < 2.2e-16 (size distribution variance, miR-29 Th17 paradigm) (DeepRNA-Reg vs dCLIP prediction region size variance, F-test)
- count 306 of 341 dCLIP TargetScan sites overlapped; +293 added (Th2 WT vs miR-23,24,27 KO TargetScan overlap (Figure 1D))
- count 65 of 78 dCLIP TargetScan miR-29 targets overlapped; +119 added (Th17 WT vs miR-29ab1 KO TargetScan overlap (Figure 1E))
- fold_change >80% more canonical seed binding sequences (DeepRNA-Reg vs dCLIP, miR-23/24/27 seeds in Th2)
- other ≥15% decrease in expression with miR-24 or miR-27 mimic vs control (high-confidence target filter) (Threshold for high-confidence miR-24/27 target gene calls)
- count 3 independent experiments (Flow cytometric CD28 MFI measurement in WT vs miR-23,24,27 KO CD4+ T cells)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
DeepRNA-Reg, a recurrent-neural-network algorithm for differential HITS-CLIP analysis, is benchmarked against dCLIP 1.7 using AGO2 HITS-CLIP data from wildtype vs. miRNA-cluster-knockout mouse Th2 and Th17 cells. Performance was assessed primarily through counts-based metrics (miRNA seed-motif abundance, TargetScan 7.2 overlap, DBE-partitioned motif enrichment) and concordance with independent gene-expression data visualised as CDF plots. Formal inferential tests were confined to F-tests comparing the variance of predicted-region size distributions and Mann-Whitney U tests comparing seed-sequence centring; CD28 protein expression was validated by flow cytometry in biological triplicates.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| F-test (variance ratio test) | Comparison of size-distribution variance of differentially enriched sites called by DeepRNA-Reg vs. dCLIP, Th2 paradigm (WT vs. miR-23,24,27 KO) | Degrees of freedom imply ~15,220 DeepRNA-Reg predictions and ~44,518 dCLIP predictions (F_{15219,44517} = 2.41, p < 2.2e-16) | not stated |
| F-test (variance ratio test) | Comparison of size-distribution variance of differentially enriched sites called by DeepRNA-Reg vs. dCLIP, Th17 paradigm (WT vs. miR-29ab1 KO) | Degrees of freedom imply ~11,447 DeepRNA-Reg predictions and ~47,169 dCLIP predictions (F_{11446,47168} = 2.56, p < 2.2e-16) | not stated |
| Mann-Whitney U test | Centring of canonical and 8-mer miRNA seed binding sequences relative to the centre of predicted regions, Th2 (miR-23, miR-27) and Th17 (miR-29) paradigms | — | not stated |
-
Variance of predicted site size distributions was compared with an F-test (variance ratio test), which assumes normally distributed populations in each group↳ Could also: Levene's test or the Brown-Forsythe test could also compare group spread — These alternatives are more robust to departures from normality; genomic region sizes are typically right-skewed, so a normality-free variance comparison could complement the F-test result
-
Seed-sequence centring within predicted regions was compared using the Mann-Whitney U test, a rank-sum test sensitive primarily to location shift↳ Could also: A two-sample Kolmogorov-Smirnov test or a permutation test on the mean/median positional distance could also compare the full distributional shape of seed positions — These approaches detect any shape difference (scale, skew, multimodality) rather than location shift alone, providing a more general test of whether the positional distributions differ between the two algorithms
-
Concordance between DBE scores and miRNA-induced expression fold-change was displayed as CDF plots stratified by DBE percentile, without a formal summary statistic↳ Could also: A rank correlation (Spearman's rho) or area under the precision-recall curve between DBE rank and expression response could also quantify this relationship — A single correlation coefficient or AUC value would provide a compact, directly comparable effect-size measure across algorithms and miRNA conditions, supplementing the visual CDF approach
-
Algorithm performance was benchmarked using seed-motif counts and TargetScan overlap as the reference positive set, with predictions treated as a binary call↳ Could also: A receiver-operating-characteristic (ROC) analysis with area under the curve (AUC), using validated or high-confidence TargetScan sites as the positive class, could also quantify discrimination across all score thresholds — ROC/AUC summarises the sensitivity-specificity trade-off across the full range of score thresholds in a single metric and is a standard approach for comparing the predictive performance of two classifiers
-
Multiple inferential tests and descriptive comparisons were conducted across two paradigms and many metrics without any stated correction for multiple comparisons↳ Could also: A Benjamini-Hochberg FDR correction applied across the family of formal tests (the two F-tests and the Mann-Whitney comparisons) could also be reported — Explicit multiplicity control clarifies which individual comparisons remain significant after accounting for the number of tests performed and is standard practice when several hypothesis tests are reported together
-
The main performance comparisons are reported as raw counts or descriptive percentages (e.g., number of TargetScan sites captured, number of seed sequences) without confidence intervals or effect sizes↳ Could also: Bootstrap confidence intervals around the difference in captured TargetScan sites or seed-motif counts could also be reported — Confidence intervals would convey the precision of the count-based comparisons and support inference about how reliably the observed differences might replicate in other datasets or conditions
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41055236 (DeepRNA-Reg)
Paper: Sekhon et al. 2025, RNA Biol. "DeepRNA-Reg: a deep-learning based approach
for comparative analysis of CLIP experiments." PMCID PMC12505516.
Code: https://github.com/AnselLab/DeepRNA-Reg (authors' own tool; MIT; single
script DeepRNAreg.py + pretrained DeepRNAreg.h5).
Data: GEO GSE273503 (Th2 / fresh-CD4 WT vs miR-23/24/27-KO; raw FASTQ in SRA
PRJNA1142021; processed files GSE273503_WT_vs_miRNA_KO.txt.gz + _miRNA_KO_vs_WT.txt.gz).
What the tool does (pipeline)
Inputs: 2 aligned BAM files (--bam1/--bam2) + a BED of loci (--bed). Steps: per-base coverage (samtools view|depth) -> MA-normalisation (linregress on log coverage) -> Savitzky-Golay denoise (window 21, poly 3) -> pretrained LSTM inference -> sign-map to {-1,0,1} -> contiguous stretches >3 nt called as differentially enriched -> AUC (Simpson) -> percentile DBE score (1-10). Outputs: enriched_cond1.csv, enriched_cond2.csv.
In scope (pipeline-derived, attempted)
- A. Model architecture & parameter count (Methods): "3 LSTM layers (150,150,50)
- Dense(1), total 312,051 parameters, loss=MAE, optimizer=Adam". Directly checkable by loading the shipped DeepRNAreg.h5. PRIMARY clean data point / fabrication check.
- B. GSE273503 processed outputs: characterise the deposited comparison tables; check whether they are DeepRNAreg/dCLIP prediction tables and whether prediction counts are consistent with reported claims (e.g. "almost twice as many"). Paper's own data.
- C. End-to-end functional run: run DeepRNAreg.py on a controlled BAM+BED input to confirm the documented pipeline executes and emits enriched_cond{1,2}.csv with DBE scores. Reproducibility-of-software check.
Out of scope (not attempted; the hard ~20%) — why
- Full re-derivation of the figure-level comparisons vs dCLIP (Fig 1-8): requires re-aligning SRA FASTQ to mm genome with unspecified CLIP params, rebuilding the 3'UTR BED, AND running dCLIP — alignment/peak parameters not specified runnably.
- F-test / KS statistics, TargetScan-overlap counts, icSHAPE/viennaRNA panels: depend on the above intermediate products + external annotation sets not pinned in the repo. Qualitative claims (">80% more", "twice as many") are not exact numbers. Recorded as out-of-scope, not as mismatches.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean software/linkage reproduction of the authors' own MIT tool: every concrete number checks out 1:1 — 312,051 parameters, LSTM 150/150/50+Dense(1), MAE/Adam from the shipped DeepRNAreg.h5, and the GSE273503 deposits (44,517 / 40,820 rows) carry the tool's exact output schema, with the 44,517 count matching the paper's Fig.2 F-test denominator df as an internal-consistency cross-check. No fabrication signal and no factual deviation on anything tested. The limitation is scope, not error: the paper's actual comparative-analysis results (figure-level dCLIP, TargetScan/icSHAPE overlaps, qualitative '>80% more' claims) were not re-derived because the alignment params and external data are not runnably pinned — partly an authors'-side specification gap. Overall a solid but partial reproduction, hence yellow on core-claim and overall.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.