Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

DeepRNA-Reg: a deep-learning based approach for comparative analysis of CLIP experiments.

RNA Biol · 2025
L1 100/100 PQI 95
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the SOFTWARE 1:1, and partially the results. DeepRNA-Reg is the authors' own MIT tool (one script + pretrained DeepRNAreg.h5). EXACT 1:1 on the concrete checkable numbers: the shipped model is precisely the reported architecture (3 LSTM 150/150/50 + Dense1, 312,051 params, MAE loss, Adam) and the GSE273503 processed deposits are this tool's own output tables (44,517 / 40,820 predictions); the recovered WT-vs-KO count 44,517 equals the paper's reported Fig.2 F-test denominator df 44,517, an internal-consistency cross-check. The pipeline also runs end-to-end and emits the documented enriched_cond{1,2}.csv (needs >=2 BED loci; multiprocess.cpu_count() must be capped on HPC nodes). NOT ATTEMPTED (the hard ~20%): re-deriving the figure-level dCLIP comparisons (Fig 1-8), TargetScan-overlap counts and icSHAPE/viennaRNA panels — these need SRA re-alignment with unspecified CLIP params plus running dCLIP and external annotation/assay data not pinned in the repo; several headline phrases ('>80% more','~twice as many') are qualitative, not exact numbers. No fabrication signal. All heavy compute on «our HPC»; only small results on «host»; raw/intermediate data kept on «infra».

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-14 ⛓ 495843fbe2ec
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a deep-learning approach (DeepRNA-Reg) outperform the current best method (dCLIP) for comparative/differential analysis of paired HITS-CLIP experiments, yielding more sensitive, precise, and biologically valid predictions of differential RBP (Ago2/miRNA) binding?

Core claims
  • DeepRNA-Reg, a recurrent neural network-based algorithm, predicts differentially enriched sites in paired HITS-CLIP datasets and outperforms dCLIP 1.7. method
  • DeepRNA-Reg identifies a significantly larger set of differential target sites containing miRNA seed binding sequences (>80% more canonical seeds for miR-23/24/27 in Th2) than dCLIP, indicating greater sensitivity. finding
  • DeepRNA-Reg predictions localize smaller genomic regions with less size variance and greater centring of seed motifs, increasing positional precision relative to dCLIP. finding
  • DeepRNA-Reg's differential binding enrichment (DBE) score shows greater concordance with miRNA-induced mRNA expression changes and with biologically relevant sequence features than dCLIP's DBE. finding
  • DeepRNA-Reg confirmed known miR-24/miR-27 targets (Ikzf1, Gata3, Aff4, Gpr174, Cnot6, Clcn3) and identified CD28 as a novel direct target of miR-24/miR-27 in Th2 cells. mechanism
  • DeepRNA-Reg predictions show greater translatability across distinct biological milieux (Th2 miR-23,24,27 and Th17 miR-29 paradigms). finding
  • DeepRNA-Reg provides an orthologous deep-learning tool addressing the scarcity of computational methods for comparative HITS-CLIP analysis. resource
Experimental setups
Assay System Perturbation Readout Platform
AGO2 HITS-CLIP (AHC) In vitro differentiated mouse WT vs miR-23,24,27 KO (Mirc11/Mirc22 deletion) Th2 cells miRNA cluster gene knock-out Ago2 RNA binding occupancy / differential enrichment at 3'UTR sites
AGO2 HITS-CLIP (AHC) In vitro differentiated mouse WT vs miR-29ab1 KO (Mirc33 conditional mutant) Th17 cells miRNA gene knock-out Ago2 RNA binding occupancy / differential enrichment at 3'UTR sites
Gene expression profiling / RNA-sequencing Dgcr8 Δ/Δ Tbx21 -/- (miRNA-deficient) Th2 cells transfected with miR-23a, miR-24, miR-27a, or control mimic oligonucleotide miRNA mimic transfection mRNA fold change in expression upon miRNA mimic vs control
Flow cytometry WT and miR-23,24,27 KO mouse CD4+ T (Th2) cells miRNA cluster gene knock-out CD28 cell surface protein expression (mean fluorescent intensity, MFI)
Computational pathway analysis (Ingenuity Pathway Analysis / TargetScan) DeepRNA-Reg high-confidence prediction set (Th2) none Intersection with IPA IL-4 upstream regulators and TargetScan 7.2 predicted target sites Ingenuity Pathway Analysis; TargetScan 7.2
Key results
  • DeepRNA-Reg Th2 prediction set contained >80% more canonical (6-8 nt) miR-23/24/27 seed binding sequences than dCLIP >80% more
  • In Th2, DeepRNA-Reg overlapped 306 of 341 dCLIP-captured TargetScan sites and added 293 additional TargetScan predictions +293 sites (306/341 overlap)
  • In Th17, DeepRNA-Reg overlapped 65 of 78 dCLIP-captured TargetScan miR-29 targets and added 119 additional predictions +119 sites (65/78 overlap)
  • DeepRNA-Reg prediction region sizes showed significantly less variance than dCLIP in both paradigms miR-23,24,27: F=2.41; miR-29: F=2.56
  • DeepRNA-Reg offers almost twice as many predictions at top 10% (DBE percentile 90-100) confidence level with similar miRNA-induced expression reduction as dCLIP ~2-fold more predictions
  • More restrictive DBE percentile subsets of DeepRNA-Reg showed strong trend of greater seed motif enrichment (trend less clear for dCLIP)
  • miR-23,24,27 KO Th2 cells express more surface CD28 than WT, consistent with direct miR-24/27 targeting
  • DeepRNA-Reg detected differential Ago2 binding at conserved canonical miR-27 and miR-24 binding motifs in Cd28 3'UTR
Key statistics
  • other F(15219,44517) = 2.41, p < 2.2e-16 (size distribution variance, miR-23,24,27 Th2 paradigm) (DeepRNA-Reg vs dCLIP prediction region size variance, F-test)
  • other F(11446,47168) = 2.56, p < 2.2e-16 (size distribution variance, miR-29 Th17 paradigm) (DeepRNA-Reg vs dCLIP prediction region size variance, F-test)
  • count 306 of 341 dCLIP TargetScan sites overlapped; +293 added (Th2 WT vs miR-23,24,27 KO TargetScan overlap (Figure 1D))
  • count 65 of 78 dCLIP TargetScan miR-29 targets overlapped; +119 added (Th17 WT vs miR-29ab1 KO TargetScan overlap (Figure 1E))
  • fold_change >80% more canonical seed binding sequences (DeepRNA-Reg vs dCLIP, miR-23/24/27 seeds in Th2)
  • other ≥15% decrease in expression with miR-24 or miR-27 mimic vs control (high-confidence target filter) (Threshold for high-confidence miR-24/27 target gene calls)
  • count 3 independent experiments (Flow cytometric CD28 MFI measurement in WT vs miR-23,24,27 KO CD4+ T cells)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

DeepRNA-Reg, a recurrent-neural-network algorithm for differential HITS-CLIP analysis, is benchmarked against dCLIP 1.7 using AGO2 HITS-CLIP data from wildtype vs. miRNA-cluster-knockout mouse Th2 and Th17 cells. Performance was assessed primarily through counts-based metrics (miRNA seed-motif abundance, TargetScan 7.2 overlap, DBE-partitioned motif enrichment) and concordance with independent gene-expression data visualised as CDF plots. Formal inferential tests were confined to F-tests comparing the variance of predicted-region size distributions and Mann-Whitney U tests comparing seed-sequence centring; CD28 protein expression was validated by flow cytometry in biological triplicates.

Replicationbiological Sample sizeTwo independent experimental paradigms (Th2 and Th17 cell types); CD28 flow cytometry validation described as 'three independent' experiments (text truncated before completion); no formal power analysis stated GroupsWildtype Th2 cells vs. miR-23,24,27 cluster KO Th2 cells; wildtype Th17 cells vs. miR-29ab1 KO Th17 cells; algorithm comparison is DeepRNA-Reg vs. dCLIP 1.7 Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
F-test (variance ratio test) Comparison of size-distribution variance of differentially enriched sites called by DeepRNA-Reg vs. dCLIP, Th2 paradigm (WT vs. miR-23,24,27 KO) Degrees of freedom imply ~15,220 DeepRNA-Reg predictions and ~44,518 dCLIP predictions (F_{15219,44517} = 2.41, p < 2.2e-16) not stated
F-test (variance ratio test) Comparison of size-distribution variance of differentially enriched sites called by DeepRNA-Reg vs. dCLIP, Th17 paradigm (WT vs. miR-29ab1 KO) Degrees of freedom imply ~11,447 DeepRNA-Reg predictions and ~47,169 dCLIP predictions (F_{11446,47168} = 2.56, p < 2.2e-16) not stated
Mann-Whitney U test Centring of canonical and 8-mer miRNA seed binding sequences relative to the centre of predicted regions, Th2 (miR-23, miR-27) and Th17 (miR-29) paradigms not stated
Approaches that could also have been used
  • Variance of predicted site size distributions was compared with an F-test (variance ratio test), which assumes normally distributed populations in each group
    Could also: Levene's test or the Brown-Forsythe test could also compare group spread — These alternatives are more robust to departures from normality; genomic region sizes are typically right-skewed, so a normality-free variance comparison could complement the F-test result
  • Seed-sequence centring within predicted regions was compared using the Mann-Whitney U test, a rank-sum test sensitive primarily to location shift
    Could also: A two-sample Kolmogorov-Smirnov test or a permutation test on the mean/median positional distance could also compare the full distributional shape of seed positions — These approaches detect any shape difference (scale, skew, multimodality) rather than location shift alone, providing a more general test of whether the positional distributions differ between the two algorithms
  • Concordance between DBE scores and miRNA-induced expression fold-change was displayed as CDF plots stratified by DBE percentile, without a formal summary statistic
    Could also: A rank correlation (Spearman's rho) or area under the precision-recall curve between DBE rank and expression response could also quantify this relationship — A single correlation coefficient or AUC value would provide a compact, directly comparable effect-size measure across algorithms and miRNA conditions, supplementing the visual CDF approach
  • Algorithm performance was benchmarked using seed-motif counts and TargetScan overlap as the reference positive set, with predictions treated as a binary call
    Could also: A receiver-operating-characteristic (ROC) analysis with area under the curve (AUC), using validated or high-confidence TargetScan sites as the positive class, could also quantify discrimination across all score thresholds — ROC/AUC summarises the sensitivity-specificity trade-off across the full range of score thresholds in a single metric and is a standard approach for comparing the predictive performance of two classifiers
  • Multiple inferential tests and descriptive comparisons were conducted across two paradigms and many metrics without any stated correction for multiple comparisons
    Could also: A Benjamini-Hochberg FDR correction applied across the family of formal tests (the two F-tests and the Mann-Whitney comparisons) could also be reported — Explicit multiplicity control clarifies which individual comparisons remain significant after accounting for the number of tests performed and is standard practice when several hypothesis tests are reported together
  • The main performance comparisons are reported as raw counts or descriptive percentages (e.g., number of TargetScan sites captured, number of seed sequences) without confidence intervals or effect sizes
    Could also: Bootstrap confidence intervals around the difference in captured TargetScan sites or seed-motif counts could also be reported — Confidence intervals would convey the precision of the count-based comparisons and support inference about how reliably the observed differences might replicate in other datasets or conditions
Software: DeepRNA-Reg · dCLIP 1.7 · TargetScan 7.2 · Ingenuity Pathway Analysis (IPA)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE116348 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE116466 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE130655 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE273503 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE77105 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41055236 (DeepRNA-Reg)

Paper: Sekhon et al. 2025, RNA Biol. "DeepRNA-Reg: a deep-learning based approach for comparative analysis of CLIP experiments." PMCID PMC12505516. Code: https://github.com/AnselLab/DeepRNA-Reg (authors' own tool; MIT; single script DeepRNAreg.py + pretrained DeepRNAreg.h5). Data: GEO GSE273503 (Th2 / fresh-CD4 WT vs miR-23/24/27-KO; raw FASTQ in SRA PRJNA1142021; processed files GSE273503_WT_vs_miRNA_KO.txt.gz + _miRNA_KO_vs_WT.txt.gz).

What the tool does (pipeline)

Inputs: 2 aligned BAM files (--bam1/--bam2) + a BED of loci (--bed). Steps: per-base coverage (samtools view|depth) -> MA-normalisation (linregress on log coverage) -> Savitzky-Golay denoise (window 21, poly 3) -> pretrained LSTM inference -> sign-map to {-1,0,1} -> contiguous stretches >3 nt called as differentially enriched -> AUC (Simpson) -> percentile DBE score (1-10). Outputs: enriched_cond1.csv, enriched_cond2.csv.

In scope (pipeline-derived, attempted)

  • A. Model architecture & parameter count (Methods): "3 LSTM layers (150,150,50)
    • Dense(1), total 312,051 parameters, loss=MAE, optimizer=Adam". Directly checkable by loading the shipped DeepRNAreg.h5. PRIMARY clean data point / fabrication check.
  • B. GSE273503 processed outputs: characterise the deposited comparison tables; check whether they are DeepRNAreg/dCLIP prediction tables and whether prediction counts are consistent with reported claims (e.g. "almost twice as many"). Paper's own data.
  • C. End-to-end functional run: run DeepRNAreg.py on a controlled BAM+BED input to confirm the documented pipeline executes and emits enriched_cond{1,2}.csv with DBE scores. Reproducibility-of-software check.

Out of scope (not attempted; the hard ~20%) — why

  • Full re-derivation of the figure-level comparisons vs dCLIP (Fig 1-8): requires re-aligning SRA FASTQ to mm genome with unspecified CLIP params, rebuilding the 3'UTR BED, AND running dCLIP — alignment/peak parameters not specified runnably.
  • F-test / KS statistics, TargetScan-overlap counts, icSHAPE/viennaRNA panels: depend on the above intermediate products + external annotation sets not pinned in the repo. Qualitative claims (">80% more", "twice as many") are not exact numbers. Recorded as out-of-scope, not as mismatches.
A_total_params
Reported
312,051 model parameters
Reproduced
312,051 (loaded from shipped DeepRNAreg.h5)
exact
A_arch_lstm
Reported
3 LSTM layers (150,150,50) + Dense(1)
Reproduced
LSTM150 / LSTM150 / LSTM50 / Dense1 (91200+180600+40200+51)
exact
A_loss_opt
Reported
loss=mean absolute error; optimizer=Adam
Reproduced
mean_absolute_error; Adam
exact
B_geo_output_linkage
Reported
DeepRNAreg differential-binding tables deposited in GSE273503
Reproduced
GSE273503 supplementary files carry DeepRNAreg.py's exact output schema; WT_vs_KO=44,517 / KO_vs_WT=40,820 predictions
exact
B_ftest_df_crosscheck
Reported
Fig.2 Th2 F-test denominator df = 44,517
Reproduced
WT_vs_miRNA_KO deposited prediction count = 44,517
exact
C_pipeline_runs
Reported
DeepRNAreg.py emits enriched_cond1.csv + enriched_cond2.csv
Reproduced
ran to completion (exit 0), both CSVs with documented schema, differential signal recovered
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

This is a clean software/linkage reproduction of the authors' own MIT tool: every concrete number checks out 1:1 — 312,051 parameters, LSTM 150/150/50+Dense(1), MAE/Adam from the shipped DeepRNAreg.h5, and the GSE273503 deposits (44,517 / 40,820 rows) carry the tool's exact output schema, with the 44,517 count matching the paper's Fig.2 F-test denominator df as an internal-consistency cross-check. No fabrication signal and no factual deviation on anything tested. The limitation is scope, not error: the paper's actual comparative-analysis results (figure-level dCLIP, TargetScan/icSHAPE overlaps, qualitative '>80% more' claims) were not re-derived because the alignment params and external data are not runnably pinned — partly an authors'-side specification gap. Overall a solid but partial reproduction, hence yellow on core-claim and overall.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

153.5 k
tokens (I/O) · 10.9 M incl. cache
22 min
runtime · 1.16 CPU-h
15.7 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine