Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Differential analysis of RNA structure probing experiments at nucleotide resolution: uncovering regulatory functions of RNA structure.

Nat Commun · 2022
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH + reproduced 1:1 on self-contained data. Registry mislabeled the code (paper is DiffScan, not diffBUM-HMM); corrected to github.com/yub18/DiffScan. The core DiffScan pipeline (init->normalize->scan) reproduces EXACTLY: the deterministic README example (t1=0 SVRs; t2 SVRs 132-150 Q=79.0 & 92-100 Q=54.9, identical to README) and the headline Flu benchmark number (top-20 = 0.600 = paper's 60% for DiffScan). Specificity claim reproduced illustratively on a constructed identical-condition null (FPR ~0.05, 0-5 false nt /99). NOT attempted: the simulation Jaccard (Fig 2) and icSHAPE application (Fig 4) - both gated on inputs absent from the public deposit (specific simulation sequences; replicate-level reactivity + a downregulation gene list), not on any method failure. GSE117840 + Flu.rda fully profiled. Verdict: strong partial - the method itself is faithfully and exactly reproducible; two figure-specific numbers need non-public supplementary inputs.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-20 ⛓ 6b2af7443a32
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Differential analysis of RNA structure probing (SP) data across conditions can reveal structurally variable regions (SVRs) at nucleotide resolution, and a normalization + scan-statistic framework (DiffScan) can detect these SVRs more accurately than existing methods across multiple SP platforms.

Core claims
  • DiffScan is a computational framework combining a Normalization module and a Scan module to identify SVRs at nucleotide resolution from SP data. method
  • DiffScan achieves superior or comparable power and accuracy in SVR detection compared to state-of-the-art methods (deltaSHAPE, dStruct, RASA, PARCEL) across simulated and benchmark datasets. finding
  • DiffScan is compatible with multiple SP platforms, including icSHAPE, DMS-seq, and SHAPE-Seq. finding
  • The Scan module uses a scan statistic Q(R) with Monte Carlo evaluation to identify SVRs of adaptively determined length while controlling family-wise error rate. method
  • Application of DiffScan to a multi-subcellular icSHAPE RNA structurome dataset suggests links between RNA structural variation and mRNA abundance, possibly mediated by RNA binding proteins such as serine/arginine rich splicing factors. finding
  • The Normalization module removes systematic bias (e.g., sequencing depth, signal-to-noise ratio differences) between conditions using a pivot-set-based linear transformation approach similar to MAnorm. method
  • False positive rates of DiffScan, dStruct, and PARCEL were below 0.05 across negative control datasets, while deltaSHAPE and RASA showed inflated false positive rates. finding
Experimental setups
Assay System Perturbation Readout Platform
Simulated SP reactivity data (SHAPE and icSHAPE empirical models) 1000 RNA transcripts selected from human HEK293 cells simulated mixtures of multiple RNA secondary structure conformations at varying proportions/signal strengths SVR detection accuracy (Jaccard index, average distance, precision/recall, specificity) three empirical reactivity models: two SHAPE-type and one icSHAPE-type
Negative control SP datasets (Control 1-6) multiple SP platforms including SHAPE-Seq and icSHAPE none (no true SVRs present between compared conditions) false positive rate of SVR detection SHAPE-Seq, icSHAPE
Benchmark SHAPE-Seq structure probing (Flu dataset, from Choudhary et al.) Bacillus cereus crcB fluoride riboswitch not specified in extracted text detection accuracy against explicitly annotated SVRs SHAPE-Seq
icSHAPE structure probing multi-subcellular compartments (e.g., nucleoplasm and cytoplasm) of human cells comparison across subcellular compartments SVRs and downstream motif enrichment analysis linked to mRNA abundance and RNA binding protein motifs (e.g., SR splicing factors) icSHAPE
Key results
  • DiffScan consistently had the largest Jaccard index between top predicted nucleotide positions and true simulated SVRs across all 9 simulation settings.
  • DiffScan consistently had the minimum average distance between predicted and true SVRs across all 9 simulation settings.
  • DiffScan achieved the best precision at the same recall rate among compared methods.
  • False positive rates for DiffScan, dStruct, and PARCEL were below 0.05 across all tested negative control datasets. <0.05
  • deltaSHAPE showed consistent inflation of false positives, covering approximately 15% of nucleotide positions across datasets. ~15%
  • RASA's error rate was slightly inflated in Control 2 but well-controlled in other negative control datasets.
  • In the Choudhary et al. simulation framework, DiffScan and dStruct outperformed deltaSHAPE, RASA, and PARCEL in Jaccard index and precision at fixed recall.
  • Simulated datasets contained 38,317 SVRs total: 12,201 single-nucleotide, 13,048 with lengths 2-5 nt, and 13,068 with lengths greater than 5 nt.
Key statistics
  • other false positive rate < 0.05 (DiffScan, dStruct, and PARCEL across negative control datasets)
  • other ~15% of nucleotide positions falsely reported (deltaSHAPE false positive rate across negative control datasets)
  • count 38,317 simulated SVRs (total simulated SVRs across 1000 sampled RNA transcripts)
  • count 12,201 single-nucleotide SVRs; 13,048 SVRs of 2-5 nt; 13,068 SVRs >5 nt (length distribution of simulated SVRs)
  • other transcript lengths 60 nt to 1,972 nt (range of simulated RNA transcript lengths)
  • other SVR lengths between 1 nt and 81 nt (range of simulated SVR lengths)
  • other simulated SVRs covered 1.6% to 67.3% of nucleotide positions (coverage of simulated SVRs within transcripts)
  • count 10 conformations per transcript (number of sampled structural conformations per RNA transcript in simulation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper presents DiffScan, a computational/statistical framework for differential analysis of RNA structure probing (SP) data, rather than a conventional wet-lab experiment with grouped biological samples. Reactivities from two conditions are first normalized (winsorization, quantile normalization across replicates, and an MAnorm-like linear-model normalization between conditions), then a two-sided Wilcoxon test is applied at each nucleotide position to quantify differential signal, and a scan statistic aggregates positional p-values across candidate regions of varying length. Significance of the scan statistic is assessed via Monte Carlo sampling, which is used to control the family-wise error rate across the many overlapping candidate regions. Performance was evaluated against other SVR-detection methods using simulated and benchmark datasets with metrics such as Jaccard index, average distance, precision-recall curves, and false-positive rate.

Replicationmixed Sample sizeSample sizes are described in terms of simulated data (e.g., 1000 randomly selected RNA transcripts, 10 sampled conformations each, 38,317 simulated SVRs) rather than as a power calculation for biological replicates; use of within-condition replicates is described as optional ("if available") Groupstwo experimental/cellular conditions per transcript (e.g., nucleoplasm vs. cytoplasm, or simulated condition A vs. B) Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesyes Multiplicity correctionMonte Carlo sampling-based evaluation of the scan statistic, described as controlling the family-wise error rate across overlapping candidate regions
Statistical tests used
Test Applied to n Assumptions
Two-sided Wilcoxon (rank-sum) test per-nucleotide-position comparison of reactivities between two conditions within a small surrounding window not explicitly stated (based on replicate reactivities within a local window) not stated (described as nonparametric, chosen to avoid assuming a specific reactivity distribution)
Scan statistic Q(R) with Monte Carlo significance evaluation region-level aggregation of positional p-values to identify structurally variable regions (SVRs) across enumerated overlapping windows null not stated
Jaccard index, average distance, precision-recall curves benchmarking DiffScan against deltaSHAPE, dStruct, RASA, and PARCEL on simulated and curated benchmark datasets 1000 simulated RNA transcripts (9 simulation settings); two curated benchmark datasets (e.g., Flu dataset) na
False positive rate assessment on negative control datasets six negative control SP datasets (Control 1-6) with no true SVRs, across SHAPE-Seq and icSHAPE platforms six negative control datasets as stated not stated
Approaches that could also have been used
  • Positional differential signal is quantified with a two-sided Wilcoxon (nonparametric rank-based) test at each nucleotide position.
    Could also: A parametric test (e.g., Welch's t-test) on the reactivity values — If reactivities within the local window were approximately normally distributed, a parametric test could offer somewhat greater power; the nonparametric choice made here avoids needing to verify that assumption, which can be advantageous when reactivity distributions vary across platforms.
  • Multiple testing across overlapping candidate regions is addressed via Monte Carlo-based control of the family-wise error rate (FWER).
    Could also: A false discovery rate (FDR) procedure, such as Benjamini-Hochberg, applied to the regional p-values — FDR control is often preferred over FWER control in large-scale scans where many regions are tested, since it can increase detection power by tolerating a controlled proportion of false positives rather than bounding the probability of any false positive.
  • Between-condition normalization uses an MAnorm-like linear model fit on a pivot set of structurally invariant positions.
    Could also: Size-factor or trimmed-mean normalization methods from RNA-seq differential expression tools (e.g., DESeq2's median-of-ratios or edgeR's TMM) — These alternative normalization schemes, also inspired by high-throughput sequencing, offer another established way to correct for between-condition scaling differences and could be compared as a sensitivity check on the choice of pivot-set-based normalization.
  • SVR detection performance is compared using Jaccard index, average distance, and precision-recall curves.
    Could also: ROC curves with area-under-the-curve (AUC) summaries — ROC-AUC is a complementary standard metric for classifier/detector comparison and could provide an additional single-number summary alongside the precision-recall-based evaluation already used.
  • A Bayesian posterior-probability method (diffBUM-HMM) was excluded from the false-positive-rate comparison because Bayesian approaches do not precisely control type I error.
    Could also: Reporting Bayesian posterior probabilities or Bayes factors alongside frequentist p-values, rather than excluding the method from that specific comparison — Bayesian and frequentist error metrics answer different questions (posterior probability of an effect vs. long-run false-positive rate); presenting both frameworks side by side, where feasible, can let readers see how each quantifies uncertainty for the same regions.
  • SVR length is determined adaptively by the scan statistic rather than fixed in advance.
    Could also: A fixed or user-specified window length, as used by comparator methods deltaSHAPE and dStruct — A fixed window can simplify interpretation and computation when there is strong prior knowledge of typical SVR length, whereas the adaptive approach used here trades that simplicity for flexibility across variable SVR lengths.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 35869080 (DiffScan)

⚠ Metadata correction (registry error)

The scaffold/manifest pinned code_url = github.com/marangiop/diff_BUM_HMM and the enrich attributed odd authors. That is wrong: marangiop/diff_BUM_HMM is the diffBUM-HMM method (Marangio et al., Genome Biology 2021, 10.1186/s13059-021-02379-y) — a competing tool, not this paper.

PMID 35869080 is the DiffScan paper:

  • Title: "Differential analysis of RNA structure probing experiments at nucleotide resolution: uncovering regulatory functions of RNA structure"
  • Authors: Bo Yu, Pan Li, Qiangfeng Cliff Zhang, Lin Hou
  • Nature Communications 13:4227 (2022), DOI 10.1038/s41467-022-31875-3
  • Code (correct): https://github.com/yub18/DiffScan (R package)
  • Data accession GSE117840 is correct — it is DiffScan's icSHAPE application data.

This room reproduces DiffScan, the actual subject of the paper.

What DiffScan is

An R package for differential analysis of RNA structure-probing (reactivity) data. Pipeline: init() (QC + segment) → normalize() (quantile/bias removal) → scan() (Monte-Carlo null + Wilcoxon scan statistic Q, FWER-controlled at alpha) → SVRs (structurally variable regions) at nucleotide resolution. scan() downloads a precalculated null Qmax_null.rds from HuggingFace on first run.

Reported result types (paper) and pipeline mapping

# Reported result Pipeline Scope
R1 README worked example: t1 → 0 SVRs; t2 → 2 SVRs at 92–100 nt (Q≈54.9) & 132–150 nt (Q≈79.0), seed=123 init→scan (deterministic) IN — self-contained, exact
R2 Flu dataset SVRs; benchmark "top-20 nt: 60% in annotated SVRs" (Fig 3; deltaSHAPE 63%, dStruct 38%) init→normalize→scan on shipped data/Flu.rda IN — data shipped
R3 Simulation: largest Jaccard across all 9 settings; lowest avg distance (Fig 2) simulate()+scan, needs ViennaRNA + 1000 seqs PARTIAL — needs OneDrive supplement + ViennaRNA
R4 Negative controls: FPR < 0.05 (Supp Fig 12) scan on identical-condition controls PARTIAL — needs supplement data
R5 icSHAPE application (GSE117840): 61 cyto-downregulated transcripts all contain SVRs, p=1.7e-3 (Fig 4) full pipeline on GSE117840 OUT/STRETCH — large, needs reactivity matrices + annotation

Out of scope (wet-lab / external): the SHAPE/icSHAPE library prep, the published ViennaRNA structure ground truth, RBP-binding annotations.

Datasets

  • Shipped in repo: data/Flu.rda (Flu SHAPE-Seq benchmark) — primary in-repo data.
  • GSE117840 (GEO): icSHAPE of HEK293 subcellular compartments — paper's application dataset → profiled in data/dataset_profile.json.
  • OneDrive supplement (1drv.ms link in README): full reproduction scripts+data (simulation seqs, negative controls). Fetch attempt for R3/R4.

Primary reproduction targets (floor → stretch)

  1. R1 exact (deterministic, shipped) — floor.
  2. R2 Flu SVRs + benchmark overlap — floor.
  3. R3/R4 simulation + negative controls if supplement fetchable — stretch.
  4. GSE117840 profile (always) + R5 application if feasible — stretch.
Figures / tables: Fig 3Fig 12Fig 2Fig 4
R1_t1
Reported
0 SVRs (README t1)
Reproduced
0 SVRs
exact
R1_t2
Reported
2 SVRs: 132-150 (Q=79.0) & 92-100 (Q=54.9), P=0
Reproduced
2 SVRs: 132-150 (Q=79.0,P=0) & 92-100 (Q=54.9,P=0)
exact
R2_flu_top20
Reported
0.60 (DiffScan top-20 in annotated SVRs, Flu)
Reproduced
12/20 = 0.600
exact
R2_flu_svr
Reported
detects annotated SVRs at nt resolution (Fig 3)
Reproduced
3 SVRs (9-28,34-51,70-74); precision 0.488, recall 0.875
partial
R4_negctrl
Reported
FPR < 0.05 (virtually no false positives)
Reproduced
constructed null: FPR 0.0505 (condA, 1 SVR) / 0.0000 (condB)
partial
R3_sim
Reported
largest Jaccard 9/9 simulation settings (Fig 2)
Reproduced
not attempted (gated on ViennaRNA + paper's 1000-seq set)
m.public.grade.uncheckable
R5_icshape
Reported
61/61 cyto-downregulated transcripts have SVRs, p=1.7e-3 (Fig 4)
Reproduced
not attempted (gated: replicate reactivity + downreg list absent from GSE117840 deposit)
m.public.grade.uncheckable

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

268.1 k
tokens (I/O) · 18.3 M incl. cache
50 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.