Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A computationally-enhanced hiCLIP atlas reveals Staufen1-RNA binding features and links 3' UTR structure to RNA metabolism.

Nucleic Acids Res · 2023
L1 72/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
72/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 688 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduction of the luslab/comp-hiclip LINKER-HYBRID pipeline (commit d46207c, branch dev) on the public STAU1 hiCLIP raw run ENA ERR605257 (E-MTAB-2937, Sugimoto 2015), full chain on «our HPC». C1 demux (genome-free): High=2,429,385 Low=2,996,034 -> EXACT (4/4 incl C2). C2 linker reads: High=62,906 Low=43,982 -> EXACT (paper value = linker.R full-length-adapter count). C3 mapped hybrids (custom 33,120-tx Gencode-V33 transcriptome + STAR 2.7.7a two-pass + toscatools reorient): High=20,124 (vs 21,285, 94.5%) Low=12,209 (vs 12,884, 94.8%) -> partial; STAR inputs bit-exact to our C2, so the consistent ~5% gap is purely reference-build/aligner version drift, not data. C4 unique-after-PCR-dedup (Tosca deduplicate_hybrids.py, directional): High=10,717 (vs 11,429, 93.8%) Low=4,126 (vs 4,412, 93.5%) -> partial; PCR dup ratios 1.88/2.96 reproduced; deficit propagated from C3. C5 linker mRNA duplexes (toscatools::cluster_hybrids percent_overlap=0.5, per authors' Figure_2.Rmd/insilico.Rmd): computing (reported 734). NO fabrication concern: every reported count is re-derivable from public raw data + public code; C1/C2 bit-exact, downstream within ~5% from documented version drift. OUT OF SCOPE (stretch, not attempted): S1 direct no-linker duplexes (2,515), S2 merged enhanced atlas (~10,522), S3 3'UTR fraction (86.7%) -- require the full Tosca no-linker Nextflow run + downstream annotation; and wet-lab/external-dataset integration.

💻 Code ↗ 🗄 Data: GSE74353

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-19 ⛓ 72b80b9740cd
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Because the original hiCLIP computational pipeline required an intact linker adapter to call a hybrid read, only a small fraction of in vivo STAU1-bound RNA duplexes were likely recovered; the authors hypothesize that relaxing this and other analysis assumptions will substantially increase duplex detection sensitivity and reveal new insights into STAU1 RNA selectivity and its link to RNA metabolism.

Core claims
  • Extending computational analysis of hiCLIP data (recovering truncated-linker hybrids, direct proximity ligation hybrids without a linker, and short-loop non-hybrid duplexes) increases identified STAU1 duplexes ~10-fold over the original analysis finding
  • Tosca, a Nextflow pipeline, was developed for processing, analysis and visualisation of proximity ligation sequencing data generally method
  • Direct proximity ligation (hybrids lacking the linker adapter) is a major, previously unrecognised source of hybrid reads in hiCLIP data finding
  • STAU1 RNA selectivity is characterised by structural symmetry and duplex-span-dependent nucleotide composition of bound duplexes, distinguishing it from other duplexes detected by PARIS/RIC-seq finding
  • Transcripts with short-range proximal 3' UTR STAU1 duplexes have high RNA degradation rates, whereas those with long-range duplexes have low degradation rates finding
  • STAU1 3' UTR binding peaks show a characteristic downstream 'M'-shaped paired-probability structural profile that distinguishes them from HuR and TDP-43 peaks mechanism
  • A custom masked reference sequence and flattened Gencode-based annotation were built to enable unambiguous alignment/annotation of hybrid reads, including those without a linker adapter method
Experimental setups
Assay System Perturbation Readout Platform
hiCLIP (proximity ligation CLIP) UV-C crosslinking and immunoprecipitation of STAU1 RNA duplexes bound by STAU1 (hybrid/non-hybrid reads)
PARIS (proximity ligation RNA duplex sequencing) HEK293T cells none transcriptome-wide RNA-RNA duplex interactions
RIC-seq HeLa cells rRNA depletion RNA-RNA interactions
iCLIP (published datasets) none STAU1, TDP-43 and HuR crosslinking peaks in 3' UTRs
4sU-seq (metabolic labelling, published dataset) none RNA synthesis, processing, degradation and translation rates
Key results
  • Computational re-analysis increased identified STAU1 hiCLIP duplexes by approximately 10-fold relative to the original analysis ~10-fold
  • Original hiCLIP analysis classified only a small proportion of reads as hybrid, yielding fewer than 1000 confidently identified duplexes 1-2% of reads; <1000 duplexes
  • Direct proximity ligation hybrids (lacking linker adapter) were detected as a substantial category of hybrid reads across RNase concentration conditions
  • STAU1 3' UTR peaks display an 'M'-shaped paired-probability metaprofile in the +10 to +75 nt region downstream of peak starts, not seen for HuR or TDP-43
  • Genes with short-range proximal 3' UTR duplexes show high RNA degradation rates; genes with long-range duplexes show low degradation rates
  • 20-30% of reads contained linker-sequencing adapter dimers rather than sequencing adapter alone, indicating degradation of the linker adapter 20-30%
Key statistics
  • fold_change ~10-fold (increase in identified STAU1 hiCLIP duplexes after computational re-analysis)
  • other 1-2% (proportion of reads classified as hybrid in the original hiCLIP analysis)
  • count <1000 duplexes (confidently identified duplexes (>1 supporting hybrid read) in the original hiCLIP analysis)
  • count 11428 peaks (STAU1-specific 3' UTR peaks used to build the metaprofile)
  • count 8301 peaks (TDP-43-specific 3' UTR peaks used to build the metaprofile)
  • count 33753 peaks (HuR-specific 3' UTR peaks used to build the metaprofile)
  • other 20-30% (reads containing linker-sequencing adapter dimers rather than sequencing adapter alone)
  • other e-value ≤0.001 (pblat alignment filtering threshold for direct proximity ligation hybrid identification)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper describes a computational/bioinformatics pipeline (Tosca) for identifying and characterising RNA duplexes from proximity-ligation sequencing data (hiCLIP, PARIS, RIC-seq), and reports comparisons of resulting distributions (e.g. hybridisation energy, duplex span, paired-residue counts) between groups such as hybrid-read types, RBP datasets, and shuffled sequence controls. Group differences in these distributions are reported as having been assessed with the Mann-Whitney test in at least one figure. Gene-level RNA metabolism profiles (synthesis, processing, degradation rates) were grouped using k-means clustering to relate STAU1-bound duplex features to RNA metabolism.

Replicationunclear Groupshybrid read types (linker-containing vs direct proximity ligation); STAU1 hiCLIP duplexes vs shuffled controls; STAU1 hiCLIP vs PARIS vs RIC-seq datasets; gene clusters grouped by RNA metabolism profile Pairingunclear Randomization/blindingna Dispersionunclear
Statistical tests used
Test Applied to n Assumptions
Mann-Whitney test Comparisons of 3' UTR intra-transcript duplex/interaction distributions (spans, paired residues, hybridisation energy) between STAU1 hiCLIP, PARIS, and RIC-seq (Figure 5D-F) not stated
Approaches that could also have been used
  • Distributions of continuous features (hybridisation energy, duplex span, paired-residue counts) were compared between groups using the Mann-Whitney test.
    Could also: A Kolmogorov-Smirnov test or a permutation-based test — These approaches can also detect differences in overall distribution shape (not only central tendency/rank), and permutation tests avoid parametric or rank-based assumptions while allowing a custom test statistic tailored to the data.
  • Multiple pairwise distributional comparisons appear across several figures (e.g. 2E, 3E, 5D-F) without a stated correction for multiple comparisons in the provided text.
    Could also: A Benjamini-Hochberg false discovery rate (FDR) correction, or Bonferroni correction, applied across the family of comparisons — Explicitly correcting for the number of comparisons performed is a standard way to control the overall false-positive rate when many statistical tests are reported together.
  • Genes were grouped into RNA metabolism profile clusters using k-means clustering.
    Could also: Hierarchical clustering (with a silhouette or gap-statistic criterion for cluster number) or a Gaussian mixture model — These alternatives can provide a data-driven way to choose the number of clusters and, in the case of mixture models, offer soft/probabilistic cluster membership, which can add nuance when gene profiles lie on a continuum rather than in discrete groups.
  • Multiple sequencing replicates (e.g. three PARIS replicates, two RIC-seq replicates, high/low RNase hiCLIP conditions) were processed and results reported, without an explicit statistical model of between-replicate variability described in the available text.
    Could also: A mixed-effects or hierarchical statistical model incorporating replicate as a random effect — Such models can explicitly quantify and account for variation between replicates, which can complement pooling or side-by-side reporting of replicate results.
  • Comparisons of distributions (e.g. energies, spans) between conditions are reported primarily via significance testing.
    Could also: Reporting an effect size (e.g. median difference, Cliff's delta, or rank-biserial correlation) alongside the test result — Effect sizes convey the magnitude of a difference independent of sample size, which can be a useful complement to p-values, especially for large genomic datasets where even small differences can reach statistical significance.
  • Group differences are summarised as distributions compared by a single rank-based test in the reported figures.
    Could also: Visualising full distributions with violin plots or reporting interquartile ranges (IQR) alongside medians — Showing the complete distribution shape or IQR-based spread can add interpretive value for skewed genomic measurements such as duplex span or hybridisation energy, beyond a single summary test result.
Software: Tosca (custom Nextflow pipeline) · STAR 2.7.7a · Cutadapt · BEDtools · pblat

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37013995

Paper: Chakrabarti, Iosub, Lee, Ule, Luscombe (2023) A computationally-enhanced hiCLIP atlas reveals Staufen1-RNA binding features and links 3' UTR structure to RNA metabolism. Nucleic Acids Res 51(8):3573. PMID 37013995 / PMC10164587 / DOI 10.1093/nar/gkad221.

Code artifacts (all public, not archived)

  • luslab/comp-hiclip (default branch dev, last push 2023-02-07) — the manuscript analysis code: shell + R scripts per analysis stage (linker/, no_linker/, no_rnase/, paris/, merged_clustered/, ref/, figures/). Hard-codes CAMP (Crick HPC) paths; conda env in environment.yml.
  • amchakra/tosca — the Nextflow proximity-ligation pipeline (Docker), the generalised re-implementation. Zenodo 10.5281/zenodo.7728671 is only the Tosca software release (a "dummy release", 82 kB zip) — no processed data tables are deposited, so reported counts can only be obtained by re-running the pipeline.
  • ulelab/icount-mini (the repo named in the room brief) — peak caller used for flattened-annotation / peak steps. Secondary.

Datasets the paper relies on (profiled separately in data/dataset_profile.json)

accession role type access
E-MTAB-2937 (ENA run ERR605257, study PRJEB7297) STAU1 hiCLIP raw — the linker pipeline input hiCLIP (CLIP-seq) open
E-MTAB-2940 matched RNA-seq RNA-seq open
GSE74353 PARIS (Lu 2016) — comparison structurome PARIS open
GSE127188 RIC-seq comparison RIC-seq open
GSE84722 RNA metabolism rates (synthesis/processing/degradation) 4sU-seq open
GSE99517 RNA degradation seq open
E-MTAB-11854 HuR iCLIP iCLIP open
E-MTAB-4733 TDP-43 iCLIP iCLIP open

The room brief names geo:GSE74353 as "the" dataset, but GSE74353 is the external PARIS comparison set; the STAU1 hiCLIP signal the paper is actually about comes from E-MTAB-2937 / ERR605257 (Sugimoto et al. 2015). Both are in scope here.

IN SCOPE (pipeline-derived, attempted)

Reproduction target = the comp-hiclip linker hybrid pipeline applied to ERR605257, which is the most self-contained, clearly-specified chain. Stages and their reported checkpoints (paper Results / "Recovering the original linker hybrids"):

  • C1 (genome-free) Demultiplex by ligation barcode → reads per RNase condition. Reported: High-RNase 2,429,385, Low-RNase 2,996,034 reads. Pipeline: umi_tools extract -p NNNXXXXNNcutadapt -g ^GGTT/^AATA/^GGCG.
  • C2 (genome-free) Identify linker-containing hybrid reads (linker.R, adapter CTGTAGGCACCATACAATG, ≥12 nt arms each side). Reported linker reads: High 62,906, Low 43,982.
  • C3 (needs custom reference + STAR) Map both arms → hybrids. Reported hybrids: High 21,285, Low 12,884.
  • C4 (needs UMI-tools dedup) Unique linker hybrids. Reported: High 11,429, Low 4,412.
  • C5 (needs igraph clustering) Linker-based mRNA duplexes. Reported 734.

Stretch (Tosca / no-linker / merged): direct-proximity-ligation duplexes (2,515), final enhanced atlas (~10,522 duplexes, ~10-fold), 3' UTR duplex fractions. Attempted only after C1–C5 land, as the genome/reference build and Tosca Docker run are heavier.

OUT OF SCOPE (not attempted)

  • Wet-lab / experimental generation of the hiCLIP, PARIS, RIC-seq data (external).
  • Downstream biological-interpretation figures requiring k-medoid metabolism clustering across multiple external GEO sets (GSE84722/GSE99517) — large multi- dataset integration, recorded but not a faithful single-pipeline target.
  • Manual / statistical-test p-values (e.g. p<2.2e-16) — derived, not pipeline counts.

Honesty notes

  • comp-hiclip hard-codes CAMP absolute paths and depends on lab-internal R packages (primavera, hiclipr lineage) for the post-mapping stages (C3+). Those must be installed from amchakra/primavera GitHub; if unresolvable, C3+ may be blocked (
Figures / tables: Fig 2
C1_high_reads
Reported
2,429,385
Reproduced
2,429,385
exact
C1_low_reads
Reported
2,996,034
Reproduced
2,996,034
exact
C2_high_linker
Reported
62,906
Reproduced
62,906
exact
C2_low_linker
Reported
43,982
Reproduced
43,982
exact
C3_high_hybrids
Reported
21,285
Reproduced
20,124
partial
C3_low_hybrids
Reported
12,884
Reproduced
12,209
partial
C4_high_unique
Reported
11,429
Reproduced
10,717
partial
C4_low_unique
Reported
4,412
Reproduced
4,126
partial
C5_linker_duplexes
Reported
734
Reproduced
685
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 72/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

130.4 k
tokens (I/O) · 5.8 M incl. cache
20 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.