RNA structure maps across mammalian cellular compartments.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to attempt, but the headline structural figure (Fig 1e: exon/intron Gini of HEK293 chromatin icSHAPE) does NOT reproduce 1:1 from the deposited data. ROOT CAUSE is a data-deposition gap: the authors' Fig 1e Gini was computed on a gene-level pre-mRNA shape.out (their script's input_path '.../hek_gene_ch_vivo/shape.out'), but GSE117840 deposited only TRANSCRIPT-LEVEL (spliced) per-compartment scores. Verified: 0/4686 (vivo) and 1/10264 (vitro) deposited chromatin entries are intron-containing pre-mRNA. Therefore the INTRON arm (Gini 0.7, the paper's main contrast) is not independently verifiable from public data. Reimplementing the authors' documented algorithm (20-nt non-overlapping windows, >=10 valid, standard Gini, protein_coding only, GENCODE v28) on the deposited mature transcripts gives the EXON-arm Gini ~0.58 for BOTH in vivo and in vitro, not the reported 0.5-vs-0.6 split. Window counts land in the same order of magnitude (~24k vivo / ~58k vitro pc-tx vs reported 18930 / 51409), supporting that the metric family is right but not exact. The read-depth=100 pipeline parameter matches Methods. NO fabrication evidence: values are plausible and on the right scale, but the specific stratified claim cannot be confirmed from what was deposited. NOT ATTEMPTED (last-20%, out of scope): replicate correlation (per-replicate data not deposited), Fig 2-6 TR/TE/CLIP/m6A analyses (need large external datasets + authors' private toolbox + undeposited intermediates), and raw FASTQ->icSHAPE base-calling (deposited .out IS the pipeline output).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 45assessed: 2026-06-16 ⛓ a9b2b4e66142
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhether mapping RNA secondary structure (structuromes) separately across distinct subcellular compartments (chromatin, nucleoplasm, cytoplasm) in mammalian cells reveals compartment-specific RNA structural changes and their regulation by transcription, translation, decay, RNA modifications, and RBP binding that whole-cell averaging obscures.
- ★ icSHAPE-based cytotopic RNA structuromes across chromatin, nucleoplasm and cytoplasm in human and mouse cells substantially expand the scope of RNA structural information beyond whole-cell data. resource
- ★ RNA structure plays a central role connecting transcription, translation and RNA decay, with structure formed at chromatin during transcription largely persisting and potentially mediating the transcription-translation link. finding
- ★ RNA structural changes are pervasive across compartments and between species, with greater structural variation in vivo than in vitro, reflecting regulation by many factors. finding
- ★ RNA modifications (m6A, pseudouridylation) and RBP binding underlie RNA structural changes; protein binding can explain most structural change sites (3392 of 5903). mechanism
- ★ Structural analysis of CLIP-seq data distinguishes m6A 'readers' (canonical YTH, HNRNPC, IGF2BP) from 'anti-readers', and validates LIN28A as an m6A anti-reader. finding
- HNRNPC acts as an m6A reader indirectly by binding a purine-rich motif that becomes unpaired and accessible upon nearby m6A modification rather than recognizing the N6-methyl group. mechanism
- icSHAPE combined with subcellular RNA fractionation enables in vivo and in vitro structure mapping per compartment as a method/resource. method
- Intronic regions are more folded in vivo than exonic regions (Gini 0.7 vs 0.5), independent of RBP binding. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| in vivo icSHAPE (RNA structure probing + sequencing) | v6.5 mouse embryonic stem (mES) cells and human HEK293 cells, fractionated into chromatin, nucleoplasm, cytoplasm | subcellular fractionation; none (chemical SHAPE probing) | per-nucleotide structural flexibility (icSHAPE reactivity score, unpaired bases) / Gini index | high-depth deep sequencing (~200 million reads per replicate) |
| in vitro icSHAPE (refolded naked RNA, control) | isolated RNA from chromatin/nucleoplasm/cytoplasm fractions of mES and HEK293 cells | RNA purification and in vitro refolding | per-nucleotide icSHAPE reactivity of refolded RNA | deep sequencing |
| RT-qPCR (fractionation quality control) | mES / HEK293 subcellular fractions | fractionation | abundance of landmark RNAs across fractions | — |
| Western blot (fractionation quality control) | mES / HEK293 subcellular fractions | fractionation | compartment-specific marker protein levels | — |
| CLIP-seq (LIN28A) and integration of published RBP CLIP-seq | human/mouse cells; published RBP binding datasets | none / RBP binding | RBP binding sites and m6A-dependent binding (reader vs anti-reader) | — |
| Integrative analysis with external transcription rate, translational efficiency, RNA half-life, m6A and pseudouridylation datasets | human and mouse transcriptome | none (computational correlation) | correlation of Gini index/icSHAPE reactivity with transcription, translation, decay, modifications | — |
- ▼ Lower 5'UTR transcriptional rate correlates with more RNA structure in chromatin fraction. r=-0.19, p=1.5e-6
- ▼ More 5'UTR RNA structure correlates with decreased translational efficiency in cytoplasm. r=-0.31, p=1.7e-48
- ▼ More-structured RNAs have shorter half-lives in nucleoplasm. r=-0.23, p=4.6e-91
- ▼ More-structured RNAs have shorter half-lives in cytoplasm. r=-0.18, p=1.1e-36
- ▲ Intronic regions are more folded in vivo than exonic regions. Gini 0.7 vs 0.5
- – Protein binding (CLIP-seq) explains most RNA structural change sites. 3392 of 5903 sites
- ▲ m6A destabilizes local RNA structure across all three fractions, with largest differences in nucleoplasm consistent with nuclear m6A deposition.
- ▲ Pseudouridylated regions show higher icSHAPE reactivity (less structured / hindered folding), predominantly in nucleus.
- correlation r=-0.31, p=1.7e-48 (5'UTR structure vs translational efficiency, cytoplasm)
- correlation r=-0.23, p=4.6e-91 (RNA structure vs half-life, nucleoplasm)
- correlation r=-0.18, p=1.1e-36 (RNA structure vs half-life, cytoplasm)
- correlation r=-0.19, p=1.5e-6 (transcriptional rate vs 5'UTR structure, chromatin)
- correlation r>0.75 (Pearson correlation across replicates for top 60% most-abundant transcripts)
- count 3392 of 5903 (RNA structural change sites explained by protein binding)
- other Gini 0.7 vs 0.5 (in vivo intron vs exon); 0.6 vs 0.6 in vitro (Gini index of structure in intron vs exon regions)
- count ~200 million reads per replicate (icSHAPE library sequencing depth)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study is a genome-wide, sequencing-based descriptive design (icSHAPE structure probing across three subcellular compartments in two species), with conclusions drawn primarily from correlation analyses between per-nucleotide structural reactivity (summarized via Gini index) and external measures of transcription rate, translational efficiency, and RNA half-life. Reproducibility was assessed with Pearson correlations between replicates, and a custom statistical method was implemented to call regions of structural change and to compare modified versus unmodified sites. Relationships were reported as correlation coefficients with associated p-values, and a mediation-versus-confounding model comparison was used to interpret the transcription–structure–translation linkage.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation coefficient (replicate reproducibility) | replicate vs replicate and across fractions for top ~60% most-abundant transcripts (Supplementary Figures 2–3) | — | not stated |
| Pearson/quantitative correlation (structure vs transcription rate) | 5'UTR structure of chromatin fraction vs transcriptional rate (Figure 2a); r=-0.19, p=1.5×10^-6 | — | not stated |
| correlation (structure vs translational efficiency) | 5'UTR structure of cytoplasmic fraction vs translation efficiency (Figure 2b); r=-0.31, p=1.7×10^-48 | — | not stated |
| correlation (structure vs RNA half-life) | nucleoplasm (r=-0.23, p=4.6×10^-91) and cytoplasm (r=-0.18, p=1.1×10^-36) (Figure 2c-d) | — | not stated |
| mediation vs confounding model comparison | interpreting transcription–RNA structure–translation relationship (Figure 2g) | — | not stated |
| custom statistical method for regions of structural variation / overlap enrichment | calling structural-change regions and overlap of HNRNPC binding with structural-variation sites (Figure 3, Supplementary Figure 9) | — | not stated |
-
Reproducibility and structure–phenotype relationships were quantified with Pearson correlation coefficients.↳ Could also: Spearman or Kendall rank correlation could also be reported alongside Pearson. — Rank-based correlations make no linearity or normality assumption and are robust to outliers and skew common in sequencing-derived measures, so they would complement the Pearson values for monotonic-but-nonlinear trends.
-
Numerous transcriptome-wide p-values were reported for correlation and overlap analyses.↳ Could also: A multiple-testing framework such as Benjamini-Hochberg FDR or a permutation-based null could also be applied across the family of genome-wide comparisons. — An explicit correction or empirical null conveys how many findings are expected by chance across many simultaneous tests and is often preferred when reporting large numbers of p-values.
-
Correlations were summarized by a point estimate of r with an associated p-value.↳ Could also: A 95% confidence interval (e.g., via Fisher z-transformation or bootstrap) around each correlation could also be presented. — Interval estimates communicate the precision of the association and the practical magnitude, which point estimates and p-values alone do not fully convey.
-
The transcription–structure–translation linkage was interpreted with a mediation-versus-confounding model comparison.↳ Could also: A formal causal-mediation analysis with bootstrapped indirect-effect estimates, or a structural-equation/partial-correlation framework, could also be used. — These standard approaches quantify the indirect (mediated) effect and its uncertainty directly, which would add an effect-size and interval to the model-selection conclusion.
-
Structural-change regions were identified with a custom statistical method and significance of overlap was asserted.↳ Could also: A permutation or region-shuffling null (e.g., bedtools-style randomization) could also be used to calibrate overlap enrichment. — An explicit empirical null provides an enrichment statistic and p-value that account for gene length, expression, and region composition, helping contextualize the observed overlap.
-
Replicate agreement was reported as a correlation threshold (r > 0.75) for abundant transcripts.↳ Could also: Concordance metrics such as the intraclass correlation coefficient or irreproducible discovery rate (IDR), commonly used for high-throughput sequencing replicates, could also be reported. — ICC/IDR jointly assess agreement in both rank and scale and quantify reproducibility of called features, complementing a Pearson-based summary.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Lower 5'UTR transcription rate correlates with more RNA structure in the chromatin fraction.other mouse human down 2019×1papers★ This paper is the founder (earliest)
-
More 5'UTR RNA structure correlates with decreased translational efficiency in cytoplasm.other mouse human down 2019×1papers★ This paper is the founder (earliest)
-
Intronic regions are more folded in vivo than exonic regions.other mouse human up 2019×1papers★ This paper is the founder (earliest)
-
m6A destabilizes local RNA structure across all fractions, with the largest effect in nucleoplasm.other mouse human down 2019×1papers★ This paper is the founder (earliest)
-
Pseudouridylated regions are less structured (higher icSHAPE reactivity), predominantly in the nucleus.other mouse human down 2019×1papers★ This paper is the founder (earliest)
-
RBP binding (CLIP-seq) explains the majority of in-vivo RNA structural change sites.other mouse human 2019×1papers★ This paper is the founder (earliest)
-
More-structured RNAs have shorter half-lives in the cytoplasm.other mouse human down 2019×1papers★ This paper is the founder (earliest)
-
More-structured RNAs have shorter half-lives in the nucleoplasm.other mouse human down 2019×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
2 downstream papers · 1 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Structure-Mediated RNA Decay by UPF1 and G3BP1. 2020 · 236 cites
- Spen links RNA-mediated endogenous retrovirus silenc... 2020 · 49 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-30886404 (Sun et al. 2019, NSMB, "RNA structure maps across mammalian cellular compartments")
- Repo: github.com/lipan6461188/RNA_Structure_Dynamics @ 2c8ff3f63c9f7afba03c7394171c3ec23a883e8b (per-figure scripts; depends on authors' personal
tools.py/ParseTrans/icSHAPEtoolbox + GAP parser github.com/lipan6461188/GAP @ 874e7a5) - Data: GEO GSE117840 — processed icSHAPE score files (.out.txt.gz), one per compartment x condition x species.
Data structure (verified, «job»-2180721)
.outformat: col1 = Ensembl transcript ID (ENST), col2 = length, col3 = RPKM, col4.. = per-nucleotide icSHAPE score (NULL = no confident score).- The deposited per-compartment files are MATURE / spliced-transcript reactivities. Checked against GENCODE v28: of HEK293 chromatin entries, 0/4686 (vivo) and 1/10264 (vitro) are intron-containing pre-mRNA (out_len == genomic span); the rest match spliced length (or are single-exon). => the deposited data has essentially no intron reactivities.
In scope (pipeline-derived, attempted)
- Figure 1e — windowed Gini index of icSHAPE reactivity, HEK293 chromatin, EXON arm. Authors' algorithm (
Figure 1/1e-RBP Gini/1.calcGINI.py): non-overlapping 20-nt windows over NULL-filtered reactivities, >=10 valid per window, standard Gini index, protein_coding transcripts only. Reproducible from deposited data for the EXON arm (mature transcript = exonic). - Pipeline parameter: read-depth threshold = 100 (Methods text).
Out of scope / not reproducible from deposited data
- Figure 1e INTRON arm (Gini ~0.7). Requires intron reactivities, which live only in the authors' gene-level pre-mRNA
shape.out(their script'sinput_path = ".../hek_gene_ch_vivo/shape.out"). That gene-level intermediate was NOT deposited in GSE117840. Not derivable. (DATA-DEPOSITION GAP.) - Replicate correlation (Pearson r > 0.75, Supp Fig 2). Per-replicate scores not deposited (only combined per-compartment scores).
- Fig 2 TR/TE/structure correlations, Fig 3-6 dynamics/RBP/m6A analyses. Require large external datasets (Ribo-seq TR/TE, CLIP-seq, m6A maps), the authors' undeposited intermediate tables, and their personal toolbox -> the hard last 20%, not attempted.
- Raw icSHAPE base-calling from FASTQ. Heavy; the deposited .out files ARE the pipeline output, so re-running the base caller adds no auditable value vs. cost. Not attempted.
Heavy-compute note
All compute on «our HPC»/«infra» («job», 2180720, 2180721, 2180724). «host» holds small results only.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The headline Fig 1e result (HEK293 chromatin icSHAPE exon/intron Gini: introns more folded, 0.7 vs 0.5 in vivo) does not reproduce 1:1 from GSE117840. Root cause is an authors-side data-deposition gap: only spliced transcript-level scores were deposited, so the intron arm is non-derivable (0/4686 vivo entries are intron-containing) and the deposited exon arm gives ~0.58 for both conditions rather than the reported 0.5/0.6 split. Values are plausible and on the right scale (window counts within ~12%), with no fabrication evidence — the central contrast is simply unverifiable from the public record, so this is graded a data-availability partial, not a discrepancy or refutation.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.