Chemical reversible crosslinking enables measurement of RNA 3D distances and alternative conformations in cells.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓No authors-side cause for any deviation
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL (well-reproduced core). SHARC-seq / CRSSANT pipeline, COMPUTE RAN on «our HPC» SLURM compute nodes. (1) ALL 4 shipped deterministic CRSSANT tests reproduced: gaptypes classification (10093->9179/339/1/1/281/0) EXACT; gapfilter splice/short-gap (352->291) byte-identical SAM EXACT; gapmcluster human-7SK TG assembly (275 alignments) byte-identical EXACT; crssant DG clustering (24 DGs) byte-identical in 3/5 runs, within-tol (clique-iteration non-determinism). 3/4 byte-identical -> strong evidence the read-processing + clustering core is faithful. (2) Headline biological claim '17 DGs on 7SK' (Supp Fig 10f) reproduced END-TO-END FROM RAW FASTQ (SRR13797237): STAR (exact README params) -> gaptypes -> gapfilter -> crssant on the 939 N-gapped reads with both arms inside 7SK -> 24-25 RN7SK DGs (5 runs); the top-17 by read support (>=5 reads) form a well-supported core consistent with the reported 17, extra 7-8 are low-support (2-5 reads) removed by a coverage/confidence threshold -> graded partial (same regime; not an exact-17 claim). (3) Dataset GSE167812: 12/12 samples present, profiled; deposit = processed read-arm tables (HeLa _anno.txt + in-vitro 1HR2 bedpe), NOT final per-RNA DG BEDPE; deepest HeLa lib dominated by CVA21 virus + rRNA. (4) Two latent bugs in shipped crssant.py line2info() documented (split '\n' vs '\t'; undefined genesdict) - dead SA-chimeric path never run by the authors' tests; NOT a fabrication signal. NOT reproduced: per-library exact %gapped for all 10 libraries (S1, partial). OUT OF SCOPE (not attempted): wet-lab crosslinking-efficiency assays, 3D distance/structural modeling (not pipeline-derived).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-19 ⛓ 06e39fa3fb0b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors test whether reversible bifunctional 2′-hydroxyl acylation crosslinkers of defined length (SHARC), combined with exonuclease trimming and proximity ligation sequencing, can measure spatial distances between nucleotides in RNA and thereby capture 3D tertiary structures and alternative conformations of RNA in living cells.
- ★ SHARC uses chemical crosslinkers of defined lengths to measure distances between nucleotides in cellular RNA method
- ★ SHARC-exo (crosslinking + exo trimming + proximity ligation + sequencing) enables transcriptome-wide RNA tertiary structure contact maps at high accuracy and precision method
- ★ SHARC data provide constraints that improve Rosetta-based RNA 3D structure modeling to near-nanometer resolution finding
- ★ Integrating SHARC-exo with other crosslinking-based methods reveals compact folding of the 7SK RNA, a regulator of transcriptional elongation finding
- ★ Aromatic dicarboxylic acid-derived crosslinkers (e.g., dipicolinic acid) achieve near-quantitative (97-99%) RNA crosslinking efficiency finding
- ★ SHARC crosslinks can be selectively reversed under mild alkaline conditions without causing RNA phosphodiester chain damage finding
- ★ Exonuclease (RNase R) trimming pinpoints crosslink sites at near-nucleotide resolution by stalling ~5 nt from the crosslinked base method
- Proteinase K/TNA extraction removes proteins prior to crosslink detection, indicating detected interactions are RNA-RNA rather than protein-mediated finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Crosslinking efficiency assay (polyacrylamide gel electrophoresis) | model self-complementary RNA 1 duplex, in vitro | treatment with 8 activated dicarboxylic acid crosslinkers | % crosslinked RNA | urea-denatured TBE PAGE gel |
| 1H NMR hydrolysis kinetics | dipicolinic acid imidazolide (DPI), in vitro | pH 7.4 buffer, room temperature, time course | hydrolysis half-life / rate constant | NMR |
| Crosslink reversal / RNA stability assay | model RNA 1 duplex crosslinked with DPI, in vitro; model dinucleotide ApA and compound 2 (DPI methyl ester) | alkaline conditions (Borate buffer pH 10.0-11.0, 37°C) | reversal efficiency and RNA degradation | 1H NMR and urea-denatured TBE PAGE gel |
| SHARC-exo (crosslinking, RNase III digestion, DD2D gel isolation, RNase R exo trimming, proximity ligation, RT, high-throughput sequencing) | HEK293T cells (ribosome, 7SL, RNase P, spliceosome, 7SK RNA) | DPI crosslinking (5, 12.5, 25 mM) with varying RNase R trimming times | gapped-read fraction, crosslinked fragment recovery %, inter-nucleotide spatial distances | high-throughput sequencing |
| PARIS psoralen crosslinking with exo trimming validation | cells (28S rRNA) | psoralen crosslinking followed by RNase R trimming | positional enrichment of uridine at trimmed 3′ ends | high-throughput sequencing |
| icSHAPE reactivity comparison | human ribosome, HEK293T cells | none (comparison to existing icSHAPE dataset) | icSHAPE signal along gapped-read arms | — |
| Proteinase K digestion / TNA RNA extraction | HEK293T cells | proteinase K treatment | protein removal efficiency | — |
- ▲ Aromatic dicarboxylic acid crosslinkers (terephthalic, isocinchomeronic, dipicolinic acids) showed 97-99% crosslinking efficiency, versus 1-24% for oxalic/succinic acid 97-99%
- ▲ DPI hydrolysis half-life at pH 7.4 was ~5 min, faster than the related SHAPE reagent NAI (~30 min half-life) ~6-fold faster
- ▲ Compound 2 (DPI methyl ester) hydrolyzed far faster than the ApA phosphodiester bond, allowing selective crosslink reversal ~1000-fold difference in rate constant (3.5×10⁻⁴ s⁻¹ vs <4.0×10⁻⁷ s⁻¹)
- – SHARC crosslinked RNA duplex was nearly fully reversed after alkaline treatment without apparent degradation
- ▲ Trimmed ribosome samples showed enrichment of single-stranded nucleotides peaking at the 5th nucleotide from the 3′ end ~1.3-fold over non-trimmed
- ▲ icSHAPE reactivity signal was enriched near the crosslink site in trimmed samples ~3.7-fold
- – SHARC-exo measured ribosome inter-nucleotide distances close to the physical crosslinker length, with distances significantly narrower than shuffled-read controls mode ~8 Å; 51% within 20 Å (unrefined), 31-49% within 20-40 Å (refined)
- ▼ Distances constrained by secondary structure (dsRNA) were predominantly short, while core and expansion-segment tertiary contacts showed progressively broader distance distributions 96.2% (dsRNA) vs 58.2% (core) vs 33.9% (ES) within 20 Å
- fold_change 97-99% crosslinking efficiency (aromatic reagents) vs 1-24% (oxalic/succinic) (crosslinking efficiency of SHARC reagents on model RNA 1 duplex)
- other DPI hydrolysis half-life ~5 min at pH 7.4 (DPI reaction kinetics measured by NMR)
- other rate constant: ApA <4.0×10⁻⁷ s⁻¹; compound 2 3.5×10⁻4 s⁻¹ (R²=0.99); ~1000-fold difference (selective crosslink reversal vs phosphodiester stability)
- pvalue p < 10⁻³⁰⁰ (Wilcoxon rank-sum test) (SHARC-exo distance distribution vs randomly shuffled reads, ribosome)
- count 51% of minimal distances within 20 Å, mode ~8 Å (ribosome cryo-EM distance validation, all gapped reads)
- count 31% within 20 Å and 49% within 40 Å (ribosome distances restricted to 5th±2 nt positions)
- count dsRNA 96.2%, core 58.2%, expansion segments 33.9% of distances within 20 Å (distances by structural category (secondary vs tertiary contacts))
- count 1.01%, 1.31%, 1.89% RNA fragments recovered as crosslinked at 5, 12.5, 25 mM DPI; 3.3-14.5% gapped reads (SHARC-exo crosslinking and library statistics in HEK293T cells)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This methods paper presents SHARC, a chemical crosslinking-sequencing approach for measuring RNA inter-nucleotide distances in cells. Quantitative benchmarking relies primarily on Wilcoxon rank-sum tests comparing observed distance distributions to shuffled null distributions, with crosslinking efficiencies characterized as mean ± SD from technical replicates. Hydrolysis kinetics are summarized by rate constants with goodness-of-fit (R²). Most structural results are reported descriptively as cumulative distance percentages and modes rather than through formal inferential tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon rank-sum (WRS) test | Comparison of minimal inter-nucleotide distance distributions between SHARC-exo gapped reads and randomly shuffled reads in the human 28S rRNA (Fig. 2e and Fig. 2g) | All gapped reads mapped to the human ribosome; exact count not stated in provided text | not stated |
| Exponential decay curve fitting (rate constant estimation) | Hydrolysis kinetics of model SHARC compound 2 and ApA dinucleotide (Fig. 1h); R²=0.99 reported for compound 2 | — | not stated |
-
p-values from Wilcoxon rank-sum tests were reported only as bounds (p < 10^-300) without accompanying effect sizes↳ Could also: A rank-biserial correlation or median difference with 95% CI could also be reported alongside the p-value — With very large n (many sequencing reads), even trivially small differences produce near-zero p-values; effect sizes quantify the magnitude of the difference independently of sample size and support comparison across experiments or datasets
-
Distance distributions were compared to a shuffled null using the Wilcoxon rank-sum test, which is sensitive to location shift↳ Could also: A Kolmogorov-Smirnov (KS) test or a permutation test on a chosen summary statistic could also be used — KS and permutation tests are sensitive to differences in the full distributional shape—including spread and tail behavior—which may be particularly informative given the explicitly described long-tailed distance distributions and heterogeneous RNA conformations
-
Crosslinking efficiency experiments used n=3 technical replicates (same RNA preparation, repeated measurements)↳ Could also: Independent biological replicates (separate RNA preparations and crosslinking reactions) could also be used — Biological replicates capture preparation-to-preparation variability and support broader inference about the method's reproducibility across experimental batches, complementing technical replicates that primarily reflect measurement precision
-
Dispersion was reported as SD from a small number of replicates (n=3 technical; n=2 biological)↳ Could also: 95% confidence intervals could also be reported alongside or instead of SD — CIs directly express uncertainty in the estimated mean and are often recommended for small-n experiments because they scale with n, whereas SD reflects only the spread of observations and stays roughly constant as n grows
-
Multiple comparisons were performed across DPI concentrations, trimming conditions, and RNA structural categories without a stated multiplicity correction↳ Could also: A Benjamini-Hochberg FDR correction or Bonferroni correction could also be applied across the family of comparisons — Applying a correction makes explicit which findings remain significant after accounting for the number of tests and clarifies the family-wise error rate, which is a standard consideration when many comparisons are reported in the same study
-
Structural subpopulations (dsRNA, core tertiary, expansion segment) were characterized descriptively by the percentage of reads falling within fixed distance thresholds↳ Could also: Mixture modeling (e.g., a mixture of log-normal or gamma components fit to the distance distributions) could also formally decompose reads into structural subpopulations — Mixture models can simultaneously estimate the proportion and characteristic distance parameters of each conformational state, providing a quantitative complement to the threshold-based summaries used and potentially better capturing the heterogeneous conformations the paper emphasizes
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35177610 (SHARC-seq / CRSSANT)
Paper: Van Damme et al. 2022, Nat Commun 13:911. "Chemical reversible crosslinking enables measurement of RNA 3D distances and alternative conformations in cells." (method = SHARC-seq)
- DOI 10.1038/s41467-022-28602-3 · PMCID PMC8854666
- Analysis code: https://github.com/zhipenglu/CRSSANT (third-party/lab tool, P16 OK)
- Data: GEO GSE167812 (SRA SRP308411); processed =
GSE167812_RAW.tar(≈113 MB, BEDPE + TXT)
The pipeline (CRSSANT), per repo README + paper Methods
- Preprocess FASTQ (adapter/barcode trim).
- STAR (v2.7.0f in paper) map reads — chimeric/softclip aware
(
--chimSegmentMin 5 --chimJunctionOverhangMin 5 --chimScoreJunctionNonGTAG 0 --chimOutType WithinBAM HardClip --outFilterMultimapNmax 10). softreverse.py— rearrange softclipped alignments, remap.gaptypes.py— classify primary alignments into 6 types: continuous / gap1 / gapm / trans / homo / bad.gapfilter.py— drop splice junctions + 1–2 nt gaps.crssant.py— cluster gap1+trans alignments into Duplex Groups (DGs) and Non-overlapping Groups (NGs). Params: t_o=0.5 (spectral)/0.1 (cliques), t_eig=5.gapmcluster.py— multi-gap → tri-segment groups (higher-order).
In scope (pipeline-derived, reproducible)
- S1 Per-library % gapped reads — paper: "3.3–14.5% of the reads are gapped" (Results; per-library in Supplementary Table 1). Reproduce by mapping the SRA libraries and counting gap1+gapm vs total primary alignments.
- S2 Duplex Groups (DGs) per RNA / per library — the deposited BEDPE files in
GSE167812_RAW.tarARE the crssant.py output. Reproduce DG assembly from the raw reads and compare DG counts to the deposited BEDPE (ground-truth pipeline output) and to any paper-stated counts (e.g. 7SK "17 DGs", Supp Fig 10f). - S3 Read/alignment classification counts (continuous/gap1/gapm/trans) per library — intermediate, checkable against Supplementary Table 1.
Reproduction strategy
Treat the deposited processed BEDPE/TXT (GSE167812_RAW.tar) as the authors' pipeline output. Two-tier:
- Tier A (cheap, profiling): download the processed tar (small) → profile the BEDPE/TXT (N DGs, columns, RNAs covered) and cross-check against paper numbers.
- Tier B (full repro, heavy → «our HPC»): download raw FASTQ for selected SRR (SRP308411) → STAR map → gaptypes → gapfilter → crssant.py → compare regenerated DG counts + %gapped against Tier-A deposited output and Supplementary Table 1.
Out of scope (not attempted)
- Wet-lab: crosslinking-efficiency gels (1.01/1.31/1.89%), DD2D recovery, RNase titration — bench measurements, not pipeline-derived.
- 3D-distance / structural-modeling claims (Å distances, 28S/18S 3D models, VARNA/ PyMOL renderings) — derived from external structures + manual modeling, not the read-processing pipeline. Noted but not reproduced.
Heavy-compute note
STAR human-genome mapping + alignment classification needs ≥30–100 GB scratch and
RAM → must run on «our HPC» (SLURM). «host» only orchestrates. As of first pass
the «our HPC» VPN tunnel is unreachable («host» ssh times out) — waiting for the
central fix before submitting Tier-B jobs; Tier-A profiling can proceed once the
tunnel is up (downloads happen on front1/«infra», never «host»).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.