MUTACLASH: identifying functional small RNA target sites using crosslinking-induced mutations.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the core. SRR3882949 = C. elegans ALG-1 iCLIP; we reproduced Table 1's mutation-frequency row by running the paper's OWN reference (mRNA_WS275.fa, shipped in the repo), the paper's OWN aligner (BWA-MEM), and the paper's OWN MD/CIGAR mutation rule (pipeline/find_deletion/bwa_find.py) with the repo's default trim params (trim_galore -q30 --length 17 --max_length 70). 53.65M raw -> 21.49M trimmed -> 9.38M primary-aligned reads on «our HPC» («job», ~7 min). DELETION reproduces to the same order of magnitude: 6.82-9.27% vs reported 5.8% (deletion-only def within ~18%), ~13x above the mRNA-seq baseline (0.026%) -> confirms the paper's central deletion claim => partial/1:1-ish. SUBSTITUTION does NOT reproduce in absolute value: ~20-26% vs 7.41%, robust to MAPQ (28% at MAPQ>=30, so not a multimapping artifact) -> the gap is explained by alignment SCOPE: we count over all transcriptome-aligned reads, whereas the paper computes over the ChiRA-defined non-hybrid TARGET set after dedup+filtering; non-target reads (rRNA/tRNA/ncRNA) force-aligned to mRNA inflate mismatches but not deletions. NOT attempted (the hard ~20%): full ChiRA chimera calling + hybrid/non-hybrid partition, UMI/PCR dedup weighting, binding-site prediction (pirScan/miRanda/RNAup), 22G-RNA regulatory effects (Fig 4,7), per-position CIM distribution (Fig 2), and the PRG-1 CLASH / mRNA-seq rows of Table 1 (different accessions). No fabrication concern: the deletion rate is derivable from shipped reference + public SRA + shipped code; the substitution mismatch is a scope/definition difference, not a non-derivable number.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 30assessed: 2026-06-16 ⛓ 8f02ea2aa745
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors hypothesize that crosslinking-induced mutations (CIMs) are present in PIWI/Argonaute CLASH data and act as single-nucleotide-resolution molecular footprints of small RNA–mRNA binding that can identify functional small RNA target sites otherwise overlooked by current tools.
- ★ CIMs are present and enriched in PIWI (piRNA) and Argonaute (miRNA) CLASH data and serve as molecular footprints of Argonaute binding on target mRNAs. finding
- ★ MUTACLASH is a bioinformatics pipeline that identifies small RNA–mRNA hybrids and locates CIMs within them, with a filtering step to avoid alignment-ambiguity artifacts. resource
- ★ CIMs are enriched at nucleotides corresponding to the center of piRNA and miRNA binding sites (deletions peak at piRNA/miRNA positions 11–12, substitutions at 9–10). finding
- ★ CIMs are enriched at nucleotides within piRNA target sites that exhibit local base-pairing mismatches. finding
- ★ For piRNA noncanonical sites with seed-region CIMs, reduced seed pairing is compensated by increased non-seed base-pairing to facilitate stable binding. mechanism
- ★ mRNAs with noncanonical binding sites and/or low hybrid abundance marked by CIMs show stronger regulatory effects than those without CIMs, indicating CIMs mark functionally relevant binding events. finding
- CIM presence shows only moderate correlation with hybrid abundance, so it provides an independent metric beyond pairing score/energy or hybrid abundance. finding
- CIMs show a uridine-derived mutation preference consistent with uridine's higher photoreactivity during UV crosslinking. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| CLASH (crosslinking, ligation, and sequencing of hybrids) | C. elegans (PRG-1 PIWI / piRNA) | UV crosslinking | piRNA–mRNA hybrid reads and CIMs (deletions/substitutions) on target mRNA | — |
| iCLIP (CLASH-like method) | C. elegans (ALG-1 Argonaute / miRNA) | UV crosslinking | miRNA–mRNA hybrid reads and CIMs on target mRNA | — |
| mRNA sequencing (RNA-seq) | C. elegans | none (no UV crosslinking) | baseline mRNA deletion/substitution frequency for comparison | — |
| Bioinformatic CIM/hybrid analysis (MUTACLASH pipeline) | C. elegans CLASH/iCLIP data sets | none | CIM positions, base-pairing ratios, hybrid abundance, regulatory effects | MUTACLASH (integrating prior hybrid-calling algorithms; miRanda, piRNA targeting scores) |
- ▲ PRG-1 CLASH non-hybrid mRNA reads showed elevated deletion and substitution frequency vs mRNA-seq 6.65% deletion / 6.76% substitution vs 0.026% / 1.4%
- ▲ ALG-1 iCLIP non-hybrid mRNA reads showed elevated mutation frequency vs mRNA-seq 5.8% deletion / 7.41% substitution
- ▲ PRG-1 CLASH hybrid reads contained CIMs above mRNA-seq baseline 2.8% deletion / 3.27% substitution
- – mRNA deletions peak at positions corresponding to piRNA nt 11–12 and substitutions at nt 9–10 (center enrichment) ~2 nt shift between deletion and substitution peaks
- – CIM-containing hybrid abundance correlates only moderately with total hybrid abundance R=0.245–0.403
- ▼ piRNA base-pairing ratio is reduced around the location of CIMs (e.g., at positions 5 and 14)
- – Seed-region CIMs associated with reduced seed pairing but increased non-seed pairing
- ▲ CIMs show uridine-derived mutation preference; U-to-C dominate U substitutions and C-to-U dominate C substitutions U-derived 42%–64%; U→C 83%–90%; C→U 84%
- fold_change Deletion 6.65% vs 0.026%; Substitution 6.76% vs 1.40% (PRG-1 CLASH non-hybrid mRNA reads vs mRNA-seq)
- count Deletion 5.8%; Substitution 7.41% (ALG-1 iCLIP non-hybrid mRNA reads)
- count Deletion 2.8%; Substitution 3.27% (PRG-1 CLASH hybrid reads)
- count Deletion 4.59%; Substitution 3.67% (ALG-1 iCLIP hybrid reads)
- count Deletion 0.026%; Substitution 1.40% (mRNA-seq non-hybrid baseline)
- correlation R = 0.245–0.403 (CIM-containing hybrid abundance vs total hybrid abundance (Supplemental Fig. S2A–D))
- other 42%–64% uridine-derived mutations (CIM nucleotide preference in PRG-1/ALG-1 CLASH)
- pvalue P<0.05 (*) and P<0.01 (**) (differences in seed/non-seed pairing ratios between hybrids with CIMs at given positions)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational bioinformatics study introducing MUTACLASH, a pipeline to identify crosslinking-induced mutations (CIMs) in CLASH/iCLIP sequencing data from C. elegans PIWI (PRG-1) and Argonaute (ALG-1) experiments. The main analyses are descriptive and comparative: mutation frequencies in CLASH reads versus mRNA-seq reads (Table 1), positional distributions of CIMs across small RNA target sites, Pearson correlations between CIM-containing hybrid abundance and total hybrid abundance, and box-plot comparisons of piRNA pairing ratios across CIM-stratified hybrid groups. Significance is reported as threshold indicators (* P<0.05; ** P<0.01) without naming the underlying test.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation (R, inferred from reported R values) | Correlation between CIM-containing hybrid abundance and total hybrid abundance per small RNA target site (Supplemental Fig. S2A–D) | Read counts per target site across ~4M PRG-1 and ~134K ALG-1 hybrid reads (exact site-level n not stated) | not stated |
| Unnamed significance test (P<0.05, P<0.01 thresholds only) | Differences in piRNA pairing ratios at seed and non-seed regions between CIM-position-stratified hybrids (Fig. 3D, 3E) | null | not stated |
| Frequency comparison (descriptive percentages, no formal test stated) | Mutation rates (deletion %, substitution %) in mRNA-seq vs PRG-1 CLASH vs ALG-1 iCLIP reads (Table 1) | mRNA-seq: 20,038,571 reads; PRG-1 non-hybrid: 4,053,251; ALG-1 non-hybrid: 15,107,177; PRG-1 hybrid: 4,017,410; ALG-1 hybrid: 133,974 | na |
| Positional distribution / count-based enrichment (descriptive, no formal test stated) | Distribution of CIM counts at each nucleotide position across piRNA and miRNA target sites (Fig. 2A, 2B) | null | na |
-
Correlations between CIM-containing hybrid abundance and total hybrid abundance are reported as Pearson R (inferred)↳ Could also: Spearman rank correlation could also be used — Hybrid read counts are count data with heavy right-skew; a rank-based measure requires no assumption of bivariate normality and is often preferred for sequencing count distributions
-
Differences in piRNA pairing ratios between CIM-stratified groups are tested with an unnamed significance test (P-value thresholds only)↳ Could also: Reporting the specific test (e.g. Wilcoxon rank-sum / Mann-Whitney U, or Kruskal-Wallis with post-hoc) by name would be standard practice — Naming the test and reporting exact P-values or test statistics allows readers to assess assumptions, reproducibility, and the magnitude of evidence; it is a standard requirement in most journals
-
Mutation frequency differences between library types (Table 1) are presented as descriptive percentages without a formal significance test↳ Could also: A chi-square test or Fisher's exact test on read counts could also be applied to each pairwise library comparison — With millions of reads the differences are visually large, but a formal test with an effect-size measure (e.g. odds ratio) would quantify the magnitude and provide a reproducible decision criterion
-
Multiple positional comparisons across up to ~21 piRNA/miRNA positions are performed without a stated correction for multiple testing↳ Could also: A Benjamini-Hochberg FDR correction or Bonferroni correction across the family of positional tests could also be applied — When significance is assessed at many positions simultaneously, an FDR or family-wise error rate correction is a standard way to control the expected proportion of false discoveries, which is especially relevant when individual P-values are reported
-
Positional enrichment of CIMs is described visually via count distributions (Fig. 2A, 2B)↳ Could also: A permutation or bootstrap test comparing the observed positional distribution to a null distribution (e.g. reads randomly sampled from non-CIM-containing hybrids) could also be used — A permutation-based approach would provide a formal statistical test of whether the observed peak positions differ from chance, complementing the visual inspection and providing a quantitative enrichment metric
-
Dispersion in box plots is shown via IQR and 5th–95th percentile whiskers; no confidence intervals are reported for group medians or proportions↳ Could also: Notched box plots or bootstrap 95% confidence intervals on group medians could also be reported — Confidence intervals on medians allow direct visual inference about whether group differences are likely to replicate, which is a common complement to or replacement for P-value thresholds, particularly for non-normally distributed data
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41330639 (MUTACLASH)
Paper: Wu WS, Lee DE, Chung CJ, Lu SY, Brown JS, Zhang D, Lee HC. "MUTACLASH: identifying functional small RNA target sites using crosslinking-induced mutations." RNA 2026. DOI 10.1261/rna.080482.125. PMCID PMC12810180.
Code: https://github.com/lu1215/MutaCLASH (commit 762accbbef8bd0d494923ebf7cef6f2b4bc066e2)
— bundles the full pipeline + the exact WS275 references + BWA + ChiRA + hyb tools.
Data (this RU): SRR3882949 = C. elegans ALG-1 iCLIP raw reads
(SRX1936747 / SAMN05382497; Illumina HiSeq 2500; 53,654,102 reads; 4.88 Gbp;
single-end). From Broughton et al. "Pairing Beyond the Seed…" (the dataset the
MUTACLASH paper re-analyses for the ALG-1 iCLIP rows of Table 1).
What MUTACLASH does (pipeline, from Methods + repo)
- Preprocess — adapter+quality trim (trim_galore v0.6.5
-q 30 --length 17 --max_length 70, auto-adapter; or cutadapt v2.10), then dedup (custom Python). - Hybrid id — ChiRA + BWA-MEM maps reads to regulator (miRNA/piRNA) + target (mRNA) references; splits chimeric (hybrid) vs single (non-hybrid) reads.
- Mutation detection — "mutated positions on mRNAs were obtained through
reading the MD tag with a custom Python script and SAMtools"
(
pipeline/find_deletion/bwa_find.py): per read, CIGARD= deletion, MD mismatch letters = substitution. - Binding prediction (pirScan / miRanda / RNAup), single-read CIMS stats, figures.
References bundled in repo: data/reference/mRNA_WS275.fa (43,040 C. elegans
WS275 transcripts), miRNA_WS275.fa (458 miRNAs).
In scope (pipeline-derived, low-hanging — what we reproduce)
Table 1 — mutation frequency of ALG-1 iCLIP, non-hybrid reads:
- reported deletion rate = 5.8 %
- reported substitution rate = 7.41 %
This is the cleanest pipeline-derived quantity tied to this RU's accession
(SRR3882949). It is a preprocessing/alignment-level statistic: the fraction of
non-chimeric reads that, after BWA-MEM alignment to the WS275 mRNA transcriptome,
carry ≥1 deletion (CIGAR D) / ≥1 substitution (MD/NM mismatch). We reproduce it by
running the paper's own aligner (BWA-MEM) on the paper's own reference
(mRNA_WS275.fa) with the paper's own trim parameters, then applying the paper's
own MD/CIGAR mutation definition.
Definition used (matches bwa_find.py): for each primary aligned read,
has_deletion = (D bases in CIGAR > 0); #substitutions = NM − (#inserted + #deleted bases) (NM/MD encode the same mismatch set the paper reads from MD),
has_substitution = (#substitutions > 0).
deletion% = reads_with_del / aligned_reads × 100; likewise substitution%.
Out of scope (the hard ~20%, not attempted — and why)
- Full ChiRA chimera calling + hybrid/non-hybrid split: the non-hybrid mutation rate is dominated by the bulk of iCLIP reads (overwhelmingly non-chimeric), so aligning all reads to mRNA is a faithful operationalisation of the "non-hybrid reads" row without the heavy/fragile ChiRA + perl-hyb stack.
- Binding-site prediction (pirScan/miRanda/RNAup), 22G-RNA regulatory effects (Figs 4,7), per-position CIM distribution (Fig 2): downstream, multi-tool, not a single pinnable number tied to this accession.
- PRG-1 CLASH / mRNA-seq rows of Table 1: different accessions, not this RU.
- Exact PCR-deduplication weighting: reported as a secondary bracket, not chased.
Possible-fabrication check
Table 1's ALG-1 row is derivable from the shipped reference + public SRA data + shipped MD-tag code → checkable. Our reproduced value is provisional; a human auditor decides the final grade.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Deletion (C1) reproduces well — 6.82–9.27% vs the reported 5.8% (Table 1), ~13x above the 0.026% mRNA-seq baseline — confirming the paper's central crosslinking-induced-deletion claim from shipped reference + public SRA + shipped code. Substitution (C2) over-counts ~3.5x (7.41% reported vs ~20–26% reproduced) but this sits on our side: we scored over the whole-transcriptome alignment instead of the paper's ChiRA-defined non-hybrid target set, and the gap is robust to MAPQ (28% at MAPQ≥30), ruling out a multimapping artifact. No fabrication concern — both numbers are derivable from the shared data/code; the substitution row simply requires the full ChiRA filtering we did not run. Net: solid, partial reproduction with an explainable scope-driven deviation.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.