Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Deep transcriptomics reveals cell-specific isoforms of pan-neuronal genes.

Nat Commun · 2025
L1 84/100 3/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

salvaged by watchdog from agreement.json (agent omitted ROOM_RESULT.json)

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-07-29

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can deep, replicate single-neuron-type transcriptomes from C. elegans reveal cell-specific alternative splicing patterns (including in pan-neuronal genes) and the regulatory factors that establish them, overcoming the low capture/sensitivity limits of conventional single-cell RNA-seq?

Core claims
  • Pan-neuronal genes (expressed in many/all neurons) harbor highly cell-specific splice variants/isoforms restricted to single or few neuron types. finding
  • Differential intron retention is widespread across neuron types, accounting for ~52% of all differential splicing — an under-appreciated source of cell-specific gene regulation. finding
  • A 'uniqueness index' algorithm aggregating significant differential splicing across all pairwise cell comparisons identifies cell-specific isoforms and candidate regulatory factors. method
  • Gene expression and alternative splicing are globally distinct and regulated orthogonally across neuronal cell types (poor correlation between the two). finding
  • Three distinct splicing factors (RNA binding proteins) are employed in vivo to control splicing in a single neuron. mechanism
  • Deep CeNGEN transcriptomes with biological replicates enable high-confidence splicing analysis even for lowly-expressed genes due to uniform gene-body coverage. resource
  • A user-friendly platform was developed for spatial transcriptomic visualization of splicing patterns at single-neuron resolution. resource
  • Different types of alternative splicing are correlated with each other across cell types (concerted regulation), except intron retention which correlates less. finding
Experimental setups
Assay System Perturbation Readout Platform
Short-read bulk RNA-seq of sorted neuron-type populations (rRNA-depleted) C. elegans isolated neuron types (46 of 118 anatomically-distinct types, CeNGEN consortium) none (neuron-type-specific fluorescent transgene labeling/sorting) differential gene expression and alternative splicing (PSI/ΔPSI) across neuron types Illumina short-read sequencing
Whole-animal polyA-selected RNA-seq (validation) C. elegans wild-type whole animals none intron retention levels to confirm cell-type observations
RT-PCR (orthogonal validation of intron retention) C. elegans whole-animal population RNA none presence of intron-skipped and intron-retained products for 5 retained introns
Transgenic fluorescent splicing reporter (GFP/RFP/BFP imaging) C. elegans neurons (pan-neuronal rgef-1 promoter; cholinergic unc-17 promoter BFP) transgene/overexpression (dgk-1 GFP/RFP splicing reporter) in vivo cell-specific splice-site selection (upstream RFP vs downstream GFP) in motor neurons
In silico differential splicing analysis (JUM) and differential expression (DESeq2) C. elegans neuron-type transcriptomes (all pairwise comparisons) none cassette exons, 5'/3' splice sites, intron retention, composite events; uniqueness index JUM, DESeq2
Key results
  • Intron retention accounts for over half of all differential splicing across neuron types 52%
  • Differential intron retention is >7-fold more prevalent than differential cassette exon splicing 7-fold
  • Total alternative splicing events detected across genes 15,515 events in 5779 genes
  • AVM vs AVL comparison yields differential cassette exons; 2070 total pairwise comparisons performed 75 cassette exons
  • unc-31 exon 3 inclusion varies continuously across neurons (AVL 78%, AVM 53%, AVG 1%); uniqueness index 23 in AVL 78% vs 53% vs 1%; index=23
  • pct-1 intron consistently spliced in most neurons but retained in CAN neuron 69% retained in CAN vs ~0%
  • mec-2/Stoml3 has highest 3' splice-site uniqueness; touch neurons (AVM/PVM) select downstream 3' splice site, others upstream
  • dgk-1 alternative 3' splice site uniquely used in excitatory motor neurons (DA/VB) but not inhibitory (VD/DD), confirmed in vivo by reporter isoforms of 67 vs 31 aa
Key statistics
  • count 15,515 alternative splicing events (total AS events in distinct genes)
  • count 2070 pairwise comparisons (all pairwise neuron-type comparisons)
  • count 75 differential cassette exons (AVM vs AVL comparison at |ΔPSI|>10%, q<0.05)
  • other 52% (fraction of differential splicing that is intron retention (5' SS 27%, 3' SS 11%))
  • fold_change 7-fold (intron retention more prevalent than cassette exon splicing)
  • other 69% retained in CAN vs ~0% elsewhere (pct-1 intron retention; n=3,4,4,3,4,4,4 replicates)
  • other ASK 72% retained vs RIC 10% (pqn-53 intron retention across neuron types)
  • other uniqueness index range -45 to +45 (=23 for unc-31 in AVL) (single neuron compared against 45 others)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study performed pairwise RNA-seq comparisons across 46 C. elegans neuron types using deep bulk transcriptomes with biological replicates generated by the CeNGEN consortium. Differential gene expression was assessed with DESeq2 and differential alternative splicing with JUM, applying ΔPSI > 10% and FDR q < 0.05 thresholds across all 2070 pairwise cell-type combinations. A custom 'uniqueness index' aggregated significant ΔPSI values across pairwise comparisons to identify cell-specific splicing events, and relationships between regulatory layers were summarized with adjusted R-squared from linear regression.

Replicationbiological Sample sizen=3–4 biological replicates per neuron type stated in figure legends; no formal power analysis or sample-size justification described Groups46 C. elegans neuron types compared in all pairwise combinations (2070 comparisons) for both gene expression and alternative splicing Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionFDR q-value (specific algorithm, e.g. Benjamini-Hochberg, not explicitly named; DESeq2 also applies its own internal FDR)
Statistical tests used
Test Applied to n Assumptions
DESeq2 (negative binomial Wald test with internal FDR correction) Pairwise differential gene expression across all 2070 combinations of 46 neuron types 46 neuron types; n=3–4 biological replicates per neuron type (exact per-comparison n not uniformly stated in text) not stated
JUM (Junction Usage Model) PSI-based differential splicing with FDR q-value Pairwise differential alternative splicing (cassette exons, intron retention, 5′ and 3′ splice sites, composite events) across all 2070 neuron-type pairs; thresholds: |ΔPSI| > 10%, q < 0.05 Minimum 5 junction-spanning reads per biological replicate; n=3–4 replicates per neuron type not stated
Custom uniqueness index (sum of all significant ΔPSI values from pairwise comparisons for one cell vs. all others) Ranking cell-specific splicing events for each of the 46 neuron types across all alternative splicing classes 45 pairwise comparisons per focal neuron (maximum index value ±45) na
Linear regression / correlation (adjusted R-squared with 95% CI) Correlating magnitude of differential splicing across types (e.g., cassette exons vs. 5′ splice sites) and gene expression vs. splicing differences, one point per cell-type pair 2070 pairwise comparisons as data points not stated
RT-PCR (qualitative band detection, orthogonal validation) Validation of five selected retained introns in whole-animal RNA null na
Approaches that could also have been used
  • Differential alternative splicing was quantified pairwise using JUM with PSI-based statistics
    Could also: rMATS, SUPPA2, or LeafCutter could also be applied for differential splicing from short-read RNA-seq — These tools implement distinct statistical models (e.g., likelihood-ratio tests, Dirichlet-multinomial) that handle read-count variability differently; using a second tool in parallel is a common robustness check and can increase confidence in shared findings
  • Replicate spread for PSI values was displayed as SEM
    Could also: SD or 95% CI could also be used to summarize biological variability across replicates — With n=3–4 replicates, SD directly conveys biological variability without scaling by sample size; 95% CI communicates estimation uncertainty; both are frequently preferred over SEM for small n because SEM shrinks with n in a way that can visually understate true spread
  • Relationships between splicing-type magnitudes and between expression and splicing were summarized with adjusted R-squared from linear regression
    Could also: Spearman rank correlation could also be used to characterize monotonic association between these quantities — Spearman correlation makes no linearity or normality assumption and is more robust to outliers, which may be relevant when ΔPSI distributions are bounded (0–100%) and potentially skewed across 2070 pairwise comparisons
  • A custom uniqueness index (summed significant ΔPSI) was used to rank cell-specific splicing events
    Could also: Published cell-type specificity metrics such as the tau index, a Jensen-Shannon divergence-based score, or a one-vs-rest linear model contrast could also rank specificity — Established metrics have known statistical properties and prior benchmarks, which can facilitate cross-study comparison and help calibrate threshold selection; they also handle missing cell-type data in defined ways
  • Intron retention events were orthogonally validated with RT-PCR on five selected events from whole-animal RNA
    Could also: Long-read RNA-seq (Nanopore or PacBio) on sorted neuronal populations could also validate isoform structure — Long reads capture complete isoform structure in a single read, which is particularly informative for composite splicing events where multiple choices occur on the same transcript; this would complement short-read PSI estimates with direct isoform-level evidence
  • No formal power analysis or sample-size justification for n=3–4 biological replicates was reported
    Could also: A simulation-based sensitivity analysis or post-hoc estimation of minimum detectable ΔPSI at the observed read depths and replicate variance could also be reported — Reporting the detectable effect size at the observed n and variance helps readers interpret negative results and calibrate confidence in the completeness of the differential splicing catalog
Software: DESeq2 · JUM (Junction Usage Model)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
7
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40379625

Paper: Wolfe Z, Liska D, Norris A. Deep transcriptomics reveals cell-specific isoforms of pan-neuronal genes. Nat Commun 16, 2025. DOI 10.1038/s41467-025-58296-2. Repo: https://github.com/xcwolfe/Differential-Expression-in-C-elegans @ 4528b8c (own code, R-markdown pipeline scripts 1–13). Data: CeNGEN bulk neuron RNA-seq, SRA PRJNA952691 (Barrett et al. 2022).

What the repo actually ships (decides reproducibility)

  • Barrett_et_al_2022_CeNGEN_bulk_RNAseq_data.csv — a count matrix (genes × replicates). VERIFIED: 31173 genes × 160 replicates spanning 41 neuron types — NOT the 46 types / 180 replicates the paper analyzed. Missing 5 types: AIM, AVL, DVC, LUA, SMB. README calls it a "sample count matrix … for all 46 cell types," but it is a 41-type/160-rep subset.
  • genenames.csv (31173 WBGene IDs), simplemine_results.csv (484 RBPs), experimental_conditions_46_neurons.csv (180 rows, the full design table), GTF + refFlat annotations.
  • R-markdown scripts 1–13 (DESeq2/PCA/UMAP, RBP search, JUM/DCC splicing analysis).
  • Does NOT ship: the full 46-type matrix, any JUM/STAR output, BAM/fastq, single-cell data. Splicing results require regenerating alignments from SRA.

In scope (pipeline-derived, reproducible from shipped data)

# Result Pipeline Script Feasibility
C1 Count matrix dimensions (genes, replicates, neuron types) data inventory repo files trivial, done
C2 RBP gene list size (484) WormBase list simplemine_results.csv trivial, done
C3 PCA separates neuron classes; PC1/PC2 variance prcomp/DESeq2 vst+plotPCA 1, 3 runs on shipped matrix
C4 DESeq2 type-vs-type differential expression → DE-gene-count matrix ( log2FC >2, padj<0.01) DESeq2 ~type
C5 # pairwise comparisons = n·(n−1) arithmetic 2 trivial

Heavy tier (gradeable headline numbers, require full SRA re-processing)

# Result (paper) Pipeline Why heavy
H1 15,515 AS events in 5779 genes STAR → JUM over 2070 pairwise comparisons author's own "~48 h / ~3300 files megaloop"; needs all 180 deep transcriptomes from SRA + JUM tool
H2 52% intron retention / 27% A5S / 11% A3S JUM ΔPSI categorisation aggregate over all comparisons
H3 352 pan-neuronal genes at median cutoff single-cell (PRJEB22693) + JUM needs non-shipped single-cell data

Out of scope (wet-lab / manual / external — not attempted)

RT-PCR validation of 5 introns, transgenic two-color splicing reporters, VISTA-SPLICE Tableau dashboard, amino-acid/reading-frame manual annotation, PhyloP conservation (author states script 13 was not in the final manuscript).

Strategy

Phase 1 (core): run the authors' DESeq2/PCA code on the shipped 41-type matrix; grade structural facts (C1,C2,C5) and the DE/PCA pipeline (C3,C4); honestly record the 41/160 vs 46/180 data-scope gap. Phase 2 (stretch): assess/attempt a scoped STAR+JUM splicing run from SRA. No fabrication; honesty over coverage.

Figures / tables: Fig 1Fig 2
C1
Reported
46
Reproduced
41 (shipped-matrix subset)
within tolerance
C2
Reported
180
Reproduced
160
within tolerance
C3
Reported
2070=46x45
Reproduced
1640=41x40 (formula exact)
within tolerance
C4
Reported
31173
Reproduced
31173
exact
C5
Reported
484
Reproduced
484
exact
C6
Reported
qualitative (no % in paper)
Reproduced
PC1=18.07%,PC2=7.64%,NN-same-type=0.562
within tolerance
C7
Reported
heatmap, no number
Reproduced
DE-gene counts per pair produced (e.g. ADLvsASG=1394)
within tolerance
H1
Reported
15515/5779
Reproduced
not reproduced (aggregate); scoped JUM attempted
m.public.grade.uncheckable
H2
Reported
52/27/11
Reproduced
scoped ADL-vs-DD JUM (qualitative IR-dominance test)
partial
H3
Reported
352
Reproduced
not reproduced (needs non-shipped single-cell data)
m.public.grade.uncheckable

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

The authors' DESeq2+PCA count-matrix pipeline reproduces cleanly on the shipped data, with C4=31173 genes and C5=484 RBPs exact and PCA/DE confirming the qualitative 'neurons separate by type' premise. The main deviation is a data-scope gap: the repo ships a 41-type/160-replicate subset (README claims all 46), so 46/180/2070 become 41/160/1640 — an input/availability issue, partly the authors' incomplete deposit. The title-level splicing claims (15515 AS events, 52/27/11%, 352 pan-neuronal genes) were not reproduced because they require ~20.4 TB SRA regeneration and non-shipped single-cell data — a feasibility/availability blocker, not a contradiction. Net: solid where computable, incomplete where data/compute were unavailable; no fabrication signal.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

399.7 k
tokens (I/O) · 62.3 M incl. cache
368 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.