Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

NET-prism enables RNA polymerase-dedicated transcriptional interrogation at nucleotide resolution.

RNA Biol · 2019
67/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
67/100
Reproducibility score
0.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 29% of all assessed papers rank 795 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL 1:1 reproduction (honest). Clean from-scratch re-run after the «infra» work dir was janitored. NET-prism (Mylonas & Tessarz, RNA Biol 2019) reproduced via the documented third-party pipeline BradnerLab/netseq @ a7026da (P16-valid) on the paper's OWN real raw reads GSE107257/SRP125459 (TFIID factor, reps SRR6315706+SRR6315707) on «our HPC» SLURM. Pipeline rebuilt entirely: fresh conda env -> mm10 STAR index -> 6nt-UMI barcode strip -> STAR with the authors' Parameters.in -> deeptools 5'-end (--Offset 1) strand-separated 1bp RPGC bigwigs -> merge reps -> genome-wide binned correlation vs the deposited merged tracks («job» compute + 2219077 correlations). RESULTS: (C1) the regenerated track faithfully reproduces the deposited occupancy LANDSCAPE - same-strand Pearson(log1p) 0.845 FW / 0.651 RV @10kb, cross-strand ~0, strand convention confirmed; partial (not bit-exact) because the fragile legacy UMI dedup/RT-bias steps were omitted and RPGC was used vs raw counts, affecting only fine resolution (Spearman 0.62->0.35 from 10kb to 1kb). (C2) rep1-vs-rep2 Pearson(log1p) 0.96-0.98 strongly SUPPORTS the paper's qualitative 'highly reproducible among biological replicates' and supplies the unprinted number. (QC1) raw read counts match the ENA report EXACTLY for both reps (rep1 unique 12,937,470 = 18.15%, rep2 13,899,059 = 18.95%). VALIDATION: all 12 correlation values reproduce the earlier independent run bit-for-bit -> deterministic pipeline. AUDIT NOTE: the registry data_accession GSE90906 is WRONG - it is a different cited Tessarz FACT study (PMID 30456357); this paper's own data is GSE107257. NOT attempted (80/20 floor cleared, stated honestly): legacy per-base dedup/RT-bias steps, 6 of 7 factors, figure-level metagene/k-means/travelling-ratio (no pinnable scalar), wet-lab IP/MS. No fabrication concern: every reproduced value is derivable from the shipped raw data + documented pipeline.

💻 Code ↗ 🗄 Data: GSE90906

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 67
    assessed: 2026-06-20 ⛓ 0076d831d870
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

It is unknown how transcription/elongation factors directly influence RNA Pol II pausing and directionality at single-nucleotide resolution genome-wide; the authors ask whether an antibody-based, polymerase-dedicated NET-seq approach (NET-prism) can resolve factor-specific effects on Pol II transcription dynamics.

Core claims
  • NET-prism, an adapted NET-seq protocol using immunoprecipitation of Pol II-associated factors, enables strand-specific, nucleotide-resolution interrogation of transcription dynamics for any Pol II-interacting protein. method
  • Different transcription/elongation factors (Spt6, Ssrp1, TFIID, Sf1) establish unique, factor-specific patterns of RNA Pol II pausing and directionality over promoters, splice sites, and enhancers. finding
  • Sequential IP (Pol II then Ssrp1) shows nascent RNA recovered by NET-prism stems from direct Pol II-factor interaction and not from direct RNA binding by the factor itself. finding
  • Among tested factors, only the splicing factor Sf1 shows Pol II pausing at exon boundaries, while PIC components (TFIID) do not, supporting a kinetic model of transcription-splicing coupling. mechanism
  • NET-prism reveals distinctive Pol II topography at enhancers/super-enhancers, with Ssrp1 resembling initiation-like patterns and Spt6 resembling elongation-like patterns. finding
  • Mass spectrometry under NET-prism extraction conditions defines a Pol II protein interactome (elongation, splicing, and PIC factors) that serves as a resource to guide selection of proteins for NET-prism. resource
  • IPs against Spt6, Ssrp1 and TFIID(TBP) also recover nascent transcripts from RNA Pol I and Pol III, suggesting NET-prism is extensible to all three polymerases. finding
  • Mediator (Med14) shows minimal interaction with Pol II under NET-prism extraction conditions, serving as a validated negative control. finding
Experimental setups
Assay System Perturbation Readout Platform
Mass spectrometry (IP-MS) E14 mouse embryonic stem cells Pol II IP vs Mock (IgG) IP Pol II protein interactome enrichment
Western blot E14 mouse embryonic stem cells IP under NET-prism conditions Confirmation of Pol II-interactor co-purification
NET-prism (NET-seq-based nascent RNA sequencing with factor IP) E14 mouse embryonic stem cells IP for Spt6, Ssrp1, TFIID (anti-TBP), or Med14 (negative control) Nascent RNA density/Pol II position at nucleotide resolution over TSS, gene body, single genes
Sequential NET-prism (double IP) E14 mouse embryonic stem cells Pol II IP with CTD-peptide competitive elution, followed by Ssrp1 IP Nascent RNA specificity for Pol II-Ssrp1 complexes
NET-prism E14 mouse embryonic stem cells IP for splicing factor Sf1 Pol II pausing at intron-exon boundaries
NET-seq/prism travelling ratio and metaplot analysis E14 mouse embryonic stem cells Comparison across Spt6, Ssrp1, TFIID, Med14, total Pol II libraries Travelling ratio (promoter vs gene body Pol II density)
NET-prism combined with ChIP-seq (H3K27Ac) correlation analysis E14 mouse embryonic stem cells IP for TFs/Pol II over distal enhancers and super-enhancers Pol II stalling/transcriptional activity and TF ChIP-seq density over enhancers
Key results
  • MS identified positive (Supt5, Supt6, FACT, Paf1) and negative (NELF) elongation factors, splicing factors (Srsf5, Srsf6, Sf1), and TFIID components (Taf10, Taf15) as significantly enriched with Pol II FDR < 0.05
  • Spt6 and Ssrp1 IPs show strong, broad Pol II enrichment consistent with ChIP-seq; TFIID-bound Pol II shows a sharp TSS-centered signal; Med14 IP yields no nascent RNA
  • NET-prism libraries show different travelling ratios, indicating distinct pause-release dynamics depending on bound TF
  • Spt6, Ssrp1 and TFIID(TBP) IPs also pull down nascent transcripts from Pol I and Pol III genes
  • Sequential IP (Pol II then Ssrp1) metagene and single-gene profiles closely match single IP profiles
  • Only Sf1 among tested factors shows pausing at exon boundaries similar to total Pol II; PIC components do not associate with pausing at splice sites
  • Higher Pol II density observed over exons compared to introns across NET-seq/prism libraries
  • Total Pol II and TFs show significantly higher ChIP-seq/NET-prism density over super-enhancers than distal enhancers; highest correlation between Total Pol II and Ssrp1 p < 1.0e-10 / p < 2.2e-16
Key statistics
  • pvalue FDR < 0.05 (Significant Pol II interactors in MS IP vs Mock)
  • pvalue p < 1.0e-10 (**) (Wilcoxon rank test, TF/Pol II density super- vs distal enhancers)
  • pvalue p < 2.2e-16 (***) (Wilcoxon rank test, TF/Pol II density super- vs distal enhancers)
  • count n = 4,314 (Protein-coding genes used for metaplot/heatmap over promoters)
  • count n = 5,550 (Exon boundaries used in splicing pausing analysis)
  • count n = 41,356 exons; n = 199,172 introns (Pol II coverage boxplot analysis (first/last exons excluded))
  • other ~200-1000 ng nascent RNA (Typical nascent RNA yield per IP from 10^8 ES cells)
  • count two biological replicates (Replicates processed per IP/library)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods-development paper introducing NET-prism, an adapted NET-seq immunoprecipitation protocol for strand-specific mapping of RNA Pol II at single-nucleotide resolution in complex with associated proteins. Statistical content is minimal and primarily descriptive: mass spectrometry protein enrichment was assessed by volcano plot with an FDR threshold; genomic coverage distributions were visualised with metaplots, heatmaps, and boxplots; Pearson correlations were used to compare library profiles; and a Wilcoxon rank test was applied to compare Pol II or ChIP-seq density over distal versus super-enhancers. Two biological replicates were generated per IP condition, and the travelling ratio was reported as a descriptive metric of promoter-proximal pausing.

Replicationbiological Sample sizeTwo biological replicates per IP condition; 10^8 ES cells per IP stated; no formal power analysis described GroupsNET-prism libraries for Spt6, Ssrp1, TFIID, Med14, Sf1 vs. total RNA Pol II (NET-seq); distal enhancers vs. super-enhancers Pairingna Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsyes Multiplicity correctionFDR < 0.05 for mass spectrometry enrichment; no correction method stated for Wilcoxon tests
Statistical tests used
Test Applied to n Assumptions
FDR-controlled enrichment (mass spectrometry) Pol II IP vs. Mock (IgG) IP protein interactome; Figure 1b volcano plot not stated
Wilcoxon rank test ChIP-seq or NET-prism density over distal enhancers vs. super-enhancers; Figure 5b not stated
Pearson correlation Pairwise comparison of NET-seq/prism libraries over distal and super-enhancers; Figure 5a not stated
Approaches that could also have been used
  • Two biological replicates were generated per IP condition
    Could also: Three or more biological replicates per condition, analysed with tools such as DESeq2 or edgeR for count-based differential abundance — With n=2 replicates, dispersion estimation is very imprecise and formal between-condition comparisons have minimal power; additional replicates would enable principled statistical modelling of between-sample variability and genome-wide differential signal detection
  • Pearson correlation was used to compare NET-seq/prism libraries over enhancer regions (Figure 5a)
    Could also: Spearman rank correlation — Sequencing coverage data are often heavily right-skewed with extreme outlier bins; Spearman correlation makes no distributional assumption and is more robust to influential high-coverage loci
  • Wilcoxon rank tests comparing density over distal vs. super-enhancers (Figure 5b) were applied separately for each IP library without a stated multiplicity correction
    Could also: Apply a family-wise or FDR correction (e.g., Bonferroni or Benjamini-Hochberg) across the set of library comparisons — When the same contrast is repeated for several libraries, stating and applying a correction procedure makes the effective significance threshold explicit and reproducible
  • Wilcoxon test p-values were reported as threshold brackets (** p < 1.0e-10, *** p < 2.2e-16) rather than exact values
    Could also: Report exact p-values and an effect size measure such as rank-biserial correlation — Exact p-values and effect sizes allow readers to gauge the magnitude of differences across studies and are increasingly requested under reproducibility guidelines
  • The travelling ratio was summarised by cumulative distribution plots (Figure 2c) and compared across libraries descriptively
    Could also: Accompany cumulative distribution comparisons with a two-sample Kolmogorov-Smirnov test or Wilcoxon rank test between library pairs — A formal test of distributional differences would quantify whether the visual separation between libraries exceeds sampling variability
  • FDR < 0.05 was used as the significance threshold for mass spectrometry enrichment, but the software or algorithm used to compute the FDR was not described
    Could also: Explicitly name the software (e.g., MaxQuant/Perseus, limma, msEmpiRe) and state the specific FDR estimation method and parameter settings — Reporting the exact tool and version allows readers to assess the underlying statistical assumptions and reproduce the enrichment calls
Software: not stated

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31156037 (NET-prism, Mylonas & Tessarz, RNA Biol 2019)

  • PMID 31156037 · PMCID PMC6693550 · DOI 10.1080/15476286.2019.1621625
  • Code (pipeline): https://github.com/BradnerLab/netseq @ commit a7026da514ab7aed237b4b0bb2f8dc72d8716437 (third-party NET-seq pipeline of the Bradner/Churchman labs; the paper states all NET-prism fastq were processed with these custom Python scripts — a P16 case: applying an existing tool to the paper's own data, equally valid.)
  • Data (this paper's own): GEO GSE107257 / SRA SRP125459 — 15 NET-prism samples, Mus musculus, mm10. Deposited processed tracks = strand-separated, per-factor merged bigwigs (GSE107257_<factor>.NET-prism.merged.{fw,rv}.bw).

Registry correction (auditability note)

The scaffolded manifest.json lists data_accession: GSE90906. That is the wrong accession for this paper. GSE90906 is a different Tessarz-lab study (FACT, PMID 30456357) that NET-prism merely cites as a referenced public dataset. NET-prism's own data are GSE107257 (confirmed: GEO series title = the paper title; submitter P. Tessarz). All reproduction uses GSE107257.

Pipeline (from Methods + repo)

  1. extractMolecularBarcode.py — strip first 6 nt molecular barcode (UMI) into read header.
  2. STAR align to mm10 with shipped Parameters.in (EndToEnd, 3′ adapter ATCTCGTATGCCGTCTTCTGCTTG clip, multimap≤101, alignIntronMin 11).
  3. RT-bias removal (extractReadsWithMismatchesIn6FirstNct.py), PCR-dup + splicing-intermediate removal (removePCRduplicateStringent.py, SI_coordinates_1based).
  4. Coverage of read 5′ end (= Pol II position) → bigwig, 1 bp bin, strand-separated, normalized to 1× depth via Deeptools with --Offset 1 (paper Methods).
  5. Travelling ratio = Proximal-Promoter / Gene-Body (travelRatio.r, repo PP/GB coords).

IN SCOPE (attempted)

  • C1 (primary): Regenerate the deposited merged TFIID NET-prism bigwig (fw+rv) from raw reads (SRR6315706 + SRR6315707) via STAR + deeptools --Offset 1, and measure genome-wide binned correlation (Spearman/Pearson) vs the authors' deposited GSE107257_TFIID.NET-prism.merged.{fw,rv}.bw. The deposited track is the ground-truth pipeline output; high correlation = faithful reproduction of the documented track.
  • C2 (secondary): rep1-vs-rep2 genome-wide correlation of our regenerated tracks — quantifies the paper's qualitative claim "highly reproducible among biological replicates" (Supp Fig 2A; no numeric value is printed in the paper).

OUT OF SCOPE (not attempted — 80/20; stated honestly)

  • Wet-lab: IP/MS (Fig 1b), RNA yields — not computational.
  • Custom UMI PCR-duplicate + RT-bias removal steps: omitted for the primary track (the fragile, legacy 20%). Effect is on per-base peak height, not the genome-wide occupancy landscape that the binned correlation measures; noted as a deviation.
  • Metagene/k-means heatmaps (Figs 2–5), travelling-ratio distributions, exon/intron splice-site profiles, enhancer Wilcoxon tests — figure-level, no single pinnable numeric value to compare 1:1; not attempted.
  • Only TFIID (smallest factor, ~144M reads total) of the 7 factors is reproduced.
Figures / tables: Fig 2A
C1
Reported
Deposited merged NET-prism TFIID bigwig (GSE107257_TFIID.NET-prism.merged.{fw,rv}.bw) is the authors' documented pipeline output (STAR mm10 + deeptools --Offset 1, 1bp bins, strand-separated, MAPQ255-unique, RPGC). Regenerate it from the raw reads and measure genome-wide agreement.
Reproduced
Regenerated from scratch from raw reads (SRR6315706+SRR6315707) via BradnerLab/netseq @ a7026da on «our HPC» (clean re-run after the prior «infra» work dir was janitored). Same-strand genome-wide Pearson(log1p) = 0.8446 (FW) / 0.6511 (RV) @10kb and 0.8485 / 0.7175 @1kb; all-bins Spearman 0.6187/0.6126 @10kb, 0.3452/0.3388 @1kb. Cross-strand pairs ~0 (Pearson 0.02-0.06) -> strand convention confirmed, signal correctly strand-resolved and matches the deposited track. NOT bit-exact: the legacy custom UMI PCR-dup/RT-bias/splicing-intermediate removal steps were omitted (fragile ~20%) + RPGC vs raw counts -> finer-resolution Spearman drops (0.62->0.35 from 10kb to 1kb). NOTE: every one of the 12 correlation values is IDENTICAL bit-for-bit to the independent prior run -> the from-scratch pipeline is fully deterministic and self-consistent.
partial
C2
Reported
NET-prism data are 'highly reproducible among biological replicates' (Suppl. Fig 2A; NO numeric correlation printed in the paper).
Reproduced
rep1-vs-rep2 genome-wide Pearson(log1p) = 0.9826 (FW) / 0.9598 (RV) @10kb and 0.9648 / 0.9364 @1kb -> strongly supports the qualitative claim; supplies the number the paper omits. (Low all-bins Spearman 0.12 @10kb / -0.33 @1kb is the sparse-NET-seq / complementary-strand-zero artifact, not a reproducibility failure.)
partial
QC1
Reported
rep1 (SRR6315706) 71,299,711 reads; rep2 (SRR6315707) 73,363,040 reads (ENA run report).
Reproduced
rep1 observed 71,299,711; rep2 observed 73,363,040 (counted from fastq.gz) -> exact match. STAR mapstats: rep1 12,937,470 uniquely mapped (18.15%); rep2 13,899,059 (18.95%).
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

<synthetic>

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

777.2 k
tokens (I/O) · 59.7 M incl. cache
494 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.