NET-prism enables RNA polymerase-dedicated transcriptional interrogation at nucleotide resolution.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL 1:1 reproduction (honest). Clean from-scratch re-run after the «infra» work dir was janitored. NET-prism (Mylonas & Tessarz, RNA Biol 2019) reproduced via the documented third-party pipeline BradnerLab/netseq @ a7026da (P16-valid) on the paper's OWN real raw reads GSE107257/SRP125459 (TFIID factor, reps SRR6315706+SRR6315707) on «our HPC» SLURM. Pipeline rebuilt entirely: fresh conda env -> mm10 STAR index -> 6nt-UMI barcode strip -> STAR with the authors' Parameters.in -> deeptools 5'-end (--Offset 1) strand-separated 1bp RPGC bigwigs -> merge reps -> genome-wide binned correlation vs the deposited merged tracks («job» compute + 2219077 correlations). RESULTS: (C1) the regenerated track faithfully reproduces the deposited occupancy LANDSCAPE - same-strand Pearson(log1p) 0.845 FW / 0.651 RV @10kb, cross-strand ~0, strand convention confirmed; partial (not bit-exact) because the fragile legacy UMI dedup/RT-bias steps were omitted and RPGC was used vs raw counts, affecting only fine resolution (Spearman 0.62->0.35 from 10kb to 1kb). (C2) rep1-vs-rep2 Pearson(log1p) 0.96-0.98 strongly SUPPORTS the paper's qualitative 'highly reproducible among biological replicates' and supplies the unprinted number. (QC1) raw read counts match the ENA report EXACTLY for both reps (rep1 unique 12,937,470 = 18.15%, rep2 13,899,059 = 18.95%). VALIDATION: all 12 correlation values reproduce the earlier independent run bit-for-bit -> deterministic pipeline. AUDIT NOTE: the registry data_accession GSE90906 is WRONG - it is a different cited Tessarz FACT study (PMID 30456357); this paper's own data is GSE107257. NOT attempted (80/20 floor cleared, stated honestly): legacy per-base dedup/RT-bias steps, 6 of 7 factors, figure-level metagene/k-means/travelling-ratio (no pinnable scalar), wet-lab IP/MS. No fabrication concern: every reproduced value is derivable from the shipped raw data + documented pipeline.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 67assessed: 2026-06-20 ⛓ 0076d831d870
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetIt is unknown how transcription/elongation factors directly influence RNA Pol II pausing and directionality at single-nucleotide resolution genome-wide; the authors ask whether an antibody-based, polymerase-dedicated NET-seq approach (NET-prism) can resolve factor-specific effects on Pol II transcription dynamics.
- ★ NET-prism, an adapted NET-seq protocol using immunoprecipitation of Pol II-associated factors, enables strand-specific, nucleotide-resolution interrogation of transcription dynamics for any Pol II-interacting protein. method
- ★ Different transcription/elongation factors (Spt6, Ssrp1, TFIID, Sf1) establish unique, factor-specific patterns of RNA Pol II pausing and directionality over promoters, splice sites, and enhancers. finding
- ★ Sequential IP (Pol II then Ssrp1) shows nascent RNA recovered by NET-prism stems from direct Pol II-factor interaction and not from direct RNA binding by the factor itself. finding
- ★ Among tested factors, only the splicing factor Sf1 shows Pol II pausing at exon boundaries, while PIC components (TFIID) do not, supporting a kinetic model of transcription-splicing coupling. mechanism
- ★ NET-prism reveals distinctive Pol II topography at enhancers/super-enhancers, with Ssrp1 resembling initiation-like patterns and Spt6 resembling elongation-like patterns. finding
- ★ Mass spectrometry under NET-prism extraction conditions defines a Pol II protein interactome (elongation, splicing, and PIC factors) that serves as a resource to guide selection of proteins for NET-prism. resource
- IPs against Spt6, Ssrp1 and TFIID(TBP) also recover nascent transcripts from RNA Pol I and Pol III, suggesting NET-prism is extensible to all three polymerases. finding
- Mediator (Med14) shows minimal interaction with Pol II under NET-prism extraction conditions, serving as a validated negative control. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Mass spectrometry (IP-MS) | E14 mouse embryonic stem cells | Pol II IP vs Mock (IgG) IP | Pol II protein interactome enrichment | — |
| Western blot | E14 mouse embryonic stem cells | IP under NET-prism conditions | Confirmation of Pol II-interactor co-purification | — |
| NET-prism (NET-seq-based nascent RNA sequencing with factor IP) | E14 mouse embryonic stem cells | IP for Spt6, Ssrp1, TFIID (anti-TBP), or Med14 (negative control) | Nascent RNA density/Pol II position at nucleotide resolution over TSS, gene body, single genes | — |
| Sequential NET-prism (double IP) | E14 mouse embryonic stem cells | Pol II IP with CTD-peptide competitive elution, followed by Ssrp1 IP | Nascent RNA specificity for Pol II-Ssrp1 complexes | — |
| NET-prism | E14 mouse embryonic stem cells | IP for splicing factor Sf1 | Pol II pausing at intron-exon boundaries | — |
| NET-seq/prism travelling ratio and metaplot analysis | E14 mouse embryonic stem cells | Comparison across Spt6, Ssrp1, TFIID, Med14, total Pol II libraries | Travelling ratio (promoter vs gene body Pol II density) | — |
| NET-prism combined with ChIP-seq (H3K27Ac) correlation analysis | E14 mouse embryonic stem cells | IP for TFs/Pol II over distal enhancers and super-enhancers | Pol II stalling/transcriptional activity and TF ChIP-seq density over enhancers | — |
- ▲ MS identified positive (Supt5, Supt6, FACT, Paf1) and negative (NELF) elongation factors, splicing factors (Srsf5, Srsf6, Sf1), and TFIID components (Taf10, Taf15) as significantly enriched with Pol II FDR < 0.05
- – Spt6 and Ssrp1 IPs show strong, broad Pol II enrichment consistent with ChIP-seq; TFIID-bound Pol II shows a sharp TSS-centered signal; Med14 IP yields no nascent RNA
- – NET-prism libraries show different travelling ratios, indicating distinct pause-release dynamics depending on bound TF
- ▲ Spt6, Ssrp1 and TFIID(TBP) IPs also pull down nascent transcripts from Pol I and Pol III genes
- – Sequential IP (Pol II then Ssrp1) metagene and single-gene profiles closely match single IP profiles
- – Only Sf1 among tested factors shows pausing at exon boundaries similar to total Pol II; PIC components do not associate with pausing at splice sites
- ▲ Higher Pol II density observed over exons compared to introns across NET-seq/prism libraries
- ▲ Total Pol II and TFs show significantly higher ChIP-seq/NET-prism density over super-enhancers than distal enhancers; highest correlation between Total Pol II and Ssrp1 p < 1.0e-10 / p < 2.2e-16
- pvalue FDR < 0.05 (Significant Pol II interactors in MS IP vs Mock)
- pvalue p < 1.0e-10 (**) (Wilcoxon rank test, TF/Pol II density super- vs distal enhancers)
- pvalue p < 2.2e-16 (***) (Wilcoxon rank test, TF/Pol II density super- vs distal enhancers)
- count n = 4,314 (Protein-coding genes used for metaplot/heatmap over promoters)
- count n = 5,550 (Exon boundaries used in splicing pausing analysis)
- count n = 41,356 exons; n = 199,172 introns (Pol II coverage boxplot analysis (first/last exons excluded))
- other ~200-1000 ng nascent RNA (Typical nascent RNA yield per IP from 10^8 ES cells)
- count two biological replicates (Replicates processed per IP/library)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a methods-development paper introducing NET-prism, an adapted NET-seq immunoprecipitation protocol for strand-specific mapping of RNA Pol II at single-nucleotide resolution in complex with associated proteins. Statistical content is minimal and primarily descriptive: mass spectrometry protein enrichment was assessed by volcano plot with an FDR threshold; genomic coverage distributions were visualised with metaplots, heatmaps, and boxplots; Pearson correlations were used to compare library profiles; and a Wilcoxon rank test was applied to compare Pol II or ChIP-seq density over distal versus super-enhancers. Two biological replicates were generated per IP condition, and the travelling ratio was reported as a descriptive metric of promoter-proximal pausing.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| FDR-controlled enrichment (mass spectrometry) | Pol II IP vs. Mock (IgG) IP protein interactome; Figure 1b volcano plot | — | not stated |
| Wilcoxon rank test | ChIP-seq or NET-prism density over distal enhancers vs. super-enhancers; Figure 5b | — | not stated |
| Pearson correlation | Pairwise comparison of NET-seq/prism libraries over distal and super-enhancers; Figure 5a | — | not stated |
-
Two biological replicates were generated per IP condition↳ Could also: Three or more biological replicates per condition, analysed with tools such as DESeq2 or edgeR for count-based differential abundance — With n=2 replicates, dispersion estimation is very imprecise and formal between-condition comparisons have minimal power; additional replicates would enable principled statistical modelling of between-sample variability and genome-wide differential signal detection
-
Pearson correlation was used to compare NET-seq/prism libraries over enhancer regions (Figure 5a)↳ Could also: Spearman rank correlation — Sequencing coverage data are often heavily right-skewed with extreme outlier bins; Spearman correlation makes no distributional assumption and is more robust to influential high-coverage loci
-
Wilcoxon rank tests comparing density over distal vs. super-enhancers (Figure 5b) were applied separately for each IP library without a stated multiplicity correction↳ Could also: Apply a family-wise or FDR correction (e.g., Bonferroni or Benjamini-Hochberg) across the set of library comparisons — When the same contrast is repeated for several libraries, stating and applying a correction procedure makes the effective significance threshold explicit and reproducible
-
Wilcoxon test p-values were reported as threshold brackets (** p < 1.0e-10, *** p < 2.2e-16) rather than exact values↳ Could also: Report exact p-values and an effect size measure such as rank-biserial correlation — Exact p-values and effect sizes allow readers to gauge the magnitude of differences across studies and are increasingly requested under reproducibility guidelines
-
The travelling ratio was summarised by cumulative distribution plots (Figure 2c) and compared across libraries descriptively↳ Could also: Accompany cumulative distribution comparisons with a two-sample Kolmogorov-Smirnov test or Wilcoxon rank test between library pairs — A formal test of distributional differences would quantify whether the visual separation between libraries exceeds sampling variability
-
FDR < 0.05 was used as the significance threshold for mass spectrometry enrichment, but the software or algorithm used to compute the FDR was not described↳ Could also: Explicitly name the software (e.g., MaxQuant/Perseus, limma, msEmpiRe) and state the specific FDR estimation method and parameter settings — Reporting the exact tool and version allows readers to assess the underlying statistical assumptions and reproduce the enrichment calls
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-31156037 (NET-prism, Mylonas & Tessarz, RNA Biol 2019)
- PMID 31156037 · PMCID PMC6693550 · DOI 10.1080/15476286.2019.1621625
- Code (pipeline): https://github.com/BradnerLab/netseq @ commit
a7026da514ab7aed237b4b0bb2f8dc72d8716437(third-party NET-seq pipeline of the Bradner/Churchman labs; the paper states all NET-prism fastq were processed with these custom Python scripts — a P16 case: applying an existing tool to the paper's own data, equally valid.) - Data (this paper's own): GEO GSE107257 / SRA SRP125459 — 15 NET-prism
samples, Mus musculus, mm10. Deposited processed tracks = strand-separated,
per-factor merged bigwigs (
GSE107257_<factor>.NET-prism.merged.{fw,rv}.bw).
Registry correction (auditability note)
The scaffolded manifest.json lists data_accession: GSE90906. That is the wrong
accession for this paper. GSE90906 is a different Tessarz-lab study (FACT, PMID
30456357) that NET-prism merely cites as a referenced public dataset. NET-prism's
own data are GSE107257 (confirmed: GEO series title = the paper title; submitter
P. Tessarz). All reproduction uses GSE107257.
Pipeline (from Methods + repo)
extractMolecularBarcode.py— strip first 6 nt molecular barcode (UMI) into read header.- STAR align to mm10 with shipped
Parameters.in(EndToEnd, 3′ adapterATCTCGTATGCCGTCTTCTGCTTGclip, multimap≤101, alignIntronMin 11). - RT-bias removal (
extractReadsWithMismatchesIn6FirstNct.py), PCR-dup + splicing-intermediate removal (removePCRduplicateStringent.py,SI_coordinates_1based). - Coverage of read 5′ end (= Pol II position) → bigwig, 1 bp bin, strand-separated,
normalized to 1× depth via Deeptools with
--Offset 1(paper Methods). - Travelling ratio = Proximal-Promoter / Gene-Body (
travelRatio.r, repo PP/GB coords).
IN SCOPE (attempted)
- C1 (primary): Regenerate the deposited merged TFIID NET-prism bigwig (fw+rv)
from raw reads (SRR6315706 + SRR6315707) via STAR + deeptools
--Offset 1, and measure genome-wide binned correlation (Spearman/Pearson) vs the authors' depositedGSE107257_TFIID.NET-prism.merged.{fw,rv}.bw. The deposited track is the ground-truth pipeline output; high correlation = faithful reproduction of the documented track. - C2 (secondary): rep1-vs-rep2 genome-wide correlation of our regenerated tracks — quantifies the paper's qualitative claim "highly reproducible among biological replicates" (Supp Fig 2A; no numeric value is printed in the paper).
OUT OF SCOPE (not attempted — 80/20; stated honestly)
- Wet-lab: IP/MS (Fig 1b), RNA yields — not computational.
- Custom UMI PCR-duplicate + RT-bias removal steps: omitted for the primary track (the fragile, legacy 20%). Effect is on per-base peak height, not the genome-wide occupancy landscape that the binned correlation measures; noted as a deviation.
- Metagene/k-means heatmaps (Figs 2–5), travelling-ratio distributions, exon/intron splice-site profiles, enhancer Wilcoxon tests — figure-level, no single pinnable numeric value to compare 1:1; not attempted.
- Only TFIID (smallest factor, ~144M reads total) of the 7 factors is reproduced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
<synthetic>Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.