Analysis of subcellular transcriptomes by RNA proximity labeling with Halo-seq.
The main results reproduced, with only marginal, non-material deviations.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL 1:1 reproduction of the quantitation->statistics half of the Halo-seq pipeline on the paper's OWN data (GSE172281; the brief's GSE116008 is the external APEX-seq comparison). Ran tximport->DESeq2 (MLE LFC, padj<0.05 & |log2FC|>=0.5, min-5-counts-all-samples, GENCODE v28) on the authors' deposited Salmon quant.sf, all on «our HPC»/«infra». RESULTS: Fibrillarin 684 enr / 445 dep vs reported 602/338 (partial, correct magnitude; confirmed unshrunken LFC since ashr->314 overshoots low); H2B 320 enr & p65 160 enr -> 'hundreds' (match); H2B-Fib enriched overlap 22 vs 20 (within-tol) with binomial P=0.099 not-significant vs reported P=0.2 not-significant (SAME conclusion). C4 LMB via Xtail 1.2.0 (built from source): 18 genes more enriched in H2B+LMB pulldown (FDR<0.1, all direction TE>0) vs reported 105 -> MISMATCH; 105 not recoverable at any reasonable FDR (v1=21,v2=45,final=18; FDR<0.3 ->40). KEY AUDIT NOTE: rnabioco/rnaroids ships ONLY the upstream Salmon/STAR/kallisto Snakemake pipeline -- NO downstream DESeq2/Xtail statistics scripts -- so exact reported counts depend on unspecified filter/threshold/annotation choices; C4's 105 is not derivable from the shipped data with the described method (possible-fabrication flag for human review). C1-C3 reproduce in direction+magnitude (no strong concern). NOT ATTEMPTED (out of scope, not pipeline-derived): wet-lab singlet-oxygen yields / dot blots / microscopy / labeling-efficiency gels; gene-class enrichment (snRNA/lncRNA/ARE/HuR) and GO lists (secondary, external annotations).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 56assessed: 2026-06-21 ⛓ b908c9da5ba1
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether a nonenzymatic, Halo-DBF-based RNA proximity labeling method (Halo-seq) can efficiently and sensitively characterize subcellular transcriptomes and reveal mechanistic insights into the RNA sequence features and regulatory factors controlling subcellular RNA localization.
- ★ Halo-seq pairs a light-activatable Halo-DBF ligand with Click chemistry to label and purify spatially defined RNA populations in living cells with high spatial specificity (~100 nm radius) method
- ★ Halo-seq displays higher RNA labeling efficiency than comparable proximity labeling methods (miniSOG2 and APEX2) finding
- ★ RNAs containing AU-rich elements are relatively enriched in the nucleus finding
- ★ Nuclear enrichment of AU-rich element-containing RNAs becomes stronger upon treatment with the nuclear export inhibitor leptomycin B finding
- ★ Halo-seq data expand the role of HuR in mediating nuclear RNA export and define a comprehensive set of HuR-dependent export transcripts mechanism
- ★ Halo-seq was used to quantify nuclear, nucleolar, and cytoplasmic transcriptomes and characterize their dynamic changes following perturbation finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA proximity labeling + RNA-seq (streptavidin pulldown vs input) | HeLa cells expressing H2B-Halo, Halo-p65, or Halo-fibrillarin | Halo-DBF + green light labeling | gene-level enrichment of pulldown vs input RNA (DESeq2) | NovaSeq (Illumina) |
| RNA dot blot | HeLa cells expressing Halo-p65 | DBF ligand present/absent, green light exposure time (0/1/5 min) | biotinylation signal via streptavidin-HRP | Sapphire molecular imager (Azure Biosystems) |
| Comparative RNA labeling assay (dot blot) | HeLa cells expressing HA-miniSOG2 vs Halo-HA | blue light (miniSOG2) vs green light (Halo-DBF) | relative RNA biotinylation efficiency | — |
| Comparative RNA labeling assay (dot blot) | HeLa cells expressing Halo-APEX2 | H2O2 (APEX2 labeling) vs Halo-DBF/light | relative RNA biotinylation efficiency | — |
| Fluorescence microscopy of Halo fusion localization | HeLa cells expressing H2B-Halo, Halo-p65, Halo-fibrillarin | none | subcellular localization via fluorescent Halo ligand | Deltavision Elite widefield fluorescence microscope |
| In situ Click imaging (Cy5-azide) | HeLa cells expressing Halo fusions | DBF/light labeling with/without Cy5 or DBF controls | subcellular location of alkynylated RNA | Deltavision Elite widefield fluorescence microscope |
| Halo-seq with drug perturbation, analyzed via Xtail | HeLa cells expressing H2B-Halo | Leptomycin B (40 ng/ml, 15 h) vs untreated | change in nuclear pulldown/input RNA ratio | NovaSeq (Illumina) |
| Comparison to external RNA-seq datasets | APEX-seq and CeFra-seq public datasets | none | cross-method comparison of subcellular RNA enrichment | — |
- ▲ Halo-seq showed higher RNA labeling efficiency than miniSOG2 and APEX2 in side-by-side comparisons
- ▲ AU-rich element-containing transcripts are enriched in the nuclear (H2B-Halo pulldown) transcriptome
- ▲ Leptomycin B treatment increases the nuclear enrichment of AU-rich element-containing RNAs relative to untreated cells
- – Halo-p65 fusion relocalizes from cytoplasm to nucleus following LMB treatment, confirming LMB activity
- ▲ Biotinylation of RNA in the Click reaction is dependent on prior addition of DBF Halo ligand
- – Typical streptavidin pulldown recovers a small fraction of input RNA, varying by Halo fusion location 0.5-5% of input RNA
- pvalue adjusted P < 0.05 (threshold for calling a gene enriched/depleted in pulldown vs input)
- fold_change absolute log2 fold change >= 0.5 (threshold for calling a gene enriched/depleted in pulldown vs input)
- count 0.5-5% (percentage of input RNA recovered by streptavidin pulldown, depending on Halo fusion)
- count 20-40 million read pairs (typical sequencing depth per RNAseq sample)
- count 25-100 μg (10 cm dish) / 300-700 μg (15 cm dish) (typical total RNA yields from Halo-seq labeling)
- other ~100 nm (estimated diffusion radius of DBF-generated singlet oxygen radicals)
- other 40 ng/ml, 15 h (leptomycin B treatment dose and duration prior to Halo-seq)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
Halo-seq characterizes spatially defined RNA populations by proximity labeling, streptavidin pulldown, and rRNA-depleted RNA-seq. Transcript abundances were quantified with Salmon, aggregated to gene level with tximport, and differential enrichment between streptavidin pulldown and input RNA was assessed with DESeq2 (adjusted P < 0.05, |log2FC| ≥ 0.5, minimum 5 counts across all samples). Changes in pulldown-to-input enrichment ratios between leptomycin B-treated and untreated conditions were identified with Xtail, a ratio-based tool originally designed for ribosome profiling data.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 Wald test with Benjamini-Hochberg FDR | Comparison of gene abundance in streptavidin pulldown versus input RNA to identify spatially enriched or depleted transcripts (nuclear, nucleolar, cytoplasmic Halo-seq experiments) | — | not stated |
| Xtail ratio-of-ratios analysis | Identification of genes whose pulldown-to-input (nuclear enrichment) ratio changed between leptomycin B-treated and untreated conditions | — | not stated |
-
The number of biological replicates used in DESeq2 analyses is not stated in the methods text↳ Could also: Explicitly reporting n biological replicates and including a brief sample-size or power justification would also be standard practice for RNA-seq studies — DESeq2's variance estimation is sensitive to replicate number; stating n allows readers to assess the reliability of dispersion estimates and the generalizability of findings
-
Xtail, a tool originally developed for ribosome profiling, was repurposed to detect condition-dependent changes in pulldown-to-input ratios↳ Could also: A DESeq2 interaction-term model (condition × sample-type) could also test whether the enrichment ratio differs by condition within a single unified statistical framework — An interaction model explicitly encodes the biological question (does nuclear enrichment change with LMB?) without requiring adaptation of an external tool, and keeps all samples in one joint variance estimate
-
DESeq2 was used to compare pulldown and input RNA abundances across all Halo-seq enrichment analyses↳ Could also: edgeR (quasi-likelihood F-test) or limma-voom could also be applied to the same count data for differential enrichment — edgeR and limma-voom are widely benchmarked alternatives with different variance-modelling assumptions; concordance across tools can increase confidence in the final gene list
-
Transcript-level abundances were quantified with Salmon (pseudo-alignment) and summarized to gene level with tximport↳ Could also: Alignment-based quantification (e.g., STAR + featureCounts or STAR + RSEM) could also be used, particularly given the paper's parallel analysis of unspliced versus spliced transcripts — Alignment-based methods retain read-level positional evidence, which can be informative when distinguishing intronic from exonic reads as done in the unspliced-transcript sub-analysis
-
Enrichment calls combined an FDR threshold (adjusted P < 0.05) with a fixed |log2FC| ≥ 0.5 cutoff↳ Could also: Reporting all FDR-significant genes ranked by effect size, without a hard fold-change floor, would also be a standard approach — Fixed fold-change thresholds are one way to balance sensitivity and specificity; reporting the full ranked list lets readers apply alternative thresholds and facilitates meta-analyses
-
A fixed minimum-count filter of 5 counts in all samples was applied prior to testing↳ Could also: Data-driven independent filtering (DESeq2's built-in results() filter on mean normalized counts) or edgeR's filterByExpr() could also be used — Automated filtering selects the count threshold that maximises discoveries at a given FDR, whereas a fixed threshold may be either too conservative or too lenient depending on library depth
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34875090 (Halo-seq, Engel et al. 2022, NAR gkab1185)
Paper in one line
Halo-seq: a singlet-oxygen RNA proximity-labeling method. HEK293T cells express a HaloTag fusion targeted to a subcellular locale (chromatin = H2B-Halo; nucleolus = Halo-fibrillarin; cytoplasm = Halo-p65). DBF + green light → local singlet oxygen → biotinylates nearby RNA → streptavidin pulldown vs input → RNA-seq → DESeq2 enrichment per compartment. LMB perturbation tested with Xtail.
CRITICAL accession note
- The room BRIEF lists
geo:GSE116008as "the data". GSE116008 is NOT this paper's data — it is the APEX-seq atlas (Fazal et al. 2019, Cell), used here only as an external comparison dataset (Halo-seq vs APEX-seq labeling). - The paper's OWN data is
GSE172281(26 Halo-seq RNA-seq samples, GSM5251724–GSM5251749). This is the dataset that drives every pipeline-derived result, so the reproduction targets GSE172281. Both are profiled.
Pipeline as described (Methods + GEO data_processing)
- cutadapt 3' adapter trim.
- Salmon (or kallisto) transcript quant against hg38 / GENCODE v28. For the
spliced/unspliced analysis a custom fasta with intron-retained + intron-removed
versions of each transcript was built by
src/add_primary_transcripts.py(the ONE script the paper cites from repo rnabioco/rnaroids). - tximport → gene-level abundances.
- DESeq2 input vs streptavidin-pulldown. Gene filter: ≥5 counts in ALL samples in that analysis. Enriched/depleted call: padj < 0.05 AND |log2FC| ≥ 0.5.
- LMB analysis: Xtail on pulldown/input ratio change, FDR<0.05 / FDR<0.1.
GEO shortcut (makes this lightweight)
GSE172281 ships per-sample Salmon *quant.sf.gz as processed supplementary
files. We therefore skip read alignment and run the described tximport→DESeq2
pipeline directly on the authors' own quant files. This is a faithful 1:1 of the
quantitation→stats half of the pipeline (the half that produces the reported gene
counts), with the upstream Salmon step taken from the authors' deposited output.
Sample design (GSE172281, all input vs pulldown/IP)
- Fibrillarin (nucleolus): Input Rep1-3 (GSM5251724-26) vs IP Rep1-3 (27-29) — 3v3
- H2B (chromatin): Input Rep1-4 (30-33) vs IP Rep1-4 (34-37) — 4v4
- H2B + LMB: Input Rep1-3 (38-40) vs IP Rep1-3 (41-43) — 3v3
- p65 (cytoplasm): Input Rep1-3 (44-46) vs IP Rep1-3 (47-49) — 3v3
IN SCOPE (pipeline-derived → attempt)
- C1 (primary): Fibrillarin pulldown — 602 enriched, 338 depleted (FDR<0.05, |log2FC|≥0.5). DESeq2. [Results, Fig 4]
- C2: H2B vs Fibrillarin significant-enriched overlap = 20 genes, overlap not significant (binomial P = 0.2). [Results]
- C3: H2B enrichment counts ("hundreds enriched and depleted") + p65 counts — quantify exact numbers (paper gives them only qualitatively for H2B/p65).
- C4 (harder): LMB — 105 genes more enriched in H2B+LMB pulldown vs untreated (FDR<0.1) via Xtail. [Results]
OUT OF SCOPE (not pipeline / not reproducible from deposit → not attempted)
- Wet-lab: singlet-oxygen yields (DBF 0.42 vs miniSOG 0.03), dot blots, microscopy, labeling-efficiency gels, 10× labeling vs APEX (assay-level, not from counts).
- Gene-class enrichment statistics (snRNA/lncRNA/ARE/HuR-CLIP Wilcoxon, HuR dose-response) — secondary, depend on external annotation sets; attempt only if core reproduces and time allows.
- GO term lists (ribosome biogenesis for fibrillarin) — qualitative; spot-check only.
- CeFra-seq / APEX-seq head-to-head localization accuracy — qualitative, no AUC given.
Compute plan
All on «our HPC»/«infra». Download quant.sf to «infra» (front1). conda env: r-base + bioconductor-{tximport,deseq2}, r-readr. GENCODE v28 GTF for tx2gene. DESeq2 is light (front1-runnable; will still submit per «infra» rules if needed).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.