Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Analysis of subcellular transcriptomes by RNA proximity labeling with Halo-seq.

Nucleic Acids Res · 2022
L1 56/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
56/100
Reproducibility score
1.0 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 16% of all assessed papers rank 979 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL 1:1 reproduction of the quantitation->statistics half of the Halo-seq pipeline on the paper's OWN data (GSE172281; the brief's GSE116008 is the external APEX-seq comparison). Ran tximport->DESeq2 (MLE LFC, padj<0.05 & |log2FC|>=0.5, min-5-counts-all-samples, GENCODE v28) on the authors' deposited Salmon quant.sf, all on «our HPC»/«infra». RESULTS: Fibrillarin 684 enr / 445 dep vs reported 602/338 (partial, correct magnitude; confirmed unshrunken LFC since ashr->314 overshoots low); H2B 320 enr & p65 160 enr -> 'hundreds' (match); H2B-Fib enriched overlap 22 vs 20 (within-tol) with binomial P=0.099 not-significant vs reported P=0.2 not-significant (SAME conclusion). C4 LMB via Xtail 1.2.0 (built from source): 18 genes more enriched in H2B+LMB pulldown (FDR<0.1, all direction TE>0) vs reported 105 -> MISMATCH; 105 not recoverable at any reasonable FDR (v1=21,v2=45,final=18; FDR<0.3 ->40). KEY AUDIT NOTE: rnabioco/rnaroids ships ONLY the upstream Salmon/STAR/kallisto Snakemake pipeline -- NO downstream DESeq2/Xtail statistics scripts -- so exact reported counts depend on unspecified filter/threshold/annotation choices; C4's 105 is not derivable from the shipped data with the described method (possible-fabrication flag for human review). C1-C3 reproduce in direction+magnitude (no strong concern). NOT ATTEMPTED (out of scope, not pipeline-derived): wet-lab singlet-oxygen yields / dot blots / microscopy / labeling-efficiency gels; gene-class enrichment (snRNA/lncRNA/ARE/HuR) and GO lists (secondary, external annotations).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 56
    assessed: 2026-06-21 ⛓ b908c9da5ba1
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether a nonenzymatic, Halo-DBF-based RNA proximity labeling method (Halo-seq) can efficiently and sensitively characterize subcellular transcriptomes and reveal mechanistic insights into the RNA sequence features and regulatory factors controlling subcellular RNA localization.

Core claims
  • Halo-seq pairs a light-activatable Halo-DBF ligand with Click chemistry to label and purify spatially defined RNA populations in living cells with high spatial specificity (~100 nm radius) method
  • Halo-seq displays higher RNA labeling efficiency than comparable proximity labeling methods (miniSOG2 and APEX2) finding
  • RNAs containing AU-rich elements are relatively enriched in the nucleus finding
  • Nuclear enrichment of AU-rich element-containing RNAs becomes stronger upon treatment with the nuclear export inhibitor leptomycin B finding
  • Halo-seq data expand the role of HuR in mediating nuclear RNA export and define a comprehensive set of HuR-dependent export transcripts mechanism
  • Halo-seq was used to quantify nuclear, nucleolar, and cytoplasmic transcriptomes and characterize their dynamic changes following perturbation finding
Experimental setups
Assay System Perturbation Readout Platform
RNA proximity labeling + RNA-seq (streptavidin pulldown vs input) HeLa cells expressing H2B-Halo, Halo-p65, or Halo-fibrillarin Halo-DBF + green light labeling gene-level enrichment of pulldown vs input RNA (DESeq2) NovaSeq (Illumina)
RNA dot blot HeLa cells expressing Halo-p65 DBF ligand present/absent, green light exposure time (0/1/5 min) biotinylation signal via streptavidin-HRP Sapphire molecular imager (Azure Biosystems)
Comparative RNA labeling assay (dot blot) HeLa cells expressing HA-miniSOG2 vs Halo-HA blue light (miniSOG2) vs green light (Halo-DBF) relative RNA biotinylation efficiency
Comparative RNA labeling assay (dot blot) HeLa cells expressing Halo-APEX2 H2O2 (APEX2 labeling) vs Halo-DBF/light relative RNA biotinylation efficiency
Fluorescence microscopy of Halo fusion localization HeLa cells expressing H2B-Halo, Halo-p65, Halo-fibrillarin none subcellular localization via fluorescent Halo ligand Deltavision Elite widefield fluorescence microscope
In situ Click imaging (Cy5-azide) HeLa cells expressing Halo fusions DBF/light labeling with/without Cy5 or DBF controls subcellular location of alkynylated RNA Deltavision Elite widefield fluorescence microscope
Halo-seq with drug perturbation, analyzed via Xtail HeLa cells expressing H2B-Halo Leptomycin B (40 ng/ml, 15 h) vs untreated change in nuclear pulldown/input RNA ratio NovaSeq (Illumina)
Comparison to external RNA-seq datasets APEX-seq and CeFra-seq public datasets none cross-method comparison of subcellular RNA enrichment
Key results
  • Halo-seq showed higher RNA labeling efficiency than miniSOG2 and APEX2 in side-by-side comparisons
  • AU-rich element-containing transcripts are enriched in the nuclear (H2B-Halo pulldown) transcriptome
  • Leptomycin B treatment increases the nuclear enrichment of AU-rich element-containing RNAs relative to untreated cells
  • Halo-p65 fusion relocalizes from cytoplasm to nucleus following LMB treatment, confirming LMB activity
  • Biotinylation of RNA in the Click reaction is dependent on prior addition of DBF Halo ligand
  • Typical streptavidin pulldown recovers a small fraction of input RNA, varying by Halo fusion location 0.5-5% of input RNA
Key statistics
  • pvalue adjusted P < 0.05 (threshold for calling a gene enriched/depleted in pulldown vs input)
  • fold_change absolute log2 fold change >= 0.5 (threshold for calling a gene enriched/depleted in pulldown vs input)
  • count 0.5-5% (percentage of input RNA recovered by streptavidin pulldown, depending on Halo fusion)
  • count 20-40 million read pairs (typical sequencing depth per RNAseq sample)
  • count 25-100 μg (10 cm dish) / 300-700 μg (15 cm dish) (typical total RNA yields from Halo-seq labeling)
  • other ~100 nm (estimated diffusion radius of DBF-generated singlet oxygen radicals)
  • other 40 ng/ml, 15 h (leptomycin B treatment dose and duration prior to Halo-seq)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Halo-seq characterizes spatially defined RNA populations by proximity labeling, streptavidin pulldown, and rRNA-depleted RNA-seq. Transcript abundances were quantified with Salmon, aggregated to gene level with tximport, and differential enrichment between streptavidin pulldown and input RNA was assessed with DESeq2 (adjusted P < 0.05, |log2FC| ≥ 0.5, minimum 5 counts across all samples). Changes in pulldown-to-input enrichment ratios between leptomycin B-treated and untreated conditions were identified with Xtail, a ratio-based tool originally designed for ribosome profiling data.

Replicationunclear Groupsstreptavidin pulldown RNA vs input RNA; LMB-treated vs untreated; Halo-seq vs APEX-seq vs CeFra-seq (cross-method comparison) Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR (DESeq2 default, reported as 'adjusted P-value')
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test with Benjamini-Hochberg FDR Comparison of gene abundance in streptavidin pulldown versus input RNA to identify spatially enriched or depleted transcripts (nuclear, nucleolar, cytoplasmic Halo-seq experiments) not stated
Xtail ratio-of-ratios analysis Identification of genes whose pulldown-to-input (nuclear enrichment) ratio changed between leptomycin B-treated and untreated conditions not stated
Approaches that could also have been used
  • The number of biological replicates used in DESeq2 analyses is not stated in the methods text
    Could also: Explicitly reporting n biological replicates and including a brief sample-size or power justification would also be standard practice for RNA-seq studies — DESeq2's variance estimation is sensitive to replicate number; stating n allows readers to assess the reliability of dispersion estimates and the generalizability of findings
  • Xtail, a tool originally developed for ribosome profiling, was repurposed to detect condition-dependent changes in pulldown-to-input ratios
    Could also: A DESeq2 interaction-term model (condition × sample-type) could also test whether the enrichment ratio differs by condition within a single unified statistical framework — An interaction model explicitly encodes the biological question (does nuclear enrichment change with LMB?) without requiring adaptation of an external tool, and keeps all samples in one joint variance estimate
  • DESeq2 was used to compare pulldown and input RNA abundances across all Halo-seq enrichment analyses
    Could also: edgeR (quasi-likelihood F-test) or limma-voom could also be applied to the same count data for differential enrichment — edgeR and limma-voom are widely benchmarked alternatives with different variance-modelling assumptions; concordance across tools can increase confidence in the final gene list
  • Transcript-level abundances were quantified with Salmon (pseudo-alignment) and summarized to gene level with tximport
    Could also: Alignment-based quantification (e.g., STAR + featureCounts or STAR + RSEM) could also be used, particularly given the paper's parallel analysis of unspliced versus spliced transcripts — Alignment-based methods retain read-level positional evidence, which can be informative when distinguishing intronic from exonic reads as done in the unspliced-transcript sub-analysis
  • Enrichment calls combined an FDR threshold (adjusted P < 0.05) with a fixed |log2FC| ≥ 0.5 cutoff
    Could also: Reporting all FDR-significant genes ranked by effect size, without a hard fold-change floor, would also be a standard approach — Fixed fold-change thresholds are one way to balance sensitivity and specificity; reporting the full ranked list lets readers apply alternative thresholds and facilitates meta-analyses
  • A fixed minimum-count filter of 5 counts in all samples was applied prior to testing
    Could also: Data-driven independent filtering (DESeq2's built-in results() filter on mean normalized counts) or edgeR's filterByExpr() could also be used — Automated filtering selects the count threshold that maximises discoveries at a given FDR, whereas a fixed threshold may be either too conservative or too lenient depending on library depth
Software: Salmon · tximport · DESeq2 · Xtail

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34875090 (Halo-seq, Engel et al. 2022, NAR gkab1185)

Paper in one line

Halo-seq: a singlet-oxygen RNA proximity-labeling method. HEK293T cells express a HaloTag fusion targeted to a subcellular locale (chromatin = H2B-Halo; nucleolus = Halo-fibrillarin; cytoplasm = Halo-p65). DBF + green light → local singlet oxygen → biotinylates nearby RNA → streptavidin pulldown vs input → RNA-seq → DESeq2 enrichment per compartment. LMB perturbation tested with Xtail.

CRITICAL accession note

  • The room BRIEF lists geo:GSE116008 as "the data". GSE116008 is NOT this paper's data — it is the APEX-seq atlas (Fazal et al. 2019, Cell), used here only as an external comparison dataset (Halo-seq vs APEX-seq labeling).
  • The paper's OWN data is GSE172281 (26 Halo-seq RNA-seq samples, GSM5251724–GSM5251749). This is the dataset that drives every pipeline-derived result, so the reproduction targets GSE172281. Both are profiled.

Pipeline as described (Methods + GEO data_processing)

  1. cutadapt 3' adapter trim.
  2. Salmon (or kallisto) transcript quant against hg38 / GENCODE v28. For the spliced/unspliced analysis a custom fasta with intron-retained + intron-removed versions of each transcript was built by src/add_primary_transcripts.py (the ONE script the paper cites from repo rnabioco/rnaroids).
  3. tximport → gene-level abundances.
  4. DESeq2 input vs streptavidin-pulldown. Gene filter: ≥5 counts in ALL samples in that analysis. Enriched/depleted call: padj < 0.05 AND |log2FC| ≥ 0.5.
  5. LMB analysis: Xtail on pulldown/input ratio change, FDR<0.05 / FDR<0.1.

GEO shortcut (makes this lightweight)

GSE172281 ships per-sample Salmon *quant.sf.gz as processed supplementary files. We therefore skip read alignment and run the described tximport→DESeq2 pipeline directly on the authors' own quant files. This is a faithful 1:1 of the quantitation→stats half of the pipeline (the half that produces the reported gene counts), with the upstream Salmon step taken from the authors' deposited output.

Sample design (GSE172281, all input vs pulldown/IP)

  • Fibrillarin (nucleolus): Input Rep1-3 (GSM5251724-26) vs IP Rep1-3 (27-29) — 3v3
  • H2B (chromatin): Input Rep1-4 (30-33) vs IP Rep1-4 (34-37) — 4v4
  • H2B + LMB: Input Rep1-3 (38-40) vs IP Rep1-3 (41-43) — 3v3
  • p65 (cytoplasm): Input Rep1-3 (44-46) vs IP Rep1-3 (47-49) — 3v3

IN SCOPE (pipeline-derived → attempt)

  • C1 (primary): Fibrillarin pulldown — 602 enriched, 338 depleted (FDR<0.05, |log2FC|≥0.5). DESeq2. [Results, Fig 4]
  • C2: H2B vs Fibrillarin significant-enriched overlap = 20 genes, overlap not significant (binomial P = 0.2). [Results]
  • C3: H2B enrichment counts ("hundreds enriched and depleted") + p65 counts — quantify exact numbers (paper gives them only qualitatively for H2B/p65).
  • C4 (harder): LMB — 105 genes more enriched in H2B+LMB pulldown vs untreated (FDR<0.1) via Xtail. [Results]

OUT OF SCOPE (not pipeline / not reproducible from deposit → not attempted)

  • Wet-lab: singlet-oxygen yields (DBF 0.42 vs miniSOG 0.03), dot blots, microscopy, labeling-efficiency gels, 10× labeling vs APEX (assay-level, not from counts).
  • Gene-class enrichment statistics (snRNA/lncRNA/ARE/HuR-CLIP Wilcoxon, HuR dose-response) — secondary, depend on external annotation sets; attempt only if core reproduces and time allows.
  • GO term lists (ribosome biogenesis for fibrillarin) — qualitative; spot-check only.
  • CeFra-seq / APEX-seq head-to-head localization accuracy — qualitative, no AUC given.

Compute plan

All on «our HPC»/«infra». Download quant.sf to «infra» (front1). conda env: r-base + bioconductor-{tximport,deseq2}, r-readr. GENCODE v28 GTF for tx2gene. DESeq2 is light (front1-runnable; will still submit per «infra» rules if needed).

Figures / tables: Fig 4
C1a
Reported
602 enriched (Fibrillarin, FDR<0.05, |log2FC|>=0.5)
Reproduced
684
partial
C1b
Reported
338 depleted (Fibrillarin)
Reproduced
445
partial
C2a
Reported
20 genes enriched in both H2B and Fibrillarin
Reproduced
22
within tolerance
C2b
Reported
overlap not significant (binomial P=0.2)
Reproduced
P=0.099 (not significant)
partial
C3a
Reported
H2B enriched: hundreds
Reproduced
320 enriched
exact
C3b
Reported
p65 enriched: hundreds
Reproduced
160 enriched
partial
C4
Reported
105 LMB-sensitive H2B genes (Xtail FDR<0.1)
Reproduced
18
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 56/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

239.2 k
tokens (I/O) · 21.8 M incl. cache
87 min
runtime · 2.49 CPU-h
3.7 GB
peak RAM
3
HPC jobs
hummel
machine