Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification of new ETV6 modulators through a high-throughput functional screening.

iScience · 2022
68/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 32% of all assessed papers rank 765 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (both in-scope pipeline results reproduced; marquee results out-of-scope by data availability). Reproduction unit = GEO GSE79373 + Cufflinks (third-party tool; paper reports no original code). C1 (deposit integrity): independently re-aligned both deposited Reh RNA-seq runs (REHWT SRR3236162, REH2-11Plenti SRR3236158) with STAR 2.7.10b -> HTSeq 2.0.5 (reverse-stranded, Ensembl GRCh37.75) and compared per-gene to the deposited GSE79373 HTSeq counts: Pearson(log1p) 0.9886/0.9885, ~84% exact per-gene counts, total counts within 0.7% -> deposit integrity CONFIRMED. C2 (the assigned Cufflinks step): Cufflinks v2.2.1 on the REHWT BAM yields 63653 gene FPKMs; the paper's FPKM<=0.21 'non-expressed' cutoff splits them into 46448 non-expressed (median deposited count 0, 76% zero) vs 17205 expressed (median deposited count 813, 5% zero) -> the cutoff is a valid, reproduced non-expressed threshold. OUT OF SCOPE (not deposited under any accession, recorded not graded): the headline shRNA functional screen (~140k shRNAs; 2858 positive -> 1241 retained -> 90 candidate genes -> 13 validated; top5 AKIRIN1/COMMD9/DYRK4/JUNB/SRP72) and the targeted-RNA-seq Table 1 ratios. The exact 2858->1241 shRNA reduction is not reproducible because it needs the undeposited shRNA->target mapping, but the Cufflinks/FPKM<=0.21 step it depends on does reproduce. Honest verdict: described well enough for the deposited Reh expression (1:1 on deposit integrity, valid on the FPKM cutoff); the paper's central screen rests on undeposited data and cannot be independently checked.

💻 Code ↗ 🗄 Data: GSE79373

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

ETV6's regulatory network remains unclear despite its role in hematological malignancies; the paper tests whether an unbiased genome-wide shRNA screen can identify novel genes that modulate ETV6's transcriptional repressive activity.

Core claims
  • A genome-wide shRNA screen in an engineered ETV6-dependent Blasticidin-sensitive pre-B ALL cell line can identify modulators of ETV6 repressive transcriptional activity method
  • 13 shRNAs were identified that knock down their target genes and induce overexpression of ETV6 transcriptional target genes finding
  • Silencing of AKIRIN1, COMMD9, DYRK4, JUNB, and SRP72 leads to abrogation of ETV6 repressive activity finding
  • The Reh EBS3tk BlastR clone 2-1 ETV6-5 cell system is a suitable, reversible model to screen for ETV6 modulators resource
  • The identified modulators have a broad, gene-specific impact on the ETV6 transcriptional network finding
  • AKIRIN1, DYRK4, SRP72, and JUNB modulate t(12;21)-associated ETV6 target genes, suggesting conservation of modulator activity across leukemic contexts finding
  • Structure-prediction-based interaction analysis can nominate additional ETV6 modulator candidates missed by single-shRNA filtering method
Experimental setups
Assay System Perturbation Readout Platform
genome-wide shRNA loss-of-function screen with deep sequencing Reh pre-B ALL cell line (EBS3tk-BlastR, ETV6-expressing, clone 2-1 ETV6-5) shRNA library (~140,000 unique shRNAs, 15 pools) shRNA enrichment/normalized counts under Blasticidin selection vs control
qRT-PCR Reh EBS3tk BlastR clones (2-1 to 2-22) lentiviral ETV6 overexpression BlastR mRNA expression
ChIP-seq Reh clone 2-1 none (endogenous ETV6 ChIP) ETV6 binding/integration site of EBS3tk BlastR
PCR (genomic) Reh clone 2-1 none confirmation of EBS3tk BlastR insertion site in NEGR1 intron 1
qRT-PCR / Western blot Reh 2-1 ETV6-5 clone shRNA knockdown of ETV6 or NAM (nicotinamide) treatment BlastR expression and ETV6 protein level
proliferation assay (cell counting) Reh 2-1 ETV6-5, shCtl-GFP, shETV6-GFP sorted cells Blasticidin dose range, ETV6 knockdown cell proliferation over 12 days
targeted RNA sequencing (custom gene panel) Reh ETV6-expressing cell lines individually transduced with candidate shRNAs (90 candidates) shRNA knockdown of individual candidate genes expression of 84-gene ETV6 target/control panel
qPCR validation Reh shDYRK4 cells shRNA knockdown of DYRK4 DYRK4 expression (knockdown efficiency)
Key results
  • ETV6 re-expression reduces BlastR expression in Reh clone 2-1 compared to pLENTI control 2.7-fold, p ≤ 0.001
  • 126,042 unique shRNAs identified by deep sequencing; 1,241 retained after threshold filtering, yielding 90 candidate ETV6 modulator genes
  • 13 shRNAs efficiently knocked down target genes and impaired ETV6 transcriptional activity
  • 5 top modulators (AKIRIN1, COMMD9, DYRK4, JUNB, SRP72) significantly impact expression of ETV6 target genes
  • 40 of 84 candidate genes confirmed as ETV6-dependent targets; 22 of these are direct ETV6 ChIP-seq targets
  • shDYRK4 knockdown efficiency validated by qPCR 73% knockdown
  • AKIRIN1, DYRK4, SRP72, and JUNB knockdown samples cluster with ETV6-null pLENTI sample based on t(12;21)-associated target gene expression; COMMD9 shows only moderate impact
  • EBS3tk BlastR insertion into NEGR1 intron 1 does not alter NEGR1 expression and NEGR1 is not significantly ETV6-regulated 53.1 vs 51.3 FPKM; logFC=-0.46, FDR=0.50
Key statistics
  • fold_change 2.7-fold reduction in BlastR expression with ETV6 vs pLENTI (qRT-PCR validation of ETV6-dependent BlastR repression in clone 2-1)
  • pvalue p ≤ 0.001 (BlastR repression by ETV6 vs pLENTI control)
  • count 126,042 unique integrated shRNAs out of ~140,000 library (genome-wide shRNA screen sequencing depth)
  • count 2,858 positive shRNAs; 1,241 retained after filtering (shRNA screen threshold selection)
  • count 81 genes targeted by ≥2 distinct shRNAs (77 by two, 4 by three) (prioritized candidate genes from screen)
  • count 90 total candidate modulator genes (81 + 9 from structure prediction) (final candidate gene list for validation)
  • count 40 of 84 candidate target genes confirmed ETV6-dependent; 22 direct ChIP-seq targets (targeted RNA-seq validation of ETV6 transcriptional network)
  • fold_change logFoldChange = -0.46, FDR = 0.50 (NEGR1 expression change upon ETV6 re-expression)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a genome-wide shRNA screen with sequential read-count thresholds (including a 2 SD cutoff relative to a control shRNA) to nominate ETV6-modulator candidates, then validated candidates with targeted RNA sequencing and qRT-PCR. Comparisons between conditions (e.g., BlastR expression across clones/treatments, shRNA pairs/trios, knockdown categories, and ETV6-target versus control-gene expression) were assessed with paired or unpaired Student's t-tests, and grouping patterns were additionally visualized with unsupervised hierarchical clustering/heatmaps. Results were reported mainly as mean ± SD with exact p-values (Table 1) or threshold-based significance markers (*, **, ***) in figures.

Replicationtechnical Sample sizeTechnical replicates (n=4) stated for qRT-PCR assays (Figure 2A, D, E); n=81 gene pairs/trios for shRNA comparisons (Figure 3E); n=31/32/22 for knockdown categories (Figure 4D); n=22 target genes vs n=62 control genes for network validation (Figure 4C, G); no formal power calculation or sample-size justification described. GroupsETV6-expressing vs control clones; shRNA-silenced vs control (SHC) samples; ETV6 target genes vs ETV6-independent control genes; knockdown-efficiency categories Pairingmixed Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionnot stated for the multiple t-tests performed across shRNA/gene comparisons; FDR (false discovery rate) was reported for one differential expression value (NEGR1 logFoldChange) drawn from a separate prior RNA-seq dataset
Statistical tests used
Test Applied to n Assumptions
Student's t-test (paired) Figure 3E: comparison of normalized counts between best and second-best shRNA of a pair/trio n = 81 gene pairs/trios not stated
Student's t-test (unpaired) Figure 4D: relative expression compared across knockdown categories (Best KD, KD, No KD) Best KD n=31, KD n=32, No KD n=22 not stated
Student's t-test Figure 4C: p values for ETV6-target vs non-target gene expression per shRNA sample 22 target genes vs 62 control genes per shRNA not stated
Student's t-test (paired) Table 1: 'Targets SHC vs shRNA (paired)' p-values for the five top modulators 22 ETV6 target genes per comparison not stated
Student's t-test Figure 2A: BlastR expression in ETV6-expressing clone vs mock/pLENTI controls n = 4 technical replicates not stated
SD-based threshold for hit calling (2 SD above SHC202 control counts) Figure 3A: selection of over-represented shRNAs in the genome-wide screen 126,042 unique integrated shRNAs screened not stated
Approaches that could also have been used
  • Technical replicate variability in qRT-PCR assays (n=4) is summarized as mean ± SD.
    Could also: Reporting the mean with a 95% confidence interval or SEM alongside SD — A CI directly conveys the precision of the estimated mean, which can be a useful complement to SD (which describes spread of the replicate measurements themselves) when communicating how confidently a difference between conditions is estimated.
  • Multiple independent Student's t-tests were performed across many shRNA and gene-target comparisons (e.g., Figure 4C, Table 1) without a stated multiple-testing correction.
    Could also: Applying a family-wise error control (Bonferroni, Holm) or a false-discovery-rate procedure (Benjamini-Hochberg) across the full set of comparisons — When many tests are run on the same dataset, an FDR or family-wise correction is a standard way to keep the expected proportion of false positives controlled across the whole comparison set, complementing the per-comparison p-values already reported.
  • Hit selection in the genome-wide shRNA screen used a fixed 2 SD threshold relative to a control shRNA to call over-represented shRNAs.
    Could also: A formal screen-analysis pipeline such as MAGeCK, RSA (redundant siRNA activity), or a permutation/empirical null-based FDR approach — These established screen-analysis frameworks incorporate replicate variability and multiple-hairpin redundancy directly into a statistical hit-calling score, which can complement a simple SD-based cutoff, particularly for genome-wide libraries with many hairpins per gene.
  • Grouping of expression profiles across shRNA samples was visualized with unsupervised hierarchical clustering/heatmaps (Figures 4H, 5A).
    Could also: Pairing the clustering with a formal group-separation statistic, such as PERMANOVA or a silhouette/bootstrap-support metric — A quantitative test of cluster separation can complement the visual heatmap pattern with a numerical measure of how distinct the identified groups are.
  • Comparisons between paired best/second-best shRNAs (Figure 3E) and paired SHC-vs-shRNA target expression (Table 1) used paired Student's t-tests.
    Could also: A non-parametric paired alternative such as the Wilcoxon signed-rank test — A non-parametric paired test can also be used when normality of the paired differences is uncertain, and would complement the parametric t-test results, especially for smaller sample sizes.
  • Differential expression for a single gene (NEGR1) was reported with a logFoldChange and an FDR value from RNA-seq data.
    Could also: Extending the same FDR-based differential expression framework (e.g., DESeq2, edgeR, or limma-voom) consistently to the targeted RNA-seq validation comparisons — Using the same model-based differential-expression approach across all RNA-seq-derived comparisons, rather than combining it with t-tests for some analyses, can offer a unified variance-modeling framework across the whole dataset.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35198911

Paper: Neveu et al. 2022, iScience — "Identification of new ETV6 modulators through a high-throughput functional screening." PMID 35198911 / PMC8851229 / DOI 10.1016/j.isci.2022.103858.

Assigned reproduction unit (from BRIEF.md):

Important up-front facts (control-plane verified)

  1. GSE79373 is REUSED data, generated for the lab's earlier paper Neveu et al. 2016, Blood "CLIC5: a novel ETV6 target gene..." (PMID 27540136). The 2022 paper's Data Availability cites GSE79373 (RNA-seq), GSE102785 (ChIP-seq), PXD031001 (proteomics).
  2. The 2022 paper's code-availability statement reads verbatim: "This paper does not report original code." The assigned GitHub link (cole-trapnell-lab/ cufflinks) is the generic third-party Cufflinks tool, NOT authors' code. Per BRIEF rule P16, applying this third-party tool to the paper's own data is a valid reproduction.

GSE79373 contents (from GEO sample table)

  • 25 RNA-seq samples, Homo sapiens, genome hg19 / Ensembl GRCh37.75.
  • 17 patient ALL samples (315_T856_T) on AB 5500 SOLiD (Lifescope align).
  • 7 Reh cell-line samples on Illumina HiSeq 2500 (STAR align): GSM2093552 REH2-11ETV6-ETS-NLS-HIS, GSM2093553 REH2-11ETV6-HIS, GSM2093554 REH2-11Plenti, GSM2093555 REH2-1ETV6-ETS-NLS-HIS, GSM2093556 REH2-1ETV6-HIS, GSM2093557 REH2-1Plenti, GSM2093558 REHWT.
  • Deposited processed files = HTSeq counts (*.counts.txt.gz), GSE79373_RAW.tar (~5.2 MB).
  • GEO data_processing: Lifescope 2.1 (SOLiD) / STAR (Illumina) → PICARD 1.107 mark-dup → GATK 3.2-2 SplitNCigar → HTSeq 0.6.1p1 counts vs Ensembl GRCh37.75.
  • Raw reads: SRA SRP071966.

2022 paper bioinformatic pipeline (verbatim from Methods)

  • shRNA screen: "aligned on a custom reference containing all the target sequences of the MISSION Human shRNA library using the Bowtie2 ... version 2.2.3"; "read counts per shRNA were obtained using BEDTools version 2.22.1"; normalized per total reads.
  • Targeted RNA-seq: raw counts normalized per total reads + 12 housekeeping genes; relative expression vs SHC002/SHC202 controls.
  • Expression filter: "shRNAs which target non-expressed genes (FPKM ≤0.21) in Reh cells."
  • "The gene expression values (FPKM) were calculated using the Cufflinks tool version 2.2.1 on hard-clipped BAM files."

In scope (pipeline-derived AND backed by deposited data)

  • R1 — Reproduce GSE79373 deposited Reh counts. Re-align a Reh Illumina sample (REHWT = GSM2093558; control REH2-11Plenti = GSM2093554) from SRA reads with STAR → HTSeq (Ensembl GRCh37.75), compare per-gene counts to the deposited *.counts.txt.gz. Verifies the deposit reproduces. (Pipeline: STAR+HTSeq.)
  • R2 — Cufflinks FPKM in Reh cells + the FPKM ≤0.21 filter. Run Cufflinks v2.2.1 on the Reh BAM (the assigned tool) to compute FPKM; characterize the expressed/non-expressed split and verify that FPKM ≤0.21 is a sensible non-expressed cutoff (this is the exact 2022 step that uses GSE79373 + Cufflinks).

Out of scope (no deposited data and/or wet-lab/manual)

  • Headline screen counts (1,241 shRNAs → 81 genes → 90 candidates → 13 validated; Fig 3–4, Table 1): the shRNA-barcode deep-sequencing data is not deposited under any accession → cannot reproduce (blocker: no_data_accession for that assay).
  • Targeted RNA-seq relative-expression (Table 1 ratios, e.g. DYRK4 3.21, SRP72 2.55): targeted-RNA-seq reads not deposited → cannot reproduce.
  • Proteomics (PXD031001), ChIP-seq (GSE102785), reporter/wet-lab assays: out of scope for this RU (different accessions / wet-lab).

Honest expected outcome

A partial reproduction: the only deposited+pipeline+named-tool intersection is the Reh-cell expression (R1/R2). The paper's marquee results are not backed by deposited data,

Figures / tables: TableFig 3
C1
Reported
GSE79373 Reh deposited HTSeq counts (GRCh37.75) reproducible from raw reads
Reproduced
REHWT Pearson(log1p)=0.9886 Spearman=0.979 exact=83.8% within1=88.2%; REH2-11Plenti Pearson=0.9885 Spearman=0.989 exact=83.4%; totals within 0.7%; reverse-stranded
within tolerance
C2
Reported
Cufflinks v2.2.1 FPKM<=0.21 = unexpressed-gene cutoff in Reh
Reproduced
Cufflinks v2.2.1 on REHWT: 63653 genes, 46448 FPKM<=0.21 (non-expressed, median deposited count 0, 76% zero) vs 17205 FPKM>0.21 (expressed, median deposited count 813, 5% zero) -> cutoff cleanly separates non-expressed from expressed
within tolerance
OOS1
Reported
shRNA screen 2858->1241->90->13 (top5 AKIRIN1,COMMD9,DYRK4,JUNB,SRP72)
Reproduced
NOT ATTEMPTED - shRNA-barcode deep-seq data not deposited under any accession
partial
OOS2
Reported
targeted-RNA-seq Table 1 relative-expression ratios
Reproduced
NOT ATTEMPTED - targeted-RNA-seq reads not deposited
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.