Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

TFAP2 paralogs facilitate chromatin access for MITF at pigmentation and cell proliferation genes.

PLoS Genet · 2022
L2 95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1 within pipeline-version tolerance). The paper's RNA-seq DESeq2 result (GSE190610 suppl xlsx) was reproduced independently from raw FASTQ via the described method: Trim Galore -> STAR (UCSC hg19 + GENCODE v19) --quantMode GeneCounts -> DESeq2 per clone vs WT, padj<0.05. 12 paired-end SRA runs aligned. Significant-DEG counts: 5657->5644 (clone4.3, 99.8%) and 5578->5494 (clone2.12, 98.5%). Per-gene log2FC concordance Pearson 0.997-0.998, sign concordance 99.9%, exemplar genes near-exact (ARMC4 7.595->7.614; CTSG -9.911->-9.920). Anti-fabrication: PASS - the deposited DE table is genuinely derivable from the raw data. The ~76% Jaccard of the exact padj<0.05 gene set reflects boundary genes flipping across tool/annotation versions, not effect-size disagreement. NOT attempted (out of scope): the repo's prep_Fisher Table-1 enrichment helper (consumes a pre-made DE table + MACS2 peaks + un-shipped gene-coordinate references + hard-coded paths -> under-determined to run 1:1), CUT&RUN/ATAC peak calling + DiffBind, and zebrafish scRNA. NOTE: prior room had staged this pipeline but was finalized mid-alignment and its «infra» workdir (~33G) was reclaimed by the janitor; this run rebuilt env+index+alignment+DESeq2 from scratch. One transient STAR-index-write failure (exit 110, not OOM/disk) on the first setup attempt; clean rerun succeeded.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ 8957b7bde0d2
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
👤 1 human curator(s) · Level L2 2026-06-15
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

TFAP2 paralogs act as pioneer/initiating transcription factors that facilitate MITF's access to chromatin at enhancers regulating pigmentation and proliferation genes, and independently repress enhancers near cell-adhesion genes.

Core claims
  • Pigmentation genes are only expressed in mitfa-expressing zebrafish melanocyte-lineage cells that also express tfap2 paralogs finding
  • TFAP2 paralogs directly activate enhancers near genes enriched for roles in pigmentation and cell proliferation mechanism
  • TFAP2 paralogs directly repress enhancers near genes enriched for roles in cell adhesion mechanism
  • TFAP2 is necessary for MITF binding and chromatin accessibility at co-activated enhancers, indicating TFAP2 acts as a pioneer factor for MITF mechanism
  • TFAP2-KO SK-MEL-28 cells proliferate less and adhere to one another more than WT cells finding
  • Tfap2a and Tfap2e act redundantly to promote melanocyte number and pigmentation onset in zebrafish embryos finding
  • TFAP2 and MITF do not appear to co-operatively repress the same enhancers finding
  • TFAP2A/B/C can bind nucleosomes and behave similarly to the canonical pioneer factor FOXA1 mechanism
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq GFP+ cells sorted from Tg(mitfa:GFP) zebrafish embryos, 28 hpf none cell-type clustering and gene expression (mitfa, tfap2 paralogs, pigmentation genes) 10x Chromium; Seurat/UMAP
CRISPR knockout SK-MEL-28 human melanoma cells TFAP2A and TFAP2C double knockout (TFAP2-KO) generation of TFAP2-KO cell line for downstream assays
ATAC-seq SK-MEL-28 WT vs TFAP2-KO cells TFAP2A/TFAP2C KO chromatin accessibility / nucleosome positioning
CUT&RUN (anti-H3K27Ac, anti-H3K4Me3, anti-H3K27Me3) SK-MEL-28 WT vs TFAP2-KO cells TFAP2A/TFAP2C KO active enhancer marks and silenced chromatin marks
CUT&RUN (anti-MITF) SK-MEL-28 WT vs TFAP2-KO cells TFAP2A/TFAP2C KO MITF chromatin binding genome-wide
CUT&RUN (anti-TFAP2A) SK-MEL-28 WT and MITF loss-of-function mutant cells MITF loss-of-function TFAP2A chromatin binding
bulk RNA-seq SK-MEL-28 WT vs TFAP2-KO human cell lines TFAP2A/TFAP2C KO differential gene expression
zebrafish genetics / live imaging tfap2a and tfap2e mutant zebrafish embryos (single and double mutants) tfap2a and/or tfap2e loss-of-function mutation number and pigmentation of dorsal melanophores at 36 hpf
Key results
  • Pigmentation genes (dct, pmel, mreg, trpm1b) are detected only in tfap2-high MIX cells, not tfap2-low MIX cells, despite both expressing mitfa
  • tfap2a;tfap2e double mutant embryos have fewer and paler melanophores than tfap2a single mutants or WT siblings
  • TFAP2-KO SK-MEL-28 cells show reduced proliferation compared to WT
  • TFAP2-KO SK-MEL-28 cells show increased cell-cell adhesion compared to WT
  • At many TFAP2A-bound chromatin elements, TFAP2 loss causes loss of enhancer activity and, at a large subset, chromatin condensation
  • MITF binding is TFAP2-dependent at TFAP2-activated enhancer subsets
  • TFAP2 inhibits enhancers near cell-adhesion genes, and at a subset of these excludes MITF binding
Key statistics
  • count 11,217 cells (GFP+ cells sequenced from Tg(mitfa:GFP) zebrafish embryos at 28 hpf)
  • count 1918 cells (sox10-positive cells re-clustered into 10 subclusters)
  • count 28 clusters comprising 11 cell types (initial clustering of GFP+ zebrafish embryo cells)
  • count n=9 (melanophore counts per embryo, tfap2a low/low mutants at 36 hpf)
  • count n=32 (melanophore counts per embryo, tfap2a low/low; tfap2e +/ui157 embryos at 36 hpf)
  • count n=10 (melanophore counts per embryo, tfap2a low/low; tfap2e ui157/ui157 double mutants at 36 hpf)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is an integrative genomics investigation combining single-cell RNA-seq (10x Chromium, 11,217 cells) of zebrafish embryos with bulk genomic assays (ATAC-seq, CUT&RUN for MITF/TFAP2A and histone marks) in SK-MEL-28 melanoma cells, plus in vivo melanophore counts in mutant zebrafish. Single-cell data were clustered and visualized with Seurat/UMAP and ordered with Monocle pseudotime, while a discrete group comparison of melanophore numbers across mutant genotypes was reported with a box plot and a Student's t-test.

Replicationmixed Sample sizeFor the in vivo comparison, per-genotype embryo counts are given (n = 9, 32, 10); for scRNA-seq, total cells sequenced (11,217) and re-clustered (1918) are stated; no formal power/sample-size calculation is described in the provided text Groupstfap2a/tfap2e mutant zebrafish genotypes; WT vs TFAP2-KO and MITF-mutant SK-MEL-28 cells Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated in provided text
Statistical tests used
Test Applied to n Assumptions
Student's t-test number of pigmented dorsal melanophores compared across tfap2a low/low, tfap2a low/low;tfap2e +/ui157, and tfap2a low/low;tfap2e ui157/ui157 embryos (Fig 1L) n = 9, 32, and 10 embryos for the three genotype groups, respectively not stated
Monocle pseudotime trajectory analysis lineage ordering of sox10-positive cells from neural crest through MIX/MX to melanophores (Fig 1B) 1918 sox10-expressing cells na
Seurat clustering / differential expression for marker identification cluster annotation and genes differentially expressed between MIX tfap2-low and tfap2-high clusters (Fig 1G, S1 Table) 11,217 total cells; 1918 re-clustered sox10+ cells not stated
Approaches that could also have been used
  • The three-genotype melanophore-count comparison was reported with the Student's t-test.
    Could also: A one-way ANOVA (or Kruskal-Wallis) with a post-hoc multiple-comparison correction such as Tukey HSD or Dunn's test could also have been used. — A single omnibus test with post-hoc correction also controls the family-wise error rate across the simultaneous group comparisons, which some readers prefer when more than two groups are involved.
  • Group spread was conveyed with box-plot quartiles and min/max whiskers.
    Could also: Reporting a 95% confidence interval for the group mean, or standard deviation, alongside the box plot could also be done. — A CI or SD complements the visual spread by directly quantifying uncertainty in the estimate, which is often valued for the small per-group n here (e.g., n = 9 and n = 10).
  • The melanophore counts were compared with a parametric Student's t-test.
    Could also: A non-parametric Mann-Whitney U test, or an exact/permutation test, could also have been applied. — A rank-based or permutation alternative makes no normality assumption and is a common choice for small-sample count data, providing a robustness check on the parametric result.
  • The exact p-value is attributed to the test but the numeric value is not shown in the excerpt.
    Could also: Reporting exact p-values together with the effect size (e.g., difference in means and its CI) could also accompany the test. — Exact p-values plus an effect-size estimate convey both the strength and the practical magnitude of the difference, aiding interpretation and downstream meta-analysis.
  • Cell-population structure was derived with Seurat clustering at fixed parameters (dims = 30, resolution = 1.2) and visualized with UMAP.
    Could also: Sensitivity analyses across a range of resolution/dimension parameters, or alternative embeddings (e.g., t-SNE) and clustering metrics (e.g., silhouette scores), could also be reported. — Parameter-sensitivity reporting helps convey how stable the identified clusters are to analytic choices, which readers often find informative for single-cell studies.
Software: Seurat · UMAP · Monocle (pseudotime) · 10x Chromium / Cell Ranger pipeline

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
38
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

HPA003259 HPA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35580127 (TFAP2 paralogs / MITF, PLoS Genet 2022)

GSE190610 is a multi-omic series: CUT&RUN (TFAP2A, MITF, H3K27Ac/Me3, H3K4Me3), ATAC-seq, and RNA-seq, all in SK-MEL-28 melanoma cells (WT vs TFAP2-KO).

Repo (github.com/ahelv/Differential_Expression @ bbf8fe2)

ONE R function prep_Fisher (compute_dysregulation.R). It is NOT a DE pipeline: it consumes a pre-computed RNA-seq DE table (ext_gene, logFC, pvalue, padj) plus a MACS2 peak bedfile and runs bedtools closest + a Fisher's Exact Test for enrichment of dysregulated genes near peaks (Table 1). It hard-codes paths (/bedtools2/bin, /Directory/) and references un-shipped globals (gene_bed, gene_list_unique, gene_bed_flank). No example data, no expected values. → Running the repo 1:1 is under-determined (the 4 inputs + gene-coordinate references are not provided) = hard 20%.

In scope (low-hanging, well-specified pipeline) — PRIMARY TARGET

RNA-seq differential expression (the input the repo consumes; the paper's own result):

  • Reported result: GSE190610_Differentially_expressed_genes.xlsx (series suppl). Two sheets, each a DESeq2 table (Gene, log2FC, pvalue, padj):
    • DEGs_4.3 = Clone 4.3 (TFAP2-KO) vs WT — 5657 genes
    • DEGs2.12 = Clone 2.12 (TFAP2-KO) vs WT — 5578 genes
  • Methods (PMC9159589): Trim Galore 0.6.3 → STAR align to hg19/GRCh37 --quantMode GeneCounts → DESeq2, adj p<0.05.
  • Data: 12 paired-end RNA-seq runs (raw FASTQ on SRA; no counts deposited): WT = SRR18982187-190; Clone2.12 = SRR18982191-194; Clone4.3 = SRR18982195-198.
  • Reproduction plan: realign per the described pipeline (STAR/GRCh37 + GENCODE v19, GeneCounts) → DESeq2 per clone vs WT → compare per-gene log2FC + significant-gene overlap against the deposited xlsx. Independent quantification = anti-fabrication check.

Out of scope (not attempted; why)

  • CUT&RUN / ATAC-seq peak calling, DiffBind, the Fisher enrichment (Table 1): the repo's inputs/reference files are not shipped (docs_insufficient for 1:1), and the multi-assay integration is the hard last ~20%.
  • Zebrafish scRNA-seq (Cell Ranger): separate large pipeline, out of scope.
DEG_4.3_count
Reported
5657 sig DEGs (Clone4.3-KO vs WT, DESeq2 padj<0.05)
Reproduced
5644 (99.8%)
within tolerance
DEG_2.12_count
Reported
5578 sig DEGs (Clone2.12-KO vs WT, DESeq2 padj<0.05)
Reproduced
5494 (98.5%)
within tolerance
DEG_4.3_top_ARMC4_log2FC
Reported
7.595
Reproduced
7.614
exact
DEG_2.12_top_CTSG_log2FC
Reported
-9.911
Reproduced
-9.920
exact
log2FC_concord_clone4.3
Reported
~1.0 expected (anti-fabrication)
Reproduced
Pearson 0.9972 / Spearman 0.9976 / sign 99.9%
exact
log2FC_concord_clone2.12
Reported
~1.0 expected (anti-fabrication)
Reproduced
Pearson 0.9982 / Spearman 0.9983 / sign 99.9%
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is an incomplete reproduction, not a discrepancy: the study is fully eligible (public code, public GSE190610 raw FASTQ on SRA, and a method specified well enough to reproduce 1:1), and a faithful Trim Galore→STAR(hg19/GENCODE19)→DESeq2 pipeline was correctly staged. However, «job» was still aligning (0/12 samples) when finalize was forced, so no reproduced values were computed — the reported DEG counts (5657 / 5578 at padj<0.05) and concordance checks remain PENDING. The shortfall is entirely on our side / operational (premature finalize), with a secondary note that the authors' public repo is a Fisher-enrichment helper rather than the DE pipeline, so the DE result had to be re-staged independently. No evidence of any author/data defect; severity is unmeasured because nothing was run to completion.

👤 Schlein Lab (curation team) L2 95/100
🔴1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

327.1 k
tokens (I/O) · 24 M incl. cache
141 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.