Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Dynamic reversal of random X-Chromosome inactivation during iPSC reprogramming.

Genome Res · 2019
L1 75/100 PQI 83
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH? Yes for the assigned public data; the methods name the dataset, build, tools and the qualitative result clearly. 1:1 vs different: PARTIAL 1:1 on what GSE106340 supports. The assigned accession GSE106340 is the EXTERNAL Schiebinger/Waddington-OT 10X reprogramming scRNA-seq (65,781 mouse cells) that the paper re-analysed -- NOT the paper's own allele-resolved data (that is GSE126229, a different accession). From GSE106340 I reproduced (a) the exact cohort size 65,781 cells [exact], and (b) the paper's central qualitative claim -- X-chromosome reactivation -- as a significant monotone rise in the mean X-linked:autosomal expression ratio across the time course (MEF 0.81 -> 2i-iPSC 1.26, ~55%; Spearman rho=0.842, p=0.0022), strongest in 2i/naive iPSCs as expected [partial: trend confirmed, no exact paper number is pinned to this accession]. Pipeline: GSE106340 normalized matrix -> per-chromosome mean-expression aggregation (numpy/pandas) + mm10 GENCODE M25 gene->chr; cell-day labels read directly from the matrix header. WHAT I DID NOT ATTEMPT (the hard 20%): the allele-specific Mus/(Mus+Cast) headline -- 156 X-linked genes, median allelic ratio -1.148 (d13)/-0.144 (iPSC), Fig 1H/2A -- because it requires the allele-resolved GSE126229 (not the assigned accession); Monocle 2.10.0 pseudotime; external ChIP-seq. The brief's code link kundajelab/atac_dnase_pipelines is a text-mining false positive (no ATAC/DNase result in this scRNA-seq paper). No fabrication concern surfaced: every reproduced value is derivable from the shipped public matrix. All heavy compute ran on «our HPC»/«infra»; «host» holds results only.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 75
    assessed: 2026-06-14 ⛓ bf2f17a2e910
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper asks when and how chromosome-wide reversal of random X-Chromosome inactivation (XCR) occurs during reprogramming of mouse somatic cells to iPSCs, and which genomic features, pluripotency transcription factors, and chromatin regulators enable or restrict the reactivation of stably silenced X-linked genes.

Core claims
  • XCR during iPSC reprogramming is hierarchical, with subsets of X-linked genes reactivating early, intermediate, late, and very late. finding
  • XCR initiates earlier than previously thought, before the onset of full pluripotency network activation and before complete Xist loss. finding
  • Early-reactivating genes are located genomically closer to genes that escape XCI than late-reactivating genes. finding
  • Early-reactivating genes show increased pluripotency transcription factor binding. finding
  • Histone deacetylases (HDACs) restrict XCR in reprogramming intermediates, and the hypoacetylated state of the Xi persists until late reprogramming stages. mechanism
  • Allelic activation of X-linked genes involves combined action of chromatin topology, pluripotency TFs, and chromatin regulators. mechanism
  • An allele-resolution inducible reprogramming mouse model (Mus Xi-GFP / Cast Xa) enables allele-specific transcriptome tracing of XCR. resource
  • Imprinted autosomal genes (Impact, Peg3) are reactivated/erased during iPSC reprogramming. finding
Experimental setups
Assay System Perturbation Readout Platform
Allele-specific full-transcript RNA-seq (Smart-seq2) Female Mus musculus musculus (Xi-GFP) x Mus musculus castaneus (Cast) MEFs, FUT4+ reprogramming intermediates, iPSCs, and ESCs OSKM (Pou5f1/Oct4, Sox2, Klf4, Myc) overexpression-induced reprogramming Allele-resolved X-linked gene expression / maternal-to-total read ratios Smart-seq2
Fluorescence-activated cell sorting (FACS) Female mouse embryonic fibroblasts (MEFs); FUT4/SSEA-1 marked reprogramming intermediates none (selection of GFP-negative Xi-GFP cells and FUT4+ intermediates) X-GFP allele status and cell surface marker FUT4/SSEA-1
Live/fluorescence and phase contrast imaging Reprogramming cells day 0–12 (Mus Xi-GFP / Cast iPSC system) OSKM reprogramming GFP fluorescence as readout of X-Chromosome reactivation
Single-cell RNA-seq (reanalysis) iPSC reprogramming cells, alternative reprogramming system and genetic background (Schiebinger et al. 2019) reprogramming Pseudotime ordering (Monocle) and X-linked gene reactivation timing per cell
RNA fluorescence in situ hybridization (RNA-FISH) ESCs none Biallelic expression of early X-linked genes
Key results
  • 11% (18/156) of informative X-linked genes reactivate as early as day 8 of reprogramming ('early' genes). 18/156 (11%)
  • X-to-autosome expression ratio progressively increases in FUT4+ intermediates starting day 10, while Chromosomes 2 and 8 do not change.
  • Average Mus/Cast allelic ratio approaches equal biallelic expression by iPSC stage, indicating completed Xi reactivation. log2 Mus/Cast median day13=-1.148, day15=-1.143, iPSCs=-0.144
  • Xist is gradually down-regulated starting day 8, followed by Tsix activation in iPSCs.
  • Reactivation of several early genes occurs in single cells still expressing high Xist, indicating XCR before complete Xist loss.
  • Reactivation of early genes precedes activation of pluripotency gene Prdm14, indicating XCR initiates before full pluripotency network activation.
  • Silenced paternal alleles of imprinted genes Impact and Peg3 become biallelically expressed during reprogramming.
  • Complete allelic information extracted for 156 X-linked genes spanning early, intermediate, late, very late, and escapee classes. 156 genes
Key statistics
  • count 18/156 (11%) (Early reactivated X-linked genes at day 8)
  • count 156 (Informative high-confidence X-linked genes with complete allelic information)
  • other log2 Mus/Cast median = -1.148 (Allelic X expression ratio at day 13)
  • other log2 Mus/Cast median = -1.143 (Allelic X expression ratio at day 15)
  • other log2 Mus/Cast median = -0.144 (Allelic X expression ratio in iPSCs (near-equal biallelic))
  • other ~24 h (Duration of imprinted Xi reversal in epiblast, contrasted with multi-day rXCI reversal)
  • other 0.15–0.85 (Maternal/total ratio range defining biallelic expression in heatmaps)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study uses allele-specific full-transcript Smart-seq2 RNA-seq across a reprogramming time course (day 2, 8, 10, 13, 15, iPSCs, ESCs) plus reanalysis of published single-cell RNA-seq data to track X-Chromosome reactivation. Results are largely reported as descriptive quantitative measures (log2-transformed normalized read counts, allelic ratios of maternal/total reads, X-to-autosome expression ratios, median allelic ratios) and visualized via PCA, heatmaps, Monocle pseudotime ordering, and a generalized additive model fit. From the provided text, formal hypothesis-testing statistics, p-values, and dispersion measures are not explicitly stated for most comparisons.

Replicationunclear Sample size156 informative high-confidence X-linked genes with complete allelic information passing SNP criteria; number of biological/technical replicates per time point not stated in the provided text Groupsreprogramming time points (day 2/8/10/13/15), iPSCs, and ESCs; maternal Mus (Xi) vs paternal Cast (Xa) alleles Pairingna Randomization/blindingnot stated Dispersionunclear
Approaches that could also have been used
  • Group differences (e.g., allelic ratios across time points, or X/A ratios between chromosomes) are presented descriptively with summary statistics such as medians.
    Could also: Pairing the descriptive summaries with formal tests (e.g., Mann-Whitney U / Wilcoxon for two-group comparisons, or Kruskal-Wallis with Dunn's post-hoc across time points) and reporting exact p-values. — Adding inferential statistics and exact p-values would quantify the strength of evidence for the observed differences and complement the descriptive trends; it is a common companion to ratio-based summaries.
  • Central tendency of allelic and X/A ratios is summarized using medians.
    Could also: Reporting an accompanying dispersion measure such as IQR, SD, or a 95% confidence interval alongside the central value. — Explicit dispersion or interval estimates convey the spread and uncertainty of the ratios, which is informative given gene- and cell-level variability.
  • Genes are categorized into reactivation classes (early, intermediate, late, very late, escapee) using fixed allelic-ratio thresholds at specific days (e.g., biallelic at day 8 = early).
    Could also: A model-based clustering or change-point/trajectory-classification approach (e.g., fitting per-gene reactivation curves and clustering their parameters) in addition to threshold-based binning. — A continuous model-based grouping can capture gradations between classes and reduce sensitivity to specific cutoff values, complementing the threshold definitions.
  • Single-cell pseudotime and class-level trends are summarized with a generalized additive model curve.
    Could also: Reporting the GAM with confidence bands, or comparing to alternative smoothers (e.g., LOESS) or mixed-effects models accounting for cell/replicate structure. — Confidence bands around the fitted curve and structure-aware models would convey fit uncertainty and account for non-independence among cells from the same sample.
  • A reprogramming time course is sampled at discrete days with allele-resolution populations.
    Could also: Explicitly stating the number of biological replicates and a sample-size/power rationale. — Describing replication and the basis for n helps readers interpret the reproducibility and statistical resolution of the time-course estimates.
Software: Monocle (single-cell pseudotime ordering) · Generalized additive model (curve fitting; tool not named)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
68
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE126229 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE69823 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-31515287

Paper: Dynamic reversal of random X-Chromosome inactivation during iPSC reprogramming. Genome Research 2019. PMID 31515287 · PMC6771397 · DOI 10.1101/gr.249706.119.

What the paper actually contains (Methods + Data availability)

  • The paper's own data = allele-resolved Smart-seq2 scRNA-seq of female M. m. musculus (X-GFP) × M. m. castaneus (Cast) reprogramming → GEO GSE126229 (NOT the accession assigned to this RU). This is where the headline allele-specific result lives: allelic ratio Mus/(Mus+Cast), 156 X-linked genes with allelic info, reactivation kinetics (Fig 1H, 2A).
  • The paper ALSO re-analyses an external dataset: Schiebinger et al. 2019 (Waddington-OT), the mouse iPSC-reprogramming 10X time course = GEO GSE106340, 65,781 cells, 22 samples (days 0–16, 2i & serum). Used for pseudotime (Monocle 2.10.0) and X:autosome expression dynamics.

Accession assigned to this RU = GSE106340 (the EXTERNAL Schiebinger data)

The brief pins geo:GSE106340. That is the public, downloadable 10X matrix (GSE106340_expression.matrix.flt.nrm.10X.txt.gz, 753 MB + GSE106340_RAW.tar, 488 MB). It is not the allele-resolved data, so the 156-gene Mus/Cast allelic ratio is not derivable from GSE106340 — that needs GSE126229.

Code availability

  • Paper ships NO own analysis code / no GitHub (verbatim data-availability statement names only GEO accessions).
  • The brief's code link kundajelab/atac_dnase_pipelines is a text-mining false positive: this is a Smart-seq2/10X scRNA-seq paper, there is no ATAC/DNase pipeline result to reproduce. Not attempted (non_pipeline for that artifact).
  • P16 path: apply a standard scRNA-seq quantification (mean expression by chromosome) to the paper's named public data GSE106340 — equally valid.

IN SCOPE (low-hanging, light compute on the assigned public data GSE106340)

  • C1 — cohort size. Total cells in the normalized matrix = reported "65,781 cells". Direct matrix dimension. Grade: exact / mismatch.
  • C2 — X-chromosome reactivation signal. Mean X-linked vs mean autosomal expression ratio (X:A) across the reprogramming time course should rise toward iPSC (the paper's central qualitative claim: silenced Xi reactivates → X output increases). Computed from the GSE106340 normalized matrix + per-sample day labels (RAW.tar / sample titles) + mm10 gene→chromosome map (GENCODE M25). Grade: partial (trend reproduces) / mismatch.

OUT OF SCOPE (the hard 20% — stated, not attempted)

  • Allele-specific Mus/Cast ratio, 156 X-linked genes, Fig 1H/2A median allelic ratios (−1.148 d13, −0.144 iPSC): require GSE126229 (different accession, not assigned) — allelic info absent from GSE106340. → data_restricted-to-other-accession.
  • Monocle 2.10.0 pseudotime ordering (heavier; X:A-vs-day is the clean proxy).
  • External ChIP-seq (GSE90893/25409/36905/69823) — out of scope.
  • atac_dnase_pipelines — not used by any reproducible result here.

Pipeline named per in-scope result

  • C1/C2: 10X/Smart-seq2 normalized expression matrix → per-chromosome mean expression aggregation (standard scRNA-seq summarization; numpy/pandas). No bespoke tool required; mm10 GENCODE M25 for gene→chr.
Figures / tables: Fig 1
C1
Reported
65,781 cells
Reproduced
65781
exact
C2
Reported
X output / X:autosome expression ratio rises across reprogramming toward iPSC (Xi reactivation, Fig 1)
Reproduced
X:A median 0.813 (MEF d0) -> 1.229 (d16 2i) -> 1.258 (2i-iPSC); Spearman rho(day vs X:A)=0.842, p=0.0022; rise strongest in 2i (naive) iPSCs
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator headless) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

No genuine discrepancy on the assigned data: C1 (65,781 cells) is an exact match and C2 reproduces the paper's qualitative X-reactivation trend significantly (X:A 0.813→1.258; rho=0.842, p=0.0022). The limitation sits on our side / data scope, not the authors': the assigned accession GSE106340 is the external Schiebinger dataset the paper re-analysed, so the paper pins no exact number to it and we confirm the claim only via a self-chosen X:A proxy. The paper's exact allele-resolved headline lives on GSE126229 (out of scope) — derivable in principle, just not from the assigned data. Severity is negligible and no fabrication concern surfaced; overall a solid, explainable yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

124.1 k
tokens (I/O) · 7.9 M incl. cache
28 min
runtime · 0.24 CPU-h
2.3 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine