Immuno-detection by sequencing enables large-scale high-dimensional phenotyping in cells.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce 1:1. Re-ran the authors' own IDseq R package (github.com/jessievb/IDseq @6cbfa5b) on the authors' own raw FASTQ (GEO GSM2671836 / SRA SRR5705251-254, ID-seq spike-in DNA-tags) with the authors' own config index files, all compute on «our HPC» (SLURM 2176463). The regenerated per-(antibody,well) unique-UMI count table matches the deposited GEO supplementary count table essentially bit-for-bit: 2 of 4 spike conditions identical, the other 2 within 4-6 UMI of ~1M; overall 624/634 cells exactly equal, Pearson=Spearman=1.000, total 7,134,140 vs deposited 7,134,130 (rel diff 1.4e-6). Regex matched 98.8% of 8.24M reads. Strong positive evidence AGAINST fabrication for this result. NOT attempted (out of scope, see scope.md): downstream dose-response/EC50, PKIS kinase-inhibitor screen, EGF time-series, signal-to-background, clustering figures - these are manual/statistical analyses not shipped in the IDseq package. Only the spike-in sample (smallest) was processed; the other 6 GEO samples use the identical pipeline and would reproduce the same way.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 96assessed: 2026-06-14 ⛓ c4007f3775f0
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a DNA-tagged antibody plus high-throughput sequencing technology (ID-seq) enable accurate, large-scale, high-dimensional measurement of many (phospho-)proteins across many samples simultaneously, and be used to dissect the role of kinases in human epidermal stem cell renewal and differentiation?
- ★ ID-seq combines antibody-based protein detection with DNA-sequencing of DNA-tagged antibodies to measure large numbers of (phospho-)proteins in many samples in parallel method
- ★ ID-seq allows precise, sensitive and specific multiplexed quantification of up to 84 (phospho-)proteins in hundreds of samples simultaneously with a four-order-of-magnitude dynamic range finding
- ★ Multiplexing does not interfere with antibody detection, as singleplex and multiplexed measurements correspond closely finding
- ★ A generalised linear mixed model accounting for negative binomial count distribution enables identification of treatment effects on each antibody signal method
- ★ Decreased mTOR signalling is associated with increased keratinocyte differentiation mechanism
- ★ Screening ~300 PKIS kinase inhibitor probes (targeting 225 kinases) uncovers 13 kinases potentially regulating epidermal renewal through distinct mechanisms finding
- ★ A 70 antibody–DNA conjugate panel covering cell cycle, apoptosis, DNA damage, epidermal differentiation and multiple signalling pathways serves as a broadly applicable resource resource
- EGFR inhibition with AG1478 induces keratinocyte differentiation with concurrent BMP and Notch pathway activation driven by changes in mRNA expression mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ID-seq (immuno-detection by sequencing via DNA-tagged antibodies) | primary human epidermal stem cells (keratinocytes) | AG1478 EGFR inhibition (10 μM, 48 h) | counts of antibody-coupled DNA barcodes quantifying (phospho-)protein levels | next-generation sequencing; dsDNA tag with 10-nt barcode + 15-nt UMI |
| ID-seq PKIS screen | human epidermal keratinocytes in 384-well plates | 294 PKIS kinase inhibitor probes (24 h) | 70-antibody (phospho-)protein phenotype profiles per probe | next-generation sequencing |
| immuno-PCR (singleplex epitope detection) | fixed cell populations | none / antibody dilutions / IgG controls | single antibody DNA-tag signal for comparison to ID-seq | — |
| in-cell-western / immunofluorescence (IF) | primary skin stem cells / fixed cells | differentiation, EGF/BMP stimulation, DNA damage (mitomycin C, hydroxyurea), AG1478, DMH1, phosphatase treatment | antibody signal validation and TGM1 differentiation marker level | — |
| colony formation assay with automated image analysis | epidermal stem cells | 18 high-PC2 PKIS probes (n=3 replicates) | colony number, colony size/distribution, TGM1 level per colony | — |
| RT-qPCR | differentiating human keratinocytes | AG1478, increasing cell density, RAPTOR siRNA silencing | mRNA levels (TGM1, PPL, ID2, HES2, BMP ligands, NOTCH receptors, RAPTOR, mTOR) | — |
| quantitative proteomics | keratinocytes | EGFR inhibition (differentiation) | TGM1 protein level | — |
| siRNA-mediated silencing | primary skin stem cells | siRNA knockdown of selected proteins / RAPTOR | epitope abundance-dependent decrease in antibody-barcode counts; differentiation marker expression | — |
- – Singleplex (immuno-PCR) and multiplexed (ID-seq) measurements of 17 antibodies are highly correlated, showing multiplexing does not interfere with detection R = 0.98 ± 0.046
- – ID-seq library preparation is highly reproducible across separate preparations and sequencing runs R = 0.98
- – Signal variability of 69 antibody–DNA conjugates was below 20% across biological replicates CV < 0.2
- – AG1478 treatment significantly altered (phospho-)protein levels, identifying 13 increased and 7 decreased proteins including upregulated TGM1 and NOTCH1 13 up / 7 down
- ▲ PC2 from PCA of screen data captures differentiation, correlating with TGM1, NOTCH1, SMAD3, Cyclin B1 and GAPDH upregulation
- ▼ High-PC2 (differentiating) probes show strong downregulation of mTOR pathway activity (phospho-mTOR, phospho-S6)
- – 15 of 18 high-PC2 probes showed a significant effect on at least one colony phenotype, validating PC2 as a marker of differentiation 15/18 probes
- – Replicate PKIS screens were highly correlated with low UMI duplicate rates, indicating high data quality R = 0.98; 1.2% UMI duplicates
- correlation R = 0.98 ± 0.046 (singleplex immuno-PCR vs multiplexed ID-seq across 17 antibodies)
- correlation R = 0.98 (reproducibility of PCR-based ID-seq library preparation)
- correlation R > 0.99 (technical replicates using nine distinct DNA tag sequences per antibody)
- other CV < 0.2 (below 20% variation) (variability of 69 antibody–DNA conjugates across 14 biological replicates)
- count 13 increased and 7 decreased (phospho-)proteins (p < 0.01, ANOVA) (AG1478 treatment effect (n = 6))
- fold_change ~75-fold signal over no-cell background (84 antibodies signal over technical noise)
- correlation R = 0.98 (replicate PKIS screens correlation; 1.2% UMI duplicate rate)
- count 294 PKIS compounds targeting 225 kinases; 70-antibody panel (PKIS screen scope in 384-well plates, 24 h treatment)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper introduces ID-seq, an antibody-DNA barcode sequencing technology, and applies it to screen ~294 kinase inhibitor probes across ~70 (phospho-)protein phenotypes in primary human keratinocytes. The primary statistical model was a generalised linear mixed model (GLM) with negative binomial error distribution and likelihood ratio testing (referred to by the authors as ANOVA) applied per antibody to quantify compound effects. Principal component analysis (PCA) was then used to aggregate correlated phenotypic measurements into interpretable biological axes, and two-sample t-tests with 1% FDR correction identified molecular features distinguishing differentiating from non-differentiating cell states. Results were reported as effect estimates and -log10 p-values on volcano plots, with boxplots using median and IQR for distribution summaries.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Generalised linear mixed model (negative binomial distribution) with likelihood ratio test, described by authors as ANOVA | Effect of AG1478 treatment and individual PKIS probes on each of ~70 antibody phenotypes (Fig. 1e and PKIS screen) | n = 6 biological replicates for AG1478 experiment; n per PKIS probe not explicitly stated | not stated |
| Two-sample t-test with 1% FDR correction | Comparison of ~70 molecular phenotypes between high-PC2 (top 10%) and low-PC2 (bottom 10%) PKIS probe groups (Fig. 3a) | Top and bottom 10% of 294 probes (~29 probes per group); exact per-group n not stated | not stated |
| Pearson correlation (r / R) | Singleplex (immuno-PCR) vs. multiplex (ID-seq) signal concordance (Fig. 1b); library preparation reproducibility (Fig. 1c); PKIS replicate screen concordance (Supplementary Fig. 13) | n = 17 antibodies for singleplex/multiplex comparison; n = 4 for insert panel example; n not stated for library reproducibility | na |
| Principal component analysis (PCA) | Aggregation of signed log10 p-values from 294 PKIS probes × 70 antibody analyses to identify axes of biological variation (Fig. 2b) | 294 PKIS probes × 70 antibody phenotypes | na |
-
A custom negative binomial GLM was implemented (Supplementary Note 3) to model antibody barcode count data↳ Could also: Use established Bioconductor packages such as DESeq2 or edgeR, which also model sequencing count data with negative binomial distributions and include built-in normalization, dispersion shrinkage, and likelihood ratio or Wald testing — These packages provide extensively peer-reviewed implementations with robust empirical Bayes dispersion estimation that can improve stability at small sample sizes (n = 6); their assumptions and normalization steps are fully documented, which facilitates comparison across studies
-
A p < 0.01 threshold was applied across ~70 simultaneous antibody-level tests in the AG1478 ANOVA analysis without a stated multiplicity correction for the panel↳ Could also: Apply Benjamini-Hochberg FDR or Bonferroni correction across the 70 simultaneous tests — With 70 tests at α = 0.01, approximately 0.7 false positives would be expected by chance; an explicit correction quantifies and bounds this rate, which is informative when interpreting the 20 significant hits reported and is consistent with the FDR approach the authors used elsewhere
-
Two-sample t-tests were used to compare high-PC2 vs. low-PC2 probe groups on antibody-derived measurements↳ Could also: Mann-Whitney U (Wilcoxon rank-sum) test, which does not assume normality of the underlying measurements — Antibody count-derived phenotype scores may not be normally distributed, particularly with small group sizes; a non-parametric alternative makes fewer distributional assumptions, at the cost of modestly reduced power when normality holds
-
Unsupervised PCA was used to construct a composite differentiation score (PC2) from the full 70-antibody panel↳ Could also: Supervised dimensionality reduction such as partial least squares discriminant analysis (PLS-DA), or sparse PCA that selects a minimal antibody subset — PCA maximizes total variance regardless of biological class structure; supervised or sparse alternatives could identify the antibody combinations most discriminative of differentiation status, potentially yielding a more interpretable and parsimonious composite score
-
Effect sizes were reported as point estimates on volcano plots without accompanying uncertainty intervals↳ Could also: Report 95% confidence intervals around each effect estimate alongside or instead of p-value thresholds — Confidence intervals convey both the direction and precision of each effect; this is particularly informative when effects are modest, as the authors note for TGM1 upregulation, and allows readers to assess practical as well as statistical significance
-
Dispersion was reported as standard deviation in some contexts (Fig. 1b) and as coefficient of variation in others (Fig. 1d), with boxplots elsewhere↳ Could also: Use a single, consistent dispersion metric — such as 95% confidence intervals or standard error of the mean — across all result summaries — Consistent use of one dispersion metric simplifies cross-figure comparison; confidence intervals are often preferred for small n because they simultaneously capture variability and sample size, making it easier to judge the precision of each estimate
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
15 of 18 high-PC2 kinase inhibitor probes significantly alter at least one colony phenotype (colony number, size distribution, or TGM1 level) in human keratinocyte colony formation assays, orthogonally validating the ID-seq differentiation axisimaging human keratinocyte 2018×1papers★ This paper is the founder (earliest)
-
ID-seq multiplex antibody-DNA barcode measurements correlate with singleplex immuno-PCR across 17 antibodies (R=0.98 ± 0.046), demonstrating that multiplexing does not interfere with individual epitope detectionother human keratinocyte 2018×1papers★ This paper is the founder (earliest)
-
Differentiation-inducing high-PC2 kinase inhibitor probes suppress mTOR pathway activity, evidenced by downregulation of phospho-MTOR and phospho-RPS6 in human keratinocytes measured by ID-seqother human keratinocyte down 2018×1papers★ This paper is the founder (earliest)
-
PC2 from PCA of a 294-compound PKIS kinase-inhibitor ID-seq screen captures a differentiation axis in human keratinocytes, with SMAD3, TGM1, NOTCH1, CCNB1, and GAPDH showing high positive loadingsother human keratinocyte up 2018×1papers★ This paper is the founder (earliest)
-
EGFR inhibition (AG1478) upregulates TGM1 and NOTCH1 as part of a mixed (phospho-)proteome response (13 proteins up, 7 down) in primary human epidermal keratinocytes, detected by ID-seqother human keratinocyte up 2018×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-29921844 (ID-seq, van Buggenum et al. 2018, Nat Commun)
- Paper: "Immuno-detection by sequencing enables large-scale high-dimensional phenotyping in cells." DOI 10.1038/s41467-018-04761-0
- Code: https://github.com/jessievb/IDseq (R package, MIT/GPL-3, default branch
master, latest commit6cbfa5b7073fa99945498f23debe11e3415cdd1b2021-06-30). GEO data_processing names the same tool as "R-package immunoSeq version 1.0.0" (the package was renamed IDseq). - Data: GEO GSE100135 / SRA SRP109587 / BioProject PRJNA390781.
The pipeline (what the repo does)
The IDseq R package turns raw single-end FASTQ into a per-(antibody, well) UMI count table:
IDseq_split_reads()— for each FASTQ, regex-extract from each read:[ACTGN]+ (UMI 15nt)(Barcode_1=antibody 10nt)(ANCHOR ATCAGTCAACAGATAAGCGA)(Barcode_2=well 10nt) [ACTGN]+. Reads without an exact match get approximate matching (aregexec, max.distance = 2). Writes split table (UMI, Barcode_1, Barcode_2, sample_folder).IDseq_umi_count()+IDseq_barcode_count()— collapse duplicate UMIs and count the number of UNIQUE UMI strings per (Barcode_1, Barcode_2, sample_folder).IDseq_barcode_match()— left-join the count table to the experiment'santibody_barcode_index.txt(on Barcode_1) andwell_barcode_index.txt(on Barcode_2) →barcode_count_matched.tsv.
IN SCOPE (attempted) — pipeline-derived, directly checkable
Re-run the authors' own IDseq package on the authors' own raw FASTQ, with the
authors' own config index files, and compare the regenerated count table to the
deposited barcode_count_matched.tsv for that GEO sample.
Target sample: GSM2671836 ("ID-seq spike-in DNA-tags", synthetic construct — smallest, cleanest sample, ~8.2M reads over 4 SRA runs SRR5705251–254 = 4 spike conditions / sample_folders spike25–28). Deposited output: 40 antibody barcodes × 4 well barcodes × 4 folders = 624 rows; total unique-UMI count = 7,134,130.
Comparison metrics (provisional grades; human decides):
- total unique-UMI count over the whole sample (one scalar) vs deposited;
- per-cell agreement: each reproduced folder optimally matched to a deposited spike folder, then Pearson/Spearman correlation and exact-equal fraction of the per-(Barcode_1,Barcode_2) counts.
OUT OF SCOPE (not attempted) — downstream / wet-lab / manual
- Dose-response curves, EC50/potency estimates, kinase-inhibitor (PKIS) screen hits, EGF time-series dynamics, signal-to-background, clustering/heatmaps (Figs 2–6): these are downstream statistical/normalisation analyses not in the shipped IDseq package and depend on manual normalisation steps not in the repo.
- Wet-lab steps (antibody–DNA conjugation, staining, validation).
- Reason for skipping: the repo ships only the FASTQ→count-table pipeline; the count table is the foundational, unambiguous pipeline output. The 80% with a clearly specified, runnable pipeline.
Hard rules compliance
- All compute on «our HPC» (SLURM via front1). Data + envs on «infra»
«path». «host» holds only small results + pointers.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Re-ran the authors' own IDseq R package (pinned commit) on the authors' own raw FASTQ (GSM2671836 / SRR5705251-254) with their config index files. The regenerated unique-UMI count table matches the deposited GEO table essentially bit-for-bit: spike26 (3,103,497) and spike28 (2,192,047) identical, all four conditions Pearson=Spearman=1.000, 624/634 cells exactly equal, total 7,134,140 vs 7,134,130 (rel diff 1.4e-6). The only differences are 4-6 UMI on two conditions from borderline approximate-match/N-containing reads at the regex boundary — immaterial. Strong positive evidence against fabrication; a clean 1:1 reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.