Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Cell Specific eQTL Analysis without Sorting Cells.

PLoS Genet · 2015
L1 90/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 79% of all assessed papers rank 212 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the publicly-checkable parts 1:1. The paper's core output is a multi-cohort genotype x expression INTERACTION eQTL meta-analysis; it is NOT reproducible from scratch on public data because 6/7 discovery cohorts (n=5863) are controlled-access and the one public accession (ArrayExpress E-TABM-1036 = DILGOM) is expression-only (no genotypes, no measured cell counts). What we reproduced on «our HPC»/«infra»: (C1) re-derived the headline interaction tallies directly from the shipped per-eQTL S3 Table using the paper's stated FDR<0.05 + interaction-Z-sign rule -> 13124 tested / 1117 significant / 909 neutrophil-mediated / 208 lymphocyte-mediated, ALL EXACT, and independently confirmed by the table's own 'Mediated by cell type' column (generic 12007, neutrophils 909, lymphocytes 208). (C2) re-implemented the neutrophil-proxy construction (S1's 58 HT12v3 probes -> mapped Array_Address_Id to ILMN via GPL6947 -> PCA) on the public DILGOM expression matrix: PC1 explains 64.3% of variance with all-positive coherent loadings, reproducing the premise that these probes form a single neutrophil axis. NOT attempted (and why): the from-scratch interaction-eQTL discovery (controlled raw data), purified-cell replication Fig 3 (wet-lab/controlled), proxy external-validation correlations R=0.75/0.81 (need measured counts absent from public data), and Crohn's enrichment S5 (depends on full discovery). Fabrication check: no discrepancy -- the abstract's tallies are fully internally consistent with the shipped supplementary data, verified two independent ways. Verdict provisional; human audit sheet in AUDIT.md.

💻 Code ↗ 🗄 Data: E-TABM-1036

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-14 ⛓ 3a61b4230861
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Cell-type specific (e.g., neutrophil- or lymphocyte-specific) cis-eQTL effects can be inferred from whole blood gene expression data, without physically sorting cells, by using predicted cell-type proportions as an interaction term in a genotype-by-environment (GxE) meta-analysis.

Core claims
  • A genotype x predicted-cell-count interaction (GxE) meta-analysis across whole blood datasets can detect neutrophil-specific cis-eQTLs without cell sorting method
  • The same approach can predict lymphocyte-specific cis-eQTLs finding
  • Predicted cell-type-specific eQTL effects replicate in independent cell-type specific datasets finding
  • Crohn's disease-associated SNPs preferentially affect gene expression within neutrophils, including at the NOD2 locus finding
  • A 58-probe expression signature, summarized via PCA, provides a proxy for neutrophil percentage in datasets lacking direct cell counts resource
  • 95% of the 58 neutrophil-correlated probes show much higher expression in purified neutrophils than in 13 other purified blood cell types (BLUEPRINT RNA-seq) finding
  • Neutrophils comprise ~60% of white blood cells but had no published eQTL dataset prior to this work due to purification difficulty finding
Experimental setups
Assay System Perturbation Readout Platform
cis-eQTL GxE meta-analysis (genotype x cell-count-proxy interaction, linear model) whole blood, human, 7 cohorts (EGCUT, InCHIANTI, Rotterdam Study, Fehrmann, SHIP-TREND, KORA F4, DILGOM) none (natural genetic variation, interaction with inferred neutrophil proportion) cell-type-mediated vs non-cell-type-mediated cis-eQTL effects on gene expression Illumina HT12v3 expression array (genotyping/array platform per cohort)
Cell-count proxy derivation via correlation + PCA of expression probes whole blood, human, EGCUT cohort (training) and SHIP-TREND (validation), both with actual neutrophil counts none neutrophil percentage proxy (first principal component of 58 probes) Illumina HT12v3 expression array
RNA-seq expression comparison across purified cell types purified neutrophils vs. 13 other purified blood cell types, BLUEPRINT epigenome project none (cell sorting/purification) relative expression level of the 58 neutrophil-correlated probes/genes RNA-seq (BLUEPRINT)
cis-eQTL lookup/annotation whole blood, human none 13,124 previously identified cis-eQTLs used as the eQTL set for cell-type-specificity testing
Key results
  • 58 Illumina HT12v3 probes identified as positively correlated with neutrophil percentage in EGCUT training data Spearman R > 0.57
  • 95% of the 58 probes show much higher expression in purified neutrophils vs. 13 other purified blood cell types 95%
  • Predicted neutrophil percentage (proxy) strongly correlates with actual neutrophil percentage in EGCUT Spearman R=0.75, Pearson R=0.76
  • Predicted neutrophil percentage (proxy) strongly correlates with actual neutrophil percentage in SHIP-TREND (independent validation) Spearman R=0.81, Pearson R=0.82
  • Actual and proxy neutrophil percentage both show weak positive correlation with age in EGCUT actual: Pearson R=0.08, P=0.02; proxy: Pearson R=0.14, P=6x10^-5
  • Neither actual nor proxy neutrophil percentage is associated with gender in EGCUT actual P=0.31; proxy P=0.11
  • Including more or fewer than 58 top probes gives similar proxy-to-actual correlations
Key statistics
  • correlation Spearman R > 0.57 (threshold for selecting 58 neutrophil-correlated probes in EGCUT)
  • correlation Spearman R = 0.75, Pearson R = 0.76 (proxy vs actual neutrophil percentage, EGCUT training cohort)
  • correlation Spearman R = 0.81, Pearson R = 0.82 (proxy vs actual neutrophil percentage, SHIP-TREND validation cohort)
  • correlation Pearson R = 0.08, P = 0.02 (actual neutrophil percentage vs age, EGCUT)
  • correlation Pearson R = 0.14, P = 6x10^-5 (proxy neutrophil percentage vs age, EGCUT)
  • pvalue P = 0.31 (actual), P = 0.11 (proxy) (neutrophil percentage vs gender association test, EGCUT)
  • count 5,863 samples (total unrelated whole blood samples across 7 discovery cohorts)
  • count 13,124 cis-eQTLs (previously discovered whole blood cis-eQTLs used as basis for cell-type-specificity testing)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is a genome-environment interaction (GxE) meta-analysis across 5,683 (reported elsewhere as 5,863) unrelated whole-blood samples from seven cohorts, designed to infer cell-type-specific cis-eQTLs without sorting cells. A cell-count proxy was built from expression probes via principal component analysis and validated against measured cell counts using correlation, and cell-type specificity was tested by fitting a linear model containing a SNP-by-proxy interaction term for a predefined set of 13,124 previously discovered cis-eQTLs. Predictions were replicated in independent cell-type-specific datasets, and disease-associated SNPs were tested for enrichment of cell-type-specific effects.

Replicationbiological Sample size5,683 unrelated whole blood samples from seven discovery cohorts (text also reports 5,863); EGCUT used as training and SHIP-TREND/independent cell-type-specific datasets used for validation/replication; no formal power calculation described Groupsgenotype groups for cis-eQTL effects, modeled as continuous interaction with inferred/measured neutrophil percentage Pairingna Randomization/blindingna Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Spearman rank correlation selecting 58 probes correlated with neutrophil percentage (R>0.57) and validating the neutrophil proxy vs. actual neutrophil percentage (EGCUT R=0.75, SHIP-TREND R=0.81) not stated
Pearson correlation proxy/neutrophil percentage vs. age (EGCUT R=0.08, P=0.02; proxy R=0.14, P=6x10^-5) and proxy vs. actual neutrophil percentage (EGCUT R=0.76, SHIP-TREND R=0.82) not stated
Student's t-test testing association of neutrophil percentage and of the proxy with gender (P=0.31; P=0.11) not stated
Linear regression with a SNP-by-cell-count-proxy interaction term distinguishing cell-type-mediated/specific from generic cis-eQTL effects across the 13,124 cis-eQTLs 5,683 whole blood samples (text also states 5,863) not stated
Principal component analysis (first PC as proxy) summarizing the 58 neutrophil-marker probes into a single neutrophil-percentage estimate per cohort na
Meta-analysis across cohorts combining the seven whole-blood cohorts to detect cell-type-specific interaction effects not stated
Approaches that could also have been used
  • Cell-type specificity was assessed by adding a single SNP-by-proxy interaction term in a linear model.
    Could also: A formal likelihood-ratio or nested-model comparison (full model with interaction vs. reduced model without it), or a mixed-effects model accounting for cohort structure, could also be used. — An explicit nested-model test makes the contribution of the interaction term transparent and a mixed model can pool cohorts while modeling between-cohort heterogeneity directly.
  • The cell-count proxy was derived as the first principal component of selected marker probes.
    Could also: Reference-based deconvolution methods (e.g., CIBERSORT-style regression or constrained least-squares with a cell-type signature matrix) could also estimate cell-type proportions. — Reference-based deconvolution yields proportions on an interpretable scale for multiple cell types simultaneously, which can complement a single-component proxy.
  • Proxy validation relied on Spearman and Pearson correlation coefficients with p-values.
    Could also: Reporting agreement metrics such as a Bland-Altman analysis, concordance correlation, or root-mean-square error with confidence intervals could also summarize prediction quality. — Agreement statistics capture systematic bias and absolute prediction error, which correlation alone does not convey, and CIs would quantify uncertainty in the estimates.
  • Probe selection used a Spearman correlation threshold (R>0.57) to pick 58 marker probes.
    Could also: A cross-validated or penalized feature-selection approach (e.g., elastic net) on the training cohort could also define the marker set. — Cross-validation gives an out-of-sample performance estimate and penalized selection can reduce sensitivity of the proxy to the specific threshold chosen.
  • Several univariate tests (correlation with age, t-test for gender) were reported with individual exact p-values.
    Could also: A single multivariable model including age and gender together, optionally with multiple-testing adjustment across the reported associations, could also be used. — A joint model accounts for covariate correlation simultaneously, and an explicit adjustment communicates the family of comparisons considered.
  • Disease-variant enrichment for cell-type-specific eQTLs is framed as a downstream test (details beyond the available excerpt).
    Could also: Standard enrichment frameworks such as permutation-based null distributions, Fisher's exact tests with FDR control, or matched-SNP background sampling could also be applied. — These approaches provide a calibrated null that accounts for properties like minor allele frequency and gene density when assessing enrichment significance.

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
171
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

1p31 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
AC034220 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
E-MTAB-264 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
E-MTAB-945 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
E-TABM-1036 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE20142 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-25955312

Paper: Westra HJ et al. (2015) Cell Specific eQTL Analysis without Sorting Cells. PLoS Genet 11(5):e1005223. PMID 25955312 · PMCID PMC4425538 · DOI 10.1371/journal.pgen.1005223. Code: https://github.com/molgenis/systemsgenetics (eQTL meta-analysis pipeline) — pinned d9f0387c89531c1fd31513b448498a024707f216 (master, not archived). Data accession (brief): ArrayExpress E-TABM-1036 = DILGOM cohort, whole-blood Illumina HT12 expression, quantile-normalized matrix Files/test.tab (310 MB, ~518 samples). PUBLIC, expression-only (no genotypes, no measured cell counts).

Method (what the paper does)

Cell-type-specific eQTLs are detected with an interaction model on whole-blood data:

  • base: Y ≈ I + β1·G + e
  • interaction: Y ≈ I + β1·G + β2·P + β3·(P·G) + e

where Y=expression, G=genotype, P=neutrophil-percentage proxy. The interaction term β3 (P·G) tests cell-type specificity; interaction Z-scores are meta-analyzed across cohorts (sample-size–weighted Z method). The neutrophil proxy P is built purely from expression: 58 HT12v3 probes correlated with measured neutrophil% (Spearman R>0.57 in EGCUT), then PC1 of those probes is used as the proxy phenotype.

In scope (pipeline-derived, attempted)

id result pipeline step reproducibility from public data
C1 Aggregate interaction tallies: 13,124 cis-eQTLs tested → 1,117 significant at FDR<0.05 (909 neutrophil-mediated/positive + 208 lymphocyte-mediated/negative) FDR threshold + sign classification applied to the shipped per-eQTL interaction table (S3 Table) YES — re-derive from shipped S3 results; a consistency/fabrication check of the headline numbers against the data behind them
C2 Neutrophil proxy is a single coherent expression axis (PC1 of the 58 S1 probes) PCA proxy construction (S1 Table probes) applied to public DILGOM expression (E-TABM-1036) PARTIAL — method re-run on the paper's own public cohort; PC1 variance-explained + sign coherence checkable, but external validation R=0.75/0.81 needs measured counts not in public data

Out of scope (not attempted — and why)

  • Full interaction-eQTL discovery from raw data (the actual β3 Z-scores in S3). Requires individual-level genotype + expression for 7 discovery cohorts (EGCUT, InCHIANTI, Rotterdam, Fehrmann, SHIP-TREND, KORA F4, DILGOM; n=5,863). Six are controlled-access (EGA/dbGaP/on-request); the one public accession (E-TABM-1036/DILGOM) is expression-only — no genotypes. → data_restricted for the full pipeline.
  • Replication in purified cells (neutrophils, CD4/CD8 T, B-cells, monocytes, LCLs; Fig 3) — wet-lab–generated / controlled data. Out of scope.
  • Proxy external-validation correlations R=0.75 (EGCUT)/0.81 (SHIP-TREND) — need measured neutrophil counts paired with expression; not in public E-TABM-1036. Not attempted.
  • Crohn's disease enrichment (S5) — depends on the full discovery output. Not attempted.

Honest framing

The full multi-cohort genotype×expression interaction meta-analysis is not reproducible from publicly obtainable data. What is reproducible and auditable: (C1) re-deriving the reported aggregate counts from the shipped supplementary result table, and (C2) re-running the neutrophil-proxy construction method on the paper's own public DILGOM expression. This is a partial reproduction by design (80/20).

Figures / tables: S2 TableS3 TableS1 Table
C1a
Reported
13124 cis-eQTLs tested
Reproduced
13124
exact
C1b
Reported
1117 interaction eQTLs at FDR<0.05
Reproduced
1117
exact
C1c
Reported
909 neutrophil-mediated
Reproduced
909
exact
C1d
Reported
208 lymphocyte-mediated
Reproduced
208
exact
C2
Reported
neutrophil proxy = PC1 of 58 S1 probes (used as proxy phenotype; ext. R=0.75/0.81)
Reproduced
PC1=64.3% var (next 5.9%), 100% same-sign loadings, mean|pairwise r|=0.63 on public DILGOM; 45/58 probes present
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

A strong partial with no fabrication: the headline interaction-eQTL tallies (13124 tested -> 1117 significant = 909 neutrophil + 208 lymphocyte) re-derived exactly from the shipped S3 table two independent ways, and the neutrophil-proxy PC1 reproduced on public DILGOM (64.3% variance, coherent loadings). The limitations are data-access, not authors' defects: 6/7 discovery cohorts are controlled-access and external proxy correlations (R=0.75/0.81) need measured counts absent from public data, so the from-scratch interaction meta-analysis and Fig 3 were not re-run. q7 is graded limited because the exact tallies are an internal-consistency re-count of the authors' own output rather than an independent regeneration from genotypes.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

165.1 k
tokens (I/O) · 9.7 M incl. cache
16 min
runtime · 0.01 CPU-h
0.5 GB
peak RAM
2
HPC jobs
hummel
machine