Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Newborn sex-specific transcriptome signatures and gestational exposure to fine particles: findings from the ENVIRONAGE birth cohort.

Environ Health · 2017
L1 71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to run the pipeline, NOT well enough to byte-reproduce the headline number. arrayQC_Module is a third-party limma wrapper; I ran the equivalent limma one-color pipeline (read.maimages green.only -> backgroundCorrect subtract -> normalizeBetweenArrays quantile) on all 146 GSE83393 Agilent FES arrays on «our HPC» («job», 1:49). Cohort claims reproduce cleanly: 146 deposited -> 142 analyzed (146-4), 76 girls/66 boys (GEO has 77/67/2). The 16,844-gene claim does NOT reproduce at the literally-stated filter: removing genes flagged in >30% of arrays yields 22,142 genes; 16,844 only emerges under a stricter, undocumented 'present in ~96% of arrays' filter (the value sits inside the reproduced sensitivity band 14,961-36,337). Not a clear fabrication -- most likely an under-described filtering step / different gene-ID namespace -- but flagged for human audit. PM2.5 differential-expression and pathway results (C4-C6) NOT attempted: the per-gene regression needs the paper's full covariate set (maternal age, gestational age, smoking, BMI, season, batch) which is not deposited in GEO; only the PM2.5 exposures are. Honest outcome: PARTIAL -- pipeline reproduced, cohort numbers reproduced, headline gene count within-band but not exact, association results out of scope.

💻 Code ↗ 🗄 Data: GSE83393

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-16 ⛓ eec170683db3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Specific transcriptome profiles in cord blood may arise in response to gestational fine particulate matter (PM2.5) exposure; the study investigates sex-specific transcriptomic responses to long- and short-term gestational PM2.5 exposure to elucidate underlying molecular mechanisms of PM2.5-induced adverse health effects.

Core claims
  • Gestational PM2.5 exposure is associated with sex-specific gene expression changes in newborn cord blood, with major differences between boys and girls. finding
  • This is the first whole genome gene expression study in cord blood to identify sex-specific pathways altered by PM2.5. finding
  • For long-term exposure, neurodevelopment and RhoA pathways were modulated in boys, while defensin expression was down-regulated in girls. mechanism
  • For short-term exposure, pathways related to synaptic transmission and mitochondrial function were altered in boys, and immune response pathways in girls. mechanism
  • Some processes were altered in both sexes (e.g. DNA damage for long-term, olfactory signaling for short-term exposure). mechanism
  • Whole genome microarray of cord blood RNA combined with spatial-temporal interpolation/dispersion PM2.5 exposure modeling enables identification of exposure-associated transcriptome signatures. method
  • The ENVIRONAGE birth cohort cord blood transcriptome dataset (142 mother-newborn pairs, 16,844 genes) is a resource for studying early life exposome effects. resource
Experimental setups
Assay System Perturbation Readout Platform
Whole genome gene expression microarray Cord blood (whole blood) of 142 mother-newborn pairs, ENVIRONAGE birth cohort Gestational ambient PM2.5 exposure (long-term annual average and short-term last month of pregnancy) Genome-wide mRNA expression levels (16,844 genes); gene-PM2.5 association via multivariable linear regression Agilent Whole Human Genome 8 × 60K microarray; Agilent DNA G2505C Scanner; cyanine-3 one-color Quick-Amp labeling
PM2.5 exposure assessment (spatial-temporal interpolation + dispersion modeling) Maternal residential addresses, study area in Flanders/South-East-Limburg, Belgium none (observational exposure estimation) Daily PM2.5 concentration (μg/m3) at high-resolution receptor grid Kriging interpolation with CORINE land cover + IFDM dispersion model
RNA isolation and quality control Cord blood collected in Tempus tubes none RNA yield and RNA Integrity Number (RIN, samples <6 excluded) Tempus Spin RNA Isolation kit; NanoDrop Spectrophotometer; Agilent 2100 Bioanalyzer
Pathway overrepresentation analysis Genes significantly (p<0.05) associated with PM2.5, per sex none (computational) Overrepresented pathways with p-value ConsensusPathDB
Gene set enrichment analysis (GSEA) Genes ranked by log2-fold change, per sex none (computational) Enrichment scores; pathways with q<0.05 and p<0.005 GSEA software (MSigDB version 5.0); Cytoscape 3.2.0 EnrichmentMap
Principal component analysis and partial correlation Significant genes (p<0.05) for long/short-term exposure, per sex none PC scores correlated with PM2.5 exposure (partial correlation R)
Key results
  • 1269 (7.5%) genes showed a significant interaction between PM2.5 and newborn sex for long-term exposure 1269 genes (7.5%)
  • Long-term PM2.5 exposure significantly associated with 1358 genes in boys and 724 genes in girls; 75 differentially expressed in both 1358 (boys), 724 (girls), 75 overlap
  • Short-term PM2.5 exposure differentially affected 432 (2.6%) genes between sexes; 1144 genes in boys and 507 in girls significant, 55 in overlap 432 (2.6%); 1144 boys, 507 girls, 55 overlap
  • Defensins pathway down-regulated in girls for long-term exposure (e.g. DEFA3, DEFB1, DEFA4 down) p=5.7E-04
  • Axon guidance pathway modulated in boys for long-term exposure (neurodevelopment) p=1.4E-02
  • PC1 significantly associated with long-term PM2.5 in girls and boys girls R=0.51 (p<0.0001); boys R=-0.40 (p=0.004)
  • 180 genes in boys and 113 genes in girls significantly associated with both long- and short-term exposure 180 (boys), 113 (girls)
  • TNF receptor signaling and T cell receptor signaling pathways overrepresented in boys for long-term exposure (immune/inflammatory) TNF p=4.8E-03; TCR p=1.8E-02
Key statistics
  • count 142 mother-child pairs (final sample) (Study population after exclusions)
  • count 16,844 genes in final dataset (Genes after preprocessing for statistical analyses)
  • mean 16.0 (range 11.8–20.6) μg/m3 (Long-term (annual average before delivery) PM2.5 exposure)
  • mean 13.3 (range 6.5–34.8) μg/m3 (Short-term (last month of pregnancy) PM2.5 exposure)
  • correlation R=-0.63, p<0.0001 (PC2 partial correlation with long-term PM2.5 in boys)
  • correlation R=0.51, p<0.0001 (PC1 partial correlation with long-term PM2.5 in girls)
  • pvalue 5.7E-04 (Defensins pathway overrepresentation in girls, long-term)
  • count 76 girls (53.5%), 66 boys (46.5%) (Newborn sex distribution)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This observational birth cohort study (ENVIRON AGE, n=142 mother-newborn pairs) used multivariable-adjusted linear regression—with a sex × PM2.5 interaction term—to test associations between gestational PM2.5 exposure (long-term: annual average; short-term: last month of pregnancy) and whole-genome gene expression (16,844 genes) in cord blood, extracting sex-stratified estimates. Significant genes (p<0.05) were submitted to overrepresentation analysis (ConsensusPathDB) and Gene Set Enrichment Analysis (GSEA with FDR correction via gene-set permutation) to identify modulated pathways. Principal component analysis was applied to significant gene sets and partial correlation coefficients between component scores and PM2.5 were reported as a secondary summary of association.

Replicationbiological Sample sizeFinal sample of 142 mother-newborn pairs after microarray quality control and exposure data exclusions; no formal a priori power calculation described GroupsBoys vs. girls; long-term (annual average) vs. short-term (last month) PM2.5 exposure; each of 16,844 genes tested separately Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFDR correction (GSEA only, gene-set permutation-based q-value); no correction stated for gene-level regression across 16,844 genes or for ConsensusPathDB overrepresentation analysis
Statistical tests used
Test Applied to n Assumptions
Multivariable-adjusted linear regression with sex × PM2.5 interaction term; sex-stratified fold changes extracted Association of each of 16,844 genes with long-term (5 μg/m³ increment) and short-term (10 μg/m³ increment) PM2.5 exposure 142 mother-newborn pairs (66 boys, 76 girls) not stated
Principal component analysis (PCA) Dimensionality reduction of genes significant at p<0.05 for each sex and each exposure window 66 boys, 76 girls (separate analyses) na
Partial correlation coefficient (R) Association between principal component scores and long-term and short-term PM2.5 exposure 66 boys, 76 girls (separate analyses) not stated
Overrepresentation analysis (ConsensusPathDB; hypergeometric-type test) Pathway enrichment of genes significantly associated with PM2.5 (p<0.05) for each sex and exposure window; threshold p<0.05 not stated
Gene Set Enrichment Analysis (GSEA) with gene-set permutation test and FDR correction Pathway enrichment using log2-fold-change-ranked gene lists for each sex and exposure window; threshold q<0.05 and p<0.005 16,844 genes ranked per analysis not stated
Single stochastic regression imputation (SAS proc MI, FCS statement) Sensitivity analysis adjusting for white blood cell counts and neutrophil percentage, missing in 31 of 142 newborns stated
Approaches that could also have been used
  • Gene-level associations with PM2.5 were declared significant using a nominal p<0.05 threshold applied simultaneously across 16,844 genes, with no stated correction for multiple comparisons at the gene level
    Could also: Benjamini-Hochberg false discovery rate (FDR) correction applied across all gene-level tests, as is standard practice in microarray studies (e.g., limma with adjusted p-values) — Controlling the FDR at, say, 5% or 10% across tens of thousands of simultaneous tests quantifies the expected proportion of false discoveries among declared significant genes; this is the dominant convention in genome-wide expression analyses and aids interpretation of large gene lists
  • Single stochastic regression imputation was used to handle missing white blood cell and neutrophil data (~22% of observations) in the sensitivity analysis
    Could also: Multiple imputation by chained equations (MICE), generating multiple completed datasets and pooling regression estimates using Rubin's rules — Multiple imputation appropriately propagates uncertainty due to missingness into standard errors and confidence intervals; single imputation treats imputed values as observed, which can underestimate variability — a distinction that becomes more material as the fraction of missing data increases
  • The relationship between continuous PM2.5 exposure and gene expression was modeled as linear across the exposure range
    Could also: Natural or restricted cubic spline terms for PM2.5 within the regression framework, or categorical quartile coding with a test for trend — Spline approaches let the data reveal non-linear exposure-response shapes without imposing proportionality; for environmental exposures, thresholds or diminishing returns are plausible, and visually inspecting fitted splines is a common complementary step
  • Pathway overrepresentation analysis (ORA) in ConsensusPathDB used a p<0.05 threshold with no stated correction for the number of pathways tested
    Could also: FDR correction (e.g., Benjamini-Hochberg) applied to ORA pathway p-values, consistent with the approach used in the GSEA branch of the same analysis — Applying a uniform multiple-testing standard across both enrichment methods — ORA and GSEA — facilitates consistent interpretation and limits the expected rate of spurious pathway findings when many pathways are evaluated simultaneously
  • Sex-specific responses were characterized by including a sex × PM2.5 interaction term in a single combined model and then extracting sex-stratified fold changes
    Could also: Fully stratified models fit separately for boys and girls, with interaction p-values explicitly reported per gene to indicate the strength of statistical support for sex-differential effects — Reporting the per-gene interaction p-value alongside stratified estimates is a common complementary approach; it clarifies which genes show statistically supported sex-differential associations rather than all genes with any sex-specific nominal significance
  • Results were summarized at the gene level using fold changes and p-values only, with no confidence intervals for the regression coefficients
    Could also: 95% confidence intervals for the fold changes (or log2-fold changes) associated with each PM2.5 increment, at least for the top reported genes — Confidence intervals convey both the direction and precision of the estimated effect; for studies with modest n (142 total, ~66–76 per sex), interval width provides important context for assessing the stability of individual gene estimates
Software: R with in-house arrayQC pipeline 2.15.3 · Agilent Feature Extraction Software 10.7.3.1 · GSEA (MSigDB) 5.0 · ConsensusPathDB (online, Max Planck Institute for Molecular Genetics) · Cytoscape with EnrichmentMap plugin 3.2.0 · SAS (proc MI)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
32
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-28583124

Paper: Winckelmans et al. 2017, Environ Health 16:52. "Newborn sex-specific transcriptome signatures and gestational exposure to fine particles: findings from the ENVIRONAGE birth cohort." PMID 28583124 / PMC5458481 / DOI 10.1186/s12940-017-0264-y.

Data: GEO GSE83393 — 146 deposited Agilent one-color FES arrays (GPL17077, Agilent SurePrint G3 Human GE v2 8x60K, Cy3 single channel), cord blood, ENVIRONAGE cohort. Raw FES .txt.gz per sample + GSE83393_RAW.tar.

Code: https://github.com/BiGCAT-UM/arrayQC_Module @ commit 350f023abdc8a749f4fddbd4b1f40ea71dd809b3 (2015-11-23). A third-party generic QC/normalization tool for Agilent/GenePix spotted arrays from the BiGCAT/Maastricht group (same dept as authors de Kok/Kleinjans). It is a thin wrapper around limma: read.maimages (green.only for one-color) → background correction → control omission → bad-spot flagging (gIsWellAboveBG) → log2 → normalizeBetweenArrays (method="quantile" for one-color). Output = normalized expression matrix + QC plots. Per brief P16, applying this tool to the paper's own data is a valid reproduction.

Methods pipeline (from paper Methods)

"Gene expression ... arrayQC (R 2.15.3): local background correction, control omission, bad spot flagging, log2 transformation and quantile normalization. ... Quality control resulted in exclusion of four newborns. ... genes with >30% flagged data were removed; for genes with multiple probes the probe with the largest interquartile range (IQR) was retained, yielding 16,844 genes."

In scope (pipeline-derived, attempted)

id result reported pipeline difficulty
C1 # arrays deposited / analyzed 146 deposited; 142 analyzed (4 excluded by QC) GEO metadata + arrayQC array-level QC easy (input verifiable; exact 4 = threshold-dependent)
C2 sex split of analyzed set 76 girls / 66 boys (142) GEO Sex: characteristic easy (metadata-derived)
C3 # genes after normalization + flag-filter + IQR collapse 16,844 genes limma one-color pipeline (= arrayQC) on the 146 FES files medium — the clean compute target

Out of scope (the hard ~20%, not attempted or only noted)

  • PM2.5 association DE counts (long-term girls 724 / boys 1358 / overlap 75; short-term 507 / 1144 / overlap 55; interaction 1269 = 7.5%): require the full per-gene regression with the paper's covariate set (maternal age, gestational age, smoking, BMI, season, batch, ...). Only PM2.5 long/short values are in GEO characteristics; the remaining covariates are NOT deposited → not faithfully reproducible. NOT attempted (would need fabricated covariates).
  • Pathway / GSEA tables (Tables 2–5): downstream of the DE lists, out of scope.
  • Which specific 4 newborns failed QC: arrayQC array-level exclusion is threshold/visual → not exactly reproducible; we corroborate the count + arithmetic only (146 deposited; 2 have no submitter sex; 146−4=142, 77−1=76 girls, 67−1=66 boys).

Plan

One «our HPC» SLURM job: download GSE83393_RAW.tar to «infra», extract 146 FES files, run limma one-color pipeline (the arrayQC steps), apply the paper's flag-filter (>30% not-well-above-BG removed) + largest-IQR-per-gene collapse, report the gene count and per-array flagged fraction. Compare to 16,844 / 142 / 76+66.

Figures / tables: Table
C1
Reported
146 arrays deposited; 142 newborns analyzed (4 excluded by QC)
Reproduced
146 FES arrays read+normalized; 4 worst-flagged arrays (50.6/47.9/47.8/46.4%) plausibly the 4 QC exclusions -> 142
within tolerance
C2
Reported
76 girls / 66 boys (n=142)
Reproduced
GEO: 77 girls / 67 boys / 2 unspecified (n=146) -> 76/66 after removing the 4
exact
C3
Reported
16844 genes after normalization + >30%-flag filter + largest-IQR collapse
Reproduced
22142 at the stated >30% rule; sweep 14961(100%-present)..19311(90%)..22142..36337(all)
partial
C4-C6
Reported
PM2.5 DE counts 724/1358/75, 507/1144/55, interaction 1269 (7.5%)
Reproduced
NOT ATTEMPTED (covariates not deposited in GEO)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

Cohort claims reproduce cleanly — 146 deposited→142 analyzed (146−4) and 76 girls/66 boys derive exactly from GEO metadata plus the 4 QC exclusions. The headline 16,844-gene count is not byte-reproducible: the paper's stated '>30% flagged' rule gives 22,142, and 16,844 only appears under an undocumented '~96%-present' filter (within the 14,961–36,337 sweep) — an authors'-side under-specification, not a fabrication signal. The central conclusion (sex-specific PM2.5 transcriptome signatures, C4–C6) could not be tested at all because the required covariates are not deposited in GEO, so the core claim is neither confirmed nor refuted. Overall a solid pipeline-level reproduction with explainable, mostly authors'-side deviations → partial.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

153.4 k
tokens (I/O) · 13.7 M incl. cache
24 min
runtime · 0.04 CPU-h
1.3 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine