Newborn sex-specific transcriptome signatures and gestational exposure to fine particles: findings from the ENVIRONAGE birth cohort.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to run the pipeline, NOT well enough to byte-reproduce the headline number. arrayQC_Module is a third-party limma wrapper; I ran the equivalent limma one-color pipeline (read.maimages green.only -> backgroundCorrect subtract -> normalizeBetweenArrays quantile) on all 146 GSE83393 Agilent FES arrays on «our HPC» («job», 1:49). Cohort claims reproduce cleanly: 146 deposited -> 142 analyzed (146-4), 76 girls/66 boys (GEO has 77/67/2). The 16,844-gene claim does NOT reproduce at the literally-stated filter: removing genes flagged in >30% of arrays yields 22,142 genes; 16,844 only emerges under a stricter, undocumented 'present in ~96% of arrays' filter (the value sits inside the reproduced sensitivity band 14,961-36,337). Not a clear fabrication -- most likely an under-described filtering step / different gene-ID namespace -- but flagged for human audit. PM2.5 differential-expression and pathway results (C4-C6) NOT attempted: the per-gene regression needs the paper's full covariate set (maternal age, gestational age, smoking, BMI, season, batch) which is not deposited in GEO; only the PM2.5 exposures are. Honest outcome: PARTIAL -- pipeline reproduced, cohort numbers reproduced, headline gene count within-band but not exact, association results out of scope.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-16 ⛓ eec170683db3
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusSpecific transcriptome profiles in cord blood may arise in response to gestational fine particulate matter (PM2.5) exposure; the study investigates sex-specific transcriptomic responses to long- and short-term gestational PM2.5 exposure to elucidate underlying molecular mechanisms of PM2.5-induced adverse health effects.
- ★ Gestational PM2.5 exposure is associated with sex-specific gene expression changes in newborn cord blood, with major differences between boys and girls. finding
- ★ This is the first whole genome gene expression study in cord blood to identify sex-specific pathways altered by PM2.5. finding
- ★ For long-term exposure, neurodevelopment and RhoA pathways were modulated in boys, while defensin expression was down-regulated in girls. mechanism
- ★ For short-term exposure, pathways related to synaptic transmission and mitochondrial function were altered in boys, and immune response pathways in girls. mechanism
- ★ Some processes were altered in both sexes (e.g. DNA damage for long-term, olfactory signaling for short-term exposure). mechanism
- Whole genome microarray of cord blood RNA combined with spatial-temporal interpolation/dispersion PM2.5 exposure modeling enables identification of exposure-associated transcriptome signatures. method
- The ENVIRONAGE birth cohort cord blood transcriptome dataset (142 mother-newborn pairs, 16,844 genes) is a resource for studying early life exposome effects. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole genome gene expression microarray | Cord blood (whole blood) of 142 mother-newborn pairs, ENVIRONAGE birth cohort | Gestational ambient PM2.5 exposure (long-term annual average and short-term last month of pregnancy) | Genome-wide mRNA expression levels (16,844 genes); gene-PM2.5 association via multivariable linear regression | Agilent Whole Human Genome 8 × 60K microarray; Agilent DNA G2505C Scanner; cyanine-3 one-color Quick-Amp labeling |
| PM2.5 exposure assessment (spatial-temporal interpolation + dispersion modeling) | Maternal residential addresses, study area in Flanders/South-East-Limburg, Belgium | none (observational exposure estimation) | Daily PM2.5 concentration (μg/m3) at high-resolution receptor grid | Kriging interpolation with CORINE land cover + IFDM dispersion model |
| RNA isolation and quality control | Cord blood collected in Tempus tubes | none | RNA yield and RNA Integrity Number (RIN, samples <6 excluded) | Tempus Spin RNA Isolation kit; NanoDrop Spectrophotometer; Agilent 2100 Bioanalyzer |
| Pathway overrepresentation analysis | Genes significantly (p<0.05) associated with PM2.5, per sex | none (computational) | Overrepresented pathways with p-value | ConsensusPathDB |
| Gene set enrichment analysis (GSEA) | Genes ranked by log2-fold change, per sex | none (computational) | Enrichment scores; pathways with q<0.05 and p<0.005 | GSEA software (MSigDB version 5.0); Cytoscape 3.2.0 EnrichmentMap |
| Principal component analysis and partial correlation | Significant genes (p<0.05) for long/short-term exposure, per sex | none | PC scores correlated with PM2.5 exposure (partial correlation R) | — |
- – 1269 (7.5%) genes showed a significant interaction between PM2.5 and newborn sex for long-term exposure 1269 genes (7.5%)
- – Long-term PM2.5 exposure significantly associated with 1358 genes in boys and 724 genes in girls; 75 differentially expressed in both 1358 (boys), 724 (girls), 75 overlap
- – Short-term PM2.5 exposure differentially affected 432 (2.6%) genes between sexes; 1144 genes in boys and 507 in girls significant, 55 in overlap 432 (2.6%); 1144 boys, 507 girls, 55 overlap
- ▼ Defensins pathway down-regulated in girls for long-term exposure (e.g. DEFA3, DEFB1, DEFA4 down) p=5.7E-04
- – Axon guidance pathway modulated in boys for long-term exposure (neurodevelopment) p=1.4E-02
- – PC1 significantly associated with long-term PM2.5 in girls and boys girls R=0.51 (p<0.0001); boys R=-0.40 (p=0.004)
- – 180 genes in boys and 113 genes in girls significantly associated with both long- and short-term exposure 180 (boys), 113 (girls)
- – TNF receptor signaling and T cell receptor signaling pathways overrepresented in boys for long-term exposure (immune/inflammatory) TNF p=4.8E-03; TCR p=1.8E-02
- count 142 mother-child pairs (final sample) (Study population after exclusions)
- count 16,844 genes in final dataset (Genes after preprocessing for statistical analyses)
- mean 16.0 (range 11.8–20.6) μg/m3 (Long-term (annual average before delivery) PM2.5 exposure)
- mean 13.3 (range 6.5–34.8) μg/m3 (Short-term (last month of pregnancy) PM2.5 exposure)
- correlation R=-0.63, p<0.0001 (PC2 partial correlation with long-term PM2.5 in boys)
- correlation R=0.51, p<0.0001 (PC1 partial correlation with long-term PM2.5 in girls)
- pvalue 5.7E-04 (Defensins pathway overrepresentation in girls, long-term)
- count 76 girls (53.5%), 66 boys (46.5%) (Newborn sex distribution)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This observational birth cohort study (ENVIRON AGE, n=142 mother-newborn pairs) used multivariable-adjusted linear regression—with a sex × PM2.5 interaction term—to test associations between gestational PM2.5 exposure (long-term: annual average; short-term: last month of pregnancy) and whole-genome gene expression (16,844 genes) in cord blood, extracting sex-stratified estimates. Significant genes (p<0.05) were submitted to overrepresentation analysis (ConsensusPathDB) and Gene Set Enrichment Analysis (GSEA with FDR correction via gene-set permutation) to identify modulated pathways. Principal component analysis was applied to significant gene sets and partial correlation coefficients between component scores and PM2.5 were reported as a secondary summary of association.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Multivariable-adjusted linear regression with sex × PM2.5 interaction term; sex-stratified fold changes extracted | Association of each of 16,844 genes with long-term (5 μg/m³ increment) and short-term (10 μg/m³ increment) PM2.5 exposure | 142 mother-newborn pairs (66 boys, 76 girls) | not stated |
| Principal component analysis (PCA) | Dimensionality reduction of genes significant at p<0.05 for each sex and each exposure window | 66 boys, 76 girls (separate analyses) | na |
| Partial correlation coefficient (R) | Association between principal component scores and long-term and short-term PM2.5 exposure | 66 boys, 76 girls (separate analyses) | not stated |
| Overrepresentation analysis (ConsensusPathDB; hypergeometric-type test) | Pathway enrichment of genes significantly associated with PM2.5 (p<0.05) for each sex and exposure window; threshold p<0.05 | — | not stated |
| Gene Set Enrichment Analysis (GSEA) with gene-set permutation test and FDR correction | Pathway enrichment using log2-fold-change-ranked gene lists for each sex and exposure window; threshold q<0.05 and p<0.005 | 16,844 genes ranked per analysis | not stated |
| Single stochastic regression imputation (SAS proc MI, FCS statement) | Sensitivity analysis adjusting for white blood cell counts and neutrophil percentage, missing in 31 of 142 newborns | — | stated |
-
Gene-level associations with PM2.5 were declared significant using a nominal p<0.05 threshold applied simultaneously across 16,844 genes, with no stated correction for multiple comparisons at the gene level↳ Could also: Benjamini-Hochberg false discovery rate (FDR) correction applied across all gene-level tests, as is standard practice in microarray studies (e.g., limma with adjusted p-values) — Controlling the FDR at, say, 5% or 10% across tens of thousands of simultaneous tests quantifies the expected proportion of false discoveries among declared significant genes; this is the dominant convention in genome-wide expression analyses and aids interpretation of large gene lists
-
Single stochastic regression imputation was used to handle missing white blood cell and neutrophil data (~22% of observations) in the sensitivity analysis↳ Could also: Multiple imputation by chained equations (MICE), generating multiple completed datasets and pooling regression estimates using Rubin's rules — Multiple imputation appropriately propagates uncertainty due to missingness into standard errors and confidence intervals; single imputation treats imputed values as observed, which can underestimate variability — a distinction that becomes more material as the fraction of missing data increases
-
The relationship between continuous PM2.5 exposure and gene expression was modeled as linear across the exposure range↳ Could also: Natural or restricted cubic spline terms for PM2.5 within the regression framework, or categorical quartile coding with a test for trend — Spline approaches let the data reveal non-linear exposure-response shapes without imposing proportionality; for environmental exposures, thresholds or diminishing returns are plausible, and visually inspecting fitted splines is a common complementary step
-
Pathway overrepresentation analysis (ORA) in ConsensusPathDB used a p<0.05 threshold with no stated correction for the number of pathways tested↳ Could also: FDR correction (e.g., Benjamini-Hochberg) applied to ORA pathway p-values, consistent with the approach used in the GSEA branch of the same analysis — Applying a uniform multiple-testing standard across both enrichment methods — ORA and GSEA — facilitates consistent interpretation and limits the expected rate of spurious pathway findings when many pathways are evaluated simultaneously
-
Sex-specific responses were characterized by including a sex × PM2.5 interaction term in a single combined model and then extracting sex-stratified fold changes↳ Could also: Fully stratified models fit separately for boys and girls, with interaction p-values explicitly reported per gene to indicate the strength of statistical support for sex-differential effects — Reporting the per-gene interaction p-value alongside stratified estimates is a common complementary approach; it clarifies which genes show statistically supported sex-differential associations rather than all genes with any sex-specific nominal significance
-
Results were summarized at the gene level using fold changes and p-values only, with no confidence intervals for the regression coefficients↳ Could also: 95% confidence intervals for the fold changes (or log2-fold changes) associated with each PM2.5 increment, at least for the top reported genes — Confidence intervals convey both the direction and precision of the estimated effect; for studies with modest n (142 total, ~66–76 per sex), interval width provides important context for assessing the stability of individual gene estimates
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Axon guidance pathway genes show differential expression in boys' newborn cord blood associated with long-term gestational PM2.5 exposure.microarray human cord-blood mixed 2017×1papers★ This paper is the founder (earliest)
-
Defensins pathway genes (DEFA3, DEFB1, DEFA4) are downregulated in girls' newborn cord blood in association with long-term gestational PM2.5 exposure.microarray human cord-blood down 2017×1papers★ This paper is the founder (earliest)
-
TNF receptor signaling and T cell receptor signaling pathways are overrepresented among genes differentially expressed in boys' newborn cord blood associated with long-term gestational PM2.5 exposure.microarray human cord-blood mixed 2017×1papers★ This paper is the founder (earliest)
-
The first principal component of PM2.5-associated genes in newborn cord blood correlates significantly with long-term PM2.5 exposure in both girls (R=0.51) and boys (R=-0.40), in opposite directions.microarray human cord-blood mixed 2017×1papers★ This paper is the founder (earliest)
-
Long-term gestational PM2.5 exposure associates with differential gene expression in newborn cord blood in a sex-specific manner, with more genes affected in boys (1358) than in girls (724).microarray human cord-blood mixed 2017×1papers★ This paper is the founder (earliest)
-
Long-term gestational PM2.5 exposure shows significant sex-specific transcriptomic effects in newborn cord blood, with 1269 genes (7.5%) exhibiting a significant PM2.5-by-sex interaction.microarray human cord-blood 2017×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-28583124
Paper: Winckelmans et al. 2017, Environ Health 16:52. "Newborn sex-specific transcriptome signatures and gestational exposure to fine particles: findings from the ENVIRONAGE birth cohort." PMID 28583124 / PMC5458481 / DOI 10.1186/s12940-017-0264-y.
Data: GEO GSE83393 — 146 deposited Agilent one-color FES arrays
(GPL17077, Agilent SurePrint G3 Human GE v2 8x60K, Cy3 single channel),
cord blood, ENVIRONAGE cohort. Raw FES .txt.gz per sample + GSE83393_RAW.tar.
Code: https://github.com/BiGCAT-UM/arrayQC_Module @ commit
350f023abdc8a749f4fddbd4b1f40ea71dd809b3 (2015-11-23). A third-party generic
QC/normalization tool for Agilent/GenePix spotted arrays from the BiGCAT/Maastricht
group (same dept as authors de Kok/Kleinjans). It is a thin wrapper around limma:
read.maimages (green.only for one-color) → background correction → control
omission → bad-spot flagging (gIsWellAboveBG) → log2 → normalizeBetweenArrays
(method="quantile" for one-color). Output = normalized expression matrix + QC plots.
Per brief P16, applying this tool to the paper's own data is a valid reproduction.
Methods pipeline (from paper Methods)
"Gene expression ... arrayQC (R 2.15.3): local background correction, control omission, bad spot flagging, log2 transformation and quantile normalization. ... Quality control resulted in exclusion of four newborns. ... genes with >30% flagged data were removed; for genes with multiple probes the probe with the largest interquartile range (IQR) was retained, yielding 16,844 genes."
In scope (pipeline-derived, attempted)
| id | result | reported | pipeline | difficulty |
|---|---|---|---|---|
| C1 | # arrays deposited / analyzed | 146 deposited; 142 analyzed (4 excluded by QC) | GEO metadata + arrayQC array-level QC | easy (input verifiable; exact 4 = threshold-dependent) |
| C2 | sex split of analyzed set | 76 girls / 66 boys (142) | GEO Sex: characteristic |
easy (metadata-derived) |
| C3 | # genes after normalization + flag-filter + IQR collapse | 16,844 genes | limma one-color pipeline (= arrayQC) on the 146 FES files | medium — the clean compute target |
Out of scope (the hard ~20%, not attempted or only noted)
- PM2.5 association DE counts (long-term girls 724 / boys 1358 / overlap 75; short-term 507 / 1144 / overlap 55; interaction 1269 = 7.5%): require the full per-gene regression with the paper's covariate set (maternal age, gestational age, smoking, BMI, season, batch, ...). Only PM2.5 long/short values are in GEO characteristics; the remaining covariates are NOT deposited → not faithfully reproducible. NOT attempted (would need fabricated covariates).
- Pathway / GSEA tables (Tables 2–5): downstream of the DE lists, out of scope.
- Which specific 4 newborns failed QC: arrayQC array-level exclusion is threshold/visual → not exactly reproducible; we corroborate the count + arithmetic only (146 deposited; 2 have no submitter sex; 146−4=142, 77−1=76 girls, 67−1=66 boys).
Plan
One «our HPC» SLURM job: download GSE83393_RAW.tar to «infra», extract 146 FES files,
run limma one-color pipeline (the arrayQC steps), apply the paper's flag-filter
(>30% not-well-above-BG removed) + largest-IQR-per-gene collapse, report the gene
count and per-array flagged fraction. Compare to 16,844 / 142 / 76+66.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Cohort claims reproduce cleanly — 146 deposited→142 analyzed (146−4) and 76 girls/66 boys derive exactly from GEO metadata plus the 4 QC exclusions. The headline 16,844-gene count is not byte-reproducible: the paper's stated '>30% flagged' rule gives 22,142, and 16,844 only appears under an undocumented '~96%-present' filter (within the 14,961–36,337 sweep) — an authors'-side under-specification, not a fabrication signal. The central conclusion (sex-specific PM2.5 transcriptome signatures, C4–C6) could not be tested at all because the required covariates are not deposited in GEO, so the core claim is neither confirmed nor refuted. Overall a solid pipeline-level reproduction with explainable, mostly authors'-side deviations → partial.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.