Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Design of a targeted blood transcriptional panel for monitoring immunological changes accompanying pregnancy.

Front Immunol · 2024
L1 73/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
73/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 664 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

CLEAN RE-RUN («job», node n109) of the prior partial reproduction, which was re-queued only because its ROOM_RESULT shipped qc_room=null (the mandatory end-of-room QC block was never filled) -- the computational result itself was sound and now reproduces BIT-FOR-BIT. DESCRIBED WELL ENOUGH to reproduce the publicly-checkable computational core 1:1. We ran the paper's own published tool (BloodGen3Module 1.8.0, Drinchai/BloodGen3Module repo; R 4.3.3, GEOquery 2.68.0) with the paper's exact parameters (per-gene t-test, FC>=1.5, p<0.1, each pregnancy timepoint vs postpartum reference) on the paper's own PUBLIC reference cohort GSE108497 (PROMISSE, Illumina HT-12 V4). Sample selection reproduced the paper to the unit: healthy pregnant controls grp_p_tp HC_P_1..HC_P_5 = n 38/37/37/35/17, exactly the paper's reported per-timepoint counts. EXACT: 382-module / 38-aggregate BloodGen3 repertoire; the 3rd-trimester fingerprint (Fig 3A) of adaptive/lymphoid down (A1 modules -79..-92%) with erythroid (A37 +100%) and neutrophil-activation (A38 +100%) up. WITHIN-TOL: M12.6 (T cells) declines through gestation, exactly -40% at first trimester (matching the quoted '-40% range') deepening to -76% by 3rd tri; M14.50 (inflammation) up to +25% by 3rd tri. HONEST NON-MATCH: M16.64 (prostaglandin) is flat in PROMISSE microarray (+3.57/0/+3.57/-3.57%), consistent with the paper's own statement that this signal is MSP-driven -- not a fabrication signal. PARTIAL CROSS-CHECK: 25 of 38 aggregates reach the 10% cutoff in the 3rd-tri PROMISSE comparison vs the paper's 22 of 38 measured on the MSP RNA-seq cohort across all timepoints -- a same-ballpark analog, not 1:1. NOT ATTEMPTED (the hard 20%): the MSP RNA-seq pipeline (PRJNA898879, raw SRA only) that drives the gene-selection arithmetic and the wet-lab validation. No fabrication detected on the reproduced outputs; all verdicts provisional pending human audit (AUDIT.md).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 73
    assessed: 2026-06-15 ⛓ adde20ff5fc8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whole blood transcript abundance patterns, analyzed through the BloodGen3 module framework, can be used to identify a targeted gene panel capable of monitoring the immune trajectory ('immune clock') of pregnancy and detecting deviations from normal pregnancy progression.

Core claims
  • A targeted panel of 176 transcripts plus 8 housekeeping genes was identified for monitoring immunological changes during pregnancy resource
  • Changes in transcript abundance occur in early stages of pregnancy, with similar patterns observed in both the MSP and PROMISSE datasets finding
  • Functional gene annotation showed significant changes in lymphoid, prostaglandin, and inflammation-associated compartments compared to postpartum controls finding
  • BloodGen3, a fixed blood transcriptional module repertoire, can be applied to analyze and visualize gene expression patterns across independently generated pregnancy datasets method
  • Gene selection combined transcript abundance in whole blood, degree of correlation with BloodGen3 modules, and knowledge of pregnancy biology method
  • Whole blood transcript profiling via minimally invasive finger-prick sampling is amenable to clinical translation and high-frequency, self-collected sampling protocols finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (mRNA sequencing) whole blood, MSP cohort pregnant women (Thailand-Myanmar border) none (longitudinal pregnancy timepoints vs postpartum) transcript/gene expression abundance Illumina HiSeq 4000 with Truseq Stranded mRNA kit; STAR alignment to hg38; bcbio pipeline
microarray gene expression profiling whole blood, PROMISSE cohort healthy pregnant controls (Canada/USA) none (longitudinal pregnancy timepoints vs postpartum) gene expression signal intensity Illumina HT-12 V4 beadchips (47,231 probes), scanned on Illumina Beadstation 500, normalized with GenomeStudio
BloodGen3 fixed module repertoire analysis whole blood transcriptome data (both MSP and PROMISSE cohorts) none module response (% of constitutive genes differentially expressed), fold-change and expression difference BloodGen3Module R package
Key results
  • 176 transcripts identified, complemented with 8 housekeeping genes, forming the targeted panel
  • Changes in transcript abundance were detectable in early pregnancy stages, with similar patterns across the MSP and PROMISSE datasets
  • Significant changes seen in lymphoid, prostaglandin, and inflammation-associated compartments relative to postpartum controls
  • 15 women with uneventful pregnancies profiled at 6 timepoints each (first/second/third trimester, delivery, 1-month and 3-month postpartum) in the MSP dataset
  • 38 healthy pregnant controls selected from the PROMISSE cohort, sampled at 5 prespecified timepoints
  • Significance for differential expression determined using fold-change and p-value cutoffs against 3-month postpartum reference FC>=1.5, p<0.1
Key statistics
  • count 176 transcripts (size of the targeted gene panel identified)
  • count 8 housekeeping genes (added to the 176-transcript panel)
  • count 15 women (MSP dataset uneventful-pregnancy participants, 6 timepoints each)
  • count 38 (healthy pregnant controls selected from PROMISSE cohort)
  • fold_change |FC| > 1.5 (cutoff for significant gene expression change vs postpartum reference)
  • pvalue p < 0.1 (significance cutoff (t-test vs postpartum reference); DESeq2 FDR < 0.1 for group-level analysis)
  • count 985 individual transcriptome profiles (basis for construction of the BloodGen3 module repertoire (16 datasets))
  • count 382 co-expressed gene sets reduced to 38 module aggregates (BloodGen3 repertoire dimension reduction via k-means clustering)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper employed two independent longitudinal whole-blood transcriptome datasets — a de novo MSP cohort (n=15 women, 6 timepoints, RNA-seq) and a publicly available PROMISSE cohort (n=38 healthy controls, 5 timepoints, Illumina microarray) — and compared gene expression at each pregnancy timepoint against a mean postpartum reference using t-tests (p<0.1, |FC|≥1.5). Module-level differential expression in the MSP dataset additionally used DESeq2 with a FDR<0.1 threshold. Transcriptome patterns were organized through the pre-defined BloodGen3 fixed co-expression module repertoire, with individual-level module responses quantified by fixed fold-change and expression-difference thresholds (|FC|>1.5, |DIFF|>10). Both datasets were analyzed independently and findings compared descriptively to identify a 176-gene targeted panel.

Replicationbiological Sample size15 women selected from MSP cohort for transcriptome profiling; 38 healthy pregnant controls selected from PROMISSE; no formal power calculation or sample-size justification described GroupsPregnancy timepoints (first trimester, second trimester, third trimester, delivery, 1-month postpartum) vs 3-month postpartum reference; independently in two cohorts from resource-limited (Thailand) and resource-rich (Canada/USA) settings Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR via DESeq2 (FDR<0.1) for module-level group-level analysis; no explicit multiplicity correction described for the per-gene t-test comparisons across the full transcript set
Statistical tests used
Test Applied to n Assumptions
t-test (each timepoint mean vs mean postpartum reference, applied per gene) Gene-level differential expression across pregnancy timepoints vs 3-month postpartum reference in both MSP and PROMISSE datasets n=15 (MSP); n=38 (PROMISSE) not stated
DESeq2 Wald test with Benjamini-Hochberg FDR Module-level group differential expression analysis, MSP dataset n=15 (MSP) not stated
Fixed fold-change and expression-difference threshold (|FC|>1.5, |DIFF|>10; no formal test) Individual sample-level BloodGen3 module response in both datasets n=1 per comparison (individual sample vs reference mean) na
Approaches that could also have been used
  • Per-gene t-tests comparing each timepoint to the postpartum reference were applied across thousands of transcripts with a nominal p<0.1 threshold; no genome-wide FDR correction is described for this step
    Could also: Benjamini-Hochberg FDR correction applied to the full set of t-test p-values across all tested transcripts (e.g., via p.adjust in R) — When testing thousands of genes simultaneously, FDR adjustment is a standard way to characterize the expected proportion of false discoveries and is directly comparable across published transcriptomic studies
  • The longitudinal repeated-measures structure (same individual profiled at up to 6 timepoints) was addressed with separate per-timepoint t-tests against a fixed reference mean
    Could also: Linear mixed-effects models (e.g., lme4 or nlme in R, or limma with duplicateCorrelation for array data) that explicitly model within-subject correlation across timepoints — Mixed-effects models account for the non-independence of repeated observations from the same individual, which can improve the accuracy of standard error estimates and statistical power in small-n longitudinal designs
  • Individual-level module responses were determined using fixed fold-change and expression-difference cutoffs (|FC|>1.5, |DIFF|>10) rather than a formal statistical test
    Could also: Empirical Bayes moderated t-statistics (e.g., limma/voom) that borrow information across genes to stabilize variance estimates for each individual-vs-reference comparison — When the reference baseline is estimated from a small number of postpartum samples, moderated statistics can produce more stable per-gene variance estimates than a simple fold-change threshold
  • The two cohorts (MSP and PROMISSE) were analyzed on separate platforms (RNA-seq vs microarray) and compared descriptively to identify overlapping gene candidates
    Could also: Formal cross-cohort integration such as Fisher's combined p-value method, fixed-effects meta-analysis of standardized effect sizes, or cross-platform normalization followed by joint modeling — Quantitative integration methods yield a single summary statistic of evidence strength across cohorts and enable formal assessment of between-cohort heterogeneity
  • Module-level functional grouping used the pre-defined BloodGen3 co-expression repertoire constructed from external datasets
    Could also: Gene Set Enrichment Analysis (GSEA) or over-representation analysis against curated gene sets (MSigDB, GO, KEGG) applied to the ranked or filtered gene list from these specific datasets — Data-driven enrichment methods provide formal enrichment scores and FDR-corrected p-values derived from this study's own data, complementing fixed-repertoire approaches and facilitating comparison with the broader literature
  • Results were reported as module response percentages and fold-change thresholds without accompanying measures of variability across participants
    Could also: Reporting mean expression with SD or 95% CI, or displaying individual-level trajectory plots alongside group summaries — Dispersion and confidence measures convey between-individual variability and estimation precision, which is particularly informative for small cohorts and helps readers interpret effect magnitudes
Software: DESeq2 (R/Bioconductor) · BloodGen3Module (R package) · bcbio rnaseq pipeline 1.2.3 · FastQC 0.11.9 · STAR 2.6.1d · Samtools 1.3 · Illumina GenomeStudio

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
6
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

NCT02797327 NCT in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJNA898879 BioProject in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38352867

Paper: Brummaier T, Rinchai D, Toufiq M, et al. (2024) Design of a targeted blood transcriptional panel for monitoring immunological changes accompanying pregnancy. Front Immunol 14:1319949. PMID 38352867 · PMCID PMC10861739 · DOI 10.3389/fimmu.2024.1319949.

Code: https://github.com/Drinchai/BloodGen3Module — a third-party / authors' Bioconductor R package implementing the fixed BloodGen3 blood transcriptional module repertoire (382 modules organised into 38 aggregates). Per the brief (P16), applying this published tool to the paper's own data is an equally valid reproduction. Installed pinned via conda bioconductor-bloodgen3module=1.8.0 (R 4.3) — exact build recorded in reproduction/environment.lock.

Data accessions:

  • PROMISSE = GEO GSE108497 — Illumina HumanHT-12 V4 microarray, the public reference cohort. 512 samples (SLE + non-SLE); the paper uses the healthy pregnant controls across 5 timepoints (<16, 16–23, 24–31, 32–40 wk EGA, and postpartum). PUBLIC, fully obtainable. ← what we reproduce.
  • MSP = PRJNA898879 — de-novo bulk RNA-seq, 88 profiles / 6 timepoints (15/15/15/15/13/15). The discovery cohort. Raw reads on SRA only (no shipped count matrix). ← out of scope (heavy: SRA→counts), see below.

What the paper does (Fig 2, four-step panel design)

  1. Pre-determined module repertoire — fixed BloodGen3 (382 modules / 38 aggregates).
  2. Module-aggregate selection — Groupcomparison (t-test) of each pregnancy timepoint vs the postpartum reference, FC ≥ 1.5, p < 0.1. An aggregate is retained if ≥1 constitutive module reaches a 10 % average response measured across all MSP samples22 of 38 aggregates retained.
  3. Module-set delineation46 module sets.
  4. Candidate transcript selection → 2 530 genes → 894 (35.3 %) removed (median count < 50 in MSP) → correlation filter (r > 0.5) → final 176 test + 8 housekeeping = 184 transcripts for the validation assay.

In scope (pipeline-derived, attempted on PUBLIC GSE108497)

id result pipeline step reproducibility
C1 Module response fingerprint, 3rd trimester (32–40 wk) vs postpartum (Fig 3A): the genome-wide module grid of % increased (red) / decreased (blue) BloodGen3Module Groupcomparison (t-test, FC≥1.5, p<0.1) + gridplot YES — run the published tool on the paper's own public reference cohort
C2 M12.6 (T-cell functions) — steady decline across gestation, "−40 % range" by 3rd trimester per-module net % response, trajectory across timepoints YES (sign + approximate magnitude; PROMISSE shown in Fig 3)
C3 M16.64 (prostaglandin, A31) — gradual increase 1st→3rd trimester per-module net % response YES (direction + trajectory)
C4 M14.50 (inflammation) — steady increase 1st trimester→delivery per-module net % response YES (direction + trajectory)
C5 PROMISSE analog of aggregate responsiveness — how many of 38 aggregates reach the 10 % average-response cut-off in the 3rd-trimester PROMISSE comparison aggregate-level averaging of module responses PARTIAL — the paper's 22/38 is measured on MSP, not PROMISSE; we report the PROMISSE figure as an explicitly-labelled analog, not a claim of equality

Out of scope (not attempted — and why)

  • The headline 22/38 aggregate selection, the 894/35.3 % gene exclusion, the 46 module sets, and the final 176+8 panel — all measured on the MSP RNA-seq cohort (PRJNA898879), which is deposited as raw SRA reads only. Producing the count matrix requires a full fastq→counts RNA-seq pipeline; per the 80/20 rule this is the hard last 20 % and is left unattempted. The numbers are internally consistent (894 = 35.3 % of 2530; 176+8 = 184) but not re-derivable here without the SRA→count step. → non_pipeline/heavy, deferred.
  • Wet-lab validation assay
Figures / tables: Fig 2Fig 3Fig 3A
C1
Reported
382 BloodGen3 modules
Reproduced
382
exact
C2
Reported
38 module aggregates
Reproduced
38
exact
C3
Reported
22 of 38 aggregates >=10% avg response (MSP cohort, all timepoints)
Reproduced
25 of 38 (GSE108497 PROMISSE, 3rd-tri vs postpartum)
partial
C4
Reported
M12.6 (T cells) steady decline, '-40% range'
Reproduced
-40,-76,-60,-76% across T1..3rd-tri
within tolerance
C5
Reported
M16.64 (prostaglandin) gradual increase
Reproduced
+3.57,0,+3.57,-3.57% (flat in PROMISSE)
did not match
C6
Reported
M14.50 (inflammation) steady increase
Reproduced
0,+41.67,+8.33,+25%
within tolerance
C7
Reported
Fig 3A fingerprint: lymphoid down, erythroid/innate up
Reproduced
A1 B/T/lymphocyte -79..-92%; A37 erythroid +100%; A38 neutrophil +100%
exact
C8
Reported
final panel 176 test + 8 HK = 184
Reproduced
NOT ATTEMPTED (MSP raw SRA, PRJNA898879, out of scope)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 73/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The publicly checkable computational core reproduced cleanly using the paper's own tool (BloodGen3Module 1.8.0) and own public cohort GSE108497: structural counts (382 modules / 38 aggregates) are exact, sample selection matched to the unit (38/37/37/35/17), and the canonical pregnancy signature (lymphoid down −79..−92%, erythroid/neutrophil +100%, M12.6 T-cell −40% at T1) reproduced robustly — so the central conclusion holds. The deviations are input-side and explainable: C3 (25/38 vs 22/38) and C5 (M16.64 flat vs reported increase) stem from substituting the PROMISSE microarray for the unavailable MSP RNA-seq cohort (raw SRA only), a difference the paper itself anticipates. The titular 176+8 panel arithmetic was not re-derivable from the shared deposit (out of scope), but is internally consistent with no fabrication signal. Overall a solid, partial reproduction with deviations attributable to our analog-cohort choice and data-availability limits rather than an authors' defect.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

487.9 k
tokens (I/O) · 57 M incl. cache
102 min
runtime · 0.05 CPU-h
2.4 GB
peak RAM
2
HPC jobs
hummel
machine