Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> EXACT 1:1 for the well-specified part. This Data Descriptor (Saarimaki et al., Sci Data 2021) deposits per-dataset normalized matrices + limma/DESeq2 DEG tables on Zenodo 10.5281/zenodo.4146981 (eUTOPIA pipeline). For the RU's named dataset GSE112780 (Affymetrix mouse lung, 28215 genes x 139 samples) I re-ran the documented limma DE step (lmFit ~0+group over the full cohort, makeContrasts per dose x timepoint ENM vs matched-time control, eBayes, topTable, threshold |logFC|>0.58 & BH<0.05) ON THE AUTHORS' DEPOSITED NORMALIZED MATRIX. All 15 contrasts reproduce EXACTLY: recomputed DEG count == deposited count for every contrast (820/755/376/238/206/110/40/39/14/5/1 and 0 for the four below-threshold contrasts; total 2604), 100% gene-list overlap, logFC correlation 1.0000 (max abs diff ~5e-14), p-value correlation 1.0000. Deposited Filtered_DEG lists are exact threshold subsets of the deposited Unfiltered_DEG tables (11/11). Collection composition matches the paper exactly: 85 microarray + 16 RNA-Seq = 101 datasets; 530 microarray contrast tables = 506 ENM-vs-control + 24 control as reported. NOT ATTEMPTED (the deliberate ~20%): raw CEL->normalized matrix (Affymetrix justRMA + ComBat batch correction) because the interactive eUTOPIA GUI's batch-variable/sample-exclusion choices are operator decisions not machine-specified; RNA-Seq FASTQ->DESeq2 (heavy); manual-curation correctness (not a pipeline). One flag for human: RNA-Seq comparison count is 36 vs paper-stated 30 (+6), a curation-count-semantics ambiguity (GSE125742 split per tissue into separate dataset dirs), NOT a fabricated pipeline value. Conclusion: the deposited DEG tables are exactly derivable from the deposited processed matrices by the documented method -> no fabrication detected in the DE step; the normalization/batch step itself was not independently re-derived.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 94assessed: 2026-06-15 ⛓ 33d9f0937705
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusExisting transcriptomics data from engineered nanomaterial (ENM) exposures are scattered, heterogeneous, and lacking standardized metadata; the paper tests whether manually curating and homogenizing these data into a unified collection (with linked ENM physicochemical characteristics) can increase their FAIRness relative to the original individual datasets.
- ★ A unified collection of 101 manually curated and homogenized transcriptomics datasets covering human, mouse, and rat ENM exposures in vitro and in vivo was compiled. resource
- ★ The curated collection exhibits a higher degree of FAIRness (Findable, Accessible, Interoperable, Reusable) than the individual original datasets composing it. finding
- ★ Each dataset was homogenized via standardized metadata curation, platform-specific preprocessing, and differential expression analysis to produce ready-for-modelling data. method
- ★ ENM physicochemical characterization data (supplier, purity, nominal/core size, hydrodynamic size, zeta potential, endotoxin, etc.) were curated and linked to each transcriptomics dataset. resource
- ★ The datasets were imported into and made publicly available through the NanoPharos database with REST API access to optimize accessibility, interoperability, and reusability. resource
- Quality assessment excluded datasets with fewer than three biological replicates, unmanageable batch effects, or non-commercial/marginal platforms. method
- Differential expression results provide full gene lists plus filtered significant DEGs using |logFC| > 0.58 and BH-adjusted p-value < 0.05. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| gene expression microarray (Agilent) | human, mouse, rat samples exposed to ENMs (in vitro and in vivo) | ENM exposure vs control | gene expression / differential expression | Agilent commercial gene expression microarrays |
| gene expression microarray (Affymetrix) | human, mouse, rat samples exposed to ENMs | ENM exposure vs control | gene expression / differential expression | Affymetrix commercial gene expression microarrays |
| gene expression microarray (Illumina BeadChip) | human, mouse, rat samples exposed to ENMs | ENM exposure vs control | gene expression / differential expression | Illumina BeadChips (illuminaHumanv3.db, illuminaHumanv4.db, illuminaRatv1.db, illuminaMousev2.db) |
| RNA-Seq | human and mouse samples exposed to ENMs | ENM exposure vs control | raw read counts / differential expression | Illumina RNA-Seq |
| ENM physicochemical characterization (TEM) | engineered nanomaterials | none | core particle size and shape | Transmission Electron Microscopy |
| ENM physicochemical characterization (DLS) | engineered nanomaterials in water and/or exposure medium | none | hydrodynamic size and zeta potential (surface charge) | Dynamic Light Scattering |
| digital data curation | GEO, ArrayExpress, ENA public repository datasets | none | homogenized metadata and preprocessed expression datasets | R (v3.5.2), eUTOPIA, GEOquery |
- – Initial repository query yielded 124 unique entries that underwent manual assessment. 124 entries
- – Final collection comprises 101 manually curated and preprocessed datasets. 101 datasets
- – Collection includes 85 preprocessed microarray-based datasets. 85 datasets
- – Microarray datasets total 506 unique ENM vs. control comparisons. 506 comparisons
- – Collection includes 16 RNA-Seq based datasets. 16 datasets
- – RNA-Seq datasets represent 23 ENM vs. control comparisons. 23 comparisons
- – 24 comparisons of non-nanoparticle compounds were used as positive/negative controls. 24 comparisons
- – Illumina microarray probes retained only if detection p-value < 0.01 in at least one sample. p < 0.01
- count 124 (initial unique repository entries identified for manual assessment)
- count 101 (manually curated and preprocessed datasets in final collection)
- count 85 (microarray-based datasets in collection)
- count 506 (unique ENM vs. control comparisons (microarray))
- count 16 (RNA-Seq based datasets)
- count 23 (ENM vs. control comparisons (RNA-Seq))
- pvalue adjusted p-value < 0.05 (Benjamini & Hochberg threshold for significant DEGs)
- fold_change |logFC| > 0.58 (threshold for significant differentially expressed genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a data curation and preprocessing descriptor, not a primary experimental study; the statistical work consists of standardized bioinformatics pipelines applied uniformly across 101 manually curated transcriptomics datasets (85 microarray, 16 RNA-Seq) from engineered nanomaterial exposures. Differential expression was assessed per pairwise comparison (ENM group vs. matched control) using limma for microarray data and DESeq2 for RNA-Seq data, with Benjamini-Hochberg-adjusted p-values and log2 fold-change thresholds applied to each dataset independently. Full gene-level statistics including fold changes and adjusted p-values are reported as output files; no aggregate inferential statistics across datasets are presented.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma linear model with empirical Bayes moderation (microarray differential expression) | All microarray datasets (Agilent, Affymetrix, Illumina BeadChip): each ENM-exposed group vs. matched control | minimum 3 biological replicates required per group; exact n varies by source dataset | not stated |
| DESeq2 Wald test (RNA-Seq differential expression, median-of-ratios normalization) | All 16 RNA-Seq datasets: each ENM-exposed group vs. matched control | minimum 3 biological replicates required per group; exact n varies by source dataset | not stated |
| Proportion test (NOISeq) for low-count filtering | RNA-Seq datasets: filter transcripts with low expression before normalization | — | not stated |
| Surrogate Variable Analysis (SVA) for unknown batch effect estimation | Microarray datasets: detection and assessment of latent technical variation | — | not stated |
| Detection p-value threshold (p < 0.01) for Illumina probe filtering | Illumina BeadChip microarray datasets: probe retention after normalization | — | na |
-
RNA-Seq differential expression was performed with DESeq2 using median-of-ratios normalization and the Wald test↳ Could also: edgeR (negative binomial GLM with likelihood ratio or quasi-likelihood F-test) or limma-voom (mean-variance trend modelling followed by limma linear model) could also have been applied — All three are widely used and benchmarked RNA-Seq DE methods; edgeR and limma-voom can behave differently from DESeq2 for small n or overdispersed libraries, and applying two methods and reporting concordant results is a common robustness check in multi-dataset collections
-
Microarray differential expression was performed with limma's empirical Bayes linear model across all platforms↳ Could also: Platform-specific moderation approaches such as RMA + SAM (Significance Analysis of Microarrays) or a mixed-effects model explicitly accounting for donor and batch as random effects could also have been used — Mixed-effects or hierarchical models directly propagate donor-level variance rather than including donor as a fixed covariate, which can be advantageous when donor numbers are small or imbalanced across groups
-
Known batch effects were corrected using ComBat from the sva package↳ Could also: limma's removeBatchEffect (for microarray) or RUVSeq (Remove Unwanted Variation, for RNA-Seq) could also have been used for batch correction — RUVSeq estimates unwanted variation from negative control genes or replicate samples and integrates directly into the count model, which some workflows prefer for RNA-Seq; removeBatchEffect is a lighter-weight alternative when the batch structure is simple and well-characterized
-
When multiple probes mapped to the same gene, the median expression value was used for summarization↳ Could also: Taking the probe with the maximum absolute signal, or the mean across probes, or a summarization method based on probe reliability scores could also have been applied — The choice of probe summarization rule can affect differential expression results for genes with many probes or probes with very different signal levels; some workflows select the most variable probe or the probe with highest mean intensity rather than the median
-
For RNA-Seq, low-count transcripts were filtered using the NOISeq proportion test↳ Could also: edgeR's filterByExpr function or a simple CPM-threshold filter (e.g., CPM > 1 in at least k samples) could also have been used — filterByExpr adapts the CPM threshold to library size and the experimental group structure, which can be useful when sample sizes and library depths vary across the 16 datasets; its filtering criterion is directly linked to the downstream testing model
-
No multiplicity correction was applied across the 500+ ENM-vs-control comparisons or across the 101 datasets in the collection↳ Could also: A cross-comparison or cross-dataset FDR procedure (e.g., pooled BH across all comparisons, or a meta-analytic p-value aggregation) could also have been applied to the collection-level output — Because each comparison is corrected independently, the collection-wide false discovery rate is not controlled; for users wishing to draw conclusions across the entire collection rather than within individual datasets, a global correction would also be applicable
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33558569
Paper: Saarimäki et al. 2021, Sci Data 8:49. "Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials." PMCID PMC7870661 · DOI 10.1038/s41597-021-00808-y
Type: Data Descriptor. The deliverable is a curated, homogeneously
re-preprocessed transcriptomics collection (101 datasets: 85 microarray + 16
RNA-Seq). Processed matrices + differential-expression (DEG) tables are deposited
on Zenodo 10.5281/zenodo.4146981 (ENM_public_data.zip 1.6 GB +
Data_characteristics.xlsx 73 kB). Code = eUTOPIA (Greco-Lab, R/Shiny
microarray preprocessing) + DESeq2 for RNA-Seq.
Pipeline described in Methods (R 3.5.2)
- Microarray (eUTOPIA): platform background correction/filtering →
normalization (Agilent: limma quantile; Affymetrix:
justRMA/affy; Illumina:lumiN/lumi) → batch assessment (PCA/HC/MDS) → batch correction (ComBat/sva) → probe→Ensembl annotation → limma DE. - RNA-Seq: FastQC → HISAT2 (GRCh38/GRCm38) → Rsubread counts → NOISeq low-count filter → DESeq2 median-of-ratios norm → DESeq2 DE.
- DEG threshold (both):
|logFC| > 0.58 AND BH adjusted p < 0.05.
In scope (pipeline-derived, attempted)
- C1 — collection composition counts. 101 datasets (85 microarray, 16
RNA-Seq); 506 microarray ENM-vs-control comparisons; 23 RNA-Seq comparisons.
Verify by enumerating
Data_characteristics.xlsx/ the deposited folder tree. (low compute, descriptive claim — confirms the deposit matches the paper.) - C2 — DEG reproduction for ONE dataset (the real pipeline test). Take the
deposited processed (normalized + batch-corrected) expression matrix for one
microarray dataset and re-run the well-specified limma DE step
(
lmFit→contrasts→eBayes→topTable, threshold |logFC|>0.58 & BH<0.05). Compare DEG count and gene overlap to the deposited DEG table for that same contrast. Prefer GSE112780 (the RU's named accession) or, per 80/20, the simplest clean 2-group microarray dataset in the deposit. This isolates the reproducible DE step from the under-documented, interactive eUTOPIA preprocessing choices.
Out of scope (not attempted — stated honestly)
- Full raw→DEG re-preprocessing of every dataset (interactive eUTOPIA Shiny; per-dataset batch-variable / sample-exclusion choices are operator decisions not machine-specified → the hard last 20%; we test the deposited processed matrix instead).
- RNA-Seq alignment (HISAT2/Rsubread/DESeq2 from FASTQ) — heavy, out of 80/20.
- Manual curation correctness (wet-lab/metadata judgement) — not a pipeline.
Grading
claims.tsv (reported vs reproduced) + agreement.json
(exact|within-tol|partial|mismatch|error). All grades provisional; human decides.
Flag any deposited value not derivable from shipped data as possible-fabrication.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This Data Descriptor reproduces 1:1 for its well-specified computational deliverable: starting from the authors' deposited normalized GSE112780 matrix, the documented limma pipeline regenerates all 15 contrasts' DEG counts exactly (total 2604, 100% gene overlap, logFC r=1.0000), and deposited Filtered_DEG lists are exact threshold subsets of the Unfiltered_DEG tables. The only deviations are dataset/comparison-counting semantics (microarray 506 vs 530=506+24 control; RNA-Seq 30 vs 36 from per-tissue splitting of GSE125742) — an input/curation-counting issue on a mix of our side and paper ambiguity, not a computational or fabrication problem. Honest boundary: the raw-CEL normalization/batch step was deliberately not re-derived (operator-chosen settings), so this confirms derivability of the DEG tables from the deposited matrices, not independent regeneration of those matrices. No fabrication concern — the perfect match is expected for deposited tables recomputed from their own deposited matrix.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.