Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials.

Sci Data · 2021
L1 94/100 PQI 94
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> EXACT 1:1 for the well-specified part. This Data Descriptor (Saarimaki et al., Sci Data 2021) deposits per-dataset normalized matrices + limma/DESeq2 DEG tables on Zenodo 10.5281/zenodo.4146981 (eUTOPIA pipeline). For the RU's named dataset GSE112780 (Affymetrix mouse lung, 28215 genes x 139 samples) I re-ran the documented limma DE step (lmFit ~0+group over the full cohort, makeContrasts per dose x timepoint ENM vs matched-time control, eBayes, topTable, threshold |logFC|>0.58 & BH<0.05) ON THE AUTHORS' DEPOSITED NORMALIZED MATRIX. All 15 contrasts reproduce EXACTLY: recomputed DEG count == deposited count for every contrast (820/755/376/238/206/110/40/39/14/5/1 and 0 for the four below-threshold contrasts; total 2604), 100% gene-list overlap, logFC correlation 1.0000 (max abs diff ~5e-14), p-value correlation 1.0000. Deposited Filtered_DEG lists are exact threshold subsets of the deposited Unfiltered_DEG tables (11/11). Collection composition matches the paper exactly: 85 microarray + 16 RNA-Seq = 101 datasets; 530 microarray contrast tables = 506 ENM-vs-control + 24 control as reported. NOT ATTEMPTED (the deliberate ~20%): raw CEL->normalized matrix (Affymetrix justRMA + ComBat batch correction) because the interactive eUTOPIA GUI's batch-variable/sample-exclusion choices are operator decisions not machine-specified; RNA-Seq FASTQ->DESeq2 (heavy); manual-curation correctness (not a pipeline). One flag for human: RNA-Seq comparison count is 36 vs paper-stated 30 (+6), a curation-count-semantics ambiguity (GSE125742 split per tissue into separate dataset dirs), NOT a fabricated pipeline value. Conclusion: the deposited DEG tables are exactly derivable from the deposited processed matrices by the documented method -> no fabrication detected in the DE step; the normalization/batch step itself was not independently re-derived.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 94
    assessed: 2026-06-15 ⛓ 33d9f0937705
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Existing transcriptomics data from engineered nanomaterial (ENM) exposures are scattered, heterogeneous, and lacking standardized metadata; the paper tests whether manually curating and homogenizing these data into a unified collection (with linked ENM physicochemical characteristics) can increase their FAIRness relative to the original individual datasets.

Core claims
  • A unified collection of 101 manually curated and homogenized transcriptomics datasets covering human, mouse, and rat ENM exposures in vitro and in vivo was compiled. resource
  • The curated collection exhibits a higher degree of FAIRness (Findable, Accessible, Interoperable, Reusable) than the individual original datasets composing it. finding
  • Each dataset was homogenized via standardized metadata curation, platform-specific preprocessing, and differential expression analysis to produce ready-for-modelling data. method
  • ENM physicochemical characterization data (supplier, purity, nominal/core size, hydrodynamic size, zeta potential, endotoxin, etc.) were curated and linked to each transcriptomics dataset. resource
  • The datasets were imported into and made publicly available through the NanoPharos database with REST API access to optimize accessibility, interoperability, and reusability. resource
  • Quality assessment excluded datasets with fewer than three biological replicates, unmanageable batch effects, or non-commercial/marginal platforms. method
  • Differential expression results provide full gene lists plus filtered significant DEGs using |logFC| > 0.58 and BH-adjusted p-value < 0.05. method
Experimental setups
Assay System Perturbation Readout Platform
gene expression microarray (Agilent) human, mouse, rat samples exposed to ENMs (in vitro and in vivo) ENM exposure vs control gene expression / differential expression Agilent commercial gene expression microarrays
gene expression microarray (Affymetrix) human, mouse, rat samples exposed to ENMs ENM exposure vs control gene expression / differential expression Affymetrix commercial gene expression microarrays
gene expression microarray (Illumina BeadChip) human, mouse, rat samples exposed to ENMs ENM exposure vs control gene expression / differential expression Illumina BeadChips (illuminaHumanv3.db, illuminaHumanv4.db, illuminaRatv1.db, illuminaMousev2.db)
RNA-Seq human and mouse samples exposed to ENMs ENM exposure vs control raw read counts / differential expression Illumina RNA-Seq
ENM physicochemical characterization (TEM) engineered nanomaterials none core particle size and shape Transmission Electron Microscopy
ENM physicochemical characterization (DLS) engineered nanomaterials in water and/or exposure medium none hydrodynamic size and zeta potential (surface charge) Dynamic Light Scattering
digital data curation GEO, ArrayExpress, ENA public repository datasets none homogenized metadata and preprocessed expression datasets R (v3.5.2), eUTOPIA, GEOquery
Key results
  • Initial repository query yielded 124 unique entries that underwent manual assessment. 124 entries
  • Final collection comprises 101 manually curated and preprocessed datasets. 101 datasets
  • Collection includes 85 preprocessed microarray-based datasets. 85 datasets
  • Microarray datasets total 506 unique ENM vs. control comparisons. 506 comparisons
  • Collection includes 16 RNA-Seq based datasets. 16 datasets
  • RNA-Seq datasets represent 23 ENM vs. control comparisons. 23 comparisons
  • 24 comparisons of non-nanoparticle compounds were used as positive/negative controls. 24 comparisons
  • Illumina microarray probes retained only if detection p-value < 0.01 in at least one sample. p < 0.01
Key statistics
  • count 124 (initial unique repository entries identified for manual assessment)
  • count 101 (manually curated and preprocessed datasets in final collection)
  • count 85 (microarray-based datasets in collection)
  • count 506 (unique ENM vs. control comparisons (microarray))
  • count 16 (RNA-Seq based datasets)
  • count 23 (ENM vs. control comparisons (RNA-Seq))
  • pvalue adjusted p-value < 0.05 (Benjamini & Hochberg threshold for significant DEGs)
  • fold_change |logFC| > 0.58 (threshold for significant differentially expressed genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a data curation and preprocessing descriptor, not a primary experimental study; the statistical work consists of standardized bioinformatics pipelines applied uniformly across 101 manually curated transcriptomics datasets (85 microarray, 16 RNA-Seq) from engineered nanomaterial exposures. Differential expression was assessed per pairwise comparison (ENM group vs. matched control) using limma for microarray data and DESeq2 for RNA-Seq data, with Benjamini-Hochberg-adjusted p-values and log2 fold-change thresholds applied to each dataset independently. Full gene-level statistics including fold changes and adjusted p-values are reported as output files; no aggregate inferential statistics across datasets are presented.

Replicationbiological Sample sizeMinimum 3 biological replicates required per experimental group as an inclusion criterion; actual n varies by source dataset and is not aggregated GroupsEach ENM-exposed experimental group (unique combination of ENM, dose, time point) vs. its matched control within each dataset; 506 microarray comparisons and 23 RNA-Seq comparisons total Pairingmixed Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
limma linear model with empirical Bayes moderation (microarray differential expression) All microarray datasets (Agilent, Affymetrix, Illumina BeadChip): each ENM-exposed group vs. matched control minimum 3 biological replicates required per group; exact n varies by source dataset not stated
DESeq2 Wald test (RNA-Seq differential expression, median-of-ratios normalization) All 16 RNA-Seq datasets: each ENM-exposed group vs. matched control minimum 3 biological replicates required per group; exact n varies by source dataset not stated
Proportion test (NOISeq) for low-count filtering RNA-Seq datasets: filter transcripts with low expression before normalization not stated
Surrogate Variable Analysis (SVA) for unknown batch effect estimation Microarray datasets: detection and assessment of latent technical variation not stated
Detection p-value threshold (p < 0.01) for Illumina probe filtering Illumina BeadChip microarray datasets: probe retention after normalization na
Approaches that could also have been used
  • RNA-Seq differential expression was performed with DESeq2 using median-of-ratios normalization and the Wald test
    Could also: edgeR (negative binomial GLM with likelihood ratio or quasi-likelihood F-test) or limma-voom (mean-variance trend modelling followed by limma linear model) could also have been applied — All three are widely used and benchmarked RNA-Seq DE methods; edgeR and limma-voom can behave differently from DESeq2 for small n or overdispersed libraries, and applying two methods and reporting concordant results is a common robustness check in multi-dataset collections
  • Microarray differential expression was performed with limma's empirical Bayes linear model across all platforms
    Could also: Platform-specific moderation approaches such as RMA + SAM (Significance Analysis of Microarrays) or a mixed-effects model explicitly accounting for donor and batch as random effects could also have been used — Mixed-effects or hierarchical models directly propagate donor-level variance rather than including donor as a fixed covariate, which can be advantageous when donor numbers are small or imbalanced across groups
  • Known batch effects were corrected using ComBat from the sva package
    Could also: limma's removeBatchEffect (for microarray) or RUVSeq (Remove Unwanted Variation, for RNA-Seq) could also have been used for batch correction — RUVSeq estimates unwanted variation from negative control genes or replicate samples and integrates directly into the count model, which some workflows prefer for RNA-Seq; removeBatchEffect is a lighter-weight alternative when the batch structure is simple and well-characterized
  • When multiple probes mapped to the same gene, the median expression value was used for summarization
    Could also: Taking the probe with the maximum absolute signal, or the mean across probes, or a summarization method based on probe reliability scores could also have been applied — The choice of probe summarization rule can affect differential expression results for genes with many probes or probes with very different signal levels; some workflows select the most variable probe or the probe with highest mean intensity rather than the median
  • For RNA-Seq, low-count transcripts were filtered using the NOISeq proportion test
    Could also: edgeR's filterByExpr function or a simple CPM-threshold filter (e.g., CPM > 1 in at least k samples) could also have been used — filterByExpr adapts the CPM threshold to library size and the experimental group structure, which can be useful when sample sizes and library depths vary across the 16 datasets; its filtering criterion is directly linked to the downstream testing model
  • No multiplicity correction was applied across the 500+ ENM-vs-control comparisons or across the 101 datasets in the collection
    Could also: A cross-comparison or cross-dataset FDR procedure (e.g., pooled BH across all comparisons, or a meta-analytic p-value aggregation) could also have been applied to the collection-level output — Because each comparison is corrected independently, the collection-wide false discovery rate is not controlled; for users wishing to draw conclusions across the entire collection rather than within individual datasets, a global correction would also be applicable
Software: R 3.5.2 · limma (Bioconductor) · DESeq2 (Bioconductor) · eUTOPIA · lumi (Bioconductor) · NOISeq (Bioconductor) 2.31.0 · Rsubread 2.2.3 · sva (Bioconductor, includes ComBat) · arrayQualityMetrics · GEOquery · HISAT2 · SAMtools 1.8-27-g0896262 · FastQC 0.11.7

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

E-MTAB-6396 ArrayExpress in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE100500 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE101992 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE103101 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE112780 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE113088 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE117056 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE122197 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE125742 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE127773 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE143717 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE14452 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE153419 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE155027 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE16727 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE17676 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE19487 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE20692 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE27212 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE29042 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE29110 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE30178 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE30180 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE30200 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE30213 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE30214 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE30215 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE35193 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE39330 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE41041 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE42066 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE42067 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE42068 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE43515 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE44294 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE45322 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE45598 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE4567 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE45868 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE46998 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE46999 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE48087 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE50176 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE51186 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE51417 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE51421 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE51636 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE51661 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE53700 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE55286 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE55349 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE56324 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE56325 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE60797 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE60798 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE60799 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE60800 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE61366 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE62253 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE62769 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE63552 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE63806 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE68036 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE7010 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE75429 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE79766 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE81564 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE81565 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE81566 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE81567 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE81568 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE81569 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE82062 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE84982 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE85711 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE86339 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE88786 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE92563 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE92899 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE92900 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE92987 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE96720 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE98236 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE99929 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33558569

Paper: Saarimäki et al. 2021, Sci Data 8:49. "Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials." PMCID PMC7870661 · DOI 10.1038/s41597-021-00808-y

Type: Data Descriptor. The deliverable is a curated, homogeneously re-preprocessed transcriptomics collection (101 datasets: 85 microarray + 16 RNA-Seq). Processed matrices + differential-expression (DEG) tables are deposited on Zenodo 10.5281/zenodo.4146981 (ENM_public_data.zip 1.6 GB + Data_characteristics.xlsx 73 kB). Code = eUTOPIA (Greco-Lab, R/Shiny microarray preprocessing) + DESeq2 for RNA-Seq.

Pipeline described in Methods (R 3.5.2)

  • Microarray (eUTOPIA): platform background correction/filtering → normalization (Agilent: limma quantile; Affymetrix: justRMA/affy; Illumina: lumiN/lumi) → batch assessment (PCA/HC/MDS) → batch correction (ComBat/sva) → probe→Ensembl annotation → limma DE.
  • RNA-Seq: FastQC → HISAT2 (GRCh38/GRCm38) → Rsubread counts → NOISeq low-count filter → DESeq2 median-of-ratios norm → DESeq2 DE.
  • DEG threshold (both): |logFC| > 0.58 AND BH adjusted p < 0.05.

In scope (pipeline-derived, attempted)

  • C1 — collection composition counts. 101 datasets (85 microarray, 16 RNA-Seq); 506 microarray ENM-vs-control comparisons; 23 RNA-Seq comparisons. Verify by enumerating Data_characteristics.xlsx / the deposited folder tree. (low compute, descriptive claim — confirms the deposit matches the paper.)
  • C2 — DEG reproduction for ONE dataset (the real pipeline test). Take the deposited processed (normalized + batch-corrected) expression matrix for one microarray dataset and re-run the well-specified limma DE step (lmFit→contrasts→eBayestopTable, threshold |logFC|>0.58 & BH<0.05). Compare DEG count and gene overlap to the deposited DEG table for that same contrast. Prefer GSE112780 (the RU's named accession) or, per 80/20, the simplest clean 2-group microarray dataset in the deposit. This isolates the reproducible DE step from the under-documented, interactive eUTOPIA preprocessing choices.

Out of scope (not attempted — stated honestly)

  • Full raw→DEG re-preprocessing of every dataset (interactive eUTOPIA Shiny; per-dataset batch-variable / sample-exclusion choices are operator decisions not machine-specified → the hard last 20%; we test the deposited processed matrix instead).
  • RNA-Seq alignment (HISAT2/Rsubread/DESeq2 from FASTQ) — heavy, out of 80/20.
  • Manual curation correctness (wet-lab/metadata judgement) — not a pipeline.

Grading

claims.tsv (reported vs reproduced) + agreement.json (exact|within-tol|partial|mismatch|error). All grades provisional; human decides. Flag any deposited value not derivable from shipped data as possible-fabrication.

Figures / tables: Table
C1a
Reported
85 microarray datasets
Reproduced
85
exact
C1b
Reported
16 RNA-Seq datasets
Reproduced
16
exact
C1c
Reported
101 datasets total
Reproduced
101
exact
C1d
Reported
506 microarray ENM-vs-control comparisons
Reproduced
530 contrast tables = 506 ENM + 24 control
within tolerance
C1e
Reported
23 RNA-Seq ENM-vs-control comparisons (+7=30)
Reproduced
36 contrast tables
did not match
C2_total
Reported
2604 DEGs over 15 GSE112780 contrasts
Reproduced
2604
exact
C2_croc12
Reported
820
Reproduced
820 (100% overlap, logFC r=1.0)
exact
C2_croc1
Reported
755
Reproduced
755
exact
C2_croc6
Reported
376
Reproduced
376
exact
C2_mw80_1
Reported
238
Reproduced
238
exact
C2_mw80_6
Reported
206
Reproduced
206
exact
C2_mw40_1
Reported
110
Reproduced
110
exact
C2_mw40_6
Reported
40
Reproduced
40
exact
C2_mw80_12
Reported
39
Reproduced
39
exact
C2_mw40_12
Reported
14
Reproduced
14
exact
C2_mw10_1
Reported
5
Reproduced
5
exact
C2_mw10_6
Reported
1
Reproduced
1
exact
C3
Reported
139 samples
Reproduced
28215 genes x 139 samples
exact
C4
Reported
Filtered_DEG = threshold(Unfiltered_DEG)
Reproduced
all 11 non-empty contrasts match
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4

This Data Descriptor reproduces 1:1 for its well-specified computational deliverable: starting from the authors' deposited normalized GSE112780 matrix, the documented limma pipeline regenerates all 15 contrasts' DEG counts exactly (total 2604, 100% gene overlap, logFC r=1.0000), and deposited Filtered_DEG lists are exact threshold subsets of the Unfiltered_DEG tables. The only deviations are dataset/comparison-counting semantics (microarray 506 vs 530=506+24 control; RNA-Seq 30 vs 36 from per-tissue splitting of GSE125742) — an input/curation-counting issue on a mix of our side and paper ambiguity, not a computational or fabrication problem. Honest boundary: the raw-CEL normalization/batch step was deliberately not re-derived (operator-chosen settings), so this confirms derivability of the DEG tables from the deposited matrices, not independent regeneration of those matrices. No fabrication concern — the perfect match is expected for deposited tables recomputed from their own deposited matrix.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

146.8 k
tokens (I/O) · 8.6 M incl. cache
18 min
runtime · 0.02 CPU-h
1.9 GB
peak RAM
5
HPC jobs
hummel
machine