PulmonDB: a curated lung disease gene expression database.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, and the method+biology reproduce 1:1 against the paper's OWN live resource (P16). PulmonDB's MySQL backend is still public (guest creds hardcoded in the AnaBVA/PulmonDB R package); I re-queried it from a «our HPC» compute node and replicated the package vignette's limma COPD-vs-IPF pipeline (R 4.3.3 + RMySQL + limma; «job»). EXACT reproductions: FOSB & CXCL2 are COPD-up/IPF-down exactly as reported, and GSE32537 recapitulates known IPF biology (MMP7/SPP1/COL1A1/MMP1 up, AGER/CAV1 down). DIFFERENT (honest mismatch, fully explained, NOT fabrication): PulmonDB is a LIVING database that grew since the 2020 freeze, so composition counts (76->75/89 GSEs near-exact; 4481->6857 contrasts; 26->34 platforms) and the headline 1781-DEG count (->9295 BH<0.05) do not reproduce as exact numbers, because the DE design now has 738 lung-biopsy contrasts vs ~164 in 2020 (~4.5x power inflates FDR-significant counts). VCAM1 was absent after filtering and FCN3 came out opposite, so those two 'similar-gene' examples were not confirmed. NOT attempted (80/20): re-running COMMAND's raw->homogenized normalization from CEL files, manual ontology curation, web/genome-browser/clustering figures, and pinning the exact 2020 DB snapshot (no versioned freeze is published).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 53assessed: 2026-06-15 ⛓ f842e4b821f7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan integrating and curating publicly available COPD and IPF transcriptomic data into a single homogenized database (PulmonDB) enable systematic comparison of these two contrasting lung diseases to identify common and distinct molecular mechanisms and generate new hypotheses?
- ★ PulmonDB is a curated, web-based gene expression database and R package integrating microarray and RNA-seq data for COPD and IPF with manually curated controlled-vocabulary annotation. resource
- ★ PulmonDB recapitulates previously reported COPD- and IPF-associated gene expression patterns, with hierarchical clustering separating disease from control samples. finding
- ★ Differential expression analysis across COPD vs IPF lung-biopsy contrasts identifies genes with opposite expression behavior between the two diseases (e.g., FOSB, CXCL2 up in COPD, down in IPF). finding
- ★ A weighted limma contrast identifies genes shared/concordantly regulated between COPD and IPF (e.g., VCAM1 up, FCN3 down in both). finding
- ★ The COMMAND>_ compendium-creation methodology, previously used for bacteria (COLOMBOS) and grapevine (VESPUCCI), is applied for the first time to human data. method
- Microarray probes were remapped to establish uniform gene annotation and data were homogenized and quality-checked across platforms. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Curated transcriptomic database integration (microarray + RNA-seq) | Human lung disease samples (COPD and IPF; lung biopsy, blood, primary cells, A549 cell line) | none (observational disease vs control) | Homogenized gene expression values / contrasts | MySQL relational database built with COMMAND>_; 26 GPL platforms (Affymetrix, Agilent, Illumina, others) |
| Hierarchical clustering / heatmap visualization of curated marker genes | IPF lung tissue biopsy samples (GSE32537, GSE21369, GSE24206, GSE94060, GSE72073, GSE35145, GSE31934) vs controls | none (IPF vs control) | Expression of 19 literature-selected IPF genes; cluster separation | Clustergrammer (PulmonDB website) |
| Hierarchical clustering / heatmap visualization of curated marker genes | COPD lung tissue biopsy samples (GSE27597, GSE37768, GSE57148, GSE8581, GSE1122) vs controls | none (COPD vs control) | Expression of 16 literature-selected COPD genes; cluster separation | Clustergrammer (PulmonDB website) |
| Differential gene expression analysis | Lung biopsy contrasts with HEALTHY/CONTROL reference (GSE52463, GSE63073, GSE1122, GSE72073, GSE24206, GSE27597, GSE29133, GSE31934, GSE37768) | none (COPD vs IPF comparison) | Differentially expressed genes between COPD and IPF | limma (R environment) |
| Weighted differential expression contrast | Lung biopsy COPD and IPF contrasts vs HEALTHY/CONTROL | none (shared-signature analysis) | Genes concordantly differentially expressed in both diseases | limma (R environment) |
| RNA-seq count retrieval (preprocessed) | Human COPD/IPF RNA-seq experiments | none | Gene and exon counts | Recount2 (Rail-RNA alignment) |
- – PulmonDB contains 76 GSEs corresponding to 4481 unique preprocessed GSM contrasts across 26 platforms (GPLs) 76 GSEs; 4481 contrasts; 26 platforms
- – Differential expression analysis between COPD and IPF identified 1781 differentially expressed genes 1781 genes
- – Hierarchical clustering of IPF marker genes separated IPF and control datasets
- – Hierarchical clustering of COPD marker genes separated patients and controls into two groups
- – Gene group I overexpressed in IPF and barely/under-expressed in COPD; group II overexpressed in COPD and underexpressed in IPF
- – FOSB and CXCL2 overexpressed in COPD and underexpressed in IPF
- – VCAM1 overexpressed and FCN3 underexpressed in both COPD and IPF relative to controls
- – Samples from the same disease group showed higher correlations and null/negative correlation with controls and the opposite disease
- count 76 GSEs (total experiments (GSEs) in PulmonDB)
- count 4481 unique preprocessed GSM contrasts (total sample contrasts in PulmonDB)
- count 26 different platforms (GPLs) (number of platforms represented)
- count 1781 differentially expressed genes (COPD vs IPF differential expression analysis)
- other 37.8% (lung biopsies as proportion of samples)
- other 33.2% (blood samples as proportion of samples)
- other 34.9% COPD, 40.5% control (30.9% healthy + 9.6% match tissue), 17.2% IPF, 1.5% other diseases (sample state composition)
- count 19 IPF genes and 16 COPD genes (literature-selected marker genes visualized)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
PulmonDB is a curated compendium integrating preprocessed microarray and RNA-seq gene expression data for COPD and IPF sourced from GEO and Recount2. The primary statistical analysis used the R package limma to perform differential expression analysis on lung-biopsy contrasts drawn from nine GEO experiments, comparing COPD and IPF samples against healthy controls and against each other via weighted contrasts. Results were reported as a total count of differentially expressed genes (1,781) and visualized as hierarchical-clustering heatmaps of the top 20 genes; no formal power analysis or per-group sample sizes are stated in the available text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma linear model with moderated t-statistics (empirical Bayes) | Differential expression between COPD and IPF across nine GEO experiments (lung biopsy contrasts vs HEALTHY/CONTROL reference); secondary weighted contrast to identify genes co-regulated in both diseases | — | not stated |
| Hierarchical clustering | Visualization of 19 known IPF genes across 7 experiments and 16 known COPD genes across 5 experiments; also applied to columns and rows of the top-20 DEG heatmaps | — | na |
-
Multiple heterogeneous GEO experiments from different platforms were pooled for differential expression analysis; explicit batch correction is not described in the available text↳ Could also: Methods such as ComBat, surrogate variable analysis (SVA), or limma's removeBatchEffect could also be applied to model and remove cross-study batch effects prior to pooling — Cross-study technical variation can dominate biological signal in multi-platform compendia; documenting an explicit batch-adjustment step and its scope aids reproducibility and interpretability
-
The DEG count (1,781 genes) is reported without a stated significance threshold or named multiple-testing correction procedure↳ Could also: Applying Benjamini-Hochberg FDR correction at a pre-specified threshold (e.g., FDR < 0.05) and reporting adjusted p-values alongside log-fold changes for each gene would also define the DEG set — Documenting the correction method and threshold makes the gene list reproducible and allows readers to calibrate confidence in individual gene-level findings
-
The 'top 20' differentially expressed genes were selected for heatmap visualization, but the ranking metric (p-value, fold-change, or a composite score) is not explicitly stated↳ Could also: A volcano plot displaying all tested genes by log-fold change versus −log10(adjusted p-value) would also convey both effect magnitude and statistical confidence simultaneously — Showing the full distribution helps readers assess how many genes cross practical versus statistical significance thresholds and contextualizes the top-20 selection within the broader landscape
-
Sample groupings were visualized using hierarchical clustering of expression heatmaps↳ Could also: Dimensionality-reduction projections such as PCA or UMAP would also reveal inter-sample relationships and potential technical confounders (e.g., platform, batch, smoking status) — Lower-dimensional projections display all samples simultaneously and can surface clustering driven by technical rather than biological variables, complementing the heatmap view
-
Data from mixed platforms (microarray and RNA-seq) were integrated and analyzed jointly with limma↳ Could also: A formal meta-analysis framework such as random-effects meta-analysis (e.g., MetaDE, REM in R) could also combine per-study effect estimates while explicitly modeling between-study heterogeneity — Meta-analytic heterogeneity statistics (e.g., I²) help distinguish findings that are robust across individual studies from those driven by a small number of experiments
-
Expression spread across samples is not reported in the main results (no SD, SEM, or CI for any gene)↳ Could also: Reporting SD or 95% CI alongside mean expression values for highlighted genes would also convey within-group variability — Measures of spread help readers assess whether group differences are consistent across individual samples or substantially influenced by outlier experiments or platforms
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- A curated collection of transcriptome datasets... L1 62/100
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Meta-analysis of gene expression profiles of l... L1 78/100
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
- Comparative profiling of skeletal muscle model... L1 64/100
- CoINcIDE: A framework for discovery of patient... L1 87/100
- A curated collection of transcriptome datasets... L1 62/100
- Meta-analysis of gene expression profiles of l... L1 78/100
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Screening of Diagnostic Biomarkers and Immune...⚑ L1 48/100 ⚑
- IRSN-23 gene diagnosis enhances breast cancer... L1 71/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 31949184 (PulmonDB: a curated lung disease gene expression database)
- Title: PulmonDB: a curated lung disease gene expression database. Villaseñor-Altamirano et al., Sci Rep 10:514 (2020).
- DOI: 10.1038/s41598-019-56339-5 · PMCID: PMC6965635
- Code: https://github.com/AnaBVA/PulmonDB (R package, last push 2020-11-16, no license)
- Live resource: http://pulmondb.liigh.unam.mx/ ; MySQL backend
«ip»:3306dbexpdata_hsapi_ipf, public guest credentials hardcoded in the R package (user=guest, `password: «redacted» Verified reachable + readable 2026-06-15.
What kind of paper this is
PulmonDB is a curated database resource, not a single analysis. Raw GEO microarray/RNA-seq experiments for COPD and IPF were re-processed into a uniform "homogenized" per-contrast log2 (M-value) compendium using the third-party tool COMMAND (COMpendia MANagement Desktop): RMA-quantile (Affymetrix) / loess (other platforms) normalization, RMA median-polish / replicate averaging summarization, probe→gene remapping to Gencode v25, manual ontology annotation. Per P16 of the brief, reproducing the paper by running its own R package against its own live database is a fully valid reproduction.
In scope (pipeline-derived, attempted)
- Database composition counts — # GSEs (76), # unique GSM contrasts (4481),
platforms/GPLs (26). Directly queryable from the live DB metadata tables.
Pipeline: COMMAND curation → MySQL schema. - COPD-vs-IPF differential expression — paper reports 1781 DE genes between
COPD and IPF. Reproduced by replicating the package vignette
vignettes/genes_IPFandCOPD.Rmd: pull homogenized M-values + curated disease/tissue annotation from the live DB, restrict to lung-biopsy COPD/IPF/CONTROL-vs-HEALTHY contrasts, NA-filter,limmadesign~0 + test2 + experiment_access_id, contrasts COPDvsIPF / COPDvsH / IPFvsH / COPDandIPFvsH,decideTests. Pipeline: limma on PulmonDB homogenized values. - Directional marker genes — FOSB & CXCL2 ("overexpressed in COPD and underexpressed in IPF"); VCAM1 & FCN3 ("consistent"/similar trend in both diseases vs control). Reproduced from the same limma fit (sign of logFC).
- GSE32537 IPF-marker sanity — GSE32537 (LGRC IPF lung biopsy) is one of the IPF validation experiments used to "recapitulate previously reported knowledge". Cross-check canonical IPF markers (MMP7, MMP1, SPP1, COMP, COL1A1) have the expected up direction in the homogenized M-values.
Out of scope (not attempted, why)
- Manual ontology curation of every sample (wet-lab/manual; not a pipeline).
- COMMAND raw→homogenized normalization from scratch — COMMAND is a desktop Java/MySQL compendium tool; re-running the full normalization of 76 GSEs from raw CEL/array files is the hard last ~20% and is not required (the homogenized output is the shipped, queryable product, which we reproduce against).
- Web-interface / Genome-browser / co-expression figures — UI features.
- The hierarchical-clustering "top 20 genes" figure beyond the DE counts/directions.
Key caveat (recorded up front): living database
The live DB has grown since the 2020 snapshot (89 GSEs total / 75 with normdata, 6857 contrasts, 34 platforms as of 2026-06-15 vs the paper's 76 / 4481 / 26). Composition counts and the exact DEG count are therefore expected to differ; the published values describe the 2020 freeze, not today's DB. We report both and grade honestly — gene directions are the robust, reproducible signal.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Method and headline biology reproduce 1:1 against PulmonDB's own live resource: FOSB (+1.428 COPD / -1.436 IPF) and CXCL2 (+0.885 / -0.865) show the reported COPD-up/IPF-down split exactly, and GSE32537 recapitulates canonical IPF markers (MMP7/SPP1/COL1A1 up, AGER down). The numeric divergences — 4481->6857 contrasts, 26->34 platforms, 1781->9295 DEGs — are fully and benignly explained by the database having grown ~4.5x in DE power since the 2020 freeze, a living-resource artifact rather than an authors' or fabrication defect (no 2020 snapshot is archived). Two minor 'similar-gene' examples (FCN3 opposite, VCAM1 absent) did not confirm. Overall: a solid reproduction whose only deviations sit on the input/version side, so the central conclusion holds but exact counts are not 1:1.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.