Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Integrated bioinformatics analysis screened the key genes and pathways of idiopathic pulmonary fibrosis.

Sci Rep · 2025
L1 93/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduces 1:1. Ran the DEG step of the authors' own 1_mRNA_process.R (limma-voom AND edgeR) on the public GEO GSE173355 PolyA count matrix on «our HPC» (SLURM 2175758). Structural counts reproduce EXACTLY (80930 protein_coding transcripts; 80930x37 matrix; 23 IPF + 14 control). edgeR DEG transcript count reproduces EXACTLY (4876; gene count 3711 vs reported 3712, off by one). limma DEG counts land within ~2.4% (5202 vs 5328 transcripts; 3994 vs 4083 genes) - a small gap consistent with limma/voom version drift (limma 3.66.0 / R 4.5.3, newer than the authors'), not a methods discrepancy. Every reported number is derivable from the shipped public data+code; no fabrication concern. NOT attempted (out of scope per 80/20, see scope.md): methylation (GSE173356/DMRcate), PPI hub genes (STRING+Cytoscape GUI), CIBERSORTx (external web service), ESTIMATE, TCGA/ELMER scripts - separate dataset/pipeline or external GUI/web tools not driven by the repo.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 93
    assessed: 2026-06-14 ⛓ 7182000d3173
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What are the key genes and pathways underlying idiopathic pulmonary fibrosis (IPF), as revealed through integrated bioinformatics analysis of gene expression, DNA methylation, and immune microenvironment data?

Core claims
  • DNA methylation negatively influences the expression of 8 genes in IPF (PDCD1LG2, TTC34, CNNM1, ADAMTS16, MKX, SLC22A3, C1orf53, CP) finding
  • The Rap1 signaling pathway, Focal adhesion, and Axon guidance are significantly enriched among both DEGs and differentially methylated sites finding
  • Immune-related scores (stromal, immune, ESTIMATE) are significantly lower in IPF than in control, with altered immune cell proportions finding
  • PPI network analysis of differentially expressed immune-related genes identified 10 hub genes and 3 core subnetworks finding
  • Integrated bioinformatics analysis of paired expression and methylation datasets combined with immune profiling identifies IPF biomarkers and pathways method
  • RT-qPCR in bleomycin-induced IPF mouse model and TGF-β-induced A549 EMT model confirmed the reliability of most bioinformatic findings finding
  • 361 differentially expressed immune-related genes (IRGs) were identified by intersecting DEGs with ImmPort IRGs resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq expression profiling (DEG analysis) human IPF and control samples (GSE173355) none differentially expressed genes/transcripts GPL24676 Illumina NovaSeq 6000
DNA methylation array (differentially methylated sites/regions) human IPF and control samples (GSE173356) none differentially methylated CpG sites and regions GPL23976 Illumina Infinium HumanMethylation850 BeadChip
immune deconvolution (ESTIMATE, CIBERSORTx, GSVA) human IPF and control samples (GSE173355) none immune/stromal/ESTIMATE scores, 22 immune cell proportions, 29 immune cell GSVA scores
PPI network analysis differentially expressed immune-related genes (human) none hub genes and core subnetworks STRING / Cytoscape (cytohubba, MCODE)
RT-qPCR validation bleomycin-induced IPF mouse model (male C57BL/6, lung tissue) intratracheal bleomycin 2.5 mg/kg vs saline relative gene expression (2^-∆∆Ct) MedChemExpress HY-17565 BLM
RT-qPCR validation A549 cells (EMT model) 10 ng/mL TGF-β1 for 72 h relative gene expression (2^-∆∆Ct); EMT markers α-SMA, Vimentin, E-cadherin Hangzhou Huaan Biotechnology antibodies
histology (HE and Masson staining) bleomycin-induced IPF mouse lung tissue intratracheal bleomycin vs saline fibrosis/tissue morphology
Key results
  • PDCD1LG2 promoter CpG cg11299543 hypermethylated leading to low PDCD1LG2 expression in IPF
  • TTC34, CNNM1, ADAMTS16, MKX, SLC22A3, C1orf53 and CP promoter CpG sites hypomethylated leading to high gene expression in IPF
  • Stromal, immune and ESTIMATE scores significantly lower in IPF than control
  • Normal group had significantly higher immune scores than IPF after adjusting for age, sex, smoking β = 304.97, 95% CI [206.8, 403.1]
  • Males exhibited significantly lower immune scores than females β = -152.54, 95% CI [-277.5, -27.6]
  • Resting dendritic cells, M0/M2 macrophages, resting mast cells, plasma cells, activated CD4 memory T cells and gamma delta T cells higher in IPF; activated mast cells, monocytes, neutrophils, activated NK cells and resting CD4 memory T cells lower in IPF
  • 4012 transcripts (3206 genes) found differentially expressed by both limma and edgeR packages
  • 4933 differentially methylated sites and 87 differentially methylated regions identified in IPF vs control
Key statistics
  • other β = 304.97, 95% CI = [206.8, 403.1], p < 0.001 (normal vs IPF immune score after adjustment)
  • other β = -152.54, 95% CI = [-277.5, -27.6], p = 0.021 (male vs female immune score)
  • pvalue P < 0.01 (stromal/immune/ESTIMATE scores IPF vs control)
  • pvalue P < 0.05 (differential immune cell infiltration IPF vs control)
  • count 5328 differentially expressed transcripts / 4083 genes (limma) (DEGs by limma, logFC>1 FDR<0.05)
  • count 4876 differentially expressed transcripts / 3712 genes (edgeR) (DEGs by edgeR)
  • count 4933 differentially methylated sites; 87 DMRs (methylation analysis, FDR<0.01, region diff >0.1)
  • count 361 differentially expressed immune-related genes; 10 hub genes; 3 core subnetworks (IRG PPI analysis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study analyzed two paired GEO datasets (GSE173355, bulk RNA expression; GSE173356, DNA methylation; 23 IPF vs 14 control samples) using parallel differential expression pipelines (limma and edgeR) and a t-test-based differential methylation pipeline, all with FDR correction. Functional enrichment (GO/KEGG via clusterProfiler and missMethyl), immune microenvironment quantification (ESTIMATE algorithm, CIBERSORTx deconvolution, GSVA scoring), mRNA-methylation association analysis (ELMER), and PPI network hub-gene extraction (STRING/Cytoscape) were integrated. Key findings were validated by RT-qPCR using the 2^-ΔΔCt method in a bleomycin mouse model and a TGF-β1-induced A549 EMT cell model.

Replicationbiological Sample size37 human samples from GEO (23 IPF, 14 control); animal and cell model sample sizes not stated in provided text GroupsIPF vs healthy control (human datasets); bleomycin-treated vs saline control (mouse); TGF-β1-treated vs untreated (A549 cells) Pairingunpaired Randomization/blindingstated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionFDR (Benjamini-Hochberg, applied through limma, edgeR, clusterProfiler, missMethyl)
Statistical tests used
Test Applied to n Assumptions
limma moderated linear model with empirical Bayes shrinkage (differential expression) DEG analysis, IPF vs control, GSE173355; Figs. 1C–D 37 (23 IPF, 14 control) not stated
edgeR negative binomial GLM (differential expression) DEG analysis, IPF vs control, GSE173355; Figs. 2C–D 37 (23 IPF, 14 control) not stated
t-test via limma (differentially methylated sites) Methylation analysis, IPF vs control, GSE173356; Figs. 3B–C 37 (23 IPF, 14 control) not stated
GSVA pathway enrichment scoring + limma differential testing (KEGG pathway scores) Differential KEGG pathway activity, IPF vs control; Figs. 1G, 2H 37 (23 IPF, 14 control) not stated
Multivariate linear regression (group + age + sex + smoking as covariates) ESTIMATE immune score, stromal score, ESTIMATE score, IPF vs control; Fig. 5B 37 (23 IPF, 14 control) not stated
Group comparison of CIBERSORTx immune cell proportions (specific test not named; P < 0.05 threshold stated) 22 immune cell type proportions, IPF vs control; Fig. 6C 37 (23 IPF, 14 control) not stated
RT-qPCR with 2^-ΔΔCt quantification (group comparison test not specified) Validation of hub/methylation-associated genes in bleomycin mouse model and A549 EMT model null not stated
Approaches that could also have been used
  • Differentially methylated sites were identified using a t-test on beta values implemented in limma
    Could also: M-value-based limma (logit-transformed beta values) or beta-binomial regression models (e.g., DSS or methylKit) are also used for Illumina methylation array data — M-values better satisfy the normality assumption of linear models, and beta-binomial approaches explicitly model the bounded [0,1] distribution of beta values, which can improve sensitivity at extremes of the methylation range
  • Hub genes were ranked solely by degree centrality in the PPI network using CytoHubba
    Could also: Betweenness centrality, closeness centrality, or the MCC (Maximal Clique Centrality) algorithm in CytoHubba could also rank nodes; taking the intersection across multiple centrality measures is another option — Different centrality metrics capture different topological roles (bottleneck hubs vs. highly connected hubs), and selecting genes consistently ranked highly across multiple metrics can increase robustness of hub identification
  • Immune microenvironment was quantified using the ESTIMATE algorithm for stromal and immune scores
    Could also: ssGSEA, MCP-counter, or xCell are complementary bulk-expression immune deconvolution/scoring approaches applicable to the same dataset — Each algorithm uses distinct gene signatures and mathematical frameworks; convergent findings across multiple methods provide stronger evidence about immune composition differences
  • DEG consensus was defined as the intersection of limma and edgeR results
    Could also: DESeq2 with Wald test (or likelihood-ratio test) is a third widely used differential expression method for count data that could also have been included in the consensus strategy — Including a third independent method increases the likelihood that the final gene list is robust across statistical modeling assumptions, a common approach in bioinformatics benchmarking studies
  • RT-qPCR validation results were presented without naming the statistical test used for group comparisons or reporting the animal/cell experiment sample sizes
    Could also: Student's t-test or Mann-Whitney U test with explicit reporting of n per group, mean or median, a dispersion measure (SD or IQR), and exact p-values is standard for RT-qPCR validation experiments — Reporting these elements allows readers to independently assess statistical power and reproducibility of the experimental validation results
  • Group comparisons of ESTIMATE scores and immune cell proportions were reported with threshold p-values (P < 0.05, P < 0.01) without accompanying dispersion measures
    Could also: Reporting SD or 95% CI alongside point estimates for boxplot comparisons, and providing exact p-values, is also standard practice — Dispersion measures allow readers to gauge within-group variability and the practical magnitude of between-group differences independently of sample size; exact p-values facilitate meta-analytic synthesis
Software: R/DESeq (PCA) R 3.6.2 · R/limma · R/edgeR · R/clusterProfiler · R/missMethyl · R/GSVA Bioconductor v2.0.5 · R/DMRcate · R/ELMER · CIBERSORTx · Cytoscape (with CytoHubba and MCODE plugins) · STRING database

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-40280949

Paper: Wu J et al. "Integrated bioinformatics analysis screened the key genes and pathways of idiopathic pulmonary fibrosis." Sci Rep 2025. DOI 10.1038/s41598-025-97037-9. Code: https://github.com/wujuan948725/RNAseq-methylation-analysis-2025 (authors' own, public, not archived; 8 R scripts). Data: GEO GSE173355 (mRNA, PolyA counts) + GSE173356 (850k methylation).

Pipeline map (which reported result comes from which pipeline)

Reported result Pipeline / script In scope?
mRNA DEGs: 5328 transcripts / 4083 genes (limma, logFC>1 & FDR<0.05) 1_mRNA_process.R — edgeR/limma-voom on GSE173355 PolyA counts YES (primary)
protein_coding transcripts = 80930; matrix 80930 × 37 (23 IPF + 14 ctrl) 1_mRNA_process.R ingest/filter YES (checkpoint)
mRNA DEGs: 4876 transcripts / 3712 genes (edgeR) 1_mRNA_process.R DEA(method='edgeR') YES (secondary, ~free)
GO/KEGG enrichment (named pathways: Rap1, Focal adhesion, PI3K-Akt…) 1_mRNA_process.R clusterProfiler partial — qualitative, depends on online annotation DB versions; not primary
Methylation: 4933 DM sites (FDR<0.01), 87 DMRs (DMRcate) 3_meth_data_analysis.R on GSE173356 OUT (separate dataset/pipeline; 80/20 — attempt only if primary is cheap)
Hub genes HRAS/MAPK3/FYN/JUN/RHOA/NFκB1/AKT1/B2M/FOS/IL6 STRING + Cytoscape cytoHubba OUT (external GUI tools, not scriptable from repo)
Immune infiltration: CIBERSORTx 22 cells / 12 different; ESTIMATE 2_immcell_DEA_KEGG.R + external CIBERSORTx OUT (CIBERSORTx is an external web service)
ELMER / TCGA-biolinks analyses (scripts 4,5) TCGA, not GSE173355 OUT (different data, exploratory)

What we attempt (80/20 — clearly-specified, low-hanging, deterministic integer outputs)

The single cleanest, fully-specified, self-contained pipeline output: the DEG counts from limma-voom on GSE173355 PolyA counts. The repo even carries the expected values inline as comments (#[1] 5328 6, unique symbols #[1] 4083), and the paper text states the same ("5328 differentially expressed transcripts, corresponding to 4083 genes"). These are exact integer counts → cleanly gradable exact/mismatch. No annotation.rda needed (rownames are ENST, not ENSG, so the DEA symbol-mapping branch is skipped; gene symbols come from target_id field 6).

Claims attempted:

  • C1 protein_coding transcript count = 80930 (and matrix 80930×37).
  • C2 limma DEGs (FDR<0.05 & |logFC|>1) = 5328 transcripts.
  • C3 unique gene symbols among DEGs = 4083.
  • C4 (secondary) edgeR DEGs = 4876 transcripts / 3712 genes.

What we do NOT attempt, and why

  • Methylation (GSE173356), hub-gene PPI (STRING/Cytoscape GUI), CIBERSORTx (external web service), ESTIMATE, TCGA/ELMER scripts — these are either a separate dataset/pipeline, external GUI/web tools not driven by the repo, or qualitative named-pathway lists whose exact reproduction depends on annotation-DB versions. Per the 80/20 rule these are out of scope for this RU.
  • GO/KEGG enrichment: deterministic in principle but the named top pathways are version-sensitive; not graded here.
C1
Reported
80930 protein_coding transcripts
Reproduced
80930
exact
C1b
Reported
matrix 80930 x 37 (23 IPF + 14 control)
Reproduced
80930x37
exact
C2
Reported
5328 limma DEG transcripts (FDR<0.05,|logFC|>1)
Reproduced
5202
within tolerance
C3
Reported
4083 genes among limma DEGs
Reproduced
3994
within tolerance
C4
Reported
4876 edgeR DEG transcripts
Reproduced
4876
exact
C4b
Reported
3712 genes among edgeR DEGs
Reproduced
3711
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

This is a near-1:1 reproduction on identical public data (GSE173355) using the authors' own 1_mRNA_process.R. Structural counts (80930×37) and the edgeR DEG transcript count (4876) reproduce exactly; the only deviation is limma DEGs landing within −2.4% (5202 vs 5328 transcripts; 3994 vs 4083 genes), fully explained by limma/voom version drift, not a methods or data defect. Every reported value is derivable from the shared data+code, so there is no fabrication concern. The yellow on q3 only reflects that the (negligible) deviation sits in the limma statistical step; coverage is partial (DEG step only), but nothing reproduced contradicts the paper.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

104.2 k
tokens (I/O) · 6.3 M incl. cache
9 min
runtime · 0.01 CPU-h
1.3 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine