Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Exploring the role and mechanism of Astragalus membranaceus and radix paeoniae rubra in idiopathic pulmonary fibrosis through network pharmacology and experimen

Sci Rep · 2023
L1 20/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
20/100
Reproducibility score
3.1 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 1% of all assessed papers rank 1166 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Network-pharmacology + GEO paper on Astragalus/Radix Paeoniae Rubra in IPF. The registry 'code' link (CJ-Chen/TBtools) is a visualization toolkit, NOT an analysis pipeline; the reproducible computational result is the GEO microarray DEG analysis described in Methods (R affy/oligo RMA + limma; |logFC|>0.5 & FDR<0.05). I reproduced that pipeline faithfully on «our HPC» for all three datasets (GSE24206 primary; GSE101286, GSE110147 secondary). The PIPELINE reproduces and runs cleanly (after rebuilding preprocessCore --disable-threading to fix an HPC pthread bug), but the HEADLINE NUMBERS DO NOT MATCH. GSE24206 at the paper's stated thresholds yields 2784 DEG genes (1747 up / 1037 down) — about 7x the reported 412 (185 up / 227 down) — and the reported down>up polarity is INVERTED vs every reproduction (up>down). A threshold scan (logFC 0.5/1/2 x FDR 0.05/0.01, raw P) and a control polarity flip both fail to reproduce 185/227; using 11 IPF arrays (one per patient) instead of 17 still gives ~2133 genes. The combined 3-dataset count (4640) matches neither the union (10231) nor the intersection (109). The 6 controls reproduce exactly and '11 IPF' is reconcilable as 11 patients (17 arrays). FLAGGED for human review as possible fabrication / undisclosed-method: the reported DEG counts are not derivable from the stated data + stated method as described (grades provisional). NOT attempted (out of scope, non_pipeline): TCMSP 32 compounds, SwissTargetPrediction 728 targets, 171/117 Venn targets, 10 STRING/Cytoscape hub genes, 173 KEGG pathways (all version-dependent web-DB/GUI steps with no shipped data/code), and the wet-lab bleomycin-mouse validation.

💻 Code ↗ 🗄 Data: GSE24206

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 20
    assessed: 2026-06-14 ⛓ 42eb3415a1a9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This study investigates the multi-target, multi-pathway molecular mechanisms by which Astragalus membranaceus (AM) and Radix paeoniae rubra (RPR) exert therapeutic effects on idiopathic pulmonary fibrosis (IPF), hypothesizing that their active components act on key IPF-associated targets and pathways to alleviate fibrosis.

Core claims
  • 117 (reported elsewhere as 171) common targets between IPF DEGs and AM/RPR drug targets identify AKT1, MAPK3, HSP90AA1, VEGFA, CASP3, JUN, HIF1A, CCND1, PTGS2, and MDM2 as key targets. finding
  • Common targets are enriched in the PI3K-AKT pathway, HIF-1 signaling pathway, apoptosis, and microRNAs in cancer. mechanism
  • Astragaloside III, (R)-Isomucronulatol, Astragaloside I, Paeoniflorin, and β-sitosterol are the main active anti-IPF components of AM and RPR. finding
  • Molecular docking shows strong binding affinity (scores -4.7 to -10.7 kcal/mol) between main active compounds and key target proteins. finding
  • AM and RPR treatment alleviates bleomycin-induced pulmonary fibrotic lung damage in rats. finding
  • AM and RPR reduce mRNA levels of AKT1, HSP90AA1, CASP3, MAPK3, and VEGFA, and protein levels of AKT1, HSP90AA1, and VEGFA. finding
  • Network pharmacology combined with molecular docking and in vivo validation can predict and confirm TCM mechanisms in IPF. method
  • Integration of GEO DEGs, TCMSP/SwissTargetPrediction targets, STRING PPI, and Metascape enrichment constitutes a reusable analysis pipeline. resource
Experimental setups
Assay System Perturbation Readout Platform
Microarray DEG analysis (bulk transcriptomics) Human IPF and control lung tissues (GSE24206, GSE101286, GSE110147) none (disease vs control) Differentially expressed genes (|logFC|>0.5, FDR<0.05) GPL570, GPL6947, GPL6244; Affy/Oligo/limma R packages
Network pharmacology target prediction Human (Homo sapiens) in silico none AM/RPR active component targets TCMSP, SwissTargetPrediction, PubChem
Protein–protein interaction network analysis Human in silico none Hub genes by degree STRING v11.5, Cytoscape v3.7.2
GO/KEGG enrichment analysis Human in silico none Enriched pathways and GO terms Metascape
Molecular docking In silico (key target proteins as receptors, components as ligands) none Binding affinity (docking score, kcal/mol) AutoDock Tools v4.2, AutoDock Vina, Chem3D v18.0, wwPDB/RCSB
Histology (H&E and Masson's trichrome) SD male rat lung tissue Bleomycin 5 mg/kg intratracheal; AM/RPR decoction 8.2 g/kg gavage Fibrotic/collagen pathology
Colorimetric assay (hydroxyproline) SD male rat lung tissue Bleomycin; AM/RPR decoction HYP concentration (OD at 558 nm) E-BC-K602-M Elabscience kit; ELISA reader
RT-qPCR and immunohistochemistry SD male rat lung tissue Bleomycin; AM/RPR decoction mRNA (2-ΔΔCt) and protein (MOD) of AKT1, HSP90AA1, CASP3, MAPK3, VEGFA, MMP7 Foregene RNA kit, NanoDrop; antibodies from Abcam/Abmart/Affinity/Immunoway/Servicebio; Image-Pro Plus 6.0
Key results
  • 4640 DEGs identified in IPF lung tissues vs normal across three datasets 4640 DEGs
  • 171 common targets between IPF DEGs and AM/RPR targets identified 171 targets
  • PPI network constructed with 166 nodes and 1030 edges yielding 10 key targets 166 nodes, 1030 edges
  • 173 KEGG pathways significantly enriched including PI3K-AKT and HIF-1 signaling 173 pathways
  • Docking scores between main compounds and key targets indicate strong binding -4.7 to -10.7 kcal/mol
  • AM and RPR reduced mRNA of AKT1, HSP90AA1, CASP3, MAPK3, VEGFA
  • AM and RPR reduced protein expression of AKT1, HSP90AA1, VEGFA
  • Key ingredient stress values in CTP network ranked top compounds Astragaloside III=19660; (R)-Isomucronulatol=17304; Astragaloside I=5622; Paeoniflorin=5112; β-sitosterol=3910
Key statistics
  • count 4640 DEGs (DEGs in IPF vs normal lung across 3 datasets (17+10+33 tissues))
  • count 171 common targets (intersection of IPF DEGs and AM/RPR targets)
  • count 728 genes (potential targets of AM and RPR from SwissTargetPrediction)
  • other docking scores -4.7 to -10.7 kcal/mol (binding affinity between main compounds and key targets)
  • count 227 down / 185 up (DEGs in GSE24206)
  • count 502 down / 571 up (DEGs in GSE101286)
  • count 1293 down / 2238 up (DEGs in GSE110147)
  • mean 67.35 (5.00), 66.75 (6.50), 62.00 (6.00) years (mean (SD) age of IPF patients in three datasets)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined computational network pharmacology (DEG identification via limma across three GEO microarray datasets, PPI network construction, GO/KEGG enrichment) with an in vivo bleomycin rat model (n=6 per group, three groups). In vivo outcome data were tested for normality with Shapiro-Wilk and analyzed with one-way ANOVA (normal distribution) or Kruskal-Wallis (non-normal distribution); results were reported as mean ± SD with threshold-based significance symbols. Molecular docking affinity scores (kcal/mol) and qPCR 2−ΔΔCt values were reported descriptively without formal inferential testing beyond the group comparisons above.

Replicationbiological Sample size6 rats per group stated; no a priori power calculation or sample size justification reported GroupsControl (saline) vs. BLM (bleomycin model) vs. HC (bleomycin + AM/RPR decoction); 3 groups Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFDR < 0.05 applied for DEG identification via limma; no post-hoc correction method stated for in vivo ANOVA or Kruskal-Wallis comparisons
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test / empirical Bayes (DEG analysis) IPF vs. control lung tissues in each of three GEO datasets (GSE24206, GSE101286, GSE110147) 17 tissues (GSE24206), 10 tissues (GSE101286), 33 tissues (GSE110147) not stated
Shapiro-Wilk normality test Pre-test for all in vivo outcome variables before selecting parametric vs. non-parametric analysis 6 per group (18 total) na
One-way ANOVA Group comparisons (Control vs. BLM vs. HC) for normally distributed in vivo outcomes (HYP colorimetric, qPCR, IHC MOD) 6 per group (18 total) not stated
Kruskal-Wallis test Group comparisons for in vivo outcomes not conforming to normal distribution 6 per group (18 total) not stated
Approaches that could also have been used
  • Three GEO microarray datasets were analyzed independently with limma and their DEG lists were intersected via Venn diagram
    Could also: A formal cross-dataset meta-analysis (e.g., using the R/MetaDE or RankProd package, or ComBat batch correction followed by joint limma analysis) could also integrate the three datasets — A joint or meta-analytic approach pools statistical power across datasets and produces a single ranked gene list with cross-study uncertainty estimates, whereas independent intersection can inflate confidence in genes that happen to pass threshold in all datasets and misses genes with consistent but sub-threshold effects in individual datasets
  • One-way ANOVA (or Kruskal-Wallis) was used for three-group in vivo comparisons, but no post-hoc test is named
    Could also: Specifying a post-hoc procedure — such as Tukey HSD (for all pairwise comparisons) or Dunnett's test (for comparisons against a single control) after ANOVA, or Dunn's test with Bonferroni or Holm correction after Kruskal-Wallis — would also control the family-wise error rate across the three pairwise contrasts — Without a named post-hoc step, it is not possible to determine which specific pairwise differences drove a significant omnibus result or how type-I error was managed across the three pairwise comparisons
  • In vivo results were reported as mean ± SD with asterisk-based significance thresholds (e.g., * P<0.05) rather than exact p-values
    Could also: Reporting exact p-values alongside mean ± SD (or mean ± SEM with n stated) would also convey the same information — Exact p-values allow readers to assess the continuous evidence gradient and to use values in downstream meta-analyses or power calculations; threshold symbols alone convey only binary significance relative to the chosen cutoff
  • Group differences were quantified only with significance symbols; no standardized effect sizes were reported
    Could also: Reporting eta-squared (η²) or partial η² for ANOVA, or rank-biserial correlation for Kruskal-Wallis, would also quantify the magnitude of group differences — Effect sizes are independent of sample size and allow comparison of biological magnitude across studies; with n=6 per group, a statistically significant result can correspond to a wide range of effect magnitudes
  • Normality was assessed with Shapiro-Wilk and the result determined the choice between ANOVA and Kruskal-Wallis, applied separately to each outcome variable
    Could also: A mixed-effects model or a permutation-based ANOVA could also be used and would be robust to non-normality without requiring a two-step test-then-choose procedure — The two-step 'test for normality then choose test' approach inflates type-I error because the normality test result is itself uncertain, especially at n=6; a single robust or permutation-based approach avoids this conditional decision
  • Molecular docking binding affinities (kcal/mol) were reported as point estimates without uncertainty quantification
    Could also: Ensemble docking across multiple protein conformations (e.g., from molecular dynamics snapshots) or reporting across multiple docking runs with different random seeds would also characterize variability in predicted binding affinity — A single rigid-receptor docking score is sensitive to the chosen crystal structure conformation; reporting a distribution of scores across conformational states gives a more complete picture of binding stability
Software: R/limma (Bioconductor) · R/affy · R/oligo · GraphPad Prism 9.4.1 · Cytoscape 3.7.2 · Metascape · AutoDock Vina / AutoDock Tools 4.2 (AutoDock Tools) · Image-Pro Plus 6.0 · TBtools 1.09876.3 · Chem3D 18.0

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
18
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37666859

Paper: Jiang H et al. (2023) Exploring the role and mechanism of Astragalus membranaceus and radix paeoniae rubra in idiopathic pulmonary fibrosis through network pharmacology and experimental verification. Sci Rep 13:14745. DOI 10.1038/s41598-023-36944-1 · PMID 37666859 · PMC10477296.

"Code" link in registry: github.com/CJ-Chen/TBtools — this is TBtools, a general-purpose bioinformatics visualization toolkit (Venn diagrams, heatmaps). It is NOT the authors' analysis pipeline; it is a soft signal. The paper's actual computational pipeline is described in Methods (R packages + web databases), not shipped as a repo.

Pipeline-derived results and their reproducibility

# Reported result Source / tool In scope? Why
GEO-DEG GSE24206: 11 IPF + 6 control; 227 down / 185 up DEGs R affy/oligo + limma, |logFC|>0.5 & FDR<0.05 YES Public GEO data + fully-specified standard pipeline (limma). Independently reproducible.
GEO-DIM GSE24206 = 11 IPF + 6 control lung tissues, GPL570 GEO metadata YES Verifiable from public GEO record.
GEO-3DS 4640 DEGs across GSE24206 + GSE101286 + GSE110147 limma on 3 sets partial/secondary The combine rule (union/intersection/meta) is not stated → can only sanity-check the per-set magnitude.
NP-compounds 32 bioactive ingredients (AM 15+4 astragalosides, RPR 13), OB≥30% & DL≥0.18 TCMSP web DB NO Version-dependent web-database query; no shipped data/code; manual curation ("4 astragalosides included despite not meeting thresholds"). non_pipeline.
NP-targets 728 compound targets SwissTargetPrediction web NO Web DB, version-dependent, no shipped table.
NP-venn 171 common targets (→117 in PPI) Venn of drug∩disease NO Derived from the non-reproducible web-DB sets above.
NP-hub 10 hub genes (AKT1, MAPK3, HSP90AA1, VEGFA, CASP3, JUN, HIF1A, CCND1, PTGS2, MDM2) STRING + Cytoscape NO Depends on the non-reproducible 117-node PPI; web-DB + GUI.
NP-kegg 173 KEGG pathways enriched; top PI3K-AKT/HIF-1/apoptosis DAVID/web enrichment NO Web-DB + version-dependent gene sets; depends on non-reproducible target list.
Wet-lab bleomycin mouse model, HE/Masson, WB, qPCR validation experimental NO Wet-lab, out of scope by design.

Decision

Reproduce the GSE24206 limma DEG analysis (GEO-DEG + GEO-DIM) as the primary in-scope pipeline result, on «our HPC», from raw CEL files via the paper's stated affy (RMA) + limma route, thresholds |logFC|>0.5 & FDR(BH)<0.05. Secondary: run the same pipeline on GSE101286 and GSE110147 and report the union magnitude vs the reported 4640 (provisional — combine rule unstated).

The network-pharmacology web-database steps are out of scope as non-reproducible pipeline results (non_pipeline / web-DB, version-dependent, no shipped artifacts). This is documented, not a quality drop of the paper.

Figures / tables: Fig 2ATableFig 3Fig 4Fig 5
C1
Reported
GSE24206 = 11 IPF + 6 control lung tissues (GPL570)
Reproduced
6 control exact; 17 IPF arrays from 11 unique patients (advanced-IPF paired upper+lower lobes) -> '11 IPF' reconcilable as 11 patients
partial
C2u
Reported
185 up-regulated DEGs in GSE24206
Reproduced
1747 up (gene-level) / 2638 up (probe); 1328 up (gene, 11-IPF subset) at stated |logFC|>0.5 & FDR<0.05
did not match
C2d
Reported
227 down-regulated DEGs in GSE24206
Reproduced
1037 down (gene-level) / 1898 down (probe); 805 down (gene, 11-IPF subset)
did not match
C3
Reported
4640 DEGs across GSE24206+GSE101286+GSE110147
Reproduced
gene-level union across 3 sets = 10231; 3-way intersection = 109 (combine rule unstated in paper)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 20/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The only reproducible computational claims — the GSE24206 DEG counts and the combined 3-dataset count that seed the entire network-pharmacology analysis — do not reproduce: standard limma at the paper's own stated thresholds yields 2784 genes (1747 up / 1037 down), ~7x the reported 412 (185 up / 227 down), with inverted up/down polarity, and the reported 4640 matches neither union (10231) nor intersection (109). The data are public and the controls reproduce exactly, so this is not a data-availability problem; the deviation sits in the authors' reported output and is not derivable from the stated data+method, with the supplied code link being a visualization tool rather than the pipeline. This is severe (magnitude + direction + an unstated combine rule) and is flagged fabrication-suspect for human confirmation against Fig 2.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

164.6 k
tokens (I/O) · 11.5 M incl. cache
22 min
runtime · 0.07 CPU-h
2.4 GB
peak RAM
4
HPC jobs
hummel
machine