Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Multi-scale integrative analyses identify THBS2+ cancer-associated fibroblasts as a key orchestrator promoting aggressiveness in early-stage lung ade

Theranostics · 2022
L1 68/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 32% of all assessed papers rank 765 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to PARTIALLY reproduce from the single open dataset pinned to this room. The paper is a large multi-scale study whose headline results (THBS2+ CAF discovery via scRNA-seq, CellChat ligand-receptor signaling, WGCNA + Cox survival) all rest on the authors' OWN unreleased single-cell and clinical-cohort data and are not reproducible from public data. Using GSE10072 (the pinned Affymetrix LUAD tumor-vs-normal microarray, 58 tumor / 49 normal), a clean limma DE on «our HPC» independently CORROBORATES the paper's central tumor-enrichment claim: THBS2 is strongly upregulated in tumor (logFC +2.09, adj.P 4.7e-22), and 7 of 8 testable THBS2+ CAF signature markers (COL3A1, COL5A2, COL1A1, COL6A3, FAP, SULF1, + THBS2) are significantly upregulated in tumor vs normal — consistent with CAF/stromal enrichment in tumor tissue. BGN was borderline (adj.P 0.088) and CTHRC1 is not on the older U133A array. This is an honest 1:1 directional corroboration of the open-data slice, NOT a reproduction of the single-cell / CellChat / survival results, which the pinned code (CellChat) cannot regenerate without the unreleased data. compute_ran=true, «our HPC» reachable throughout; healthy=true (run completed as intended).

💻 Code ↗ 🗄 Data: GSE10072

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 68
    assessed: 2026-06-20 ⛓ 54fc3933c122
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether a specific molecular biomarker and its cellular source in the tumor microenvironment can explain and predict the aggressive, high-risk phenotype seen in a subset of pathologically node-negative (pN0), early-stage lung adenocarcinoma patients who have poor post-surgical outcomes despite curative resection.

Core claims
  • THBS2 is a tumor size-independent biomarker that robustly predicts post-surgical OS and RFS in multiple independent early-stage LUAD cohorts finding
  • THBS2 is exclusively derived from a specific CAF subset distinct from CAFs defined by classical markers finding
  • THBS2 is preferentially secreted via exosomes in early-stage LUAD tumors with high aggressiveness finding
  • THBS2-high early-stage LUAD is characterized by suppressed antitumor immunity, decreased immune infiltrates, and increased immune exhaustion markers finding
  • THBS2+ CAFs mainly interact with B cells, CD8+ T lymphocytes, and macrophages within the tumor microenvironment finding
  • High THBS2 expression predicts poor response to immunotherapy and short post-treatment survival finding
  • THBS2 recombinant protein suppresses ex vivo T cell proliferation and promotes in vivo LUAD tumor growth and distant micro-metastasis mechanism
  • WGCNA applied to TCGA LUAD transcriptomic and survival data identifies gene modules correlated with RFS/OS to nominate biomarker candidates method
Experimental setups
Assay System Perturbation Readout Platform
WGCNA / bulk transcriptomics TCGA pN0/M0-stage LUAD tumor tissue none gene co-expression modules correlated with RFS and OS
scRNA-seq paired primary human LUAD tumor and matched normal lung tissue (Huadong Hospital) none cell type identity and CAF subclustering 10X Genomics Chromium; Illumina NovaSeq 6000
GSEA (hallmark gene sets) THBS2+ vs THBS2- CAFs (scRNA-seq derived) none differentially enriched pathways
Cell-cell interaction network analysis (CellChat) scRNA-seq data from lung cancer samples none ligand-receptor interactions among cell types CellChat R toolkit v1.1.2
Exosome isolation and characterization (TEM, NTA, Western blot) human LUAD tumor tissue, normal adjacent tissue, and plasma (patients and healthy donors) none exosome morphology, size, and THBS2/marker protein levels SEC + ultracentrifugation; ZetaView PMX 110; H-7650 TEM
Transwell migration assay LUAD cell line THBS2 recombinant protein (50 ng/mL) vs PBS cell migration Corning Transwell plates
Ex vivo T cell proliferation assay T cells THBS2 recombinant protein T cell proliferation
In vivo tumor growth/metastasis assay LUAD mouse model THBS2 recombinant protein tumor growth and distant micro-metastasis
Key results
  • THBS2 high expression predicts poor OS and RFS in multiple independent early-stage LUAD cohorts, independent of tumor size
  • THBS2 is exclusively expressed by a distinct CAF subset not captured by classical CAF markers
  • THBS2 is preferentially secreted via exosomes in tumors with high aggressiveness
  • Plasma THBS2 levels associate with short recurrence-free survival
  • THBS2-high LUAD shows decreased immune cell infiltrates and increased immune exhaustion markers
  • High THBS2 expression predicts poor immunotherapy response and shorter post-treatment survival
  • THBS2 recombinant protein suppresses ex vivo T cell proliferation
  • THBS2 recombinant protein promotes in vivo LUAD tumor growth and distant micro-metastasis
Key statistics
  • count 43,779 (total number of cells clustered and visualized by UMAP in scRNA-seq analysis)
  • count 332 CAFs categorized into 10 clusters (number of cancer-associated fibroblasts identified and subclustered)
  • other soft-thresholding power = 12, scale-free R2 ≈ 0.9, cut height = 0.25, minimal module size = 30 (WGCNA parameters used to construct gene co-expression network)
  • count top 30 intramodular connected genes (definition of hub genes in WGCNA module)
  • count N = 5 tumor, N = 5 normal adjacent tissue (tissue samples used for exosome isolation)
  • count N = 5 pN0M0-LUAD, N = 5 healthy controls (plasma samples used for exosome isolation and comparison)
  • other Transwell assay results from three independent experiments (replication of migration assay)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper uses a multi-omic, cohort-based discovery-and-validation design: WGCNA (module eigengenes correlated with RFS/OS via Pearson correlation) applied to TCGA LUAD transcriptomic data to nominate a prognostic gene, followed by validation of THBS2 across independent transcriptomic, proteomic, and single-cell cohorts. Functional characterization used GO/KEGG/Reactome and GSEA pathway enrichment, Seurat-based single-cell clustering/marker identification, and CellChat ligand-receptor interaction scoring, alongside in vitro (e.g., Transwell, three independent experiments) and in vivo functional assays and RECIST v1.1-based clinical response classification. The provided text does not include a dedicated 'Statistical Analysis' subsection detailing specific inferential group-comparison tests, survival-analysis methods, or multiple-testing correction, so several standard reporting fields below are marked as not stated.

Replicationmixed Sample sizeSample sizes are stated for specific sub-analyses (e.g., two paired tumor/normal samples for scRNA-seq yielding 43,779 cells; n=5 per group for tissue and plasma exosome comparisons; three independent experiments for the Transwell assay), but no formal power/sample-size calculation is described, and the TCGA training cohort size used for WGCNA is not stated in the provided text. GroupsTHBS2+ vs THBS2- CAFs; tumor vs matched adjacent normal tissue/lung; LUAD patient plasma vs healthy donor plasma; PBS- vs THBS2 recombinant protein-treated cells; immunotherapy responders vs non-responders (by RECIST) Pairingmixed Randomization/blindingnot stated Dispersionunclear
Statistical tests used
Test Applied to n Assumptions
Pearson correlation (WGCNA module eigengene vs. clinical trait) Correlating co-expression module eigengenes with RFS and OS in the TCGA LUAD training cohort not stated
Gene Ontology / KEGG / Reactome pathway enrichment analysis Functional characterization of the RFS/OS-associated WGCNA module not stated
Gene set enrichment analysis (GSEA, hallmark gene sets) Differentially expressed genes between THBS2+ and THBS2- cancer-associated fibroblasts 332 CAFs (categorized into 10 clusters) not stated
Ligand-receptor interaction scoring (CellChat) Cell-cell interaction network among cell types in the tumor microenvironment 43,779 cells (scRNA-seq) not stated
Approaches that could also have been used
  • Module-trait relationships in WGCNA were assessed by correlating module eigengenes with RFS and OS using Pearson correlation.
    Could also: A Cox proportional-hazards model, or a log-rank test comparing eigengene-defined groups, could also be used — RFS and OS are time-to-event outcomes with censoring; methods designed for censored survival data explicitly incorporate follow-up time and censoring status, which a direct correlation with a continuous trait value does not.
  • Pathway enrichment (GO, KEGG, Reactome) and GSEA were run with default clusterProfiler/GSEABase settings.
    Could also: Reporting the specific multiple-testing correction applied within these tools (e.g., Benjamini-Hochberg FDR) and/or complementary gene-set testing methods such as fgsea or camera could also be used — Making the correction method explicit helps readers directly compare the false-discovery control across the several enrichment analyses performed.
  • Differential expression between THBS2+ and THBS2- CAFs was derived from single-cell data and used as GSEA input.
    Could also: A pseudobulk approach (aggregating counts per biological replicate before applying DESeq2 or edgeR) could also be used — Pseudobulk methods are designed to account for biological replicate-level variance in scRNA-seq comparisons, which can complement cell-level differential expression testing.
  • Functional assay results (e.g., Transwell migration) are reported as derived from three independent experiments, without the comparison test specified in the provided text.
    Could also: Reporting a specific paired or unpaired test (e.g., Student's t-test) or, for small n, a non-parametric alternative such as the Mann-Whitney U test, together with a dispersion measure (SD or 95% CI), could also be used — With n=3 replicates, explicitly naming the test and showing spread helps convey both statistical and practical significance.
  • Exosome/plasma comparisons (e.g., LUAD vs. NAT tissue, patients vs. healthy donors) used small, equal group sizes (n=5) with the comparative test not detailed in the provided text.
    Could also: A non-parametric test such as the Mann-Whitney U or Wilcoxon rank-sum test could also be considered as an alternative to a t-test — With small sample sizes, distributional (normality) assumptions are harder to verify, and non-parametric tests avoid relying on that assumption.
  • WGCNA network parameters (soft-thresholding power=12, cut height=0.25, minimum module size=30) were fixed based on approximate scale-free topology fit (R²≈0.9).
    Could also: A sensitivity analysis across a range of power/cut-height values, or a module-stability assessment (e.g., bootstrapping/resampling), could also be presented — Showing results are stable across nearby parameter choices can complement a single fixed-parameter network construction.
Software: R/WGCNA · R/clusterProfiler 4.0.5 · GSEABase 1.54.0 · R/Seurat 3.2.0 (R 3.6.3) · Scrublet 3.7.3 · Cell Ranger 4.0.0 · CellChat 1.1.2 · clustree (R package)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35547750

Paper: Yang H et al. Multi-scale integrative analyses identify THBS2+ cancer-associated fibroblasts as a key orchestrator promoting aggressiveness in early-stage lung adenocarcinoma. Theranostics 2022. PMID 35547750 · PMCID PMC9065207 · DOI 10.7150/thno.69590.

Pinned code: https://github.com/sqjin/CellChat (third-party scRNA-seq cell-cell communication tool) Pinned data: GEO GSE10072 (Affymetrix HG-U133A microarray, 107 samples; 58 LUAD tumor vs 49 non-tumor lung)

This is a large multi-scale paper (TCGA WGCNA, 9 external microarray validation cohorts, scRNA-seq of 2 own surgical cases, CellChat ligand-receptor analysis, proteomics/RPPA, IHC/mIHC, exosome assays, mouse xenografts). The vast majority of figures rest on the authors' OWN unreleased scRNA-seq, wet-lab, and clinical-cohort data. Per the brief we scope to pipeline-derived results reproducible from the pinned open data (GSE10072).

IN SCOPE (pipeline-derived, reproducible from GSE10072)

The paper uses GSE10072 explicitly as an external transcriptomic validation cohort for tumor-vs-normal comparison ("transcriptomic data of early-stage LUAD and matched normal lung but without survival data"). The reproducible, low-hanging pipeline output is a differential-expression (limma) analysis of LUAD tumor vs normal on GSE10072, testing the direction/significance of the paper's central markers:

  • R1 — THBS2 upregulation in tumor. THBS2 is the paper's titular gene; it is reported as a tumor/CAF-enriched, aggressiveness-promoting marker. Reproducible claim: THBS2 is significantly upregulated in LUAD tumor vs normal in GSE10072 (direction + adj.P).
  • R2 — THBS2+ CAF top-10 marker signature direction. Paper's top upregulated genes in THBS2+ CAFs (Fig. of CAF markers): THBS2, COL3A1, BGN, COL5A2, COL1A1, COL6A3, FAP, CTHRC1, SULF1. Reproducible claim: these stromal/CAF markers are concordantly upregulated in tumor vs normal microarray (direction + adj.P per gene). This is a coherence check of the CAF signature against bulk tumor-vs-normal expression.

Pipeline named: limma on normalized GSE10072 series-matrix expression (microarray canonical path; method-card 100%-clean for arrays). Grouping from GEO sample titles ("Lung Tumor_" vs "Normal Lung_").

OUT OF SCOPE (not pipeline-derivable from pinned open data)

  • CellChat ligand–receptor analysis — requires the authors' OWN scRNA-seq (43,779 cells, 2 surgical cases); no GEO/accession provided for it. CellChat cannot be run on a bulk microarray. Not reproducible from GSE10072 → data unavailable.
  • scRNA-seq / Seurat clustering, 7 CAF subclusters, 332→10 clusters — own unreleased data.
  • WGCNA on TCGA LUAD, Cox survival / KM, hazard ratios — separate cohorts; GSE10072 has no survival data (paper states this explicitly). Not attempted here (different accession).
  • QuanTIseq immune infiltration, GSEA, proteomics/RPPA, IHC/mIHC, exosome assays, xenografts — wet-lab or external-cohort / different-accession; out of scope.

Verdict frame

A clean limma reproduction confirming THBS2 + CAF-marker upregulation in GSE10072 tumor-vs-normal is an honest partial reproduction: it independently corroborates the paper's central tumor-enrichment claim for THBS2 using the one open dataset pinned to this room, while the headline CAF/CellChat results depend on unreleased single-cell data and cannot be reproduced.

Figures / tables: figure
R1
Reported
THBS2 upregulated in LUAD tumor vs normal (GSE10072 used as tumor-vs-normal cohort; no exact value published)
Reproduced
logFC=+2.086, t=13.01, adj.P=4.67e-22 (Tumor n=58 vs Normal n=49)
within tolerance
R2
Reported
THBS2+ CAF top markers (THBS2,COL3A1,BGN,COL5A2,COL1A1,COL6A3,FAP,CTHRC1,SULF1) upregulated
Reproduced
7/8 testable markers significantly up in tumor (adj.P<0.05); BGN +0.20 n.s. (adj.P=0.088); CTHRC1 absent from HG-U133A
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 68/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

225 k
tokens (I/O) · 13.5 M incl. cache
58 min
runtime · 0 CPU-h
0.4 GB
peak RAM
1
HPC jobs
hummel
machine