Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

In vivo induction of activin A-producing alveolar macrophages supports the progression of lung cell carcinoma.

Nat Commun · 2023
L1 96/100 PQI 87
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
How its reproducibility compares
96/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 91% of all assessed papers rank 92 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1 reproduction of the IN-SCOPE part. GSE193913 deposits bulk RNA-seq FPKM of 3 sorted alveolar-macrophage populations (R1 control AMs=Vehicle_mid, R2 tumor AMs=Tumor_mid, R3 TAMs=Tumor_high); these FPKM tables ARE the paper's pipeline output (TopHat2->FPKM, mm10). From them, all Fig.2b/2c conclusions reproduce cleanly: the central claim that activin-A gene Inhba is specifically up only in R2 (32 FPKM vs 4/10 in R1/R3); every AM marker high in R1&R2; every TAM marker high in R3 (these markers also self-validate the sample->population mapping); and PCA separates the three groups. NOT attempted (out of scope / not feasible, not down-ranking the paper): the scRNA-seq Fig.3 analysis (authors' Scanpy/scVelo repo) because its input loom/CellRanger data is NOT deposited under GSE193913; raw-read re-alignment (no fastq under this accession; FPKM is the canonical deposited output); all wet-lab assays (qPCR/ELISA/follistatin/KO mice, not pipeline-derived). Grades provisional - paper shows heatmap/PCA with no in-text numbers, so comparison is qualitative pattern-match (every named gene matches direction). No fabrication signal.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 96
    assessed: 2026-06-15 ⛓ ffaf473d05a9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Do residential alveolar macrophages (AMs), which are abundant in lung cancer tissue, play an active pathological role in supporting lung cancer progression, and if so, through what secreted molecular mechanism?

Core claims
  • Alveolar macrophages support lung cancer cell proliferation and contribute to unfavourable outcomes. finding
  • INHBA is specifically upregulated in AMs under tumor-bearing conditions, leading to secretion of activin A (the INHBA homodimer). mechanism
  • Activin A secreted by tumor-bearing AMs promotes lung cancer cell proliferation; follistatin (an activin A antagonist) inhibits it. mechanism
  • scRNA-seq identifies a tumor-specific subset of INHBA-high AMs (clusters 1, 4, 8) distinct from INHBA-expressing AMs in normal lungs. finding
  • Postnatal deletion/depletion of AMs or INHBA/activin A limits tumor growth in experimental models. finding
  • INHBA/activin A expression in AMs is induced via a MyD88-JNK dependent pathway. mechanism
  • An orthotopic murine lung cancer model (LLC implanted into left lung) enables study of AM-cancer interactions in their native microenvironment. method
  • Tumor-derived DAMPs convert MARCO-high AMs to MARCO-low cells with upregulated Inhba in vitro. finding
Experimental setups
Assay System Perturbation Readout Platform
Immunohistochemistry (CD163, TTF-1) Human lung cancer patient tissue (normal vs cancer areas) none CD163+ macrophage accumulation/proportion in tissue
Cell count / proliferation assay LLC lung carcinoma cells cultured with AM cell line (AMJ2-C11) supernatant AM-conditioned media; Inhba shRNA knockdown cancer cell number and doubling time
Flow cytometry Mouse lung cells (WT C57BL/6, control- vs clodronate-liposome, tumor-bearing) clodronate liposome (CDL) macrophage depletion CD45+ autofluorescence+ AM and TAM populations (F4/80, Siglec-F, CD11b, CD11c)
Orthotopic lung tumor model / tumor volume measurement C57BL/6 mice with LLC-tdTomato (also KLN205 in DBA/2; Csf2 KO mice) CDL depletion, Csf2 knockout, follistatin treatment tumor volume, metastasis to contralateral lung
Bulk RNA-sequencing Sorted lung AMs (R1 control, R2 tumor-bearing) and TAMs (R3) from orthotopic model tumor-bearing vs control transcriptome / differential gene expression (PCA, marker heatmaps)
RT-PCR / qPCR Sorted mouse AMs, CD45+ cells, tumor cells tumor-bearing vs control; DAMP coculture Inhba (and Inhbb, Inha) expression
ELISA AMs sorted from control or tumor-bearing mice tumor-bearing vs control activin A concentration
WST-1 cell proliferation assay LLC lung cancer cells recombinant murine activin A dose treatment cell viability/proliferation over days 0,1,4
Single-cell RNA-sequencing (UMAP, RNA-velocity) Sorted AMs from control and orthotopic tumor-bearing mice (13,413 cells) tumor-bearing vs control cluster identity, Inhba expression, differentiation dynamics
Key results
  • CD163+ macrophage population significantly increased in cancer tissue vs normal lung
  • LLC cell number significantly increased when cultured with AM cell supernatant
  • Tumor volumes significantly smaller in CDL-treated (AM-depleted) mice than control-liposome mice
  • Inhba specifically upregulated only in R2 (tumor-bearing AMs), among top 15 upregulated genes from R1 to R2
  • Activin A preferentially produced/elevated in tumor-bearing AMs vs control by ELISA
  • Recombinant activin A significantly increased LLC proliferation; Inhba knockdown attenuated AM supernatant's pro-proliferative effect
  • Follistatin treatment significantly reduced tumor volume in orthotopic model
  • Inhba-high AMs present in tumor-specific clusters 1,4,8 but only cluster 7 in control; MARCO-low DAMP-cocultured AMs show significant Inhba upregulation
Key statistics
  • count n = 10 patients (CD163+ macrophage proportion in normal vs cancer lung areas)
  • count n = 9 (Ctrl), n = 10 (CDL) (tumor volume comparison Ctrl vs CDL-treated mice)
  • count 13,413 (AM single-cell transcriptomes analyzed, clustered into 14 subgroups)
  • count n = 6 per group (activin A ELISA in AMs from control vs tumor-bearing mice)
  • count n = 3 (control), n = 4 (follistatin) (tumor volume in follistatin vs PBS-treated tumor-bearing mice)
  • count n = 3 mice for R1, R2, R3 (RNA-Seq populations for PCA)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined human histopathology, an orthotopic murine lung cancer model, and functional in vitro/in vivo experiments with bulk RNA-seq and single-cell RNA-seq. Two-group comparisons were assessed with two-tailed t-tests (paired or unpaired), and multi-group comparisons with one-way ANOVA followed by Bonferroni's post hoc test. Quantitative results were reported as mean ± s.e.m. with small replicate numbers, and high-dimensional data were analyzed with PCA, UMAP clustering, and RNA-velocity.

Replicationmixed Sample sizePer-panel sample sizes given as exact n (e.g., n = 3–10 wells/mice/patients); no formal power/sample-size calculation described GroupsControl vs tumor-bearing AMs; Ctrl vs CDL/Csf2-KO/follistatin; AM subpopulations Pairingmixed Randomization/blindingnot stated DispersionSEM Effect sizesno Confidence intervalsno Multiplicity correctionBonferroni post hoc (within one-way ANOVA panels)
Statistical tests used
Test Applied to n Assumptions
paired two-tailed Student's t-test Fig 1b, proportion of CD163+ macrophages in normal vs cancer areas n = 10 patients not stated
unpaired two-tailed Student's t-test Fig 1c (LLC cell count ± AM supernatant, n = 3/group), Fig 1f (tumor volume Ctrl vs CDL, n = 9 vs 10), Fig 2e (activin A ELISA, n = 6/group), Fig 2i (tumor volume control vs follistatin, n = 3 vs 4) varies by panel as listed not stated
one-way ANOVA with Bonferroni's post hoc test Fig 2d (Inhba RT-PCR across cell populations, n = 3/group), Fig 2g (LLC count with shRNA AM conditioned media, n = 3/group), Fig 3h (Inhba RT-PCR ± DAMPs, n = 3/group) n = 3 per group not stated
Principal component analysis (PCA) Fig 2b, RNA-seq of R1/R2/R3 populations n = 3 mice per population na
UMAP clustering (hierarchical, 14 subgroups) Fig 3a, scRNA-seq of 13,413 AM transcriptomes 13,413 cells na
RNA-velocity (scVelo) and Gaussian mixture model binarization of Inhba expression Fig 3d–f, differentiation dynamics and Inhba+ classification na
Approaches that could also have been used
  • Dispersion was summarized as mean ± s.e.m. across panels, often with small n (e.g., n = 3).
    Could also: Reporting the standard deviation or a 95% confidence interval, and overlaying individual data points (which several figures already do). — SD or a CI conveys the spread or precision of the estimate directly and is often preferred for small samples, complementing the SEM.
  • Several related two-group comparisons were each evaluated with separate two-tailed t-tests across multiple panels.
    Could also: A single ANOVA (or mixed-effects model) encompassing the related groups with a post-hoc correction, as was done in the ANOVA panels. — Bundling related comparisons into one model controls the family-wise error rate across the full set and yields a unified estimate of variance.
  • Group comparisons with small n (e.g., n = 3) used parametric t-tests/ANOVA.
    Could also: Nonparametric tests such as Mann-Whitney U or Kruskal-Wallis, or explicit checks/statements of normality and variance assumptions. — Nonparametric approaches relax distributional assumptions that are hard to verify at small n, while stating assumption checks documents that the parametric choice was appropriate.
  • Differential expression and marker selection for bulk RNA-seq and scRNA-seq were presented via PCA, heatmaps, and ranked gene lists.
    Could also: Reporting model-based differential-expression statistics with multiplicity control (e.g., DESeq2/edgeR or limma with Benjamini-Hochberg FDR). — An explicit FDR-controlled framework quantifies significance for the many-gene comparisons and standardizes thresholds across the transcriptomic analyses.
  • Tumor-volume and proliferation outcomes were compared at endpoints between groups.
    Could also: Reporting effect sizes (e.g., mean differences with CIs) alongside the p-values. — Effect sizes with intervals communicate the magnitude and precision of the biological difference, not only whether it reached significance.
  • Randomization and blinding of group allocation/assessment were not described in the available text.
    Could also: Explicitly stating randomization and blinded outcome assessment procedures. — Documenting these design elements helps readers gauge how allocation and measurement bias were addressed in the in vivo experiments.
Software: scVelo (RNA-velocity) · UMAP

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
34
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE154826 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE193913 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36650150

Paper: Taniguchi et al., In vivo induction of activin A-producing alveolar macrophages supports the progression of lung cell carcinoma. Nat Commun 2023. DOI 10.1038/s41467-022-35701-8.

Two distinct computational datasets in this paper

Part Data Pipeline Deposited? In scope?
A. Bulk RNA-seq of sorted AMs (Fig.2b,2c) GSE193913 — 9 FPKM tables (3×R1, 3×R2, 3×R3) TopHat2.1.1+Bowtie2+SAMtools→FPKM (mm10) YES — FPKM tables public (1.2 MB) YES — REPRODUCED
B. scRNA-seq of AMs (Fig.3) "13,413 AM transcriptomes", CellRanger+Velocyto loom authors' Scanpy/scVelo notebooks (repo) scRNA raw/loom NOT in GSE193913; not resolvable to a public accession NO — data not deposited under this accession
Wet-lab (qPCR/ELISA/proliferation/KO mice) Fig.2d-h,3g-h,4 manual/bench n/a NO — out of scope (not pipeline-derived)

In-scope reproduction (Part A — the low-hanging, fully-specified pipeline output)

The deposited FPKM tables ARE the pipeline output (paper computed FPKM with TopHat→Cufflinks on mm10; raw fastq would be in SRA, large). We reproduce the downstream Fig.2 analysis that turns those FPKM values into the paper's conclusions:

  • C1 (Fig.2c / Results): Inhba (activin A subunit) is specifically upregulated only in R2 (AMs in tumor-bearing condition), not R1 (control AMs) nor R3 (TAMs). This is the paper's central molecular claim.
  • C2 (Fig.2c): AM lineage markers Pparg, Mrc1, Marco, Siglecf, Siglec1 are high in R1 & R2 (confirming both are AMs).
  • C3 (Fig.2c): TAM markers Ccr2, Cx3cr1, Tgfb3, Ly6c2 are high in R3.
  • C4 (Fig.2b): PCA cleanly separates the three populations R1/R2/R3.

Sample → population mapping (self-validated by C2/C3 markers)

  • R1 (AM control) = Vehicle_mid_{1,2,3} (GSM5823545-47)
  • R2 (AM tumor) = Tumor_mid_{1,2,3} (GSM5823548-50)
  • R3 (TAM tumor) = Tumor_high_{1,2,3} (GSM5823551-53)

Explicitly NOT attempted (80/20)

  • Re-running TopHat/Cufflinks from raw reads (no fastq under this accession; FPKM is the deposited, canonical pipeline output — re-alignment would not change the Fig.2 conclusions and is the hard, low-value 20%).
  • The scRNA-seq / scVelo analysis (Fig.3): its input loom/CellRanger data is not deposited under GSE193913, so it is not reproducible from public data here.
  • All wet-lab assays (qPCR, ELISA, follistatin, KO mice) — not pipeline-derived.
Figures / tables: Fig.2cFig.2b
C1
Reported
Inhba specifically upregulated only in R2 (AMs in tumor-bearing), not R1 (control) or R3 (TAMs)
Reproduced
Inhba FPKM means R1=3.97, R2=32.22, R3=9.75 -> uniquely highest in R2 (8.1x control, 3.3x TAM)
exact
C2
Reported
AM markers Pparg/Mrc1/Marco/Siglecf/Siglec1 high in R1 & R2
Reproduced
all 5 high in R1&R2, low in R3 (e.g. Marco 220/162/6.6; Siglecf 201/172/0.65)
exact
C3
Reported
TAM markers Ccr2/Cx3cr1/Tgfb3/Ly6c2 high in R3
Reproduced
all 4 specifically high in R3 (Cx3cr1 0.4/0.6/69; Ccr2 1.6/4.2/38; Ly6c2 5.2/7.8/56)
exact
C4
Reported
PCA clearly distinguishes the three populations (Fig.2b)
Reproduced
PCA separates all 3 groups; PC1=70.8% var; centroid dist R1-R2=38.6, R1-R3=118.7, R2-R3=104.2
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 96/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

The in-scope bulk RNA-seq analysis (Fig.2b/2c) reproduces 1:1 from the deposited GSE193913 FPKM tables, which are the authors' own pipeline output: Inhba is uniquely high in tumor AMs (R2=32.22 vs R1=3.97, R3=9.75), all AM and TAM markers follow the reported direction, and PCA separates the three populations. The only caveat is on our/data side, not the authors': the paper shows heatmap/PCA with no printed numbers, so comparison is qualitative pattern-match (q2 yellow), and Fig.3 scRNA-seq + wet-lab assays were not attempted because their inputs are not deposited under this accession. No deviation, no severity, no fabrication concern — the central conclusion holds fully.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

97.4 k
tokens (I/O) · 5.9 M incl. cache
12 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine