Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Transcriptional firing represses bactericidal activity in cystic fibrosis airway neutrophils.

Cell Rep Med · 2021
L1 35/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
35/100
Reproducibility score
2.2 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 3% of all assessed papers rank 1140 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH; DIFFERENT QUANTITATIVELY. Full HISAT2 2.2.1 (GRCh38 Ensembl-104) -> samtools -> featureCounts 2.0.6 (-p --countReadPairs, unstranded) -> DESeq2 1.42.0 (padj<0.01) pipeline ran to completion on «our HPC» compute nodes (env built in-«job»; align array 2243066 md5-verified 14 PE runs, mean 89.9% alignment; count+DE 2243067; sensitivity «job»). QUALITATIVE biology reproduces cleanly: C3 PCA separates Blood/LTB4/CFASN on PC1 (90% var); C4 granule effectors ELANE/MPO/ARG1 sit at the detection floor in CFASN while MMP9 stays orders of magnitude higher, declining Blood->LTB4->CFASN. QUANTITATIVE DE gene counts (C1/C2/C5/C6) are consistently 1.4x-4.3x HIGHER than reported and the CFASN-vs-LTB4 up/down asymmetry reverses (paper up>down 3417>1731; ours down>up 6442<7361). A threshold sensitivity sweep shows no single plausible filter recovers 3417/1731 exactly, but an undocumented expression prefilter qualitatively explains both effects: baseMean>50 restores the paper's up>down direction (5157>4612) and shrinks counts -- consistent with an expressed-gene prefilter and/or different annotation universe not stated in the available Methods. No processed count matrix was deposited, so exact gene-set sizes are unrecoverable. Verdict PARTIAL: qualitative claims reproduced, quantitative DE-count claims a documented mismatch flagged for human review (possible reporting/method discrepancy, NOT a confirmed fabrication). NOT attempted: wet-lab assays, proprietary Qlucore/STEM renders, GSEA exact q-values, in vivo microarray (out of scope per scope.md).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ a48ee937daba
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The authors test whether the adaptive functional changes seen in CF airway neutrophils (the 'GRIM' fate, including reduced bacterial clearance) depend on de novo transcription and translation occurring after neutrophil recruitment into the CF airway microenvironment.

Core claims
  • Neutrophils recruited to CF airways undergo rapid, broad transcriptional firing that upregulates anabolic genes and downregulates antimicrobial genes finding
  • Newly transcribed RNAs in CF airway neutrophils are mirrored by appearance of corresponding proteins, confirming active translation finding
  • The transcriptional blocker α-amanitin restores expression of key antimicrobial genes and increases bactericidal capacity of CF airway neutrophils in vitro and in short-term sputum cultures ex vivo finding
  • Acquisition of the GRIM fate requires a 'two-hit' mechanism: transepithelial migration plus subsequent exposure to the CF airway (CFASN) milieu mechanism
  • Accumulation of granule effector proteins (NE, MPO, arginase-I) in CF airway fluid and neutrophils likely results from uptake of extracellular protein rather than de novo transcription, since their transcripts are undetectable mechanism
  • CF airway neutrophils show markedly increased total RNA content compared to matched blood neutrophils, both in vivo and in the in vitro transmigration model finding
  • GRIM fate transcriptional and functional adaptation develops progressively over a 1-6 h time course after transmigration finding
  • Transcriptional blockade by α-amanitin reduces primary granule release (CD63, NE) and extracellular vesicle secretion in CFASN-transmigrated neutrophils finding
Experimental setups
Assay System Perturbation Readout Platform
flow cytometry (surface CD63/CD16, total RNA dye) blood and sputum neutrophils from CF patients (N=7) none (in vivo disease state) CD63, CD16 surface expression, total RNA content
in vitro transmigration model across airway epithelium healthy donor and CF blood neutrophils transmigrated toward CFASN or LTB4 CFASN or LTB4 chemoattractant exposure CD63/CD16 expression, RNA content
RNA content quantification transmigrated neutrophils (blood, CFASN, LTB4) CFASN vs LTB4 transmigration total RNA quantity bioanalyzer
bulk RNA sequencing / PCA / GSEA blood, CFASN-transmigrated, and LTB4-transmigrated neutrophils (N=5 experiments; N=3 subjects for kinetics) CFASN vs LTB4 transmigration; time course 1,2,4,6 h differential gene expression, pathway enrichment (hallmark, reactome, KEGG)
multiplex qPCR blood and airway (sputum) neutrophils sorted from CF patients in vivo none (in vivo) transcript levels of NE, MPO, arginase-I, MMP9 Fluidigm
quantitative proteomics blood, CFASN-transmigrated, and LTB4-transmigrated neutrophils CFASN vs LTB4 transmigration protein abundance/unique proteins per condition, pathway enrichment mass spectrometry
transcriptional blockade experiment (RNA-seq, flow cytometry) blood neutrophils transmigrated 2 h into CFASN, then treated 8 h with α-amanitin α-amanitin (RNA polymerase II/III inhibitor) gene expression changes, CD63, extracellular NE, EV release
bactericidal co-culture assay CFASN-transmigrated neutrophils in vitro and primary CF sputum neutrophils ex vivo α-amanitin treatment vs untreated, plus Pseudomonas aeruginosa challenge bacterial count/killing after 30 min co-culture
Key results
  • Airway neutrophils showed a median 3.5-fold increase in total RNA content compared to matched blood neutrophils in vivo 3.5-fold
  • CFASN-transmigrated neutrophils showed >10-fold increase in RNA content vs blood; LTB4-transmigrated showed only 2-fold increase >10-fold (CFASN); 2-fold (LTB4)
  • 639 genes were upregulated and 560 downregulated concordantly both in vivo and in vitro out of 2,010 in vivo differential genes
  • Transcripts for effector proteins NE, MPO, and arginase-I were below detection limit in both in vitro (CFASN/LTB4) and in vivo airway GRIM neutrophils
  • Incubation in CFASN alone or transmigration toward COPD ASN failed to induce the full transcriptional burst seen with CFASN transmigration
  • α-amanitin treatment upregulated 20% of genes that were significantly downregulated at 6 h post-transmigration, mostly secretory vesicle and secondary/ficolin-rich granule proteins 20% of downregulated genes recovered
  • α-amanitin reduced surface CD63, extracellular NE, and extracellular vesicle release in CFASN-transmigrated neutrophils
  • α-amanitin increased killing of Pseudomonas aeruginosa by CFASN-transmigrated neutrophils in vitro and by primary CF sputum neutrophils ex vivo
Key statistics
  • fold_change median 3.5-fold (increase in total RNA content, airway vs blood neutrophils in vivo)
  • fold_change >10-fold (increase in RNA content, CFASN-transmigrated vs blood neutrophils in vitro)
  • fold_change 2-fold (increase in RNA content, LTB4-transmigrated vs blood neutrophils in vitro)
  • count 2,010 genes differential in vivo; 639 upregulated and 560 downregulated concordant in vitro (overlap of in vivo and in vitro differential gene sets)
  • count 1,407 upregulated and 1,221 downregulated shared between CFASN and LTB4 vs blood; 3,602 uniquely upregulated and 3,148 uniquely downregulated in CFASN (genes changed at least 4-fold (|log2|>2) vs blood)
  • count 3,417 upregulated and 1,731 downregulated genes (CFASN-transmigrated vs LTB4-transmigrated neutrophils)
  • count 769 downregulated and 1,735 upregulated genes (log2 FC >2 or <-2, p<0.01) (significantly different genes, CFASN vs LTB4 among blood-differential gene subsets)
  • pvalue ∗p<0.05, ∗∗p<0.01, ∗∗∗p<0.001 (Wilcoxon matched-pairs signed rank test); GSEA FDR q<5% (statistical significance thresholds used throughout figures)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a paired within-subject design comparing matched blood and CF airway (sputum) neutrophils in vivo (N=7 adult patients) and an in vitro transmigration model (N=3–5 independent experiments). Primary between-condition comparisons of flow cytometry and functional outcomes were made with the Wilcoxon matched-pairs signed rank test, with results expressed as median and interquartile range. Transcriptomic changes were characterized by combined fold-change and nominal p-value thresholds (|log2FC| > 2, p < 0.01), with pathway-level conclusions drawn from GSEA and GO term enrichment using an FDR q < 5% threshold.

Replicationmixed Sample sizeN=7 adult CF patients (in vivo); N=3–5 independent experiments with healthy donor blood neutrophils (in vitro); no formal power calculation mentioned GroupsBlood vs. CF airway (sputum) neutrophils; blood vs. CFASN-transmigrated vs. LTB4-transmigrated neutrophils; α-amanitin-treated vs. untreated CFASN neutrophils; kinetic time points (1, 2, 4, 6 h) vs. blood Pairingpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFalse discovery rate (FDR) q < 5%
Statistical tests used
Test Applied to n Assumptions
Wilcoxon matched-pairs signed rank test Matched blood vs. airway neutrophil comparisons for CD63, CD16, RNA content, NE release, extracellular vesicles, and P. aeruginosa killing (Figures 1B, 3A–B, 4A–F) N=7 patients (in vivo); N=3–5 independent experiments (in vitro) not stated
Gene Set Enrichment Analysis (GSEA) with FDR q < 5% Pathway enrichment of transcriptomic signatures using Hallmark, Reactome, and KEGG gene sets (Figures 2C, 2D, 3D, S5C) N=5 independent experiments (CFASN vs. LTB4); N=3 subjects (kinetic time course) not stated
Gene Ontology (GO) term enrichment analysis with FDR q < 5% Pathway analysis of genes commonly regulated in vivo and in vitro (Figure 1F) null not stated
RNA-seq differential expression using |log2FC| > 2 and p < 0.01 thresholds Gene-level identification of up- and downregulated genes between blood, LTB4-, and CFASN-transmigrated neutrophils and across time points (Figures 2D, 3C, S1B–C) N=5 independent experiments (CFASN vs. LTB4); N=3 subjects (time course) not stated
Principal component analysis (PCA) Visualization of transcriptomic separation across blood, LTB4-, and CFASN-transmigrated neutrophil conditions (Figure 2A) N=5 independent experiments na
Approaches that could also have been used
  • Individual RNA-seq gene-level comparisons used a combined threshold of |log2FC| > 2 and nominal p < 0.01 to define differentially expressed genes
    Could also: A genome-wide FDR-adjusted p-value (e.g., Benjamini-Hochberg adjusted q < 0.05) applied across all tested genes could also be used, as implemented in standard pipelines such as DESeq2 or edgeR — Applying FDR adjustment at the gene level controls the expected proportion of false discoveries across the full tested set and complements the fold-change filter; it is also the most commonly expected reporting standard for RNA-seq studies
  • Flow cytometry and functional assay outcomes were reported as median and interquartile range with asterisk-based significance thresholds
    Could also: Reporting exact p-values alongside 95% confidence intervals on the median difference or effect size (e.g., Hodges-Lehmann estimator for the Wilcoxon test) could also be used — Exact p-values allow readers to calibrate evidence strength for their own purposes and facilitate future meta-analyses; confidence intervals on effect estimates communicate both magnitude and precision, which is particularly informative at small n
  • Paired Wilcoxon signed rank tests were used for all within-subject comparisons across experiments of N=3–5
    Could also: A linear mixed-effects model with subject as a random effect could also account for the paired structure while enabling covariate adjustment and providing effect estimates with standard errors — Mixed-effects models accommodate unbalanced designs, can incorporate additional predictors (e.g., patient age, lung function), and yield interpretable effect estimates alongside confidence intervals rather than only a hypothesis test result
  • Pathway enrichment was conducted with GSEA and GO term over-representation analysis
    Could also: Over-representation analysis (ORA) using a hypergeometric or Fisher's exact test applied to a discrete gene list could also be used as an independent approach — GSEA uses the full ranked gene list while ORA operates on a defined threshold-based set; applying both methods can confirm robustness of pathway findings because the two approaches make different assumptions about the signal structure in the data
  • No formal sample size or power calculation was described for the in vitro experiments
    Could also: A prospective power calculation anchored to a primary outcome (e.g., expected difference in P. aeruginosa killing) or a post-hoc sensitivity analysis could also be reported — Reporting power or sensitivity analyses contextualizes the study's ability to detect differences of a given magnitude and helps readers interpret non-significant comparisons
  • The specific RNA-seq differential expression pipeline (normalization method, statistical model, software) is not stated in the text
    Could also: Explicit reporting of the complete pipeline — e.g., DESeq2 with variance-stabilizing normalization, edgeR with TMM normalization, or limma-voom — is standard practice and could be included in Methods — Specifying the pipeline enables reproducibility and allows readers to assess how modeling choices (e.g., handling of low-count genes, dispersion estimation with small n) may influence the reported gene lists
Software: Fluidigm (multiplex qPCR platform) · Bioanalyzer (RNA quantification)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33948572

Paper: Margaroli et al. 2021, Cell Rep Med — "Transcriptional firing represses bactericidal activity in cystic fibrosis airway neutrophils." DOI 10.1016/j.xcrm.2021.100239.

Data: GEO GSE167069 / SRA SRP307042 / BioProject PRJNA702819. Bulk RNA-seq, paired-end 100 bp, Illumina HiSeq 2500, Homo sapiens. 14 samples:

  • Blood neutrophils (healthy donors): 4 — HD1,2,5,6
  • 10 h transmigrated + CFASN: 5 — HD1,2,3,5,6
  • 10 h transmigrated + LTB4: 5 — HD1,2,3,5,6

Pipeline described in Methods (in scope):

  • Aligner: HISAT2 → GRCh38 (paper writes "GRCh39" = typo; no such build).
  • Counting: featureCounts (Rsubread).
  • DE: DESeq2, significance = adjusted p (BH) < 0.01.
  • Library: TruSeq RNA Single Indexes Set B; ~20 M reads/sample target.

Repo cited: https://github.com/DaehwanKimLab/hisat2 — this is the HISAT2 tool itself (third-party aligner), NOT author analysis code. Per brief P16, applying the named third-party tool to the paper's own data is a valid reproduction.

In-scope reproducible results (pipeline-derived)

id reported location pipeline
C1 3,417 genes upregulated in CFASN vs LTB4 (padj<0.01) Results / DE HISAT2→featureCounts→DESeq2
C2 1,731 genes downregulated in CFASN vs LTB4 (padj<0.01) Results / DE same
C3 PCA separates blood / LTB4 / CFASN (N=5) Fig (PCA, Qlucore) counts→PCA (DESeq2 vst)
C4 Granule effectors (NE/MPO/ARG1) transcripts below detection; MMP9 detectable Results counts inspection

Out of scope (wet-lab / manual / external tools — not attempted)

  • Flow cytometry, total RNA content fold-changes (Fig 1B), bacterial killing assays.
  • PCA exact rendering (Qlucore Omix Explorer v3.3 — proprietary; reproduce equivalent PCA in R).
  • STEM kinetics (N=3) — STEM proprietary-ish; optional stretch.
  • GSEA / GO enrichment exact q-values (stretch; depends on ranked list).
  • In vivo microarray (N=7 CF blood/sputum) — separate dataset, not GSE167069.

Primary target (quick minimum ~80%)

C1 + C2: regenerate the CFASN-vs-LTB4 DESeq2 DE gene counts at padj<0.01. Then push to C3 (PCA) and C4 (effector transcript detection).

C1
Reported
3,417 genes upregulated in CFASN vs LTB4 (padj<0.01)
Reproduced
6442 (~group) / 6561 (~donor+group)
did not match
C2
Reported
1,731 genes downregulated in CFASN vs LTB4 (padj<0.01)
Reproduced
7361 (~group) / 7438 (~donor+group)
did not match
C3
Reported
PCA shows distinct profile for each condition (Blood/LTB4/CFASN)
Reproduced
DESeq2 vst PCA: PC1=90% var cleanly separates the 3 groups (Blood -76..-106, LTB4 -17..+25, CFASN +66..+74)
within tolerance
C4
Reported
NE(ELANE)/MPO/ARG1 transcripts below detection; MMP9 detectable
Reproduced
CFASN mean norm counts: ELANE=1.9, MPO=3.5, ARG1=20.6 (near floor) vs MMP9=546 (detectable); effectors decline Blood->LTB4->CFASN
within tolerance
C5
Reported
1,407 up & 1,221 down shared in both LTB4 & CFASN vs blood
Reproduced
3065 up & 2588 down shared
did not match
C6
Reported
3,602 up & 3,148 down uniquely in CFASN vs blood
Reproduced
4905 up & 4760 down unique (direction reproduced, magnitude high)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 35/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

201.7 k
tokens (I/O) · 14.2 M incl. cache
106 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.