Expression of neurofibromin 1 in colorectal cancer and cetuximab resistance.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for a faithful 1:1 of the RNA-seq transcript claims. The paper is mostly wet-lab; its two PIPELINE-DERIVED transcript claims come from GSE183984, whose processed FPKM matrix the authors ship on GEO. Using that matrix directly (stdlib Python on «our HPC», no aligner re-run needed), BOTH reproduced in direction and significance: (C1) NF1 is ~0.80x lower post- vs pre-cetuximab (p=0.0012, matching 'slightly lower'); (C2) NF1 shows no significant association with response (post non-PD vs PD p=0.33, matching 'not associated'). NOT attempted: all wet-lab assays (non-pipeline); the NF1 mutation-frequency claim (1.8%) because its vcf2maf input VCFs are private (data_restricted); the full FASTQ->STAR->RSEM->DESeq2 re-run (unnecessary - authors ship the pipeline's FPKM output, the correct input for a between-sample transcript comparison, and re-aligning adds no evidentiary value, the hard ~20%). Caveat: paper used n=111 post-QC RAS/BRAF-WT; GEO ships 113 and we used all 113; direction robust. Grades provisional - claims are qualitative so agreement is directional, not byte-exact; a human reviewer signs off in AUDIT.md.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 73assessed: 2026-06-15 ⛓ 2e7c9803f248
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study investigates whether neurofibromin 1 (NF1) expression level is associated with sensitivity/resistance to the anti-EGFR antibody cetuximab in colorectal cancer, testing whether modifying NF1 expression alters cetuximab response and whether NF1 could serve as a biomarker in RAS/BRAF V600 wild-type tumors.
- ★ NF1 is highly expressed in cetuximab-sensitive CRC cell lines and minimally expressed in cetuximab-resistant ones finding
- ★ siRNA knockdown of NF1 in sensitive cell lines enhances MEK/ERK phosphorylation and reduces cetuximab-induced apoptosis mechanism
- ★ NF1 overexpression (NF1-GRD) renders resistant cell lines KM12C and SW480 more susceptible to cetuximab-induced apoptosis finding
- ★ Modification of NF1 expression can affect cetuximab sensitivity in CRC cell lines finding
- ★ Pre-treatment NF1 expression levels in patient tumors were not associated with cetuximab response, limiting NF1 as a clinical biomarker finding
- Post-cetuximab-treatment tumor samples showed slightly lower NF1 transcript levels than pre-treatment samples finding
- Inactivating NF1 mutations are rare (1.8%) in CRC patients and generally not associated with NF1 protein expression finding
- NF1 negatively regulates RAS-MAPK signaling via GTPase-activating activity mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Western blot | CRC cell lines NCI-H508, Caco-2, KM12C, SW480 | none/baseline and cetuximab; NF1 siRNA knockdown; NF1-GRD overexpression | NF1, p-MEK1/2, MEK1/2, p-ERK1/2, ERK1/2, cleaved caspase-3, cleaved PARP protein levels | LuminoGraph II (ATTO); Cell Signaling/Abcam antibodies; ImageJ densitometry |
| Cell growth assay (crystal violet) | CRC cell lines NCI-H508, Caco-2, KM12C, SW480 | cetuximab 0/50/100/200 µg/ml | relative proliferation (OD 595 nm) | Sunrise microplate reader (Tecan) |
| Colony formation assay | CRC cell lines | cetuximab 0/50/100/200 µg/ml | colony number (>200 µm) | Oxford Optronix GelCount |
| siRNA knockdown / plasmid transfection | cetuximab-sensitive (knockdown) and cetuximab-resistant KM12C/SW480 (overexpression) CRC cell lines | NF1-siRNA / NF1-GRD plasmid + cetuximab 100 µg/ml | NF1 expression, apoptosis markers, cell number | Lipofectamine 3000 |
| DAPI nuclear staining | siRNA- and plasmid-transfected CRC cell lines | NF1 knockdown / overexpression + cetuximab | apoptotic bodies / nuclear condensation count | EVOS FL Auto fluorescence microscope; ImageJ |
| RT-qPCR | CRC cell lines | siRNA/plasmid transfection | NF1 transcript level normalized to GAPDH | CFX Connect Real-Time PCR (Bio-Rad); EvaGreen qPCR |
| Bulk RNA sequencing | FFPE tumor samples from 92 mCRC patients (113 samples), RAS/BRAF V600 wild-type | cetuximab treatment (pre- vs post-treatment) | NF1 transcript (FPKM), differential expression, GSEA pathway enrichment | Illumina HiSeq 2500; TruSeq RNA Access Library Prep; STAR/RSEM/DESeq2 |
| Targeted DNA sequencing (NGS) / IHC | clinical CRC patient genomic database (1,449 patients) | none (diagnostic profiling) | NF1 mutation frequency and NF1 protein expression | Genomic Laboratory Information System, Asan Medical Center |
- – NF1 highly expressed in sensitive lines (NCI-H508, Caco-2), little expression in resistant lines (KM12C, SW480)
- ▲ NF1 knockdown enhanced phosphorylation of MEK and ERK
- ▼ NF1 knockdown reduced apoptosis (fewer apoptotic bodies, less cleaved caspase and PARP)
- ▲ NF1-GRD overexpression increased cetuximab-induced apoptosis in KM12C and SW480
- – Pre-treatment NF1 expression not associated with cetuximab response in 111 RAS/BRAF WT tumors
- ▼ Post-treatment samples showed slightly lower NF1 transcript than pre-treatment
- – Inactivating NF1 mutations rare in CRC patients 1.8%
- – Biallelic inactivation of NF1 observed in small subset of cases 0.5%
- count 1.8% (frequency of inactivating NF1 mutations in CRC patients)
- count 0.5% (cases with biallelic inactivation of NF1)
- count 111 (RAS and BRAF V600 wild-type tumor samples analyzed by RNA-seq for cetuximab response)
- count 92 patients with 113 samples (patients meeting selection criteria from biomarker program)
- count 2,589 (participants enrolled in biomarker discovery program Sept 2011-March 2018)
- count 1,449 (CRC patients screened in NGS genomic database March 2017-May 2020)
- count 123,416,623 (mean total reads in RNA-seq)
- count 35,361,303 (average reads per sample in RNA-seq)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combined in vitro cell line experiments (four CRC lines, two designated sensitive and two resistant to cetuximab) with RNA sequencing of 111 analyzable FFPE tumor samples from 92 patients with RAS/BRAF-wild-type metastatic CRC treated with cetuximab. Gene expression was quantified with RSEM, normalized using DESeq2, and differential expression between response groups was summarized as log2 fold changes; pathway enrichment was assessed by GSEA via clusterProfiler. NF1 mutation frequency was estimated descriptively from an NGS database of 1,449 CRC patients.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 differential gene expression analysis (Wald test implied by package default; not explicitly named) | NF1 and global gene expression compared across cetuximab response groups (pre-Tx CR/PR, pre-Tx SD/PD, post-Tx nPD, post-Tx PD) in tumor samples | 111 samples (113 collected minus 2 that failed library QC) | stated |
| Gene set enrichment analysis (GSEA) with permutation-based p-values | NF1- and EGFR-related KEGG pathway enrichment from ranked differentially expressed gene list | 111 | not stated |
| Descriptive frequency count (proportion) | NF1 inactivating mutation frequency and biallelic inactivation rate in CRC patients from NGS database | 1449 | na |
| Densitometry (ImageJ) for semi-quantitative band comparison | Western blot quantification of NF1, p-MEK, p-ERK, caspase, and PARP across cell lines and siRNA/plasmid transfection conditions | — | not stated |
| Image-based cell counting (ImageJ particle analysis) | Relative cell number comparisons between siRNA/plasmid-transfected and control cells; apoptotic body counts by DAPI staining | — | not stated |
-
Cell line assays (western blotting, cell counting) compared cetuximab-sensitive and -resistant lines using densitometry and image-based counts without a stated formal statistical test or reported number of independent biological replicates↳ Could also: A two-sample t-test or Mann-Whitney U test applied across a stated number of independent biological replicates could also have been used, with the chosen test depending on distributional assumptions — Formal testing across biological replicates with stated n allows quantification of uncertainty and reproducibility, and makes it possible to assess whether observed differences exceed expected biological variability
-
NF1 transcript levels across four patient groups were compared using DESeq2 pairwise contrasts without an explicit omnibus test across all groups simultaneously↳ Could also: A one-way ANOVA (or Kruskal-Wallis for non-parametric data) followed by a post-hoc correction (e.g., Tukey HSD or Dunn's test) could also have been applied to normalized expression values — An omnibus test first controls the family-wise error rate across all group comparisons before proceeding to post-hoc contrasts, which is a common approach when more than two groups are compared simultaneously
-
Pre-treatment and post-treatment samples from overlapping patients were compared across response groups without explicitly modeling the within-patient pairing↳ Could also: A paired analysis (e.g., Wilcoxon signed-rank test or a linear mixed-effects model with patient as a random effect) could also have been used for patients contributing both pre- and post-treatment samples — Pairing removes between-subject variability and can increase statistical power when the same patient contributes samples to more than one group, as appears to be the case here for at least a subset of participants
-
GSEA permutation-based p-values were calculated across NF1- and EGFR-related KEGG pathways without a stated correction for testing multiple pathways↳ Could also: Benjamini-Hochberg FDR-adjusted p-values (q-values) are also commonly reported alongside raw GSEA p-values when multiple pathways are tested — When many pathways are tested simultaneously, unadjusted p-values increase the expected number of false positives; FDR adjustment is standard practice in pathway enrichment analyses
-
edgeR was not considered as an alternative to DESeq2 for the RNA-seq differential expression analysis of FFPE-derived samples↳ Could also: edgeR (negative binomial model with empirical Bayes dispersion) or limma-voom (precision-weighted linear model after variance stabilization) could also have been used for count-based DGE — Both are widely benchmarked alternatives to DESeq2 for small-to-moderate sample sizes; limma-voom can be particularly well-suited for FFPE-derived RNA where input RNA quality may introduce additional variance
-
Effect sizes for DGE comparisons were reported as log2 fold changes without accompanying confidence intervals↳ Could also: 95% confidence intervals around log2 fold changes could also have been reported; DESeq2 provides these natively, including via its lfcShrink function for shrunken estimates — Confidence intervals convey both the magnitude and the precision of an estimated effect, which can be informative when interpreting whether a fold change is biologically meaningful given the sample sizes available
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 34779495 (NF1 in CRC & cetuximab resistance)
Oncology Reports 2021;47(2):8226. PMCID PMC8611403. DOI 10.3892/or.2021.8226. Data: GEO GSE183984 (RNA-seq, colon cancer pre/post cetuximab). Listed code: github.com/mskcc/vcf2maf (a third-party VCF→MAF converter, used for the targeted-DNA-seq mutation part — see below).
Pipeline-derived results (candidates)
The paper is predominantly wet-lab (CRC cell lines, western blot, RT-qPCR, siRNA knockdown, GAP-domain overexpression, colony/viability/apoptosis assays, IHC). Those are OUT OF SCOPE (manual/experimental, not a bioinformatic pipeline).
Two computational sub-pipelines exist:
-
Patient RNA-seq (GSE183984) — FFPE tumors, TruSeq RNA Access, HiSeq 2500, pipeline: FASTQC → Trim Galore → STAR (GRCh38) → RSEM → DESeq2 (FPKM) → GSEA (clusterProfiler). The authors ship the processed matrices on GEO:
GSE183984_ASAN_RNASEQ_FPKM_ensg.csv.gzand..._raw_counts_ensg.csv.gzplus a series matrix withtime point(pre/post) andtreatment response(pre-Tx / post-Tx non-PD / post-tx PD). IN SCOPE — the downstream transcript-level claims are directly reproducible from the shipped FPKM matrix + sample labels (honest 1:1 on the paper's own data). -
Targeted DNA-seq / NF1 mutation profiling (the part
vcf2mafbelongs to): variant calling on the Asan Medical Center Genomic LIS, filtered against an in-house panel of normals and KRGDB, annotated with VEP v79 → vcf2maf. OUT OF SCOPE / not reproducible: the input VCFs/BAMs are private institutional patient data ("provided upon request from the corresponding author"); no public accession. The NF1-mutation frequency claim (1.8% inactivating mutations) is therefore not derivable from shipped artifacts.
In-scope claims attempted (from GSE183984 FPKM)
- C1 "tumor samples obtained after cetuximab treatment displayed slightly lower NF1 transcript levels compared with those in the pre-treatment samples" (Discussion; cohort = 111 RAS/BRAF-WT samples).
- C2 "pre-treatment NF1 expression levels were not associated with the cetuximab response" (same paragraph).
Not attempted (and why)
- All wet-lab assays — not pipeline-derived.
- NF1 mutation frequency (1.8%) — input data private (
data_restricted). - Full FASTQ→STAR→RSEM→DESeq2 re-run — raw reads not needed; authors ship the processed FPKM/count matrices (the pipeline's output), which is the correct, faithful input for the reported between-sample transcript comparison. Re-running the aligner is the hard ~20% with no extra evidentiary value here.
- GSEA of NF1/EGFR KEGG pathways — qualitative ("pathways enriched"), no pinnable numeric target reported; not attempted.
Caveats carried into grading: paper used n=111 (post-QC, RAS/BRAF-WT subset); GEO ships 113; we used all 113 (76 pre / 37 post). C2 response is operationalized as post-treatment non-PD vs PD because GEO labels response only on post samples.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The two pipeline-derived RNA-seq claims reproduce cleanly on the authors' own GEO-shipped FPKM matrix: NF1 is significantly lower post-cetuximab (p=0.0012, matching 'slightly lower') and unassociated with response (p=0.327, matching 'not associated'). Deviations are minor and on our/data-availability side, not the authors': qualitative endpoints allow only directional matching, the cohort is 113 vs the paper's 111 (exclusion list not deposited), and the M1 1.8% mutation frequency is uncheckable because its input VCFs are private. No fabrication signal — the checkable claims are fully derivable and hold.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.