Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Expression of neurofibromin 1 in colorectal cancer and cetuximab resistance.

Oncol Rep · 2021
L1 73/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
73/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 664 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for a faithful 1:1 of the RNA-seq transcript claims. The paper is mostly wet-lab; its two PIPELINE-DERIVED transcript claims come from GSE183984, whose processed FPKM matrix the authors ship on GEO. Using that matrix directly (stdlib Python on «our HPC», no aligner re-run needed), BOTH reproduced in direction and significance: (C1) NF1 is ~0.80x lower post- vs pre-cetuximab (p=0.0012, matching 'slightly lower'); (C2) NF1 shows no significant association with response (post non-PD vs PD p=0.33, matching 'not associated'). NOT attempted: all wet-lab assays (non-pipeline); the NF1 mutation-frequency claim (1.8%) because its vcf2maf input VCFs are private (data_restricted); the full FASTQ->STAR->RSEM->DESeq2 re-run (unnecessary - authors ship the pipeline's FPKM output, the correct input for a between-sample transcript comparison, and re-aligning adds no evidentiary value, the hard ~20%). Caveat: paper used n=111 post-QC RAS/BRAF-WT; GEO ships 113 and we used all 113; direction robust. Grades provisional - claims are qualitative so agreement is directional, not byte-exact; a human reviewer signs off in AUDIT.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 73
    assessed: 2026-06-15 ⛓ 2e7c9803f248
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study investigates whether neurofibromin 1 (NF1) expression level is associated with sensitivity/resistance to the anti-EGFR antibody cetuximab in colorectal cancer, testing whether modifying NF1 expression alters cetuximab response and whether NF1 could serve as a biomarker in RAS/BRAF V600 wild-type tumors.

Core claims
  • NF1 is highly expressed in cetuximab-sensitive CRC cell lines and minimally expressed in cetuximab-resistant ones finding
  • siRNA knockdown of NF1 in sensitive cell lines enhances MEK/ERK phosphorylation and reduces cetuximab-induced apoptosis mechanism
  • NF1 overexpression (NF1-GRD) renders resistant cell lines KM12C and SW480 more susceptible to cetuximab-induced apoptosis finding
  • Modification of NF1 expression can affect cetuximab sensitivity in CRC cell lines finding
  • Pre-treatment NF1 expression levels in patient tumors were not associated with cetuximab response, limiting NF1 as a clinical biomarker finding
  • Post-cetuximab-treatment tumor samples showed slightly lower NF1 transcript levels than pre-treatment samples finding
  • Inactivating NF1 mutations are rare (1.8%) in CRC patients and generally not associated with NF1 protein expression finding
  • NF1 negatively regulates RAS-MAPK signaling via GTPase-activating activity mechanism
Experimental setups
Assay System Perturbation Readout Platform
Western blot CRC cell lines NCI-H508, Caco-2, KM12C, SW480 none/baseline and cetuximab; NF1 siRNA knockdown; NF1-GRD overexpression NF1, p-MEK1/2, MEK1/2, p-ERK1/2, ERK1/2, cleaved caspase-3, cleaved PARP protein levels LuminoGraph II (ATTO); Cell Signaling/Abcam antibodies; ImageJ densitometry
Cell growth assay (crystal violet) CRC cell lines NCI-H508, Caco-2, KM12C, SW480 cetuximab 0/50/100/200 µg/ml relative proliferation (OD 595 nm) Sunrise microplate reader (Tecan)
Colony formation assay CRC cell lines cetuximab 0/50/100/200 µg/ml colony number (>200 µm) Oxford Optronix GelCount
siRNA knockdown / plasmid transfection cetuximab-sensitive (knockdown) and cetuximab-resistant KM12C/SW480 (overexpression) CRC cell lines NF1-siRNA / NF1-GRD plasmid + cetuximab 100 µg/ml NF1 expression, apoptosis markers, cell number Lipofectamine 3000
DAPI nuclear staining siRNA- and plasmid-transfected CRC cell lines NF1 knockdown / overexpression + cetuximab apoptotic bodies / nuclear condensation count EVOS FL Auto fluorescence microscope; ImageJ
RT-qPCR CRC cell lines siRNA/plasmid transfection NF1 transcript level normalized to GAPDH CFX Connect Real-Time PCR (Bio-Rad); EvaGreen qPCR
Bulk RNA sequencing FFPE tumor samples from 92 mCRC patients (113 samples), RAS/BRAF V600 wild-type cetuximab treatment (pre- vs post-treatment) NF1 transcript (FPKM), differential expression, GSEA pathway enrichment Illumina HiSeq 2500; TruSeq RNA Access Library Prep; STAR/RSEM/DESeq2
Targeted DNA sequencing (NGS) / IHC clinical CRC patient genomic database (1,449 patients) none (diagnostic profiling) NF1 mutation frequency and NF1 protein expression Genomic Laboratory Information System, Asan Medical Center
Key results
  • NF1 highly expressed in sensitive lines (NCI-H508, Caco-2), little expression in resistant lines (KM12C, SW480)
  • NF1 knockdown enhanced phosphorylation of MEK and ERK
  • NF1 knockdown reduced apoptosis (fewer apoptotic bodies, less cleaved caspase and PARP)
  • NF1-GRD overexpression increased cetuximab-induced apoptosis in KM12C and SW480
  • Pre-treatment NF1 expression not associated with cetuximab response in 111 RAS/BRAF WT tumors
  • Post-treatment samples showed slightly lower NF1 transcript than pre-treatment
  • Inactivating NF1 mutations rare in CRC patients 1.8%
  • Biallelic inactivation of NF1 observed in small subset of cases 0.5%
Key statistics
  • count 1.8% (frequency of inactivating NF1 mutations in CRC patients)
  • count 0.5% (cases with biallelic inactivation of NF1)
  • count 111 (RAS and BRAF V600 wild-type tumor samples analyzed by RNA-seq for cetuximab response)
  • count 92 patients with 113 samples (patients meeting selection criteria from biomarker program)
  • count 2,589 (participants enrolled in biomarker discovery program Sept 2011-March 2018)
  • count 1,449 (CRC patients screened in NGS genomic database March 2017-May 2020)
  • count 123,416,623 (mean total reads in RNA-seq)
  • count 35,361,303 (average reads per sample in RNA-seq)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined in vitro cell line experiments (four CRC lines, two designated sensitive and two resistant to cetuximab) with RNA sequencing of 111 analyzable FFPE tumor samples from 92 patients with RAS/BRAF-wild-type metastatic CRC treated with cetuximab. Gene expression was quantified with RSEM, normalized using DESeq2, and differential expression between response groups was summarized as log2 fold changes; pathway enrichment was assessed by GSEA via clusterProfiler. NF1 mutation frequency was estimated descriptively from an NGS database of 1,449 CRC patients.

Replicationmixed Sample size4 CRC cell lines (biological replicates per condition not stated); 92 patients with 111 analyzable tumor samples across 4 response groups; 1,449 CRC patients in NGS database GroupsCetuximab-sensitive (NCI-H508, Caco-2) vs. cetuximab-resistant (KM12C, SW480) cell lines; pre-Tx CR/PR (n=59) vs. pre-Tx SD/PD (n=16) vs. post-Tx nPD (n=16) vs. post-Tx PD (n=20) tumor samples Pairingmixed Randomization/blindingnot stated Dispersionnone Effect sizesyes Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
DESeq2 differential gene expression analysis (Wald test implied by package default; not explicitly named) NF1 and global gene expression compared across cetuximab response groups (pre-Tx CR/PR, pre-Tx SD/PD, post-Tx nPD, post-Tx PD) in tumor samples 111 samples (113 collected minus 2 that failed library QC) stated
Gene set enrichment analysis (GSEA) with permutation-based p-values NF1- and EGFR-related KEGG pathway enrichment from ranked differentially expressed gene list 111 not stated
Descriptive frequency count (proportion) NF1 inactivating mutation frequency and biallelic inactivation rate in CRC patients from NGS database 1449 na
Densitometry (ImageJ) for semi-quantitative band comparison Western blot quantification of NF1, p-MEK, p-ERK, caspase, and PARP across cell lines and siRNA/plasmid transfection conditions not stated
Image-based cell counting (ImageJ particle analysis) Relative cell number comparisons between siRNA/plasmid-transfected and control cells; apoptotic body counts by DAPI staining not stated
Approaches that could also have been used
  • Cell line assays (western blotting, cell counting) compared cetuximab-sensitive and -resistant lines using densitometry and image-based counts without a stated formal statistical test or reported number of independent biological replicates
    Could also: A two-sample t-test or Mann-Whitney U test applied across a stated number of independent biological replicates could also have been used, with the chosen test depending on distributional assumptions — Formal testing across biological replicates with stated n allows quantification of uncertainty and reproducibility, and makes it possible to assess whether observed differences exceed expected biological variability
  • NF1 transcript levels across four patient groups were compared using DESeq2 pairwise contrasts without an explicit omnibus test across all groups simultaneously
    Could also: A one-way ANOVA (or Kruskal-Wallis for non-parametric data) followed by a post-hoc correction (e.g., Tukey HSD or Dunn's test) could also have been applied to normalized expression values — An omnibus test first controls the family-wise error rate across all group comparisons before proceeding to post-hoc contrasts, which is a common approach when more than two groups are compared simultaneously
  • Pre-treatment and post-treatment samples from overlapping patients were compared across response groups without explicitly modeling the within-patient pairing
    Could also: A paired analysis (e.g., Wilcoxon signed-rank test or a linear mixed-effects model with patient as a random effect) could also have been used for patients contributing both pre- and post-treatment samples — Pairing removes between-subject variability and can increase statistical power when the same patient contributes samples to more than one group, as appears to be the case here for at least a subset of participants
  • GSEA permutation-based p-values were calculated across NF1- and EGFR-related KEGG pathways without a stated correction for testing multiple pathways
    Could also: Benjamini-Hochberg FDR-adjusted p-values (q-values) are also commonly reported alongside raw GSEA p-values when multiple pathways are tested — When many pathways are tested simultaneously, unadjusted p-values increase the expected number of false positives; FDR adjustment is standard practice in pathway enrichment analyses
  • edgeR was not considered as an alternative to DESeq2 for the RNA-seq differential expression analysis of FFPE-derived samples
    Could also: edgeR (negative binomial model with empirical Bayes dispersion) or limma-voom (precision-weighted linear model after variance stabilization) could also have been used for count-based DGE — Both are widely benchmarked alternatives to DESeq2 for small-to-moderate sample sizes; limma-voom can be particularly well-suited for FFPE-derived RNA where input RNA quality may introduce additional variance
  • Effect sizes for DGE comparisons were reported as log2 fold changes without accompanying confidence intervals
    Could also: 95% confidence intervals around log2 fold changes could also have been reported; DESeq2 provides these natively, including via its lfcShrink function for shrunken estimates — Confidence intervals convey both the magnitude and the precision of an estimated effect, which can be informative when interpreting whether a fold change is biologically meaningful given the sample sizes available
Software: DESeq2 (R/Bioconductor) 1.20.0 (reported as BiocManager package version; may refer to Bioconductor release) · RSEM 1.2.23 · STAR aligner 2.6.0 · clusterProfiler (R) · FASTQC 0.11.8 · Trim Galore 0.4.5 · ImageJ 1.53a · Oxford Optronix GelCount 1.1.2.0 · ImageSaver 6 (ATTO Corporation) 2.7.2

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 34779495 (NF1 in CRC & cetuximab resistance)

Oncology Reports 2021;47(2):8226. PMCID PMC8611403. DOI 10.3892/or.2021.8226. Data: GEO GSE183984 (RNA-seq, colon cancer pre/post cetuximab). Listed code: github.com/mskcc/vcf2maf (a third-party VCF→MAF converter, used for the targeted-DNA-seq mutation part — see below).

Pipeline-derived results (candidates)

The paper is predominantly wet-lab (CRC cell lines, western blot, RT-qPCR, siRNA knockdown, GAP-domain overexpression, colony/viability/apoptosis assays, IHC). Those are OUT OF SCOPE (manual/experimental, not a bioinformatic pipeline).

Two computational sub-pipelines exist:

  1. Patient RNA-seq (GSE183984) — FFPE tumors, TruSeq RNA Access, HiSeq 2500, pipeline: FASTQC → Trim Galore → STAR (GRCh38) → RSEM → DESeq2 (FPKM) → GSEA (clusterProfiler). The authors ship the processed matrices on GEO: GSE183984_ASAN_RNASEQ_FPKM_ensg.csv.gz and ..._raw_counts_ensg.csv.gz plus a series matrix with time point (pre/post) and treatment response (pre-Tx / post-Tx non-PD / post-tx PD). IN SCOPE — the downstream transcript-level claims are directly reproducible from the shipped FPKM matrix + sample labels (honest 1:1 on the paper's own data).

  2. Targeted DNA-seq / NF1 mutation profiling (the part vcf2maf belongs to): variant calling on the Asan Medical Center Genomic LIS, filtered against an in-house panel of normals and KRGDB, annotated with VEP v79 → vcf2maf. OUT OF SCOPE / not reproducible: the input VCFs/BAMs are private institutional patient data ("provided upon request from the corresponding author"); no public accession. The NF1-mutation frequency claim (1.8% inactivating mutations) is therefore not derivable from shipped artifacts.

In-scope claims attempted (from GSE183984 FPKM)

  • C1 "tumor samples obtained after cetuximab treatment displayed slightly lower NF1 transcript levels compared with those in the pre-treatment samples" (Discussion; cohort = 111 RAS/BRAF-WT samples).
  • C2 "pre-treatment NF1 expression levels were not associated with the cetuximab response" (same paragraph).

Not attempted (and why)

  • All wet-lab assays — not pipeline-derived.
  • NF1 mutation frequency (1.8%) — input data private (data_restricted).
  • Full FASTQ→STAR→RSEM→DESeq2 re-run — raw reads not needed; authors ship the processed FPKM/count matrices (the pipeline's output), which is the correct, faithful input for the reported between-sample transcript comparison. Re-running the aligner is the hard ~20% with no extra evidentiary value here.
  • GSEA of NF1/EGFR KEGG pathways — qualitative ("pathways enriched"), no pinnable numeric target reported; not attempted.

Caveats carried into grading: paper used n=111 (post-QC, RAS/BRAF-WT subset); GEO ships 113; we used all 113 (76 pre / 37 post). C2 response is operationalized as post-treatment non-PD vs PD because GEO labels response only on post samples.

C1
Reported
NF1 transcript levels slightly lower in post-cetuximab vs pre-treatment tumors
Reproduced
pre median FPKM 238.73 (n=76) vs post 192.00 (n=37); post lower, 0.80x (log2FC -0.314); Mann-Whitney p=0.0012
within tolerance
C2
Reported
Pre-treatment NF1 expression not associated with cetuximab response
Reproduced
post non-PD median 180.55 (n=16) vs post PD 201.98 (n=21); Mann-Whitney p=0.327 (not significant)
within tolerance
M1
Reported
1.8% inactivating NF1 mutations (targeted DNA-seq + vcf2maf)
Reproduced
not attempted - input VCFs are private Asan Medical Center data (on request), not in GSE183984
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 73/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

The two pipeline-derived RNA-seq claims reproduce cleanly on the authors' own GEO-shipped FPKM matrix: NF1 is significantly lower post-cetuximab (p=0.0012, matching 'slightly lower') and unassociated with response (p=0.327, matching 'not associated'). Deviations are minor and on our/data-availability side, not the authors': qualitative endpoints allow only directional matching, the cohort is 113 vs the paper's 111 (exclusion list not deposited), and the M1 1.8% mutation frequency is uncheckable because its input VCFs are private. No fabrication signal — the checkable claims are fully derivable and hold.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

92.3 k
tokens (I/O) · 5.4 M incl. cache
11 min
runtime · 0 CPU-h
0 GB
peak RAM
2
HPC jobs
hummel
machine