Cis-Regulation of the CFTR Gene in Pancreatic Cells.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Paper: Blotas et al 2025 IJMS, cis-regulation of CFTR. The single clean in-scope pipeline result C1 = ATAC-seq peaks within the CFTR TAD (Capan-1), reported as 6 interior peaks (2 upstream / 1 promoter / 3 at 3'), hg19. Reproduced with the cited third-party Galaxy workflow iwc-workflows/atacseq v0.12 tool chain (cutadapt Nextera -> bowtie2 --very-sensitive --dovetail -X1000 hg19 -> MAPQ30/concordant/no-chrM filter -> Picard MarkDuplicates -> bedtools bamtobed -> macs2 -q0.05 --nomodel --shift -100 --extsize 200 --call-summits --keep-dup all) on GSE284414 (the ATAC subseries; the brief's geo:GSE284199 is the 4C subseries -- correction documented in scope.md). OUTCOME = PARTIAL: the pipeline ran cleanly end-to-end and the qualitative result reproduces (a strong CFTR-promoter peak in every sample, plus accessible chromatin upstream and at the 3' end), but the literal headline count '6' is NOT 1:1 reproducible -- standard MACS2 q0.05 yields ~17 unique peaks in the TAD (Capan-1 merged), ~3x the reported 6. The '6' reads as a curated candidate-regulatory-element set (no intragenic category; 6 elements taken to luciferase), and its selection rule is underspecified in the Methods -- flagged low-severity as 'not directly derivable from the shipped data/code' (curation gap, NOT fabrication; a human reviewer should check the supplement). Three blockers solved this session: (1) Picard MarkDuplicates NPE on bowtie2 BAM lacking @RG -> AddOrReplaceReadGroups; (2) MACS2 ImportError 'undefined symbol: __log_finite' (glibc>=2.31 dropped finite-math aliases) -> LD_PRELOAD shim; (3) front1 /tmp full (50M tmpfs) -> TMPDIR/-pipe on «infra». All 3 subseries profiled (4/9/7 runs, all N reported=observed). Bonus: Caco-2 ATAC also processed (21/22/32 unique TAD peaks rep1/rep2/merged). NOT attempted (out of scope): 4C (GSE284199), CUT&RUN (GSE284200), Jaspar motifs, luciferase (wet-lab).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-21 ⛓ 525cc845775c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether 3D chromatin architecture and cis-regulatory elements (CREs) around the CFTR gene establish its tissue-specific expression in pancreatic cells, potentially explaining clinical heterogeneity in cystic fibrosis and CFTR-related disorders.
- ★ Multiple active CREs exist upstream and downstream of the CFTR gene in pancreatic (Capan-1) cells, identified via ATAC-seq, CUT&RUN-seq (H3K27ac), 4C-seq, and the ABC model finding
- ★ The -44 kb, -35 kb, +15.6 kb, and +37.7 kb regions act cooperatively, sharing common predicted transcription factor binding motifs, and jointly enhance promoter activity beyond individual effects finding
- ★ Active CREs regulating CFTR exist outside the CFTR TAD, including a region near the LSM8 gene at +507.6 kb showing chromatin interaction and silencer activity finding
- ★ CRISPR/Cas9-mediated homozygous deletion of the endogenous -44 kb CRE reduces CFTR expression, validating its enhancer function finding
- ★ CRE usage and activity is tissue-specific, differing between pancreatic (Capan-1) and intestinal (Caco-2) cells, with intron 12 and intron 24 OCRs acting as strong enhancers only in intestinal cells finding
- The activity-by-contact (ABC) model can be used to computationally link candidate CREs to the CFTR promoter using ATAC-seq, H3K27ac, and Hi-C-type data method
- HNF1B transcription factor motifs are present in active cCREs except intron 18, which showed no luciferase activity mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ATAC-seq | Capan-1 pancreatic ductal cells | none | genome-wide open/accessible chromatin regions | — |
| CUT&RUN-seq (H3K27ac) | Capan-1 pancreatic ductal cells | none | active enhancer histone mark peaks | — |
| 4C-seq | Capan-1 pancreatic ductal cells | none (CFTR promoter used as bait) | chromatin interaction frequency at CFTR locus, analyzed with peakC | — |
| Luciferase reporter assay | Capan-1 cells (co-transfection with pCMV-beta-galactosidase control) | candidate CRE constructs cloned upstream of minimal CFTR promoter in pGL3-Basic vector | luciferase activity (fold change) | pGL3-Basic vector |
| CRISPR/Cas9 knockout (VLP delivery) | Capan-1 cells | deletion of endogenous -44 kb CRE via flanking sgRNAs | CFTR mRNA expression by RT-qPCR; deletion confirmed by PCR and Sanger sequencing | virus-like particles (VLPs) |
| ATAC-seq | Caco-2 intestinal cells | none | open chromatin regions for cross-tissue comparison | — |
| ATAC-seq | HepG2 cells (publicly available data) | none | open chromatin regions in a CFTR-non-expressing cell line | — |
| 4C-seq | Caco-2 and HepG2 cells | none (CFTR promoter bait) | chromatin interaction profiles compared across cell types | — |
- – Six reliable ATAC/H3K27ac consensus peaks identified within the CFTR TAD (two upstream, one promoter, three at 3' end)
- – ABC model identified five links to the CFTR promoter, including -44 kb, -35 kb, -114 bp, and +15.6 kb regions
- – 4C-seq peakC analysis found significant interactions at -80.1 kb boundary, ~-60 kb, +15.6 kb, intron 26, and +37.7 kb
- – Luciferase activity increased with -44 kb region and intron 26 OCR; -35 kb and +37.7 kb showed minor silencing; combination of all four cCREs gave the largest increase 1.7-fold (-44kb); 1.8-fold (intron26); -1.39-fold (-35kb/+37.7kb); 3.2-fold (4-CRE combo)
- ▲ Intron 23 region showed significant enhancer activity in luciferase assay 3.5-fold
- ▼ Homozygous CRISPR deletion of -44 kb region reduced CFTR expression relative to untreated cells; heterozygous deletion had no effect 15% reduction
- ▼ +507.6 kb region (near LSM8, outside TAD) showed silencer activity in luciferase assay; CTCF site and LSM8 promoter region had no effect -1.32-fold
- ▲ Intron 12 and combined intron 12+24 regions strongly increased luciferase activity in Caco-2 cells but had no effect in Capan-1 cells 15-fold (intron12); 50-fold (intron12+24) in Caco-2
- fold_change 1.7-fold (-44 kb region luciferase activity vs promoter alone in Capan-1)
- fold_change 1.8-fold (intron 26 OCR luciferase activity vs promoter alone in Capan-1)
- fold_change -1.39-fold (-35 kb and +37.7 kb regions minor silencing effect in Capan-1)
- fold_change 3.2-fold (combination of -44kb, -35kb, +15.6kb, +37.7kb cCREs luciferase activity)
- fold_change 3.5-fold (intron 23 region luciferase activity in Capan-1)
- fold_change 15% reduction (CFTR expression after homozygous CRISPR deletion of -44 kb region (clone 1) vs untreated Capan-1)
- fold_change 15-fold (intron 12 luciferase activity increase in Caco-2 cells)
- fold_change 50-fold (combined intron 12 and intron 24 luciferase activity increase in Caco-2 cells)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper uses genomic profiling (ATAC-seq, CUT&RUN-seq for H3K27ac, 4C-seq) combined with a computational activity-by-contact (ABC) model to nominate candidate cis-regulatory elements, then functionally tests candidates with luciferase reporter assays (reported as fold-change relative to a promoter-only construct) and validates one element with CRISPR/Cas9 deletion followed by RT-qPCR (reported as percent change relative to untreated cells). Findings are described narratively with fold-change values and one explicit statement that a comparison was 'not significantly different,' but the provided text does not name a specific statistical test, replicate structure, or software package used for these comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| not stated (a statistical comparison is implied by the phrase 'not significantly different') | Luciferase assay, Figure 4 — combined −44kb+−35kb construct vs. −44kb alone | — | not stated |
-
Luciferase and RT-qPCR results are summarized as single fold-change or percent-change values without a stated measure of dispersion or replicate count in the text.↳ Could also: Reporting mean ± SD or SEM (or a 95% CI) alongside individual replicate data points — Showing dispersion and the number of independent replicates alongside the point estimate helps convey the precision of a fold-change or percent-change measurement, which is especially informative for small-n reporter assays.
-
A comparison between the combined −44kb+−35kb construct and the −44kb-alone construct is described as 'not significantly different' without naming the statistical test used.↳ Could also: Explicitly naming the test (e.g., a two-tailed t-test or one-way ANOVA with post-hoc comparison) and reporting the associated p-value or test statistic — Stating the specific test and exact p-value allows readers to evaluate the assumptions of the test and reproduce the significance assessment.
-
Multiple related luciferase constructs (single CREs, paired combinations, and a four-CRE combination) are compared against the promoter-only baseline across several figures.↳ Could also: A one-way or two-way ANOVA with a post-hoc multiple-comparisons correction (e.g., Tukey HSD, Dunnett's test, or Benjamini-Hochberg FDR) — When many related constructs are compared to the same baseline, an omnibus test with a correction for multiple comparisons is a standard way to control the family-wise error rate across that comparison set.
-
The endogenous validation of the −44kb enhancer relies on RT-qPCR comparing one homozygous deletion clone to untreated Capan-1 cells.↳ Could also: Testing multiple independently derived homozygous clones with a statistical comparison (e.g., paired or unpaired t-test) across biological replicates — Using several independent clones and a formal statistical comparison can help distinguish an on-target enhancer effect from clone-specific variation.
-
4C-seq interaction peaks are identified using peakC without a stated significance threshold or multiple-testing correction in the text.↳ Could also: Reporting the FDR or p-value threshold used by peakC (or an alternative 4C/Hi-C peak-calling tool with explicit multiple-testing control) — Stating the significance threshold used for peak calling clarifies how candidate interactions were distinguished from background noise across the many genomic bins tested.
-
Enhancer-promoter links are nominated using the ABC computational model, which integrates chromatin accessibility, H3K27ac, and Hi-C-derived contact data.↳ Could also: Cross-validating links with an additional enhancer-prediction approach, such as co-accessibility analysis from single-cell ATAC-seq or an independent Hi-C-based statistical model — Comparing predictions from more than one enhancer-linking method can provide converging evidence for candidate CRE-promoter relationships in a given cell type.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40332394
Paper: Cis-Regulation of the CFTR Gene in Pancreatic Cells. Blotas et al., Int J Mol Sci 2025. PMID 40332394 · PMCID PMC12027686 · DOI 10.3390/ijms26083788.
Assays in the paper (all on Capan-1 / Caco-2 / HepG2 cell lines):
- ATAC-seq — open-chromatin mapping. Pipeline: iwc-workflows/atacseq v0.12 (Galaxy), default params, mapped to hg19, peaks called with MACS2.
- CUT&RUN-seq (H3K27ac, CTCF, H3K27me3) — pipeline iwc-workflows/cutandrun v0.8.
- 4C-seq — pipeline Pipe4C v1.1.4 + peakC v0.2, hg19.
- Luciferase reporter assays — wet-lab.
GEO accession mapping (resolved from NCBI/ENA, 2026-06-15)
The brief's RU pairs geo:GSE284199 with code:iwc-workflows/atacseq, but those
two do not correspond to the same assay. The paper deposited a SuperSeries split
into three subseries:
| GEO | assay | pipeline (code) | samples | BioProject |
|---|---|---|---|---|
| GSE284199 | 4C-seq | Pipe4C + peakC | 9 (GSM8679690–98) | PRJNA1197587 |
| GSE284200 | CUT&RUN | iwc cutandrun v0.8 | 7 (GSM8679699–705) | PRJNA1197588 |
| GSE284414 | ATAC-seq | iwc atacseq v0.12 | 4 (GSM8683820–23) | PRJNA1198966 |
→ The cited code (iwc-workflows/atacseq) is the ATAC-seq pipeline, whose data live
in GSE284414, not GSE284199. We therefore reproduce the ATAC-seq result using
GSE284414 + the atacseq workflow (the internally-consistent code↔data pairing).
This correction is recorded in code/code.json and data/data.json. (Per brief
rule 2, applying the cited third-party Galaxy tool to the paper's own data is a
fully valid reproduction.)
ATAC-seq runs (GSE284414 / PRJNA1198966, ENA, paired-end)
| run | sample | reads (PE) |
|---|---|---|
| SRR31736540 | ATAC_Capan-1_rep1 | 37.1 M |
| SRR31736539 | ATAC_Capan-1_rep2 | 35.7 M |
| SRR31736538 | ATAC_Caco-2_rep1 | 39.2 M |
| SRR31736537 | ATAC_Caco-2_rep2 | 37.0 M |
IN SCOPE (pipeline-derived, attempted)
- C1 — peaks within the CFTR TAD. Reported (Results / Fig.): "Six peaks were identified within the TAD [chr7:117,039,878–117,356,812]. Two were located upstream of the promoter, one corresponded to the CFTR promoter, and three were located at the 3′ end." This is a clean, objective, MACS2-pipeline-derived count + locations. Reproduction = run iwc atacseq v0.12-equivalent (cutadapt Nextera trim → bowtie2 --dovetail -X1000 → hg19 → MAPQ30/concordant/no-chrM filter → Picard dedup → BAM→BED → MACS2) on the Capan-1 ATAC reads, then count MACS2 peaks falling in chr7:117,039,878–117,356,812 and classify them (upstream / promoter / 3′).
OUT OF SCOPE (not attempted, with reason)
- Luciferase fold-changes (−44 kb 1.7×, intron 23 3.5×, introns 12/24 15×/50× in Caco-2, etc.) — wet-lab reporter assays, non_pipeline.
- 4C-seq interaction maps / ABC links (GSE284199, Pipe4C+peakC) and CUT&RUN H3K27ac peaks (GSE284200) — separate pipelines/data; the hard ~20%, deferred per the 80/20 rule. The ATAC peak count is the single cleanest pipeline number.
- Jaspar 2024 TFBS motif mapping (HNF1B etc.) — depends on the cCRE set + external motif DB; downstream/manual, deferred.
80/20 decision
The ATAC-seq "6 peaks in the CFTR TAD" is the lowest-hanging, fully-specified, objectively-checkable pipeline output (one integer + element locations). We reproduce that with the cited workflow's tool chain on the paper's own Capan-1 ATAC data, all on «our HPC»/«infra». We do not attempt the 4C-seq, CUT&RUN, motif, or luciferase results.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The pipeline reproduced cleanly on the authors' own Capan-1 ATAC data (GSE284414) and the qualitative spatial claim — accessible chromatin at the CFTR promoter, upstream, and 3' end — holds, with the promoter peak (chr7:117,118,870-117,120,235) reproduced in every sample. However the headline count of 6 is not 1:1 reproducible: the standard MACS2 q0.05 pipeline yields ~17 unique TAD peaks (~3x). The gap is on the authors' side — an underspecified curation/selection step that picks 6 candidate regulatory elements (the same 6 taken to luciferase), not fabrication. Overall: solid, partial reproduction with an explainable, well-documented deviation → yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.