A methyl-sensitive element induces bidirectional transcription in TATA-less CpG island-associated promoters.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL. Mahpour et al. 2018 (PLoS One) is predominantly a wet-lab paper; its computational layer is thin. The ONLY deposited code (github.com/AminMahpour/pyRibbon, commit fe653da) is a 1.1KB DNA-ribbon VISUALISER used solely for Fig 2F (Sanger traces of CRISPR clones). The actual quantitative claims come from third-party tools (Homer, Bowtie/Bowtie2, 'Wigman') NOT in the repo. REPRODUCED 1:1 (grade exact, provisional): the deposited Ribbon.py runs under a pinned conda env (pycairo 1.29.0 / cairo 1.18.4) on «our HPC» and renders its shipped test.fasta DETERMINISTICALLY - two independent runs gave a byte-identical PDF (357fab48...) and pixel-identical 93x30 raster (713abeb4..., AE=0), geometry exactly matching Ribbon.py's spec. This confirms the deposited artifact is functional+deterministic, but it reproduces the TOOL on its OWN example, NOT the paper's Fig 2F (whose input - CRISPR-clone Sanger reads - is not deposited); the shipped out.png is a different 178x30 README sample, so no pixel-match to it is claimed. NOT ATTEMPTED (described, but not exactly reproducible from shipped materials): the headline numbers motif7=1408 / motif10=413 copies (C2), >400 DNase-footprint coincidences (C3), 93%/7% promoter split (C4) - all require the discovered Homer motif PWM matrices + log-odds thresholds, which are absent from paper/SI/repo (class docs_insufficient); and the conservation analysis (C6) names a tool 'Wigman' that is not resolvable (class env_unresolvable). We deliberately did NOT run a guessed-PWM scan, which would compare a different motif against the authors' numbers and risk a misleading match/mismatch. No fabrication observed - the unverifiable numbers are an auditability gap, not evidence of fabrication. Accession note: GSE81322 = Brocks et al. 2017 (PMID 28604729) CAGE-seq, legitimately reused and correctly cited.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 63assessed: 2026-06-16 ⛓ 21f1d2db3fb7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors hypothesize that a specific, independently-acting DNA promoter element drives transcription in TATA-less CpG island-associated promoters; they test whether the 'CGCG element' (consensus TCTCGCGAGA) is sufficient to drive bidirectional transcription and whether its methylation abrogates this activity.
- ★ The CGCG element (consensus TCTCGCGAGA) is a conserved 10-bp motif enriched in TATA-less, CpG island-associated promoters of ribosomal protein and housekeeping genes. finding
- ★ The CGCG element is sufficient on its own to act as a promoter element driving bidirectional transcription via divergent start sites. finding
- ★ Methylation of the CGCG element abrogates its associated promoter activity. mechanism
- ★ The CGCG element employs RNA Pol II to activate gene expression and resides in DNase-accessible regions. mechanism
- ★ When coincident with a TATA-box, transcription becomes directional but remains CGCG-dependent. finding
- Novel bidirectional luciferase (LuBiDi) and fluorescence (pmCGFP) reporter constructs were developed to assay bidirectional promoter activity. method
- The CGCG element is evolutionarily conserved among vertebrates and enriched in promoters of housekeeping genes controlling RNA metabolism and translation, and of long non-coding RNAs. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Motif discovery (Homer findMotifgenomewide) | K562 cell line genome (hg38) | none | enriched motifs in DNase-accessible CpG islands | Homer software suite |
| Luciferase reporter assay (Dual Luciferase) | HEK293T cells | 1-3 copies of CGCG element cloned into pGL2-basic | firefly/Renilla luciferase bioluminescence | pGL2-basic (Promega), pRL-TK |
| Luciferase reporter assay with transcription inhibition | HEK293T cells | alpha-amanitin 5 ug/ml vs PBS; SV40 or CGCG construct | firefly/Renilla luciferase activity | pGL2-pro; Santa Cruz alpha-amanitin |
| Bidirectional luciferase reporter assay (LuBiDi) | HEK293T cells | 0-3 copies of TCTCGCGAGA | bidirectional firefly and Renilla luciferase activity | pRL-Null/pGL2-Basic derived; pSELECT-zeo-SEAP normalization |
| Bidirectional fluorescence reporter / fluorescent microscopy & live imaging | HEK293T cells | 0-2 copies of TCTCGCGAGA in pmCGFP | eGFP and mCherry fluorescence | Nikon Eclipse TE2000-E; CMV-BFP control |
| In vitro CpG methylation reporter assay | HEK293T cells / pCpGfree-basic-Lucia plasmid | M.SssI methyltransferase + SAM vs mock | Lucia luciferase promoter activity; BstUI/NheI digestion | pCpGfree-basic-Lucia (Invivogen), M.SssI (NEB) |
| qRT-PCR | HEK293T cells | 0,1,2,4 copies of CGCG element in LuBiDi | transcript levels normalized to HPRT1 | iTaq SYBR Green, ABI-7900 |
| Chromatin immunoprecipitation (ChIP-qPCR) | HEK293T cells | reporter DNA transfection | RNA Pol II occupancy at reporter promoter | anti-Pol II (Santa Cruz N-10); Bioruptor |
- ▲ CGCG element drives luciferase reporter activity from a promoterless construct, increasing with copy number
- ▲ CGCG element drives bidirectional transcription via divergent TSS in LuBiDi/pmCGFP reporters
- ▼ Methylation of the CGCG element abolishes promoter/reporter activity
- ▼ alpha-amanitin treatment reduces CGCG-driven luciferase activity, indicating RNA Pol II dependence
- ▲ Pol II ChIP shows enrichment at CGCG-containing reporter promoter
- – CGCG element is associated with bidirectional transcription only in CGI-associated promoter contexts per GRO-Cap and Start-seq analysis
- count ~30k human CGIs analyzed (genome-wide CGI motif discovery)
- count ~50 percent of human promoters associated with a CpG island (prior estimate of CGI-associated promoters)
- other p<0.05 (*), p<0.01 (**), p<0.001 (***), p<0.0001 (****) (significance thresholds, Student t-test)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combined genome-wide bioinformatic analyses (motif discovery, genomic annotation, conservation, and GRO-Cap/Start-seq analysis) with cell-based reporter assays (luciferase, fluorescence, qRT-PCR, ChIP) to characterize the CGCG promoter element. Cell-based quantitative comparisons were analyzed with Student's t-test in GraphPad Prism 7, with significance reported as threshold symbols. Gene Ontology enrichment used the GOrilla platform with a defined ratio-of-ratios score. Error bars throughout represent standard deviation.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Student's t-test (direction and pairing not specified) | Luciferase reporter assays and other cell-based quantitative comparisons (stated as default for all analyses unless noted otherwise) | — | not stated |
| Gene Ontology enrichment score (b/n)/(B/N) | GO enrichment analysis of genes with CGCG-element promoters vs CpG island-associated background gene set | — | not stated |
| FIMO position-weight-matrix scan with p-value threshold 0.001 | Identification of CGCG motif occurrences within ±1 kbp windows around TSSs for Start-seq analysis | — | na |
-
Comparisons across multiple CGCG-copy-number conditions (0, 1, 2, 3, 4 copies) were each handled with individual Student's t-tests↳ Could also: One-way ANOVA followed by a post-hoc test (e.g., Dunnett's test vs control, or Tukey HSD for all pairs) — An ANOVA framework tests all groups jointly and a post-hoc correction accounts for the family-wise error rate that accumulates when many pairwise comparisons are made from a single experiment
-
P-values are reported as threshold symbols (*, **, ***, ****) rather than exact values↳ Could also: Report exact p-values (e.g., p = 0.018) alongside or instead of symbols — Exact p-values allow readers to gauge the strength of evidence directly, are required by many journals, and facilitate downstream meta-analysis or comparison across studies
-
Student's t-test was applied to reporter assay data without a stated assessment of normality or variance equality↳ Could also: A non-parametric alternative such as the Mann-Whitney U (Wilcoxon rank-sum) test, or a Welch's t-test if variances differ — With the small sample sizes typical of transient-transfection experiments, distributional assumptions are difficult to verify; non-parametric tests make no normality assumption, and Welch's correction handles unequal variances
-
Dispersion is reported as standard deviation (SD)↳ Could also: Report 95% confidence intervals alongside or instead of SD — Confidence intervals convey both variability and estimation precision, making effect magnitudes easier to interpret, and are increasingly preferred over SD for small-n cell-biology data
-
No multiplicity correction is described for the many statistical tests performed across the paper's reporter and genomic comparisons↳ Could also: Benjamini-Hochberg false discovery rate (FDR) correction or Bonferroni correction applied across the family of comparisons — Applying a correction such as BH-FDR limits the expected proportion of false-positive findings when many tests are evaluated, which is especially relevant for GO enrichment analysis and multi-condition reporter screens
-
Effect sizes are not reported alongside significance tests↳ Could also: Report fold-change and/or Cohen's d (or partial eta-squared if ANOVA is used) alongside p-values — Effect sizes communicate the practical magnitude of differences independently of sample size, complementing p-values which reflect both effect size and n
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
16 downstream papers · 3 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Genome-wide mapping of autonomous promoter activity... 2017 · 124 cites
- Nascent RNA sequencing analysis provides insights in... 2018 · 67 cites
- Identification of the human DPR core promoter elemen... 2020 · 48 cites
- Bidirectional transcription initiation marks accessi... 2017 · 46 cites
- Transcriptionally active enhancers in human cancer c... 2021 · 42 cites
- Identification of active miRNA promoters from nuclea... 2017 · 32 cites
- DNMT and HDAC inhibitors induce cryptic transcriptio... 2017 · 244 cites
- DNMT and HDAC inhibition induces immunogenic neoanti... 2023 · 60 cites
- The landscape of cryptic antisense transcription in... 2023 · 9 cites
- Pan-cancer transcriptome analysis reveals widespread... 2024 · 4 cites
- B-type lamins maintain transcriptional homeostasis b... 2026 · 0 cites
- Bidirectional Transcription Arises from Two Distinct... 2015 · 199 cites
- Exon-Mediated Activation of Transcription Starts. 2019 · 91 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — pmid-30332484
Paper: Mahpour A, Scruggs BS, Smiraglia D, Ouchi T, Gelman IH. A methyl-sensitive element induces bidirectional transcription in TATA-less CpG island-associated promoters. PLoS One 2018. PMID 30332484 · PMCID PMC6192621 · DOI 10.1371/journal.pone.0205608
Designated code: https://github.com/AminMahpour/pyRibbon (GPL-3, last push 2017-03-22,
commit fe653da88a61845b86f27164527dd6d5cb631981, 17 KB).
Designated data (per brief): GEO GSE81322.
1. What pyRibbon actually is
Ribbon.py (1166 bytes) is a visualisation helper, not an analysis pipeline.
It reads a FASTA sequence and draws one 1-px colour-coded rectangle per nucleotide
into a PDF (A=indigo, T=yellow, G=green, C=red), with tick marks every 50 bp. It
produces no quantitative output.
Per the paper's Methods: "Ribbon sequences were produced using the pyRibbon software which we deposited in https://github.com/AminMahpour/pyRibbon/." It was used to render Figure 2F — a visualisation of Sanger sequencing of CRISPR-edited genomic deletions in the authors' own clones.
2. Computational results reported in the paper, and their pipelines
| # | Reported result (Fig/Table) | Pipeline / tool named | In repo? | Reproducible? |
|---|---|---|---|---|
| C1 | Fig 2F ribbon visualisation of CRISPR-clone Sanger traces | pyRibbon (the repo) | ✅ yes | input data (clone Sanger reads) NOT deposited → figure not reproducible; only the tool is |
| C2 | motif 7 = 1408 copies, motif 10 = 413 copies in human (hg38) CGIs | Homer findMotifsGenomeWide + scanMotifGenomeWide v4.8 |
❌ no | motif PWM matrices + log-odds thresholds NOT provided in paper or SI → not exactly reproducible |
| C3 | >400 motif-10 incidences coincide with DNase-seq footprints (S1A) | Homer annotatePeaks + Bedtools |
❌ no | depends on C2 motif + ambiguous DNase accession (paper cites ENCFF867JRG for both WGBS and DNase — likely a typo) |
| C4 | 93% unidirectional vs 7% bidirectional promoters (Table 2) | Homer annotatePeaks |
❌ no | depends on C2 motif set |
| C5 | CGCG element within 50 bp of TSS (Fig 1B, 5E) | Homer + Start-seq (GSE62151) Bowtie 0.12.8 mm9 | ❌ no | depends on C2 motif set |
| C6 | PhyloP/PhastCons conservation in ±50 bp windows | "Wigman" software | ❌ no | "Wigman" is not identifiable as a public/citable tool → env_unresolvable |
| — | CAGE-seq alignment (GSE81322, hg38, Bowtie2) | Bowtie2 | ❌ no | GSE81322 is reused data from a different study (see §4); no specific pinnable derived number reported from it |
3. In-scope vs out-of-scope decision
In scope (attempted):
- C1 — pyRibbon functional/deterministic reproduction. This is the only
designated-code artifact. We run the deposited
Ribbon.pyon its shipped input (test.fasta) on «our HPC» and verify it (a) executes under a pinned env and (b) renders deterministically (two independent runs → pixel-identical raster) with the geometry the script specifies. This is a genuine 1:1 reproduction of the deposited code's behaviour. It does not reproduce the paper's actual Fig 2F, because the figure's input (CRISPR-clone Sanger reads) is not public.
Out of scope / not exactly reproducible (documented, not fabricated):
- C2–C5 (Homer motif counts/annotations). The headline quantitative numbers
(1408 / 413 copies, 93%/7%) require the exact discovered motif PWM matrices
and log-odds thresholds, which are not provided in the paper, SI, or any
repo. Homer v4.8 de-novo motif numbering ("#7", "#10") is run-order-dependent and
not reproducible from the text alone. Re-running de-novo discovery with a guessed
PWM/threshold would compare an apples-to-oranges motif against the authors'
numbers and could masquerade as a (mis)match — we therefore do not report a
fabricated comparison number for these.
drop_reasonclass for this slice:docs_insufficient(motif matrices/thresholds ab
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is predominantly a wet-lab paper whose only deposited code, pyRibbon, reproduced 1:1 and deterministically on its own test.fasta (byte-identical PDF, AE=0 raster) — but that artifact is just the Fig 2F ribbon visualiser, not the analysis behind the headline numbers. Every quantitative claim (motif7=1408/motif10=413, Table 2's 93%/7%, >400 DNase coincidences, conservation) is not derivable from shared materials because the Homer PWM matrices + thresholds were never deposited and the named tool 'Wigman' is unresolvable. The gap is on the authors'/documentation side (auditability), with no evidence of fabrication and no measured deviation. Verdict: partial — solid on the one deposited tool, but the central computational claims are unverifiable as published.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.