Corpus 1,280 assessed · 1,181 scored · 646 reproduced ≥75 · 170 flagged ·∅ 74/100
← New search

Transcriptional landscape of repetitive elements in normal and cancer human cells.

BMC Genomics · 2014
L1 76/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
76/100
Reproducibility score
at the mean
vs. all fields · 1181 studies
🎯 Scores higher than 48% of all assessed papers rank 588 of 1181 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Partial, largely successful pipeline-derived reproduction using the authors' own RepEnrich tool (not a reimplementation) against hg19+RepeatMasker. Two datasets were fully processed: (1) GSE18184's Pol III-machinery ChIP-seq/RNA-seq subset (11/11 samples), which qualitatively reproduces the paper's Pol III-at-tRNA-genes positive-control biology (7/8 ChIP samples 1.35x-9.87x enriched vs Input; Brf2's lack of enrichment matches known promoter-type specificity) plus expected bulk-Alu non-enrichment and an RNA-seq-vs-ChIP srpRNA sanity check -- all graded within-tol since a simple fractional-ratio metric stands in for the paper's own GLM/FDR test. (2) ERP000550's 28-sample (14-pair) prostate tumor/normal cohort, which reproduces the qualitative core finding (L1 dominant among significantly changed subfamilies, ~100% tumor-overexpressed) but not the paper's precise magnitudes (346 vs 475 significant subfamilies; 89/120 vs 99/107 L1 subfamilies; ~1.4x vs paper's reported 2-4x fold-change) -- graded partial, attributable largely to substituting a paired Wilcoxon+BH-FDR test for the paper's original paired edgeR GLM (R/edgeR not confirmed available) plus no patient sub-grouping/length-binning. Two claims were explicitly NOT attempted and are not counted as reproduced or failed: the paper's Pol II ChIP-seq/LTR cancer-vs-normal and snRNA cross-cell-line-binding claims (no Pol II ChIP-seq samples downloaded in this room) and the TNM clinical-stage correlation (no linked per-patient clinical metadata available from public ENA records). A pre-existing, disclosed data-quality issue affects 66% of minor/rare repeat-family bowtie indices (filesystem race during setup) but does not affect any family used in the graded claims. No results were fabricated to avoid a drop; unattempted items are recorded as such rather than guessed.

💻 Code ↗ 🗄 Data: GSE20309

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-07
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-07
no human curator yet
Last updated
2026-08-07

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can genome-wide transcriptional regulation of repetitive elements be quantified despite the ambiguity of multi-mapping high-throughput sequencing reads, and are retrotransposons transcriptionally more active in cancer/transformed cells than in normal cells?

Core claims
  • RepEnrich, a computational method that uses all mapping reads (uniquely mapping plus multi-mapping reads assigned to repetitive element subfamily assemblies/pseudogenomes), quantifies genome-wide repetitive element enrichment method
  • The fractional counting strategy (reads shared across N subfamilies counted as 1/Ns) provides the least biased and least variable estimate of true repetitive element abundance and is therefore the RepEnrich default method
  • Many human LTR retrotransposons are transcriptionally active in a cell line-specific manner finding
  • Cancer-derived cell lines display increased RNA Polymerase II binding to retrotransposons compared with cell lines derived from normal tissue finding
  • L1 retrotransposon RNA expression is significantly higher in prostate tumors than in normal matched controls finding
  • Increased retrotransposon transcription in transformed cells may explain somatic retrotransposition events reported in several cancers mechanism
  • snRNAs show the most shared Pol II binding across cell lines, and tRNAs and 5S rRNA show ubiquitous Pol III binding, whereas transposable elements show cell-line-restricted polymerase binding finding
  • A subset of repetitive elements, predominantly tRNAs, is co-occupied by RNA Pol II and Pol III within the same cell line finding
Experimental setups
Assay System Perturbation Readout Platform
in silico simulated ChIP-seq and input data (Hidden Markov Model-based ChIP-seq read simulator developed by the authors) whole human chromosomes (e.g. chromosome 19) simulated enrichment of specified repetitive element families (L1, Alu, SVA) vs input counts per million mapping reads (CPM) / average log2CPM per repetitive element subfamily compared to known true abundance; differential enrichment calls custom HMM simulator; Bowtie1 alignment; RepEnrich; EdgeR GLM (negative binomial)
RNA Pol II ChIP-seq (antibody not distinguishing active/inactive enzyme) human cell lines: K562 (CML), HeLa (adenocarcinoma), GM12878 (EBV-immortalized lymphoblastoid), IMR-90 fibroblasts, HUVEC endothelial cells, peripheral blood-derived erythroblasts (PBDE) none (ChIP vs input comparison) log2 fold change and FDR of repetitive element subfamily read enrichment (ChIP vs input)
ChIP-seq for RNA Pol II phosphorylated on serine 2 (Pol II S2; active elongating enzyme) human cell lines (IMR-90, K562, HeLa, GM12878 panel) none (ChIP vs input) percent of repetitive element subfamilies with significant positive enrichment (FDR <0.05, Log2FC >0)
RNA Pol III ChIP-seq IMR-90 fibroblasts, K562, HeLa, GM12878 none (ChIP vs input) significant positive enrichment of repetitive element subfamilies; overlap with Pol II enrichment
ChIP-seq for TFIIIB (Pol III-associated transcription factor complex subunits) human cell lines none (ChIP vs input) repetitive element binding as supporting evidence of Pol III occupancy
ChIP-seq for chromatin activation and repression marks human cell lines (public ENCODE/GEO/ENA datasets) none repetitive element enrichment of histone marks
RNA-seq prostate tumor tissue from prostate cancer patients with normal-matched controls none (tumor vs normal-matched comparison) repetitive element / L1 retrotransposon RNA expression levels; overexpressed transposable elements
Genome browser inspection of uniquely mapping reads at individual loci human cell lines (tRNA and snRNA genes, including Pol III-transcribed U6) none visual confirmation of Pol II and Pol III co-occupancy at or near the same gene
Key results
  • Fractional counting deviated least from true abundance across all subfamilies and was closest to true abundance in multidimensional scaling of average log2CPM vectors
  • Unique counting over- or under-estimated true abundance with the greatest variance and consistently underestimated SINEs; it was most affected by read coverage and performed poorly at lower coverage
  • Total counting performed better than unique counting overall but consistently overestimated SINEs and SVA elements
  • R-squared of estimated vs true abundance was consistently close to 1 only for fractional counting, and varied widely between 0 and 1 for unique counting R-squared close to 1 (fractional) vs 0–1 range (unique)
  • In SVA-, L1- and Alu-enrichment simulations, fractional counting recovered the most benchmark differentially enriched elements and returned the fewest false positives
  • On real K562 RNA Pol II ChIP-seq data, fractional counting identified more Pol II-enriched repetitive elements than unique counting
  • 89 repetitive element subfamilies were co-occupied by Pol II and Pol III within the same cell line, the majority being tRNAs 89 subfamilies
  • Transposable elements rarely showed polymerase binding consistent across all cell lines, showing significant Pol II or Pol III binding in only one or a few cell lines, partly explained by higher expression in transformed vs normal cell lines
Key statistics
  • count 89 repetitive elements co-enriched for RNA Pol II and RNA Pol III within the same cell line (Pol II / Pol III co-occupancy overlap (Figure 3D))
  • other FDR <0.05 with Log2FC >0 (significance threshold for positive enrichment of repetitive element subfamilies in ChIP vs input GLM comparisons)
  • other ~55% of the human genome is repetitive DNA (more recent estimates as high as two-thirds) (background genome composition)
  • other ~45% of genomic DNA is transposable elements; ~10% is the four minor repeat categories (background genome composition)
  • other 20 M reads simulated (Figure 2) and 2 M reads in two L1-enrichment simulations on two different chromosomes (simulated ChIP-seq/input in triplicate; 50 bp single-end reads, chromosome 19)
  • other significantly higher levels of L1 retrotransposon RNA expression in prostate tumors vs normal-matched controls (no numeric value stated in text provided) (prostate cancer RNA-seq)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper describes a computational method (RepEnrich) for quantifying sequencing reads mapping to repetitive genomic elements, and validates it using simulated ChIP-seq data compared against a known ground truth. Differential enrichment between ChIP and input samples (and between conditions/cell lines) was assessed with a generalized linear model (GLM) fit to a negative binomial distribution, computed via EdgeR, with significance reported using a false discovery rate (FDR) threshold of 0.05 and results expressed as log2 fold-change (Log2FC). Method performance was additionally evaluated using R-squared values and multidimensional scaling (MDS) of Euclidean distances between estimated and true abundances.

Replicationmixed Sample sizeSimulated ChIP-seq/input data were generated in triplicate for the HMM-based simulations; real ChIP-seq/RNA-seq datasets were drawn from ENCODE, GEO, and ENA without an explicit stated per-sample replicate count in this excerpt GroupsChIP-seq vs input; cancer-derived vs normal cell lines; prostate tumor vs normal-matched tissue Pairingunclear Randomization/blindingnot stated Dispersionunclear Effect sizesyes Multiplicity correctionFalse discovery rate (FDR) thresholding (FDR < 0.05); specific correction procedure not named in this excerpt but is consistent with the Benjamini-Hochberg method commonly used by EdgeR
Statistical tests used
Test Applied to n Assumptions
Generalized linear model (GLM) fit to a negative binomial distribution, computed with EdgeR Differential enrichment between ChIP-seq and input samples, and across cell lines, for Pol II/Pol III binding and RNA-seq comparisons (Figure 3) not stated
R-squared (coefficient of determination) Agreement between RepEnrich-estimated abundance and true (simulated) abundance across counting strategies (Additional file 1: Figures S5A, S6A) simulations with 2M reads, triplicate simulated samples not stated
Multidimensional scaling (MDS) of Euclidean distances Comparison of unique, total, fractional, and true CPM vectors (Figure 2D, Additional file 1: Figure S4D) not stated
Approaches that could also have been used
  • Differential enrichment between ChIP/input and between conditions was modeled with a GLM fit to a negative binomial distribution via EdgeR.
    Could also: DESeq2's Wald or likelihood-ratio test, which also models count data with a negative binomial distribution — DESeq2 uses a related but distinct dispersion-shrinkage approach and is another widely used standard for count-based differential enrichment/expression analysis; comparing results across tools can illustrate robustness to modeling choices.
  • Significance across many repetitive element subfamilies was reported using an FDR < 0.05 threshold.
    Could also: Explicitly naming and reporting the multiple-testing correction procedure (e.g., Benjamini-Hochberg) alongside the FDR cutoff, or reporting q-values directly — Making the correction method explicit alongside the threshold helps readers assess how family-wise or false-discovery error was controlled across the large number of repetitive element comparisons.
  • Agreement between RepEnrich-estimated abundance and true simulated abundance was assessed using R-squared from scatterplots against the y = x line.
    Could also: A Bland-Altman plot or concordance correlation coefficient (CCC) — These approaches directly quantify agreement and systematic bias between two measurements of the same quantity, which can complement R-squared (a measure of correlation but not necessarily agreement).
  • The similarity of unique, total, fractional, and true CPM vectors was visualized using multidimensional scaling (MDS) of Euclidean distances.
    Could also: Principal component analysis (PCA) — PCA is a commonly used complementary ordination method for visualizing sample or method similarity and can be used alongside MDS to check consistency of the observed clustering pattern.
  • Higher L1 retrotransposon RNA expression was reported in prostate tumors compared to normal-matched controls.
    Could also: A paired statistical test such as the Wilcoxon signed-rank test or a paired t-test, given the matched-sample design — Explicitly using a paired test (and stating it as such) leverages the matched tumor/normal structure of the samples and is a standard approach for matched-pair comparisons; the excerpt does not specify which test, if any, was used for this particular comparison.
  • Method performance differences (fractional vs. unique vs. total counting) were reported primarily via fold-change and R-squared summaries.
    Could also: Reporting confidence intervals or standard errors around the estimated abundance/fold-change values — Interval estimates convey the precision of the estimates in addition to their point values, which can add information for readers evaluating the counting strategies.
Software: EdgeR · Bowtie1 (aligner) · RepeatMasker (annotation)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

polIII_trna_enrichment_gse18184
Reported
Pol III machinery (Pol3/Brf1/Brf2/TFIIIC) ChIP-seq is enriched at tRNA genes vs Input, as part of the paper's genome-wide GLM/negative-binomial per-subfamily significance test (Fig 3, FDR<0.05, log2FC>0); used here as the paper's expected-biology positive control since it does not restate a single per-sample tRNA-class number.
Reproduced
RepEnrich tRNA-class fractional-read enrichment vs Input: GSM509047(Pol3/HEK293T)=7.77x, GSM509048(Pol3/HFF)=2.02x, GSM509049(Brf1/HeLa)=4.01x, GSM509050(Brf2/HeLa)=0.74x, GSM509052(Pol3/HeLa Rep1)=2.89x, GSM509053(Pol3/HeLa Rep2)=9.87x, GSM509056(TFIIIC/HeLa)=3.21x, GSM509057(Pol3/Jurkat)=1.35x. 7/8 samples show clear positive enrichment (1.35x-9.87x); Brf2 shows none (0.74x), which matches known Pol III promoter-type biology (Brf2-TFIIIB is specific to TATA-containing type-3 promoters such as U6/7SL, not the type-2 internal promoters used by tRNA genes, which recruit Brf1-TFIIIB instead) rather than indicating a reproduction failure.
within tolerance
sine_aluY_bulk_enrichment_gse18184
Reported
Paper's introduction/background states Alu elements are 'thought to be transcribed by RNA polymerase III' when retrotransposons are expressed -- motivating context, not a specific quantitative bulk-enrichment claim for aggregated genomic Alu copies.
Reproduced
SINE-class and AluY-repname fractional enrichment vs Input clusters near 1.0 (0.80x-1.11x) across all 8 ChIP samples -- no detectable bulk enrichment at this aggregation level. Expected: RepEnrich sums reads across all ~1.1M genomic Alu copies, while only a small, specific subset of loci are actively Pol III-occupied at any time, diluting locus-level enrichment below detection when averaged over the whole subfamily (consistent with independent literature, e.g. PMC4333407, which needed locus-resolved methods to find the rare active Alu loci).
within tolerance
srpRNA_positive_control_gse18184
Reported
Not a specific quantitative paper claim; used here as an internal pipeline sanity check that RepEnrich correctly captures a known highly-expressed Pol III ncRNA product (7SL/SRP RNA).
Reproduced
srpRNA-class fractional read share: Input(HeLa)=0.069%, CappedRNAseq(HeLa)=0.305%, TotalRNAseq(HeLa)=1.768% -- 4x-25x higher in RNA-seq than in Input/ChIP DNA, as expected since RNA-seq directly captures the abundant mature transcript while ChIP/Input measure genomic DNA representation.
within tolerance
l1_differential_expression_tumor_normal_erp000550
Reported
Fig 5A/5C, Fig 6: 475 repeat subfamilies significant at FDR<0.05 (paired edgeR GLM with patient as covariate, 14 tumor/normal prostate pairs); 99/107 tested L1 subfamilies significant, 97/99 (98%) of those tumor-overexpressed; 2-4 fold change across element length in 'patient group 1' (Fig 6); TNM clinical stage correlates with L1-overexpression grouping (p=0.04, Mann-Whitney U).
Reproduced
Using RepEnrich.py (hg19+RepeatMasker) + CPM normalization + paired Wilcoxon signed-rank test + Benjamini-Hochberg FDR (all 14 pairs, all subfamilies, no patient sub-grouping or length-binning): 1363/3861 subfamilies testable; 346 significant at FDR<0.05 (vs paper's 475, ~73%); 120 L1 subfamilies testable, 89 significant (vs paper's 99/107); of significant L1 subfamilies, 89/89 (100%) tumor-overexpressed (vs paper's 97/99=98% -- qualitatively matches); median L1 tumor/normal fold-change across all pairs ~1.4x (below the paper's reported 2-4x range, though not a strict apples-to-apples comparison since the paper's figure is for a specific patient subgroup with length-binning, not performed here).
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 76/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

This is one of the stronger reproduction cases: the authors' own RepEnrich code was cloned and run against the same hg19+RepeatMasker reference, and both central claims survive — L1 subfamilies are significantly tumor-overexpressed in prostate (89/89 = 100% tumor-directed vs the paper's 97/99 = 98%) and Pol III machinery is enriched at tRNA genes in 7/8 ChIP samples (1.35x-9.87x), with the lone exception (Brf2 = 0.74x) matching known type-3-promoter biology rather than signalling failure. The quantitative shortfalls — 346 vs 475 significant subfamilies, 89 vs 99 significant L1 subfamilies, median fold change ~1.4x vs the paper's 2-4x — sit squarely on our side: a rank-based Wilcoxon was substituted for the paper's paired edgeR NB-GLM, no patient sub-grouping or length-binning was performed, and a setup race corrupted 66% of minor-family indices, cutting the testable universe to 1363/3861. Two further claims (Pol II/LTR cancer-vs-normal; the TNM p=0.04 correlation) went untested purely because the Pol II GSMs were not downloaded and clinical metadata is not in the anonymized ENA records — an availability/scope gap, not an authors' defect. Overall: solid, direction- and sign-consistent reproduction with fully explainable deviations, so yellow rather than green on severity and overall, but no derivability or core-claim concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.