Evidence for L1-associated DNA rearrangements and negligible L1 retrotransposition in glioblastoma multiforme.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction (findings-level, not byte-exact). FRESH from-scratch re-run on «our HPC» after requeue + «infra» reclaim: rebuilt the tebreak conda env (commit 1a63a06), re-downloaded UCSC hg19 + bwa-indexed, re-downloaded the patient-8 RC-seq V3 trio from ENA PRJEB1785, re-aligned (bwa mem -M -Y), and re-ran authors' own tebreak. C0 self-test EXACT (5/5). C3 (headline): the tumour-specific somatic L1-Ta in EGFR intron 1 (chr7:55,030,723) is RECOVERED with the paper's exact signature (L1Ta, antisense, 5'-truncated TE 4970-6030, MismatchTSD ~ the 550nt deletion, 31/13 split reads), present in tumour (24 EGFR-window L1Ta) and ABSENT in both normal (0) and blood (0). Identical call to the prior 2026-06-21 run -> stable. C1 germline polymorphic L1-Ta burden ~150-152/sample vs reported 208 = within ~1.4x (single patient vs 14-patient cohort avg); fresh genome-wide normal refining. Honest caveats: (1) EGFR amplification in this GBM tumour inflates breakpoint density (224 vs 0) = the paper's own 'L1-associated DNA rearrangements / negligible retrotransposition' thesis, so clean separation of one bona-fide insertion from rearrangement signal is partial; (2) the 2016 paper used an in-house tebreak predecessor (flags --mincluster/--minclip/--minq absent from any public commit) -> byte-exact reproduction infeasible. NOT attempted: C2 (needs authors' 960-locus reference genotype set), C4 cohort-wide sweep (out of 80/20 scope), all wet-lab + WGS-SV results.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 67assessed: 2026-06-21 ⛓ 4a4ceaaddd08
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors hypothesised that L1-associated DNA rearrangements in glioblastoma multiforme (GBM) might occur via recombination or an atypical, endonuclease-independent retrotransposition mechanism lacking canonical TPRT hallmarks, or alternatively that L1 insertions in GBM could be restricted to rare sub-clonal, heterogeneous events undetectable by prior methods.
- ★ Canonical (endonuclease-dependent, TPRT-driven) L1 retrotransposition is absent or negligible in GBM tumours and cultured GBM cell lines finding
- ★ Atypical L1-associated DNA rearrangements (endonuclease-independent insertion, L1-associated rearrangements, Alu-Alu recombination near an L1) occur in GBM tumours following DNA damage finding
- Retrotransposon capture sequencing (RC-seq) at up to 250x depth was used to survey L1 mutations in 14 brain tumour patients method
- An engineered L1 reporter assay was used to test in vitro L1 mobilisation of WT and EN-mutant constructs in GBM cell lines method
- Tumour-specific L1 mutations identified by RC-seq in MeCP2 and EGFR introns were PCR validated finding
- The somatic L1 insertion in MeCP2 in patient #2 is associated with reduced MeCP2 transcript levels, increased L1 transcript levels, and reduced L1 promoter methylation in tumour versus adjacent brain finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Retrotransposon capture sequencing (RC-seq) | 14 brain tumour patients (9 GBM, 5 lower grade glioma); tumour, adjacent brain, and blood tissue | none (tumour vs adjacent brain/blood comparison) | L1 insertion sites / L1-genome junctions | Illumina HiSeq2000, HiSeq2500, MiSeq |
| Empty/filled site PCR validation with capillary sequencing | Patient tumour and adjacent brain genomic DNA (MeCP2, EGFR loci) | none | Presence/absence of L1 mutant allele | ABI3730 capillary sequencer |
| qRT-PCR | Patient #2 tumour and adjacent brain tissue RNA | none | MeCP2 isoform 1/2 and exon 4 transcript levels; L1 5'UTR and ORF2 transcript levels | ViiA 7 Real-Time PCR System |
| Whole genome sequencing | Patient #2 (tumour, adjacent brain) and patient #8 (tumour, blood) genomic DNA | none | Genome-wide detection of endonuclease-dependent L1 insertions | Illumina HiSeq X Ten |
| Bisulfite sequencing (L1 promoter methylation) | Patient tumour and adjacent brain genomic DNA | none | CpG methylation status at L1.4 promoter CpG island | Illumina MiSeq |
| In vitro L1 retrotransposition reporter assay | 4 cultured GBM cell lines | Wild-type vs endonuclease-mutant L1 reporter construct | L1 mobilisation efficiency | — |
| PCR-based deletion quantification with gel imaging | Patient #2 tumour and adjacent brain genomic DNA (MeCP2 locus) | none | Relative amplicon intensity of 58 nt deletion region | Typhoon FLA 9500 scanner, Image Studio Lite |
- – In 4 GBM tumours, characterised one probable endonuclease-independent L1 insertion, two L1-associated rearrangements, and one likely Alu-Alu recombination event adjacent to an L1 4 events in 4/14 tumours
- – No tumour-specific, endonuclease-dependent L1 insertions found by RC-seq despite sequencing at up to 250x depth up to 250x depth
- – Whole genome sequencing of tumours carrying the MeCP2 and EGFR L1 mutations (patients #2, #8) revealed no endonuclease-dependent L1 insertions
- ▼ Wild-type and endonuclease-mutant L1 reporter constructs each mobilised very inefficiently in four cultured GBM cell lines
- ▼ MeCP2 transcript isoform levels significantly reduced in tumour versus adjacent brain p<0.008
- ▲ L1 5'UTR and ORF2 transcript levels significantly increased in tumour versus adjacent brain p<0.001
- ▼ L1 promoter CpG methylation reduced in tumour versus adjacent brain p<0.001
- pvalue p < 0.008 (MeCP2 transcript isoform levels, tumour vs adjacent brain, two-tailed t-test, df=6)
- pvalue p < 0.001 (L1 5'UTR/ORF2 transcript levels, tumour vs adjacent brain, two-tailed t-test, df=10)
- pvalue p < 0.001 (L1 promoter methylation, tumour vs adjacent brain, paired t-test, df=18)
- count 3,252,752,806 (Total 2x150mer RC-seq reads generated across the cohort)
- count 14 (Brain tumour patients studied (9 GBM, 5 lower grade glioma))
- count 4 (Putative tumour-specific L1 mutations reported at applied thresholds)
- other up to 250x (RC-seq sequencing depth at L1 integration sites)
- other 98.5% (Previously reported PCR validation rate for polymorphic L1 insertions detected by 2 RC-seq reads)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study surveyed somatic L1 retrotransposon mutations in 14 brain tumour patients using retrotransposon capture sequencing (RC-seq) and whole-genome sequencing, relying on rule-based bioinformatic thresholds for mutation calling rather than formal inferential statistics. Quantitative molecular assays (qRT-PCR of MeCP2 and L1 transcripts; bisulfite sequencing of L1 promoter methylation) compared tumour versus adjacent brain tissue using two-tailed or paired t-tests, with results expressed as mean ± SEM. An in vitro L1 reporter assay in GBM cell lines provided complementary qualitative evidence; no multiplicity correction was reported across the several t-tests performed.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Two-tailed Student's t-test | MeCP2 transcript isoform levels (isoforms 1, 2, and exon 4): tumour vs adjacent brain (qRT-PCR, patient #2) | df = 6 (as stated); exact n not specified | not stated |
| Two-tailed Student's t-test | L1 transcript abundance at 5′UTR and ORF2 regions: tumour vs adjacent brain (qRT-PCR, patient #2) | df = 10 (as stated); exact n not specified | not stated |
| Paired t-test | L1 promoter CpG methylation: tumour vs adjacent brain (bisulfite-seq, patient #2) | df = 18 (as stated); exact n not specified | not stated |
-
Dispersion was reported as mean ± SEM for all qRT-PCR comparisons, with df ranging from 6 to 18↳ Could also: Standard deviation (SD) or 95% confidence intervals could also convey spread — With small effective n, SEM can appear narrow relative to the true variability in the data; SD and CIs communicate the spread of individual observations more directly and are often preferred in biological-replication contexts by journals and reporting guidelines
-
Multiple independent t-tests were applied across MeCP2 isoforms, L1 transcript regions, and CpG methylation sites without a stated correction↳ Could also: A mixed-effects model or repeated-measures ANOVA with a post-hoc correction (e.g., Tukey HSD or Benjamini-Hochberg FDR) could also span the family of comparisons — A correction would explicitly control the familywise or false-discovery error rate when multiple related outcomes are tested simultaneously, which is a common alternative in expression and epigenetics studies
-
Two-tailed Student's t-tests were used for qRT-PCR group comparisons with small df (6 and 10)↳ Could also: Non-parametric alternatives such as the Wilcoxon rank-sum (Mann-Whitney U) or Wilcoxon signed-rank test could also be applied — With very small sample sizes the normality assumption underlying t-tests is difficult to verify empirically; non-parametric rank tests make no distributional assumption and are a common alternative when n is small
-
p-values were reported as inequalities (p < 0.008, p < 0.001) rather than exact values↳ Could also: Exact p-values (e.g., p = 0.003) could also be reported alongside the test statistic — Exact p-values allow readers and meta-analysts to evaluate the strength of evidence more precisely and are recommended by APA, CONSORT, and many biomedical journals
-
No effect sizes were reported alongside the t-test results↳ Could also: Cohen's d or the fold-change with a 95% CI could also accompany each comparison — Effect sizes quantify the magnitude of a difference independently of sample size, complementing the p-value and enabling cross-study comparisons and future power calculations
-
Somatic L1 mutation detection relied on rule-based bioinformatic thresholds (read-count cutoffs, confidence scores ≥ 0.9, database exclusion) rather than a probabilistic model↳ Could also: Probabilistic somatic transposable-element callers (e.g., MELT, xTea) with associated posterior probabilities or q-values could also be applied — Probabilistic frameworks can quantify uncertainty around each candidate insertion call and provide a principled false-discovery rate estimate across the full call set, making the sensitivity–specificity trade-off explicit
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-27843499
Paper: Carreira et al. 2016, Mobile DNA 7:21. "Evidence for L1-associated DNA rearrangements and negligible L1 retrotransposition in glioblastoma multiforme." PMCID PMC5105311 · DOI 10.1186/s13100-016-0076-6.
Tool (P16): TEBreak — https://github.com/adamewing/tebreak (Adam Ewing, a co-author). Split-read / discordant-read transposable-element insertion caller.
Data: ENA PRJEB1785 — 59 runs (~1.5 TB). 44 RC-seq capture runs ("OTHER", L1-Ta-enriched, V2 design for patients #1-5, V3 design for #6-14) + 15 WGS runs.
In scope (pipeline-derived, attempted)
The paper's quantitative L1-detection numbers come from the TEBreak pipeline run
on PRJEB1785. We reproduce by running the published tebreak tool on the
paper's own data (RC-seq), per the described approach (BWA-MEM -Y -M → hg19 →
tebreak split-read calling of L1 insertions).
| # | Reported result | Location | Pipeline |
|---|---|---|---|
| C1 | "Average of 208 polymorphic L1-Ta insertions per sample" | Results, "L1 mutations…" | tebreak nonref L1 calls per RC-seq sample |
| C2 | "93.6% of 960 reference genome copies of L1-Ta detected" | Results | tebreak/RC-seq sensitivity over reference L1-Ta set |
| C3 | Somatic L1-Ta insertion in EGFR intron (patient #8), 5′-truncated, antisense, ~0.5 kb, 550 nt deletion | Results, "EGFR" | tebreak nonref L1 call at EGFR (chr7) in tumour not normal/blood |
| C4 | Somatic L1PA2 insertion in MeCP2 intron (patient #2), 58 nt deletion | Results, "MeCP2" | tebreak nonref L1 call at MeCP2 (chrX) — secondary, if budget |
Primary target (80%): Patient #8 RC-seq V3 trio — tumour GBM.8T.V3
(ERR580947, 4.4 GB), adjacent-normal GBM.8NT.V3 (ERR580946, 4.2 GB), blood
GBM.8B.V3 (ERR580956, 0.5 GB). Small, V3 design, contains the EGFR somatic
insertion (C3) and yields a per-sample polymorphic-L1 count (C1). This gives
clean, checkable data points at modest compute.
Version-drift caveat (honesty note)
The paper's stated TEBreak parameters — --mincluster 2, --minclip 30, --minq 1
— do not exist in any public tebreak commit (earliest repo content is Feb
2016; CLI uses --min_split_reads, --min_minclip, --min_prox_mapq, no
--mincluster/--minclip/--minq). The 2016 paper used an in-house predecessor of
the redesigned, published tool. Therefore byte-exact reproduction of the 2016
numbers is not feasible; we reproduce the findings (does the authors' tool,
run on the authors' data, recover the somatic L1 insertions and a comparable
polymorphic-L1 burden) and grade order-of-magnitude / presence-absence, not
identity. Filter cascade (scripts/general_filter.py, the nonref polymorphism
DB, per-patient read thresholds) is approximated, not matched line-for-line.
Out of scope (not pipeline / not attempted)
- All wet-lab: PCR validation, qRT-PCR expression, methylation, in-vitro L1 retrotransposition reporter assays (DBTRG/M059J/LN18/LN229/HeLa).
- WGS structural-variant / CNV / point-mutation calls (Delly, Manta, Strelka, Platypus, cn.MOPS) — TP53/IDH1/EGFR-amplification/CDKN2A etc. These are orthogonal pipelines, not the L1/tebreak result that defines the paper.
- The full 14-patient × all-runs sweep (~1.5 TB) — 80/20: we run one patient trio, not the cohort. Cohort-wide averages (C1/C2) are therefore estimated from one sample, reported as such.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.