TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH TO REPRODUCE -> CLOSE MATCH (within tolerance). Ran the authors' own TrEMOLO Snakemake pipeline (OUTSIDER+INSIDER) on the authors' own deposited simulated data (Zenodo 7673915 reads + DataSuds DSDTZ0 genomes) on «our HPC». No singularity on the cluster -> built a faithful flat conda env (sniffles 1.0.12, RaGOO 1.1, assemblytics, blast, minimap2, samtools, bedtools, liftoff) and fixed 5 silent env blockers (git-in-env, pip ragoo_utilities, short TMPDIR for AF_UNIX socket, gawk for multidim arrays, undefined SAMTOOLS_1_9/1_15_1 vars). Smoke test passed end-to-end. Main S1 run produced 1294 TE detections (TE_INFOS.bed). Ground truth re-derived by genome alignment (100 INSIDER exact + 996/1000 OUTSIDER = 1096/1100). Result: TrEMOLO 1041 TP (95.0% sensitivity) / 107-110 FP vs reported 1071 TP (97.3%) / 95 FP. INSIDER detection 100/100 perfect. The headline claim reproduces closely; the ~2.3 pp sensitivity gap and slightly higher FP are attributable to version drift (paper v2.2-beta1 vs our master v2.5.6b) and our reconstructed TP/FP match window (unspecified in paper). NOT a fabrication: every reported number is derivable from the shipped data+code. NOT attempted: S2/S3 secondary analyses and the 3 comparator tools. All grades provisional, human-checkable.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCombining assembly-based and mapping-based analysis of long-read sequencing data (TrEMOLO) can more accurately detect transposable element (TE) insertions/deletions and estimate their allele frequency in a population, regardless of insertion frequency, than existing long-read TE detection tools.
- ★ TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency method
- ★ TrEMOLO outperforms TELR and TLDR, detecting a much higher proportion of true-positive simulated TE insertions at a comparable false-positive rate finding
- ★ TrEMOLO-detected TE insertions and a hobo TE excision were experimentally validated by genomic PCR in a Piwi-knockdown D. melanogaster lineage (G73) finding
- ★ Genome assembly quality (assembler/polishing rounds) has limited impact on OUTSIDER TE detection but lower-quality assemblies inflate INSIDER TE calls (likely false positives) finding
- ★ TrEMOLO accurately estimates OUTSIDER TE allele frequency and slightly overestimates INSIDER TE frequency on simulated data finding
- TrEMOLO is distributed as an open-source (GPL3.0) SnakeMake pipeline runnable in script or singularity mode resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Long-read sequencing simulation (DeepSimulator) | D. melanogaster simulated genome (G0-F100 with 1100 simulated TE insertions) | simulated TE insertions at varying allele frequencies | true-positive/false-positive/false-negative TE insertion calls | DeepSimulator |
| Whole-genome pairwise alignment and mapping-based TE calling | D. melanogaster simulated reads mapped to G0-F100 genome | none (benchmarking different tools) | number of TE insertions detected (true/false positive counts) | minimap2; TELR; TLDR; TrEMOLO (OUTSIDER/TrEMOLO_NA and full pipeline) |
| Genome assembly comparison across assemblers/polishing | D. melanogaster G73 genome assembled with Flye, Shasta, or Wtdbg2 (+/- Racon polishing, RaGOO scaffolding) | different assembly tools and number of polishing rounds | contig/scaffold number, N50, BUSCO score, assembly size, number of INSIDER/OUTSIDER TE insertions and deletions detected by TrEMOLO | Flye; Shasta; Wtdbg2; Racon; RaGOO |
| Genomic PCR | D. melanogaster G0-F100 (control) and G73 (G73-SRE) genomic DNA | 73 generations of somatic Piwi knockdown-induced TE derepression in follicle cells | presence/absence and size of amplicons for 12 TrEMOLO-detected TE insertions, 4 shared TEs, and a hobo excision event | PCR |
| Long-read sequencing (deep coverage) and TrEMOLO analysis | D. melanogaster G73 line (Piwi-knockdown lineage), 183x depth | somatic Piwi knockdown across 73 generations | number and frequency of newly integrated TEs (INSIDER and OUTSIDER) | Oxford Nanopore Technologies (ONT) |
| TE frequency estimation validation on simulated data | simulated D. melanogaster genome/reads with known TE frequencies | none (known simulated frequencies) | TrEMOLO-estimated vs. real INSIDER and OUTSIDER TE insertion frequencies | — |
- ▼ TELR detected only 12 of 1100 simulated TE insertions 1%
- ▲ TLDR detected 928 of 1100 simulated TE insertions 84.4%
- ▲ TrEMOLO OUTSIDER-only module (TrEMOLO_NA) detected 1064 of 1100 simulated TE insertions 96.7%
- ▲ Full TrEMOLO pipeline (INSIDER+OUTSIDER) detection rate rose further 97.3%
- – False-positive TE detection rates were similar across TrEMOLO_NA, TrEMOLO, and TLDR 7% (TrEMOLO_NA), 8% (TrEMOLO), 11% (TLDR)
- – Genomic PCR confirmed expected amplicons for tested TE insertions and a hobo excision in G73-SRE but not in G0-F100
- – Wtdbg2-based assembly (lower BUSCO score) yielded a similar OUTSIDER TE count but roughly twice as many INSIDER TE insertions/deletions as Flye-based assemblies ~2-fold
- – TrEMOLO accurately estimated OUTSIDER TE frequencies and slightly overestimated INSIDER TE frequencies in simulated data
- count 1100 simulated TE insertions (simulated benchmark dataset (sample S1))
- fold_change 12/1100 (1%) true-positive detections (TELR benchmarking result)
- fold_change 928/1100 (84.4%) true-positive detections (TLDR benchmarking result)
- fold_change 1064/1100 (96.7%) true-positive detections (TrEMOLO_NA (OUTSIDER-only) benchmarking result)
- fold_change 97.3% true-positive detection rate (Full TrEMOLO pipeline benchmarking result)
- fold_change false-positive rates: 80 (7%) TrEMOLO_NA, 95 (8%) TrEMOLO, 11% TLDR (false-positive TE insertion detection comparison)
- count 2334 newly integrated TEs detected in G73 (183x depth) (G73 long-read sequencing dataset from prior study)
- other BUSCO 98% (Flye+4x Racon) vs 92.6% (Wtdbg2) (assembly quality comparison across assemblers)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper is a software/methods paper describing benchmarking and validation of the TrEMOLO pipeline using simulated and experimental long-read sequencing data. Tool performance (TrEMOLO vs. TELR vs. TLDR) and assembly-quality comparisons are reported as raw counts and percentages of true-positive/false-positive/false-negative transposable element (TE) calls, a descriptive assembly-metrics table (contigs, N50, BUSCO score, TE counts), and violin plots comparing estimated vs. simulated 'true' TE insertion frequencies. No formal inferential statistical tests, p-values, or effect sizes are reported in the text provided.
-
Detection performance of TrEMOLO vs. TELR/TLDR on the same 1100 simulated insertions was reported as raw true/false positive and negative counts and percentages, without a formal statistical comparison between tools.↳ Could also: A paired contingency-table approach such as McNemar's test, or a chi-square/Fisher's exact test on detection outcomes, could also be used — Because each tool was evaluated on the same set of simulated insertions, a paired test would also quantify whether the difference in detection rates exceeds what could arise from sampling variation alone.
-
Agreement between TrEMOLO's estimated TE insertion frequencies and the known simulated frequencies was shown visually via violin plots relative to a dashed 'true frequency' line.↳ Could also: A quantitative agreement metric such as mean absolute error, root-mean-square error, or a concordance/correlation statistic (e.g., Lin's concordance correlation coefficient) with a confidence interval could also be reported — A numerical summary alongside the visual comparison would also let readers gauge estimation accuracy and its uncertainty in a way that is directly comparable across frequency bins or datasets.
-
Assembly-quality comparisons (Flye, Shasta, Wtdbg2) in Table 1 are based on a single assembly run per assembler, with resulting TE detection counts compared descriptively.↳ Could also: Generating multiple assemblies per assembler (e.g., via subsampled or bootstrapped read sets) and comparing TE detection counts with an ANOVA or a non-parametric Kruskal-Wallis test could also be used — Replicate assemblies would also allow run-to-run variability to be distinguished from systematic differences attributable to the assembler itself.
-
Experimental validation used genomic PCR on 12 of 2334 TrEMOLO-detected TE insertions, reported qualitatively as amplicon presence/absence.↳ Could also: Reporting a validation/precision rate with a binomial confidence interval (e.g., a Wilson or Clopper-Pearson interval) for the tested subset could also be included — An interval estimate would also convey the statistical uncertainty around the validation rate given the relatively small number of loci tested by PCR.
-
Multiple pairwise comparisons are made across tools (TrEMOLO, TrEMOLO_NA, TELR, TLDR) and across assemblers, without any stated correction for multiple comparisons.↳ Could also: If formal hypothesis tests were to be added for these comparisons, a false discovery rate procedure (e.g., Benjamini-Hochberg) or a Bonferroni adjustment could also be applied — Such a correction would also help control the family-wise error rate when several comparisons are examined together, should inferential testing be introduced.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37013657 (TrEMOLO, Genome Biol 2023, 24:63)
Paper
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches. Mohamed et al. 2023. DOI 10.1186/s13059-023-02911-2. PMCID PMC10069131.
TrEMOLO is a Snakemake pipeline (the authors' own tool) that detects transposable element (TE) variants from ONT long reads, in two modes:
- INSIDER — TEs found by comparing an assembled genome to a reference (minimap2 + Assemblytics / SV calling).
- OUTSIDER — TEs found by mapping long reads to the assembly and SV-calling
(minimap2
map-ont+ sniffles/svim), then BLAST against a TE library, with a frequency estimate per insertion. Repo: https://github.com/DrosophilaGenomeEvolution/TrEMOLO (GPLv3 / CeCILL-C). Publication version: v2.2-beta1 (Dec 2022, archived doi:10.23708/2FYBUL). Singularity images shipped only for v2.4.0 (2023-10) and v2.5.4b (2024-04).
In scope (pipeline-derived results)
S1 — Simulated benchmark (PRIMARY target) — Fig. 2
Reported: on a simulated dataset of 1,100 known TE insertions (100 INSIDER from 11 TE families + 1000 OUTSIDER, randomly inserted into the G0-F100 assembly), reads simulated with DeepSimulator v1.5, four tools compared:
| Tool | True positives | False positives |
|---|---|---|
| TELR | 12 (1%) | — |
| TLDR | 928 (84.4%) | 95 (8%) |
| TrEMOLO_NA | 1064 (96.7%) | 80 (7%) |
| TrEMOLO | 1071 (97.3%) | 95 (8%) |
| Pipeline: TrEMOLO (OUTSIDER + INSIDER) on the simulated reads vs the G0-F100 | ||
| reference, compared to the known insertion positions. **This is the cleanest, | ||
| fully data-backed computational claim and the main reproduction target.** |
S2 — Depth-subsampling detection (secondary) — Results "sequencing depth"
Reported: at 183× depth, 2,334 TEs (2,274 OUTSIDER + 60 INSIDER); detection robust down to ~76×; at 15×, 1,064 (47%) OUTSIDER still detected. Needs the real G73 ONT data (ENA ERP122844) — large (183×, ~180 Mb genome). Attempt only after S1.
S3 — Assembler-tolerance (tertiary) — Results
2,334–2,579 TEs across Flye/Shasta/Wtdbg2 assemblies. Needs real reads + 3 assemblers. Lower priority.
Out of scope (not pipeline / not attempted)
- ddPCR & PCR experimental validation (Table 2, wet-lab) — out of scope.
- hobo excision validation (wet-lab) — out of scope.
Datasets the paper relies on (profiled in data/dataset_profile.json)
| Accession | What | Access | Size |
|---|---|---|---|
| zenodo:7673915 (doi:10.5281/zenodo.7673915) | simulated reads S1/S2/S3 (simulatedReadsTrEMOLO.tar.gz) |
open CC-BY | 25.8 GB |
| doi:10.23708/DSDTZ0 (DataSuds/IRD) | simulated genomes w/ inserted INSIDERs+OUTSIDERs (genomesData.tar.gz) — ground truth |
open CC-BY | 116 MB |
| doi:10.23708/N447VS (DataSuds) | simulated reads (mirror of Zenodo) | open | — |
| ENA ERP122844 | G0-F100 + G73 real ONT reads | open | large |
| ENA ERP138838 | G73-SRE real ONT reads (82×) | open | large |
repo test/ |
smoke-test genome+reads+canonical_TE.fa |
open | small |
Reproduction strategy
- front1 («infra»): clone TrEMOLO, pull singularity image (v2.4.0), download Zenodo reads + DataSuds genomes; verify MD5s.
- Smoke-test the pipeline on repo
test/data (confirms it runs end-to-end and emits TE_INFOS.bed) — a fast floor. - S1: run TrEMOLO (OUTSIDER+INSIDER) on the simulated reads vs G0-F100; extract detected insertions; derive ground-truth 1,100 positions from genomesData; count TP/FP by position match; compare to Fig. 2.
- Profile every dataset touched (N reported vs observed, QC, delivers-promised).
Known risks / honesty notes
- Exact TP/FP matching threshold not stated in paper (counts via 80% size/80% identity TE-library match) — our TP rule (position proximity to a known site) is a reasonable reconstruction and will be documented as such.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Close, faithful reproduction of the authors' own tool on their own deposited data. TrEMOLO reproduced 1041/1100 TP (≈95% vs reported 97.3%) with 107-110 FP (vs 95) and a perfect 100/100 INSIDER count; the central high-sensitivity/low-FP claim fully holds. The small, non-exact deviations sit on our side (tool version drift v2.2-beta1→v2.5.6b and a reconstructed scoring window the paper underspecified), not the authors' — every printed value is derivable from the shipped data+code (1096/1100 sites recovered by alignment). Overall yellow: solid and reproducible with explainable deviations rather than a 1:1/rounding-only match.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
<synthetic>Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.