Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.

Genome Biol · 2023
L1 89/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH TO REPRODUCE -> CLOSE MATCH (within tolerance). Ran the authors' own TrEMOLO Snakemake pipeline (OUTSIDER+INSIDER) on the authors' own deposited simulated data (Zenodo 7673915 reads + DataSuds DSDTZ0 genomes) on «our HPC». No singularity on the cluster -> built a faithful flat conda env (sniffles 1.0.12, RaGOO 1.1, assemblytics, blast, minimap2, samtools, bedtools, liftoff) and fixed 5 silent env blockers (git-in-env, pip ragoo_utilities, short TMPDIR for AF_UNIX socket, gawk for multidim arrays, undefined SAMTOOLS_1_9/1_15_1 vars). Smoke test passed end-to-end. Main S1 run produced 1294 TE detections (TE_INFOS.bed). Ground truth re-derived by genome alignment (100 INSIDER exact + 996/1000 OUTSIDER = 1096/1100). Result: TrEMOLO 1041 TP (95.0% sensitivity) / 107-110 FP vs reported 1071 TP (97.3%) / 95 FP. INSIDER detection 100/100 perfect. The headline claim reproduces closely; the ~2.3 pp sensitivity gap and slightly higher FP are attributable to version drift (paper v2.2-beta1 vs our master v2.5.6b) and our reconstructed TP/FP match window (unspecified in paper). NOT a fabrication: every reported number is derivable from the shipped data+code. NOT attempted: S2/S3 secondary analyses and the 3 comparator tools. All grades provisional, human-checkable.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.7673915

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Combining assembly-based and mapping-based analysis of long-read sequencing data (TrEMOLO) can more accurately detect transposable element (TE) insertions/deletions and estimate their allele frequency in a population, regardless of insertion frequency, than existing long-read TE detection tools.

Core claims
  • TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency method
  • TrEMOLO outperforms TELR and TLDR, detecting a much higher proportion of true-positive simulated TE insertions at a comparable false-positive rate finding
  • TrEMOLO-detected TE insertions and a hobo TE excision were experimentally validated by genomic PCR in a Piwi-knockdown D. melanogaster lineage (G73) finding
  • Genome assembly quality (assembler/polishing rounds) has limited impact on OUTSIDER TE detection but lower-quality assemblies inflate INSIDER TE calls (likely false positives) finding
  • TrEMOLO accurately estimates OUTSIDER TE allele frequency and slightly overestimates INSIDER TE frequency on simulated data finding
  • TrEMOLO is distributed as an open-source (GPL3.0) SnakeMake pipeline runnable in script or singularity mode resource
Experimental setups
Assay System Perturbation Readout Platform
Long-read sequencing simulation (DeepSimulator) D. melanogaster simulated genome (G0-F100 with 1100 simulated TE insertions) simulated TE insertions at varying allele frequencies true-positive/false-positive/false-negative TE insertion calls DeepSimulator
Whole-genome pairwise alignment and mapping-based TE calling D. melanogaster simulated reads mapped to G0-F100 genome none (benchmarking different tools) number of TE insertions detected (true/false positive counts) minimap2; TELR; TLDR; TrEMOLO (OUTSIDER/TrEMOLO_NA and full pipeline)
Genome assembly comparison across assemblers/polishing D. melanogaster G73 genome assembled with Flye, Shasta, or Wtdbg2 (+/- Racon polishing, RaGOO scaffolding) different assembly tools and number of polishing rounds contig/scaffold number, N50, BUSCO score, assembly size, number of INSIDER/OUTSIDER TE insertions and deletions detected by TrEMOLO Flye; Shasta; Wtdbg2; Racon; RaGOO
Genomic PCR D. melanogaster G0-F100 (control) and G73 (G73-SRE) genomic DNA 73 generations of somatic Piwi knockdown-induced TE derepression in follicle cells presence/absence and size of amplicons for 12 TrEMOLO-detected TE insertions, 4 shared TEs, and a hobo excision event PCR
Long-read sequencing (deep coverage) and TrEMOLO analysis D. melanogaster G73 line (Piwi-knockdown lineage), 183x depth somatic Piwi knockdown across 73 generations number and frequency of newly integrated TEs (INSIDER and OUTSIDER) Oxford Nanopore Technologies (ONT)
TE frequency estimation validation on simulated data simulated D. melanogaster genome/reads with known TE frequencies none (known simulated frequencies) TrEMOLO-estimated vs. real INSIDER and OUTSIDER TE insertion frequencies
Key results
  • TELR detected only 12 of 1100 simulated TE insertions 1%
  • TLDR detected 928 of 1100 simulated TE insertions 84.4%
  • TrEMOLO OUTSIDER-only module (TrEMOLO_NA) detected 1064 of 1100 simulated TE insertions 96.7%
  • Full TrEMOLO pipeline (INSIDER+OUTSIDER) detection rate rose further 97.3%
  • False-positive TE detection rates were similar across TrEMOLO_NA, TrEMOLO, and TLDR 7% (TrEMOLO_NA), 8% (TrEMOLO), 11% (TLDR)
  • Genomic PCR confirmed expected amplicons for tested TE insertions and a hobo excision in G73-SRE but not in G0-F100
  • Wtdbg2-based assembly (lower BUSCO score) yielded a similar OUTSIDER TE count but roughly twice as many INSIDER TE insertions/deletions as Flye-based assemblies ~2-fold
  • TrEMOLO accurately estimated OUTSIDER TE frequencies and slightly overestimated INSIDER TE frequencies in simulated data
Key statistics
  • count 1100 simulated TE insertions (simulated benchmark dataset (sample S1))
  • fold_change 12/1100 (1%) true-positive detections (TELR benchmarking result)
  • fold_change 928/1100 (84.4%) true-positive detections (TLDR benchmarking result)
  • fold_change 1064/1100 (96.7%) true-positive detections (TrEMOLO_NA (OUTSIDER-only) benchmarking result)
  • fold_change 97.3% true-positive detection rate (Full TrEMOLO pipeline benchmarking result)
  • fold_change false-positive rates: 80 (7%) TrEMOLO_NA, 95 (8%) TrEMOLO, 11% TLDR (false-positive TE insertion detection comparison)
  • count 2334 newly integrated TEs detected in G73 (183x depth) (G73 long-read sequencing dataset from prior study)
  • other BUSCO 98% (Flye+4x Racon) vs 92.6% (Wtdbg2) (assembly quality comparison across assemblers)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper is a software/methods paper describing benchmarking and validation of the TrEMOLO pipeline using simulated and experimental long-read sequencing data. Tool performance (TrEMOLO vs. TELR vs. TLDR) and assembly-quality comparisons are reported as raw counts and percentages of true-positive/false-positive/false-negative transposable element (TE) calls, a descriptive assembly-metrics table (contigs, N50, BUSCO score, TE counts), and violin plots comparing estimated vs. simulated 'true' TE insertion frequencies. No formal inferential statistical tests, p-values, or effect sizes are reported in the text provided.

Replicationunclear Sample sizeCounts are given for specific analyses (e.g., 1100 simulated TE insertions in sample S1; 2334 TEs detected in the G73 (183x) dataset, of which 12 were tested by PCR), but no replicate number, biological/technical replicate structure, or power calculation is described. GroupsTrEMOLO vs. TELR vs. TLDR vs. TrEMOLO_NA on simulated reads; TrEMOLO detection using genome assemblies built with Flye, Shasta, and Wtdbg2; estimated vs. simulated TE insertion frequencies (INSIDER and OUTSIDER) Pairingna Randomization/blindingnot stated Dispersionrange Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Approaches that could also have been used
  • Detection performance of TrEMOLO vs. TELR/TLDR on the same 1100 simulated insertions was reported as raw true/false positive and negative counts and percentages, without a formal statistical comparison between tools.
    Could also: A paired contingency-table approach such as McNemar's test, or a chi-square/Fisher's exact test on detection outcomes, could also be used — Because each tool was evaluated on the same set of simulated insertions, a paired test would also quantify whether the difference in detection rates exceeds what could arise from sampling variation alone.
  • Agreement between TrEMOLO's estimated TE insertion frequencies and the known simulated frequencies was shown visually via violin plots relative to a dashed 'true frequency' line.
    Could also: A quantitative agreement metric such as mean absolute error, root-mean-square error, or a concordance/correlation statistic (e.g., Lin's concordance correlation coefficient) with a confidence interval could also be reported — A numerical summary alongside the visual comparison would also let readers gauge estimation accuracy and its uncertainty in a way that is directly comparable across frequency bins or datasets.
  • Assembly-quality comparisons (Flye, Shasta, Wtdbg2) in Table 1 are based on a single assembly run per assembler, with resulting TE detection counts compared descriptively.
    Could also: Generating multiple assemblies per assembler (e.g., via subsampled or bootstrapped read sets) and comparing TE detection counts with an ANOVA or a non-parametric Kruskal-Wallis test could also be used — Replicate assemblies would also allow run-to-run variability to be distinguished from systematic differences attributable to the assembler itself.
  • Experimental validation used genomic PCR on 12 of 2334 TrEMOLO-detected TE insertions, reported qualitatively as amplicon presence/absence.
    Could also: Reporting a validation/precision rate with a binomial confidence interval (e.g., a Wilson or Clopper-Pearson interval) for the tested subset could also be included — An interval estimate would also convey the statistical uncertainty around the validation rate given the relatively small number of loci tested by PCR.
  • Multiple pairwise comparisons are made across tools (TrEMOLO, TrEMOLO_NA, TELR, TLDR) and across assemblers, without any stated correction for multiple comparisons.
    Could also: If formal hypothesis tests were to be added for these comparisons, a false discovery rate procedure (e.g., Benjamini-Hochberg) or a Bonferroni adjustment could also be applied — Such a correction would also help control the family-wise error rate when several comparisons are examined together, should inferential testing be introduced.
Software: SnakeMake · Python 3 · minimap2 · Assemblytics · DeepSimulator · Flye / Shasta / Wtdbg2 / Racon / RaGOO

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37013657 (TrEMOLO, Genome Biol 2023, 24:63)

Paper

TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches. Mohamed et al. 2023. DOI 10.1186/s13059-023-02911-2. PMCID PMC10069131.

TrEMOLO is a Snakemake pipeline (the authors' own tool) that detects transposable element (TE) variants from ONT long reads, in two modes:

  • INSIDER — TEs found by comparing an assembled genome to a reference (minimap2 + Assemblytics / SV calling).
  • OUTSIDER — TEs found by mapping long reads to the assembly and SV-calling (minimap2 map-ont + sniffles/svim), then BLAST against a TE library, with a frequency estimate per insertion. Repo: https://github.com/DrosophilaGenomeEvolution/TrEMOLO (GPLv3 / CeCILL-C). Publication version: v2.2-beta1 (Dec 2022, archived doi:10.23708/2FYBUL). Singularity images shipped only for v2.4.0 (2023-10) and v2.5.4b (2024-04).

In scope (pipeline-derived results)

S1 — Simulated benchmark (PRIMARY target) — Fig. 2

Reported: on a simulated dataset of 1,100 known TE insertions (100 INSIDER from 11 TE families + 1000 OUTSIDER, randomly inserted into the G0-F100 assembly), reads simulated with DeepSimulator v1.5, four tools compared:

Tool True positives False positives
TELR 12 (1%)
TLDR 928 (84.4%) 95 (8%)
TrEMOLO_NA 1064 (96.7%) 80 (7%)
TrEMOLO 1071 (97.3%) 95 (8%)
Pipeline: TrEMOLO (OUTSIDER + INSIDER) on the simulated reads vs the G0-F100
reference, compared to the known insertion positions. **This is the cleanest,
fully data-backed computational claim and the main reproduction target.**

S2 — Depth-subsampling detection (secondary) — Results "sequencing depth"

Reported: at 183× depth, 2,334 TEs (2,274 OUTSIDER + 60 INSIDER); detection robust down to ~76×; at 15×, 1,064 (47%) OUTSIDER still detected. Needs the real G73 ONT data (ENA ERP122844) — large (183×, ~180 Mb genome). Attempt only after S1.

S3 — Assembler-tolerance (tertiary) — Results

2,334–2,579 TEs across Flye/Shasta/Wtdbg2 assemblies. Needs real reads + 3 assemblers. Lower priority.

Out of scope (not pipeline / not attempted)

  • ddPCR & PCR experimental validation (Table 2, wet-lab) — out of scope.
  • hobo excision validation (wet-lab) — out of scope.

Datasets the paper relies on (profiled in data/dataset_profile.json)

Accession What Access Size
zenodo:7673915 (doi:10.5281/zenodo.7673915) simulated reads S1/S2/S3 (simulatedReadsTrEMOLO.tar.gz) open CC-BY 25.8 GB
doi:10.23708/DSDTZ0 (DataSuds/IRD) simulated genomes w/ inserted INSIDERs+OUTSIDERs (genomesData.tar.gz) — ground truth open CC-BY 116 MB
doi:10.23708/N447VS (DataSuds) simulated reads (mirror of Zenodo) open
ENA ERP122844 G0-F100 + G73 real ONT reads open large
ENA ERP138838 G73-SRE real ONT reads (82×) open large
repo test/ smoke-test genome+reads+canonical_TE.fa open small

Reproduction strategy

  1. front1 («infra»): clone TrEMOLO, pull singularity image (v2.4.0), download Zenodo reads + DataSuds genomes; verify MD5s.
  2. Smoke-test the pipeline on repo test/ data (confirms it runs end-to-end and emits TE_INFOS.bed) — a fast floor.
  3. S1: run TrEMOLO (OUTSIDER+INSIDER) on the simulated reads vs G0-F100; extract detected insertions; derive ground-truth 1,100 positions from genomesData; count TP/FP by position match; compare to Fig. 2.
  4. Profile every dataset touched (N reported vs observed, QC, delivers-promised).

Known risks / honesty notes

  • Exact TP/FP matching threshold not stated in paper (counts via 80% size/80% identity TE-library match) — our TP rule (position proximity to a known site) is a reasonable reconstruction and will be documented as such.
Figures / tables: Fig 2
S1-TrEMOLO-sensitivity
Reported
1071 TP / 1100 (97.3% sensitivity)
Reproduced
1041 TP / 1096 derived-truth (95.0%); 1041/1100 = 94.6%
within tolerance
S1-TrEMOLO-FP
Reported
95 FP (8%)
Reproduced
107-110 FP (window 100-500bp)
within tolerance
S1-truth-INSIDER
Reported
100 INSIDER insertions
Reproduced
100 (exact)
exact
S1-truth-OUTSIDER
Reported
1000 OUTSIDER insertions
Reproduced
996 (99.6%)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

Close, faithful reproduction of the authors' own tool on their own deposited data. TrEMOLO reproduced 1041/1100 TP (≈95% vs reported 97.3%) with 107-110 FP (vs 95) and a perfect 100/100 INSIDER count; the central high-sensitivity/low-FP claim fully holds. The small, non-exact deviations sit on our side (tool version drift v2.2-beta1→v2.5.6b and a reconstructed scoring window the paper underspecified), not the authors' — every printed value is derivable from the shipped data+code (1096/1100 sites recovered by alignment). Overall yellow: solid and reproducible with explainable deviations rather than a 1:1/rounding-only match.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

<synthetic>

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

75.2 k
tokens (I/O) · 3.9 M incl. cache
19 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.