Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

TEMP: a computational method for analyzing transposable element polymorphism in populations.

Nucleic Acids Res · 2014
L1 61/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
61/100
Reproducibility score
0.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 22% of all assessed papers rank 906 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the paper's central pipeline-derived claim (Fig. 3c: a pogo transposon insertion in the 42AB piRNA cluster, chr2R ~2,378,892-2,378,894, absent in w1, near-fixed in Har, and elevated in the F1 hybrid above the ~48% Mendelian null) using the TEMP pipeline (github.com/JialiUMassWengLab/TEMP, the paper's own tool) run end-to-end on the paper's SRA data (SRP007937) via BWA + dm3 on «our HPC»/SLURM. Genomic coordinates match essentially exactly for all three samples. Frequencies are directionally correct in all three cases but numerically below the paper's reported values: Har observed ~90-92% vs reported 96.77% (within-tol), w1 observed ~0% vs reported 0% (exact/absence confirmed), F1 observed ~73-79% vs reported 88.24% (partial -- still clearly above the 48% Mendelian expectation supporting the paper's selection claim, but a larger gap than Har). The most likely driver of the frequency deflation is a reference-database deviation: the paper specifies a Repbase v17.07 TE family-consensus for TEMP's -r flag, but Repbase is registration-gated, so this reproduction used FlyBase's dmel-all-transposon.fasta (release r6.68), a per-genomic-instance annotation. This fragments read support for one true pogo insertion across dozens of near-identical FBti reference copies (confirmed via header inspection and repeated MD5 hashes), diluting per-reference Frequency values without affecting breakpoint position. Minor secondary deviation: BWA 0.7.18 used instead of the paper's stated 0.6.1-r104 (TEMP requires an old bwa aln/sampe workflow that is version-tolerant here). NOT attempted in this pass: the F1 21-day and backcross 21-day WGS timepoints (SRR333512/513/514, in-scope but not run due to budget), all small-RNA/piRNA runs (out of scope -- feed a separate wet-lab/piRNA-pathway analysis, not TEMP), and repair of the TEMP repo's own bundled self-test (TEMP_Absence.sh, 0/5 matches on shipped test data due to a bedtools tokenizing mismatch -- documented as a graded but non-blocking finding since only TEMP_Insertion.sh was needed for the reproduced claims). No fabricated values: every reported number is read directly from the .insertion.refined.bp.summary files listed as source_file per claim.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-01
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-01
no human curator yet
Last updated
2026-08-01

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a computational method detect transposable element insertions and excisions—present at a wide range of population frequencies—from pooled high-throughput sequencing data, and accurately estimate their population frequencies and junction positions? The authors propose TEMP as such a method and test whether it outperforms existing structural-variation/TE callers.

Core claims
  • TEMP combines pair-end (discordant) read and split (soft-clipped) read information to identify both presence and absence of TE insertions in genomic DNA from heterogeneous/pooled samples. method
  • TEMP estimates population frequency of a transposition event as T/(T+R), the ratio of supporting read pairs to supporting plus non-supporting (junction-spanning) read pairs. method
  • TEMP pinpoints junctions of high-frequency transposition events at nucleotide (base-pair) resolution using soft-clipped reads. method
  • On simulated data, TEMP outperforms PoPoolationTE, RetroSeq, VariationHunter and GASVPro. finding
  • TEMP performs well on whole-genome human data from the 1000 Genomes Project. finding
  • TEMP applied to Drosophila reveals TE frequencies in a wild population, inheritance patterns of TEs during hybrid dysgenesis, sequence signatures of TE insertion, and possible molecular effects such as altered gene expression and piRNA production. finding
  • TEMP is freely available open-source software at github.com/JialiUMassWengLab/TEMP.git and http://zlab.umassmed.edu/TEMP/. resource
  • TEMP requires a curated library of transposon consensus sequences and cannot identify transposition events de novo. method
Experimental setups
Assay System Perturbation Readout Platform
Simulated paired-end whole-genome DNA sequencing (in silico benchmark) Drosophila melanogaster reference genome dm3, chromosome arm 2L 50 simulated TE insertions and 50 simulated excisions per experiment; simulated reads mixed with reference-genome reads at defined ratios to set population frequencies; depths 5X/10X/20X/40X Recovery of simulated insertions/excisions (correct TE family, direction, interval containing true junction) and junction accuracy within 5 nt RSVSim v1.1.1 for variant simulation; pIRS v1.1.0 Illumina paired-end read simulator (-l 90 -m 500 -v 50 -e 0.0001 -a 0 -g 0); BWA v0.6.1-r104 aln (-n 3 -l 100 -R 10000)
Comparative benchmarking of TE/SV callers on pooled simulated data Five independently simulated D. melanogaster chromosome 2L arms combined (each 5X, apparent 25X), repeated 20 times Simulated insertions/excisions as above Detection performance of TEMP versus PoPoolationTE, RetroSeq, VariationHunter and GASVPro PoPoolationTE v1.02; RetroSeq; VariationHunter CommonLaw v0.04 (mrfast v2.6.0.1, -min 400 -max 600 -e 3); GASVPro-HQ (2013 Oct release, top 250 predictions by log-likelihood ratio)
Whole-genome paired-end DNA sequencing (pooled/merged human genomes) Human; four 1000 Genomes individuals NA18517, NA19240, NA12156, NA12878 merged to mimic pooled sequencing; GRCh37/hg19 reference none Predicted presence/absence of TE insertions with frequency >=20% and >8 supporting reads, compared with DGV structural variants 1000 Genomes BAM files; DGV (database of genomic variants)
Genomic deep sequencing (pooled population) for hybrid dysgenesis analysis D. melanogaster w1 strain, Harwich strain, and w1 x Har 2–4 day F1 population Hybrid dysgenic cross (genetic cross) TE population frequencies per strain and frequency change FC = F − (H + W)/2 for parental transposons (frequency >10% in at least one parental strain) NCBI SRA SRP007937
Small RNA sequencing D. melanogaster hybrid dysgenesis populations (w1, Harwich, F1) Hybrid dysgenic cross piRNA production; junction-spanning small RNA reads >=21 nt mapping perfectly across the genome–transposon junction; piRNA cluster annotation (141 clusters, chrX_TAS excluded) NCBI SRA SRP007937; processing as in Khurana et al.
Whole-genome DNA sequencing of inbred lines (TE insertion discovery and genomic distribution) 53 Drosophila Genetic Reference Panel (DGRP) inbred lines; dm3 reference none (natural/wild-derived variation) TE insertions per line and their distribution across promoters (2 kb upstream of TSS), exons, intron/UTR, intergenic (>2 kb from genes) and piRNA clusters; binomial test enrichment/depletion per TE family NCBI SRA; BWA aln (3 mismatches); FlyBase Release 5.45 annotation; Repbase v17.07; UCSC RepeatMasker; Benjamini–Hochberg correction
RNA-seq (gene expression quantification) D. melanogaster DGRP lines RAL-362, RAL-765, RAL-517 and four F1 progeny populations (7 datasets) Line-specific TE insertions (frequency >20% in promoter/intron/exon/UTR of one line only); F1 crosses Gene expression in FPKM; genes with >2-fold higher or lower expression (pseudo count 0.5 FPKM) in the insertion-bearing line versus the two lines lacking the insertion Tophat v2.0.8b (default parameters); Cufflinks v2.1.1 (default parameters)
Sequence motif and nucleotide composition analysis of TE integration sites D. melanogaster; all Drosophila genomic sequencing datasets analyzed by TEMP none Target site duplication (TSD) length from the difference between junction coordinates on the two strands; enriched palindromic motifs and mono-/dinucleotide composition in 30 bp around junctions versus 100 bp flanking background MEME (-dna -mod zoops -nmotifs 5 -minw 4 -maxw 15 -pal)
Key results
  • TEMP outperforms PoPoolationTE, RetroSeq, VariationHunter and GASVPro on simulated pooled data (results summarized in Table 1).
  • TEMP performs well on whole-genome human data derived from the 1000 Genomes Project, with predictions compared against previously reported insertions/deletions in DGV.
  • TEMP detected 14 363 non-redundant TE insertion events with junctions detected on both strands across all Drosophila genomic sequencing data, enabling TSD length estimation and motif analysis. 14 363 insertions
  • 11 311 insertions with frequency >80% in at least one DGRP inbred line were used to profile TE distribution across promoters, exons, intron/UTR, intergenic regions and piRNA clusters; individual TE families showed significant enrichment or depletion. 11 311 insertions; q < 0.15
  • 48 genes were identified whose expression differed >2-fold in the single line carrying a line-specific TE insertion (frequency >20%) in the promoter, intron, exon or UTR relative to the two lines lacking the insertion. >2-fold; 48 genes
  • TEMP was used to characterize TE frequencies in a wild D. melanogaster population and inheritance patterns of TEs during hybrid dysgenesis, including junction-spanning piRNA reads.
Key statistics
  • count 50 insertions and 50 excisions per simulation experiment, repeated 100 times, yielding 5000 simulated insertions and 5000 simulated excisions (Simulation design on D. melanogaster chromosome arm 2L)
  • count Five simulated 2L arms at 5X each, apparent coverage 25X, process repeated 20 times (Datasets used to compare TEMP with PoPoolationTE, RetroSeq, VariationHunter and GASVPro)
  • count 14 363 non-redundant insertion events with junctions on both strands (TE insertion sequence-signature/TSD analysis across all Drosophila genomic data)
  • count 11 311 insertions with frequency >80% in at least one inbred line (DGRP TE insertion genomic-distribution analysis)
  • count 53 DGRP inbred lines; 50 lines with >20X coverage, 3 lines (RAL-362, RAL-765, RAL-517) with <20X (Drosophila Genetic Reference Panel genomic sequencing datasets)
  • count 141 piRNA clusters (excluding chrX_TAS) occupying 4 924 944 bp of dm3 (piRNA cluster annotation from Brennecke et al. used in hybrid dysgenesis and DGRP analyses)
  • pvalue q-value < 0.15 (binomial test with Benjamini–Hochberg correction) (Threshold for reporting TE family enrichment/depletion in genomic features)
  • fold_change >2-fold higher or lower expression (pseudo count 0.5 FPKM), yielding 48 genes (RNA-seq comparison of insertion-bearing line versus two lines lacking the insertion)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper primarily describes a computational method (TEMP) whose performance was evaluated using simulation-based recovery-rate comparisons against four other algorithms (PoPoolationTE, RetroSeq, VariationHunter, GASVPro), with results summarized as counts/tables rather than through inferential statistics. One explicit inferential test is reported: a binomial test (with Benjamini-Hochberg multiple-testing correction) used to assess enrichment/depletion of TE families across five genomic features. A separate analysis of TE insertions and gene expression change used a fixed fold-change threshold rather than a formal statistical test.

Replicationunclear Sample sizeSimulation comparisons were repeated 100 times (yielding 5000 simulated insertions/excisions) for TEMP's own accuracy evaluation, and 20 times for the five-algorithm comparison; the TE-family enrichment analysis was based on 11,311 insertions from 53 inbred lines; no formal power/sample-size justification was described GroupsTEMP vs. PoPoolationTE, RetroSeq, VariationHunter and GASVPro on simulated data; TE family insertion counts across five genomic feature categories; parental vs. F1 transposon frequencies in a hybrid dysgenesis cross; gene expression levels among inbred lines with vs. without a TE insertion Pairingunclear Randomization/blindingnot stated Dispersionunclear Multiplicity correctionBenjamini-Hochberg FDR procedure
Statistical tests used
Test Applied to n Assumptions
binomial test assessing enrichment or depletion of each TE family in each of five genomic features (promoters, exons, intron/UTR regions, intergenic regions, piRNA clusters) 11,311 TE insertions with frequency >80% in at least one of 53 DGRP inbred lines not stated
Approaches that could also have been used
  • TE family enrichment/depletion across genomic features was assessed with a binomial test corrected by the Benjamini-Hochberg FDR procedure.
    Could also: A Fisher's exact test, hypergeometric test, or an overdispersion-aware model (e.g., negative-binomial-based enrichment test) — These alternatives are also standard for categorical enrichment testing and can additionally accommodate overdispersion or non-independence among genomic counts, which some genomic enrichment analyses exhibit.
  • Genes potentially affected by TE insertions were identified using a fixed fold-change threshold (>2-fold change in FPKM, with a pseudocount) between lines with and without an insertion, rather than a formal significance test.
    Could also: A model-based differential expression test such as DESeq2 or edgeR (negative binomial models), or a nonparametric test like Mann-Whitney U across replicate expression values, combined with multiple-testing correction — Such approaches would also yield p-values or q-values alongside the fold-change estimate, helping to convey the statistical confidence of an expression difference in addition to its magnitude.
  • Detection performance of TEMP relative to PoPoolationTE, RetroSeq, VariationHunter and GASVPro on simulated data was compared using recovery counts/tables summarized across repeated simulations.
    Could also: A paired statistical test across simulation replicates, such as McNemar's test for paired binary detection outcomes or a paired proportion test — This could also provide a formal significance estimate for the observed differences in detection rates between algorithms across the repeated simulation runs, complementing the descriptive comparison.
  • Variability of detection accuracy across the repeated simulation replicates (100 or 20 repeats) is not reported with an explicit dispersion statistic in the excerpted text.
    Could also: Reporting SD, SEM, or a 95% confidence interval of the recovery rate across replicate simulations — This would also communicate how consistent detection performance was across independent simulation runs, in addition to the point estimates shown.
  • Multiple testing correction for the TE family/genomic feature binomial tests used the Benjamini-Hochberg FDR method.
    Could also: A Bonferroni correction — Bonferroni is also a widely used alternative that controls the family-wise error rate rather than the false discovery rate, offering a more conservative threshold that some researchers prefer for smaller test families.
  • Several analyses used dichotomized frequency cutoffs (e.g., >10% for parental transposons, >20% for insertions linked to expression change, >80% for distribution analysis) rather than modeling frequency as a continuous variable.
    Could also: A regression-based approach (e.g., logistic or linear regression) treating insertion frequency as a continuous predictor — This could also make use of the full range of frequency information rather than collapsing it at a chosen threshold, potentially capturing dose-response relationships between frequency and outcome.
Software: BWA (aln algorithm) 0.6.1-r104 · RSVSim 1.1.1 · pIRS 1.1.0 · PoPoolationTE 1.02 · RetroSeq · VariationHunter CommonLaw 0.04 · GASVPro-HQ 2013 Oct Release · mrfast 2.6.0.1 · MEME · TopHat 2.0.8b · Cufflinks 2.1.1

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

har_pogo_42AB_frequency
Reported
Har (parental strain): pogo TE insertion in 42AB piRNA cluster, chr2R ~2,378,892-2,378,894, frequency 96.77% (Fig. 3c)
Reproduced
TEMP_Insertion.sh on SRR333515+SRR333516 (Har WGS) reports the same breakpoint (Junction1=2378894, Junction2=2378896) across ~30 reference rows (fragmented by FBti reference ID, see dataset note), with Frequency tightly clustered 0.90-0.92 (e.g. 0.9038, 0.9074, 0.9111, 0.9123, 0.9170)
within tolerance
w1_pogo_42AB_absence
Reported
w1 (parental strain): pogo insertion absent at the 42AB locus, 0% (Fig. 3c)
Reproduced
TEMP_Insertion.sh on SRR333517+SRR333518 (w1 WGS) shows no signal at the real breakpoint (2378894/2378896); the only nearby row is an unrelated weak singleton (Frequency=0.05) at a different junction (2378136), consistent with absence
exact
f1_pogo_42AB_frequency
Reported
F1 hybrid (w1xHar, 2-4 day): pogo insertion frequency 88.24%, elevated above the ~48% Mendelian expectation (Fig. 3c), interpreted as positive selection
Reproduced
TEMP_Insertion.sh on SRR333511 (F1 2-4day WGS) shows the same breakpoint (Junction1=2378896) with Frequency 0.73-0.79 across rows (0.7273 singleton-class, 0.7857/0.7692 2p-class) -- clearly elevated above ~0.48 Mendelian expectation and directionally consistent with the paper's selection claim, but the absolute frequency is 9-15 points below the reported 88.24%
partial
temp_absence_repo_selftest
Reported
TEMP repo's own bundled test dataset (testrun) is expected to reproduce reference deletion calls via TEMP_Absence.sh (repo self-test, not a paper figure)
Reproduced
TEMP_Absence.sh run on the shipped test data produced 0/5 matches; traced to a bedtools intersect tokenizing/field-count mismatch inside the script against the bedtools version available in the build env
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 61/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Reproduction ran the authors' own tool (TEMP) on the authors' own public data (SRP007937) and recovered the Fig. 3c breakpoint essentially exactly (Junction1=2378894/2378896 vs reported chr2R ~2,378,892-2,378,894) in all three samples. The qualitative claim is fully confirmed — pogo absent in w1 (0%), near-fixed in Har (~0.90-0.92), and elevated in the F1 hybrid to ~0.73-0.79, well above the ~48% Mendelian null that underpins the positive-selection interpretation. The numeric shortfall (Har 96.77% -> ~91%; F1 88.24% -> ~76%) is on our side: Repbase v17.07 is registration-gated, so FlyBase's per-instance dmel-all-transposon.fasta r6.68 was substituted, fragmenting read support for one insertion across ~30 near-identical FBti rows and diluting per-reference frequency without moving the breakpoint. Two honest secondary gaps: the shipped TEMP_Absence.sh self-test fails 0/5 against modern bedtools (authors' code rot, non-blocking here), and three in-scope WGS runs (SRR333512/513/514) were never processed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.