Human Retrotransposons and Effective Computational Detection Methods for Next-Generation Sequencing Data.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
▸Reproduction agent’s raw note
DROP (non_pipeline). PMID 36295018 (Lee, Min, Mun, Han; Life/MDPI 2022, PMC9605557) is an explicitly labeled REVIEW article: no Materials & Methods, no Code Availability, Data Availability = 'data are included in the manuscript'. The authors ran no pipeline and deposited no data, so there are ZERO pipeline-derived results of theirs to reproduce. The harvested artifacts are text-mining false positives: github.com/tk2/RetroSeq is one of ~8 third-party tools the review tabulates (Table 3, cited ref [119] = the original RetroSeq paper, Keane et al. 2013), and SRS228129 is one individual ('an unreported ancestor') in a 7-person panel of a study the review summarizes. Every numeric value in the paper (RetroSeq >90% / 97% Alu / 83% L1; MELT >99%; xTea >90%; 17%/11%/0.2% composition) is a CITATION of prior literature, not an output these authors generated. The P16 third-party-tool rule does not rescue it: the review has no own data, and the cited RetroSeq figures are SENSITIVITIES that require a gold-standard MEI truth set the paper does not provide, so running RetroSeq on SRS228129 would reproduce a different publication (Keane 2013) and could not be graded against any value in THIS paper (also no_expected_result). NOT ATTEMPTED on «our HPC»: no compute justified for a non_pipeline drop, and a forced ungradeable run would violate 'do not fabricate to avoid a drop'. Dataset SRS228129 was still profiled in the same pass (metadata level: 4 runs, ~865M reads / ~173 Gbp / ~89 GB fastq, open access, complete metadata, grade B provisional); files not downloaded.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-19 ⛓ ca5932259bc7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-19
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThis review surveys human non-LTR retrotransposons (LINE-1, Alu, SVA), next-generation sequencing platforms, and computational methods, aiming to guide researchers on how retrotransposons affect the human genome and which bioinformatics tools to use for detecting retrotransposon insertions depending on research purpose.
- ★ Transposable elements make up nearly 45% of the human genome, vastly exceeding the ~1.5% that is protein-coding. finding
- ★ Non-LTR retrotransposons (LINE-1, Alu, SVA) remain active in humans and drive genomic diversity, genetic alteration, and disease via insertions. mechanism
- ★ Various computational tools (e.g. RetroSeq) have been developed to detect non-reference retrotransposon insertions from NGS data. resource
- ★ L1 elements mobilize via target-primed reverse transcription (TPRT), producing 5' truncations, a 3' poly(A) tail, and flanking target site duplications. mechanism
- ★ Alu and SVA are non-autonomous and use L1 machinery in trans to retrotranspose. mechanism
- ★ Third-generation long-read sequencers (PacBio, Nanopore) offer longer reads better suited to detecting retrotransposon-derived structural variants than short-read platforms. method
- Shin et al. (2019) found non-reference L1Hs insertion sites had 41.15% GC content, implying L1 elements integrate randomly rather than preferring AT-rich regions. finding
- The first disease-causing novel L1 insertion was reported in 1988 in factor VIII (X-linked) in a hemophilia A patient by Kazazian and colleagues. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| short-read whole-genome sequencing (paired-end) | human genome | none | non-reference TE insertions via discordant mate pairs | Illumina (RetroSeq) |
| whole-genome sequencing / variant calling comparison | normal Korean tissue and lung tumor tissue | none | SNVs, insertions, deletions (indels) | Illumina NovaSeq 6000, MGISEQ-2000, DNBSEQ-T7 |
| synthetic long-read sequencing (TruSeq) | Drosophila melanogaster | none | identification of annotated transposable elements | Illumina TruSeq synthetic long-reads |
| single-tube long fragment read (stLFR) sequencing | long human genomic DNA (10–350 kb) | Tn5 transposome co-barcoding | co-barcoded subfragments for TE detection | MGISEQ-2000 |
| single-molecule real-time (SMRT) long-read sequencing | human genome | none | structural variants including retrotransposons | PacBio RSII / Sequel |
| nanopore long-read sequencing | genomes / cancer / plant viruses | none | structural variants, pathogen detection | Oxford Nanopore (ONT) |
| GC-content analysis of insertion sites | human genome (non-reference L1Hs insertions) | none | GC content of insertion regions | — |
- – TEs constitute nearly 45% of the human genome vs ~1.5% protein-coding 45% vs 1.5%
- – Non-reference L1Hs insertion sites showed GC content suggesting random integration 41.15% GC
- ▼ L1 insertion regions previously reported lower GC content than overall genome 36–38% vs 41%
- – TruSeq synthetic long-reads correctly identified annotated TEs in Drosophila 77.8%
- – TruSeq synthetic long-reads achieve long lengths with low error 1.5–18.5 kb, ~0.03% error/base
- ▲ DNBSEQ-T7 detected slightly more indels than NovaSeq 6000
- – MGISEQ-2000 detected a loss missed by NextSeq500 101–133 bp loss
- – Nanopore average error rate 5–13%
- other ~45% of human genome is TEs (TE genome fraction)
- other ~17% of human genome is L1 (LINE-1 genome fraction)
- count >500,000 L1 copies (L1 copy number, propagated ~150 million years ago)
- count >1.1 million Alu copies (Alu interspersed copies in humans)
- count ~3000 SVA elements (SVA copies, ~0.2% of genome)
- other 41.15% GC content (non-reference L1Hs insertion sites (Shin et al. 2019))
- other 77.8% (annotated Drosophila TEs identified by TruSeq long-reads)
- other 0.1% (proportion of human genetic diseases caused by novel Alu insertions)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a narrative review article summarizing existing literature on human retrotransposons and computational detection methods for next-generation sequencing data. It does not present original experimental data, statistical comparisons, or hypothesis tests; it synthesizes findings and tools reported in prior publications.
-
The paper is structured as a qualitative narrative synthesis of prior studies on retrotransposon biology and detection tools, without a formal systematic review protocol.↳ Could also: A systematic review approach with predefined search strategy, inclusion/exclusion criteria, and PRISMA-style reporting could also be used — A systematic framework can increase transparency and reproducibility of the literature synthesis process, which may be useful when comparing performance of many computational tools.
-
Performance figures for sequencing platforms and detection tools (e.g., error rates, detection percentages) are cited from individual source studies without pooled quantitative synthesis.↳ Could also: A meta-analysis or quantitative benchmarking comparison across tools/platforms could also be conducted — Pooling comparable metrics (e.g., sensitivity/specificity of TE detection tools) across studies with a common statistical framework could allow formal comparison of methods, where sample sizes and variability permit.
-
Claims about GC content differences in L1 insertion regions across studies (e.g., 36–38% vs. 41.15%) are presented descriptively without a formal statistical test of difference.↳ Could also: A formal statistical comparison (e.g., chi-square test or two-proportion z-test) of GC content distributions between studies could also be used — This would allow a quantitative assessment of whether observed differences in GC content exceed what would be expected by chance, supplementing the narrative description.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope analysis — pmid-36295018
Title: Human Retrotransposons and Effective Computational Detection Methods for Next-Generation Sequencing Data. Authors: Lee H, Min JW, Mun S, Han K. Life (Basel) 2022. DOI 10.3390/life12101583. PMCID: PMC9605557.
Article type — determinative
This publication is an explicitly labeled REVIEW article (MDPI Life, Special-issue review). Verified directly from the full text:
- Abstract opens descriptively ("Transposable elements (TEs) are classified into two classes according to their mobilization mechanism…"), i.e. a survey, not a study with a hypothesis/result.
- No Materials & Methods section exists. The paper describes no original computational analysis run by these authors.
- No Code Availability statement. (No code was written or run by the authors.)
- Data Availability statement: "The data used to support this study are included in the manuscript." → no data was deposited or analyzed by the authors.
- Author Contributions list only conceptualization / investigation / writing / visualization / supervision — no software, no formal analysis, no data curation.
What the harvested artifacts actually are
The RU was created by text-mining a code link and a data accession out of the body of this review. Neither is an artifact the review's authors produced:
- Code
https://github.com/tk2/RetroSeq— RetroSeq is one of ~8 third-party detection tools the review tabulates (Table 3). It is the tool from the ORIGINAL RetroSeq paper, Keane, Wong & Adams 2013 (Bioinformatics), cited as reference [119]. The review did not run it. - Data
SRS228129— appears in exactly one sentence: "Seven people were selected as follows: a trio of Yoruban (NA18506, NA18507, and NA18508) and CEU (NA12891, NA12892, and NA12878) and an unreported ancestor (SRS228129)." This describes the sample panel used in a study the review summarizes (the RetroSeq/1000G mobile-element evaluation), not data the review authors generated, deposited, or analyzed. SRS228129 = ENA samplesnyder_test_sample(SAMN00672450), human WGS.
Candidate numeric claims in the paper, and why none are in scope
Every number in this review is a citation of prior literature, not a pipeline-derived result of this paper:
| Candidate value | Location | Origin |
|---|---|---|
| RetroSeq sensitivity ">90%" | Table 3 | cited from ref [119] (Keane 2013) |
| "average sensitivity of 97% and 83% for detecting Alu and L1 elements" | RetroSeq paragraph | cited from ref [119] |
| MELT >99%, xTea >90%, AluMine >98%, etc. | Table 3 | each from its own cited paper |
| "17% of L1, 11% of Alu, 0.2% of SVA" composition | Fig 1 / text | genome-annotation figure, cited |
In-scope (pipeline-derived results produced by THIS paper): NONE.
Why running RetroSeq on SRS228129 does NOT rescue this RU
The brief's P16 rule (applying a third-party tool to the paper's own data is a valid reproduction) does not apply here, because:
- The review has no own data — it deposited nothing and ran nothing. SRS228129 belongs to the original RetroSeq evaluation (a different publication), not to this review.
- The only RetroSeq numbers the review states are sensitivities (97% Alu / 83% L1). A sensitivity is TP / (TP+FN) against a gold-standard truth set of known mobile-element insertions for that sample. This review provides no truth set, and none is derivable from it. Reproducing those numbers means reproducing Keane et al. 2013, not this paper.
- No specific reported value in this paper is tied to these authors running
RetroSeq on SRS228129. Running RetroSeq would yield a call count, but the paper
reports no such count → nothing to compare 1:1 (
no_expected_result).
Decision
Drop — non_pipeline (text-mining false positive: a review article harvested
as if it shipped a reproducible pipeline). Secondary: no_expected_result — no
v
No individual results have been recorded for this entry yet.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
PMID 36295018 is an explicitly labeled review article (Lee et al., Life 2022) with no Materials & Methods, no Code Availability, and no deposited data — the harvested github.com/tk2/RetroSeq and SRS228129 are entities it merely cites/mentions, so there are zero original pipeline results to reproduce. Every number (RetroSeq >90% / 97% Alu / 83% L1, MELT >99%, 17%/11%/0.2% composition) is a literature citation, making q1/q2 red (no comparable input/endpoint) but not an authors' defect. The non-reproducibility lies on our harvesting side (a text-mining false positive), there is no factual deviation in the authors' claims (no fabrication concern), and nothing in scope is refuted — hence the correct outcome is a non_pipeline drop, graded yellow overall rather than critical.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.