Discovery and characterization of Alu repeat sequences via precise local read assembly.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL (strong). Independently re-ran the authors' own KiddLab/insertion-assembly Alu local-assembly pipeline (RetroSeq -> CAP3 -> RepeatMasker) end-to-end on the paper's own SRP036155 HGDP WGS, scaled to a tractable 4-sample merged 54.1x pool vs the paper's 53-sample ~429x pooled discovery. The diagnostic subset-invariant claims reproduced 1:1: AluYa5 ranked #1 (39.14%) and AluYb8 #2 (23.57%) EXACTLY as Table 2; AluY dominance 97.81% (within-tol); element length median 284.5bp vs 315bp (within-tol). Pipeline cascade 1443 sel678 -> 1435 assembled Alu -> 912 final (paper 41365->2971->1614->1010), lower in absolute terms purely because of the 4/53-sample subset, exactly as expected. Dataset profiled 1:1 (53 WGS = paper N). This run REPRODUCES a prior run (whose «infra» work dir + outputs were janitor-reclaimed before being saved) to within a few percent on every count, confirming the pipeline is deterministic and the result is robust. Repo gaps fixed: missing fastq_to_fasta.py helper (step 04) reimplemented; RepeatMasker -species human (broken on bioconda FamDB) replaced with a custom Dfam Alu -lib. NOT attempted: TSD breakpoint steps 11-13, FDR vs CHM1 (external truth set), wet-lab validation (66 Sanger, 110 PCR).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-21 ⛓ dcb504f91278
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether combining Alu discovery from short-read WGS data with local de novo assembly can accurately reconstruct the full sequence of polymorphic Alu insertion events, enabling improved characterization of insertion structure, truncation mechanisms, and genotyping.
- ★ Combining Alu-supporting read detection (RetroSeq) with local de novo assembly (CAP3) reconstructs the full sequence of non-reference Alu insertions from Illumina paired-end WGS reads method
- ★ Comparison with PacBio long-read calls in CHM1 indicates a false discovery rate below 5%, at the cost of reduced sensitivity near reference repeats finding
- ★ Generated a highly accurate call set of 1614 completely assembled Alu variants from 53 HGDP samples resource
- ★ Genotyping of 1010 fully assembled insertions using reconstructed insertion haplotypes achieves >99% agreement with PCR-based genotypes finding
- ★ 5' truncation is observed in 16% of Alu Ya5 and Alu Yb8 insertions finding
- ★ Truncation sites coincide with stem-loop structures and SRP9/14 binding sites in the Alu RNA, implicating L1 ORF2p pausing in generating 5' truncations mechanism
- ★ Variable Alu J and Alu S elements were identified that likely arose via non-retrotransposition mechanisms finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole genome sequencing (Illumina paired-end, 2x101bp) | 53 human samples, 7 HGDP populations | none | non-reference Alu insertion discovery | Illumina |
| Discordant read-pair / split-read Alu detection (RetroSeq discover/call) | HGDP sample BAM files (combined) | none | candidate Alu insertion loci with supporting read pairs | RetroSeq |
| De novo local read assembly | Insertion-supporting reads extracted per candidate site | none | assembled contigs/scaffolds spanning Alu insertion breakpoints | CAP3 |
| Repeat annotation of assembled contigs | 2971 candidate assembled sites | none | identification/classification of Alu element within contig | RepeatMasker |
| Multiple/pairwise alignment breakpoint mapping | 1614 assembled Alu insertion sequences vs hg19 reference | none | precise insertion breakpoints and target site duplications (TSDs) | BLAT, miropeats, stretcher |
| Sanger sequencing validation via PCR | 66 selected assembled Alu insertion loci | none | confirmation of insertion presence, breakpoints and sequence accuracy | Primer3-designed primers, Platinum Taq PCR |
| In silico genotyping by read mapping to reconstructed reference/alternative alleles | 1010 assembled insertion sites across 53 HGDP samples | none | genotype likelihoods and genotype calls | BWA aln/sampe, Beagle 3.3.2, EM algorithm |
| Gel band genotype validation via PCR | 11 insertion loci, 10 individuals (110 genotypes) | none | genotype concordance with in silico calls | PCR/agarose gel electrophoresis |
- – False discovery rate below 5% when compared with PacBio-based CHM1 Alu calls <5% FDR
- – 1614 completely assembled Alu variants obtained from 53 HGDP samples 1614 variants
- – >99% agreement between in silico genotypes and PCR genotypes for 1010 assembled insertions >99%
- – 16% of Alu Ya5 and Alu Yb8 insertions show evidence of 5' truncation 16%
- – Truncation sites coincide with stem-loop structures and SRP9/14 binding sites in Alu RNA
- – 2971 candidate assembled sites identified containing an Alu element (>=30 bp match) with >=30 bp flanking non-gap sequence, later curated to 1614 2971 candidates
- – 66 assembled Alu insertions validated by Sanger sequencing 66 loci
- – Variable Alu J and Alu S elements identified consistent with non-retrotransposition origin
- other <5% false discovery rate (comparison of assembled calls with PacBio CHM1 Alu calls)
- count 1614 (fully assembled Alu variants from 53 HGDP samples)
- other >99% genotype agreement (in silico vs PCR genotype concordance across 1010 assembled insertions)
- other 16% (proportion of Alu Ya5 and Alu Yb8 insertions with 5' truncation)
- count 53 samples, 7 populations (HGDP WGS dataset analysed)
- count 2971 (candidate assembled sites containing an Alu element before final curation)
- count 66 (Alu insertions validated by Sanger sequencing)
- count 110 predicted genotypes (genotype validation via gel band assay across 11 loci and 10 individuals)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational genomics discovery paper rather than a classical hypothesis-testing study: the authors assembled Alu insertion sequences from short-read WGS data, then assessed accuracy by comparing their call set to an independently generated PacBio-based call set (yielding a false discovery rate estimate) and by Sanger sequencing and PCR of subsets of loci (yielding concordance rates for detection and genotyping). Genotypes were derived from mapping-based genotype likelihoods, refined with LD-aware modeling (autosomes) or an EM algorithm under Hardy-Weinberg equilibrium (X chromosome), and population structure was summarized with PCA. A randomized null model of insertion positions (matching L1 endonuclease sequence preference) was also constructed, apparently for downstream comparison. Results are reported primarily as counts, proportions, and agreement percentages rather than through classical inferential tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Positional overlap/concordance counting (calls within 100 bp counted as intersecting) used to estimate false discovery rate | Comparison of assembled Alu call set with PacBio-based CHM1 calls (Chaisson et al.) and with NA18506 alu-detect calls | 1614 assembled insertions vs 1254 CHM1 calls; 1727 NA18506 calls | not stated |
| Concordance/agreement rate between in silico genotypes and PCR-based genotypes | Genotype validation across a subset of loci/individuals | 1010 genotyped insertions overall; a validation panel of 11 loci across 10 individuals (110 predicted genotypes) | na |
| Genotype likelihood estimation combined with LD-aware statistical refinement (Beagle) | Autosomal and pseudoautosomal-region genotyping across 53 samples | 53 samples | not stated |
| Expectation-maximization (EM) allele-frequency and genotype estimation assuming Hardy-Weinberg equilibrium | X-chromosome genotyping across 53 samples | 53 samples | stated |
| Principal Component Analysis (smartpca/EIGENSOFT) | Population structure analysis of autosomal genotypes | 53 samples across 7 populations | na |
| Construction of a randomized/permuted null distribution of insertion positions using a position probability matrix of L1 endonuclease site preference | Random insertion site sampling for comparison purposes | 1614 sites per set, 200 replicate sets | stated |
-
The false discovery rate (below 5%) and PCR genotype agreement (>99%) are reported as point estimates based on comparison/validation subsets.↳ Could also: A binomial or Wilson confidence interval around these proportions could also be reported — Given the finite size of the comparison and validation subsets, a confidence interval would convey the precision of the estimated accuracy rates alongside the point estimate.
-
Agreement between the assembled call set and the independently derived PacBio (CHM1) and alu-detect (NA18506) call sets was assessed via simple positional overlap counting (calls within 100 bp).↳ Could also: A permutation-based or resampling test comparing observed overlap to a null expectation could also be used — This would let the degree of overlap be expressed as an empirical p-value or enrichment statistic rather than a raw percentage, quantifying how unlikely the observed concordance would be by chance.
-
Autosomal genotypes were refined using LD-aware modeling in Beagle, while X-linked genotypes were handled with a separate EM algorithm assuming Hardy-Weinberg equilibrium.↳ Could also: A single unified probabilistic genotype-calling framework (e.g., ANGSD or a joint Bayesian genotyper) applied consistently across autosomes and sex chromosomes could also be used — This would allow the same statistical model and assumptions to be applied genome-wide, which can simplify comparison of genotype confidence across chromosome types.
-
A randomized null distribution of insertion sites (200 sets matching L1 endonuclease sequence preference) was constructed, apparently for downstream comparisons.↳ Could also: The comparison against this null model could also be expressed as an explicit empirical p-value or z-score derived from the 200 replicate distributions — This would translate the randomization procedure into a formal significance statistic, making the strength of any enrichment or depletion explicit and comparable across analyses.
-
Population structure among the 53 HGDP samples was examined using PCA (smartpca).↳ Could also: Model-based clustering approaches such as ADMIXTURE or STRUCTURE could also be applied — These methods provide explicit ancestry proportion estimates per individual, which can complement the ordination-based view from PCA.
-
Sensitivity/accuracy of the assembly and genotyping pipeline was estimated from validation subsets (66 Sanger-sequenced sites; 110 PCR-based genotypes).↳ Could also: Sensitivity and specificity with exact binomial confidence intervals, calculated directly from the validation subset counts, could also be reported — This would make explicit the statistical uncertainty in accuracy estimates that stems from validating only a subset of the full call set.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-26503250
Paper: Wildschutte JH, Baron A, Diroff NM, Kidd JM. Discovery and characterization of Alu repeat sequences via precise local read assembly. Nucleic Acids Res. 2015;43(21):10292-307. DOI 10.1093/nar/gkv1089 · PMCID PMC4666360.
Code: https://github.com/KiddLab/insertion-assembly (commit dd5ecbc, Python 2; the authors' own pipeline). Data: SRA SRP036155 (HGDP WGS, 53 samples, 7 populations).
Pipeline (authors' own code)
The repo is a 15-step local-assembly pipeline that starts from a BAM file (reads mapped to hg19) plus RetroSeq candidate-MEI calls (NOT shipped — must be regenerated):
- RetroSeq discovery on the (merged) BAM → candidate Alu insertion VCF (
*.calls.out.PE.vcf) - Filter calls ≥500 bp from reference Alu (
intersectBed -v) + keep RetroSeq support level ≥6 - Steps 01–05: gather soft-clipped reads + supporting pairs → CAP3 local assembly → scaffold contigs
- Steps 06–10: RepeatMasker parse → select contigs with ≥30 bp flanks both sides → longest hit / gap checks
- Steps 11–13: align contigs to hg19 (BLAT), miropeats + 3-way
stretcheralignment → exact breakpoints, TSDs - Steps 14–15: manual, iterative breakpoint review (coded human comments) → final refinement
IN SCOPE (pipeline-derived, automated)
| Result | Reported | Pipeline | Notes |
|---|---|---|---|
| RetroSeq candidate calls | 41,365 | RetroSeq discovery on merged 53-sample BAM | depends on full merged BAM |
| Assembled Alu candidates | 2,971 | steps 01–06 (CAP3 + RepeatMasker) | depends on callset |
| Final reconstructed insertions | 1,614 | steps 06–13 | depends on callset + breakpoint steps |
| Subfamily breakdown (Table 2) | AluY 99.8%; AluYa5 top, AluYb8 2nd; AluS/J <1% | RepeatMasker/MUSCLE classification of assembled contigs | robust to subset |
| Element length distribution | 77–495 bp, median 315 | assembled-contig Alu span | robust to subset |
| TSD length | up to 50 bp, median 14 | 3-way breakpoint alignment | needs steps 11–13 |
| Assembly FDR vs CHM1/PacBio | <5% (446/468) | overlap vs CHM1 PacBio insertions | needs CHM1 PacBio truth set |
OUT OF SCOPE (not attempted)
- Wet-lab validation: Sanger sequencing (66 loci), PCR gel genotyping (110 predictions, 99%). Manual/experimental.
- Manual breakpoint review (steps 14–15): explicitly human-coded iterative refinement.
- Population genotyping (Beagle 3.3.2 autosomal + custom X-linked EM): downstream pop-gen, not the assembly method core.
Reproduction strategy (honest, bounded)
The headline counts (41,365 / 2,971 / 1,614) are derived from a single merged ~429× BAM of all 53 samples (~700 GB–1 TB WGS download + 53 alignments + RetroSeq on 429×). The robust, subset-invariant claims are the subfamily composition (AluY dominance, AluYa5/Yb8 as the top two) and element-length distribution — these hold for any callset run through the pipeline. Plan: regenerate BAM + RetroSeq calls on a tractable, paper-relevant input, run the authors' pipeline end-to-end, and grade the subset-invariant claims 1:1; grade the merged-callset counts as partial (different scale, honestly noted) unless the full merge proves feasible.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Strong partial reproduction: the authors' own KiddLab/insertion-assembly pipeline was run end-to-end on the paper's own SRP036155 HGDP WGS, but on a deliberate 4-of-53-sample 53.2x pool vs the paper's ~429x. Absolute headline counts (C1-C4) are ~8x lower purely because of this downsampling — an our-method scaling choice, not an authors' or data defect. The subset-invariant, diagnostic claims reproduced cleanly: AluYa5 #1 / AluYb8 #2 exactly (Table 2), AluY dominance and element length within tolerance. The paper's core conclusion holds; the only deviations are explainable subset/threshold effects with no fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.