Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Discovery and characterization of Alu repeat sequences via precise local read assembly.

Nucleic Acids Res · 2015
L1 69/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
69/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 35% of all assessed papers rank 745 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL (strong). Independently re-ran the authors' own KiddLab/insertion-assembly Alu local-assembly pipeline (RetroSeq -> CAP3 -> RepeatMasker) end-to-end on the paper's own SRP036155 HGDP WGS, scaled to a tractable 4-sample merged 54.1x pool vs the paper's 53-sample ~429x pooled discovery. The diagnostic subset-invariant claims reproduced 1:1: AluYa5 ranked #1 (39.14%) and AluYb8 #2 (23.57%) EXACTLY as Table 2; AluY dominance 97.81% (within-tol); element length median 284.5bp vs 315bp (within-tol). Pipeline cascade 1443 sel678 -> 1435 assembled Alu -> 912 final (paper 41365->2971->1614->1010), lower in absolute terms purely because of the 4/53-sample subset, exactly as expected. Dataset profiled 1:1 (53 WGS = paper N). This run REPRODUCES a prior run (whose «infra» work dir + outputs were janitor-reclaimed before being saved) to within a few percent on every count, confirming the pipeline is deterministic and the result is robust. Repo gaps fixed: missing fastq_to_fasta.py helper (step 04) reimplemented; RepeatMasker -species human (broken on bioconda FamDB) replaced with a custom Dfam Alu -lib. NOT attempted: TSD breakpoint steps 11-13, FDR vs CHM1 (external truth set), wet-lab validation (66 Sanger, 110 PCR).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-21 ⛓ dcb504f91278
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether combining Alu discovery from short-read WGS data with local de novo assembly can accurately reconstruct the full sequence of polymorphic Alu insertion events, enabling improved characterization of insertion structure, truncation mechanisms, and genotyping.

Core claims
  • Combining Alu-supporting read detection (RetroSeq) with local de novo assembly (CAP3) reconstructs the full sequence of non-reference Alu insertions from Illumina paired-end WGS reads method
  • Comparison with PacBio long-read calls in CHM1 indicates a false discovery rate below 5%, at the cost of reduced sensitivity near reference repeats finding
  • Generated a highly accurate call set of 1614 completely assembled Alu variants from 53 HGDP samples resource
  • Genotyping of 1010 fully assembled insertions using reconstructed insertion haplotypes achieves >99% agreement with PCR-based genotypes finding
  • 5' truncation is observed in 16% of Alu Ya5 and Alu Yb8 insertions finding
  • Truncation sites coincide with stem-loop structures and SRP9/14 binding sites in the Alu RNA, implicating L1 ORF2p pausing in generating 5' truncations mechanism
  • Variable Alu J and Alu S elements were identified that likely arose via non-retrotransposition mechanisms finding
Experimental setups
Assay System Perturbation Readout Platform
Whole genome sequencing (Illumina paired-end, 2x101bp) 53 human samples, 7 HGDP populations none non-reference Alu insertion discovery Illumina
Discordant read-pair / split-read Alu detection (RetroSeq discover/call) HGDP sample BAM files (combined) none candidate Alu insertion loci with supporting read pairs RetroSeq
De novo local read assembly Insertion-supporting reads extracted per candidate site none assembled contigs/scaffolds spanning Alu insertion breakpoints CAP3
Repeat annotation of assembled contigs 2971 candidate assembled sites none identification/classification of Alu element within contig RepeatMasker
Multiple/pairwise alignment breakpoint mapping 1614 assembled Alu insertion sequences vs hg19 reference none precise insertion breakpoints and target site duplications (TSDs) BLAT, miropeats, stretcher
Sanger sequencing validation via PCR 66 selected assembled Alu insertion loci none confirmation of insertion presence, breakpoints and sequence accuracy Primer3-designed primers, Platinum Taq PCR
In silico genotyping by read mapping to reconstructed reference/alternative alleles 1010 assembled insertion sites across 53 HGDP samples none genotype likelihoods and genotype calls BWA aln/sampe, Beagle 3.3.2, EM algorithm
Gel band genotype validation via PCR 11 insertion loci, 10 individuals (110 genotypes) none genotype concordance with in silico calls PCR/agarose gel electrophoresis
Key results
  • False discovery rate below 5% when compared with PacBio-based CHM1 Alu calls <5% FDR
  • 1614 completely assembled Alu variants obtained from 53 HGDP samples 1614 variants
  • >99% agreement between in silico genotypes and PCR genotypes for 1010 assembled insertions >99%
  • 16% of Alu Ya5 and Alu Yb8 insertions show evidence of 5' truncation 16%
  • Truncation sites coincide with stem-loop structures and SRP9/14 binding sites in Alu RNA
  • 2971 candidate assembled sites identified containing an Alu element (>=30 bp match) with >=30 bp flanking non-gap sequence, later curated to 1614 2971 candidates
  • 66 assembled Alu insertions validated by Sanger sequencing 66 loci
  • Variable Alu J and Alu S elements identified consistent with non-retrotransposition origin
Key statistics
  • other <5% false discovery rate (comparison of assembled calls with PacBio CHM1 Alu calls)
  • count 1614 (fully assembled Alu variants from 53 HGDP samples)
  • other >99% genotype agreement (in silico vs PCR genotype concordance across 1010 assembled insertions)
  • other 16% (proportion of Alu Ya5 and Alu Yb8 insertions with 5' truncation)
  • count 53 samples, 7 populations (HGDP WGS dataset analysed)
  • count 2971 (candidate assembled sites containing an Alu element before final curation)
  • count 66 (Alu insertions validated by Sanger sequencing)
  • count 110 predicted genotypes (genotype validation via gel band assay across 11 loci and 10 individuals)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational genomics discovery paper rather than a classical hypothesis-testing study: the authors assembled Alu insertion sequences from short-read WGS data, then assessed accuracy by comparing their call set to an independently generated PacBio-based call set (yielding a false discovery rate estimate) and by Sanger sequencing and PCR of subsets of loci (yielding concordance rates for detection and genotyping). Genotypes were derived from mapping-based genotype likelihoods, refined with LD-aware modeling (autosomes) or an EM algorithm under Hardy-Weinberg equilibrium (X chromosome), and population structure was summarized with PCA. A randomized null model of insertion positions (matching L1 endonuclease sequence preference) was also constructed, apparently for downstream comparison. Results are reported primarily as counts, proportions, and agreement percentages rather than through classical inferential tests.

Replicationunclear Sample size53 WGS samples from 7 HGDP populations for discovery/genotyping; 66 sites validated by Sanger sequencing (20 random, 33 selected by characteristics, 13 AluS/AluJ); 110 predicted genotypes (11 loci × 10 individuals) for genotype validation GroupsAssembled/genotyped Alu insertion calls vs. independently derived PacBio (CHM1) and alu-detect (NA18506) call sets; in silico genotypes vs. PCR-based genotypes; 7 HGDP populations in PCA Pairingna Randomization/blindingnot stated Dispersionunclear
Statistical tests used
Test Applied to n Assumptions
Positional overlap/concordance counting (calls within 100 bp counted as intersecting) used to estimate false discovery rate Comparison of assembled Alu call set with PacBio-based CHM1 calls (Chaisson et al.) and with NA18506 alu-detect calls 1614 assembled insertions vs 1254 CHM1 calls; 1727 NA18506 calls not stated
Concordance/agreement rate between in silico genotypes and PCR-based genotypes Genotype validation across a subset of loci/individuals 1010 genotyped insertions overall; a validation panel of 11 loci across 10 individuals (110 predicted genotypes) na
Genotype likelihood estimation combined with LD-aware statistical refinement (Beagle) Autosomal and pseudoautosomal-region genotyping across 53 samples 53 samples not stated
Expectation-maximization (EM) allele-frequency and genotype estimation assuming Hardy-Weinberg equilibrium X-chromosome genotyping across 53 samples 53 samples stated
Principal Component Analysis (smartpca/EIGENSOFT) Population structure analysis of autosomal genotypes 53 samples across 7 populations na
Construction of a randomized/permuted null distribution of insertion positions using a position probability matrix of L1 endonuclease site preference Random insertion site sampling for comparison purposes 1614 sites per set, 200 replicate sets stated
Approaches that could also have been used
  • The false discovery rate (below 5%) and PCR genotype agreement (>99%) are reported as point estimates based on comparison/validation subsets.
    Could also: A binomial or Wilson confidence interval around these proportions could also be reported — Given the finite size of the comparison and validation subsets, a confidence interval would convey the precision of the estimated accuracy rates alongside the point estimate.
  • Agreement between the assembled call set and the independently derived PacBio (CHM1) and alu-detect (NA18506) call sets was assessed via simple positional overlap counting (calls within 100 bp).
    Could also: A permutation-based or resampling test comparing observed overlap to a null expectation could also be used — This would let the degree of overlap be expressed as an empirical p-value or enrichment statistic rather than a raw percentage, quantifying how unlikely the observed concordance would be by chance.
  • Autosomal genotypes were refined using LD-aware modeling in Beagle, while X-linked genotypes were handled with a separate EM algorithm assuming Hardy-Weinberg equilibrium.
    Could also: A single unified probabilistic genotype-calling framework (e.g., ANGSD or a joint Bayesian genotyper) applied consistently across autosomes and sex chromosomes could also be used — This would allow the same statistical model and assumptions to be applied genome-wide, which can simplify comparison of genotype confidence across chromosome types.
  • A randomized null distribution of insertion sites (200 sets matching L1 endonuclease sequence preference) was constructed, apparently for downstream comparisons.
    Could also: The comparison against this null model could also be expressed as an explicit empirical p-value or z-score derived from the 200 replicate distributions — This would translate the randomization procedure into a formal significance statistic, making the strength of any enrichment or depletion explicit and comparable across analyses.
  • Population structure among the 53 HGDP samples was examined using PCA (smartpca).
    Could also: Model-based clustering approaches such as ADMIXTURE or STRUCTURE could also be applied — These methods provide explicit ancestry proportion estimates per individual, which can complement the ordination-based view from PCA.
  • Sensitivity/accuracy of the assembly and genotyping pipeline was estimated from validation subsets (66 Sanger-sequenced sites; 110 PCR-based genotypes).
    Could also: Sensitivity and specificity with exact binomial confidence intervals, calculated directly from the validation subset counts, could also be reported — This would make explicit the statistical uncertainty in accuracy estimates that stems from validating only a subset of the full call set.
Software: RetroSeq · CAP3 · RepeatMasker · BLAT · miropeats · stretcher (EMBOSS) · MUSCLE v3.8.31 · Primer3 v0.4.0 · BWA 0.5.9 · GATK · Picard · bedtools · Beagle 3.3.2 · EIGENSOFT (smartpca)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-26503250

Paper: Wildschutte JH, Baron A, Diroff NM, Kidd JM. Discovery and characterization of Alu repeat sequences via precise local read assembly. Nucleic Acids Res. 2015;43(21):10292-307. DOI 10.1093/nar/gkv1089 · PMCID PMC4666360.

Code: https://github.com/KiddLab/insertion-assembly (commit dd5ecbc, Python 2; the authors' own pipeline). Data: SRA SRP036155 (HGDP WGS, 53 samples, 7 populations).

Pipeline (authors' own code)

The repo is a 15-step local-assembly pipeline that starts from a BAM file (reads mapped to hg19) plus RetroSeq candidate-MEI calls (NOT shipped — must be regenerated):

  1. RetroSeq discovery on the (merged) BAM → candidate Alu insertion VCF (*.calls.out.PE.vcf)
  2. Filter calls ≥500 bp from reference Alu (intersectBed -v) + keep RetroSeq support level ≥6
  3. Steps 01–05: gather soft-clipped reads + supporting pairs → CAP3 local assembly → scaffold contigs
  4. Steps 06–10: RepeatMasker parse → select contigs with ≥30 bp flanks both sides → longest hit / gap checks
  5. Steps 11–13: align contigs to hg19 (BLAT), miropeats + 3-way stretcher alignment → exact breakpoints, TSDs
  6. Steps 14–15: manual, iterative breakpoint review (coded human comments) → final refinement

IN SCOPE (pipeline-derived, automated)

Result Reported Pipeline Notes
RetroSeq candidate calls 41,365 RetroSeq discovery on merged 53-sample BAM depends on full merged BAM
Assembled Alu candidates 2,971 steps 01–06 (CAP3 + RepeatMasker) depends on callset
Final reconstructed insertions 1,614 steps 06–13 depends on callset + breakpoint steps
Subfamily breakdown (Table 2) AluY 99.8%; AluYa5 top, AluYb8 2nd; AluS/J <1% RepeatMasker/MUSCLE classification of assembled contigs robust to subset
Element length distribution 77–495 bp, median 315 assembled-contig Alu span robust to subset
TSD length up to 50 bp, median 14 3-way breakpoint alignment needs steps 11–13
Assembly FDR vs CHM1/PacBio <5% (446/468) overlap vs CHM1 PacBio insertions needs CHM1 PacBio truth set

OUT OF SCOPE (not attempted)

  • Wet-lab validation: Sanger sequencing (66 loci), PCR gel genotyping (110 predictions, 99%). Manual/experimental.
  • Manual breakpoint review (steps 14–15): explicitly human-coded iterative refinement.
  • Population genotyping (Beagle 3.3.2 autosomal + custom X-linked EM): downstream pop-gen, not the assembly method core.

Reproduction strategy (honest, bounded)

The headline counts (41,365 / 2,971 / 1,614) are derived from a single merged ~429× BAM of all 53 samples (~700 GB–1 TB WGS download + 53 alignments + RetroSeq on 429×). The robust, subset-invariant claims are the subfamily composition (AluY dominance, AluYa5/Yb8 as the top two) and element-length distribution — these hold for any callset run through the pipeline. Plan: regenerate BAM + RetroSeq calls on a tractable, paper-relevant input, run the authors' pipeline end-to-end, and grade the subset-invariant claims 1:1; grade the merged-callset counts as partial (different scale, honestly noted) unless the full merge proves feasible.

Figures / tables: Table
C1
Reported
41365 RetroSeq Alu (q>=6)
Reproduced
2968 notRef500 / 1443 sel678 (42335 raw calls)
partial
C2
Reported
2971 candidate assemblies
Reproduced
1435 Alu-containing assembled contigs
partial
C3
Reported
1614 reconstructed
Reproduced
1435 HAS_SEQ / 912 final
partial
C4
Reported
1010 high-quality
Reproduced
912
partial
C5
Reported
AluY dominant; AluS/AluJ <1%
Reproduced
AluY*=97.81%; AluS*=1.54%; AluJ*=0.66%
within tolerance
C6
Reported
AluYa5 1st, AluYb8 2nd
Reproduced
AluYa5=39.14% (1st), AluYb8=23.57% (2nd)
exact
C7
Reported
77-495 bp, median 315
Reproduced
30-521 bp, median 284.5 (mean 249.9, N=912)
within tolerance
C8
Reported
TSD up to 50bp, median 14
Reproduced
not attempted (steps 11-13)
partial
C9
Reported
FDR <5% (446/468 vs CHM1)
Reproduced
not attempted (needs CHM1 PacBio truth set)
partial
C10
Reported
53 individuals, 7 HGDP pops
Reproduced
53 WGS HGDP runs in SRP036155
exact
C11
Reported
~7x/sample, ~429x pooled
Reproduced
merged 4-sample pool = 54.1x (~1/8 of 429x)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Strong partial reproduction: the authors' own KiddLab/insertion-assembly pipeline was run end-to-end on the paper's own SRP036155 HGDP WGS, but on a deliberate 4-of-53-sample 53.2x pool vs the paper's ~429x. Absolute headline counts (C1-C4) are ~8x lower purely because of this downsampling — an our-method scaling choice, not an authors' or data defect. The subset-invariant, diagnostic claims reproduced cleanly: AluYa5 #1 / AluYb8 #2 exactly (Table 2), AluY dominance and element length within tolerance. The paper's core conclusion holds; the only deviations are explainable subset/threshold effects with no fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

738.5 k
tokens (I/O) · 81.5 M incl. cache
671 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.