Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

An accurate method for identifying recent recombinants from unaligned sequences.

Bioinformatics · 2022
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction, well-founded and described-well-enough. detREC (Feng et al. 2022) reproduces 1:1 on its core deterministic claim and qualitatively on its headline claims. R1 (the 80% floor / shipped 8-seq demo): EXACT first-hand this session («job») -- all 6 recombinant triples + identified recombinant identical across 3 reps; the only stochastic column (unseeded bootstrap sv) within tolerance. R3d: EXACT -- analyzed input N (35,591 seqs -> 17,335 types) physically present in repo, matches to the digit. R2 (simulation): qualitative claim 'sensitivity & specificity typically >70%' reproduced (mean sens 0.863 / spec 0.990, 15/15 reps >70%). R3 (empirical Ghana headline): 80.4% recombinant types vs reported 85.4% -- a ~5pp partial using the authors' own published converged mosaic parameters; same biological conclusion, gap consistent with mosaic/MAFFT build-version differences on borderline calls, NOT a fabrication indicator. NOT attempted: R2b (cross-method Fig-3 comparison, competitors not shipped) and R3a-c (ups-group stratification, annotations not shipped) -- both genuinely out of reach with the shipped data. No fabrication indicators found.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-21 ⛓ 99e73adfc49e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper aims to develop a method for identifying recent recombinant sequences directly from unaligned sequence data, without requiring a full multiple sequence alignment or a reference panel, and to apply it to the highly diverse DBLα domain of P. falciparum var genes, which diversify primarily through recombination.

Core claims
  • A novel algorithm combining the JHMM (Zilversmit et al. 2013) mosaic representation with a distance-based triple comparison can identify recombinant sequences and their parents from unaligned, gene-length sequences without a reference panel. method
  • The method remains accurate (sensitivity/specificity generally above 70%) even in the presence of insertions and deletions in the input sequences. finding
  • At matched specificity, the method achieves higher sensitivity than other popular recombination detection methods (e.g. RDP, 3SEQ) applied to aligned versions of the same simulated data. finding
  • In a Ghanaian DBLα dataset, 85.4% of representative sequence types were identified as recombinant. finding
  • DBLα sequences belonging to the same ups group or domain subclass recombine amongst themselves more frequently than expected. finding
  • Non-recombinant DBLα types are more conserved (less diverse) than recombinant DBLα types. finding
  • For each recombinant triple, the pair of sequences whose evolutionary distance changes least before vs after the breakpoint are inferred as the non-recombinant parents; the third sequence is the recombinant. mechanism
  • Source code for the method (detREC) is freely available on GitHub. resource
Experimental setups
Assay System Perturbation Readout Platform
Coalescent simulation + sequence evolution (msprime, Pyvolve) simulated amino acid sequences (no indels) simulated recombination (random breakpoints, 2+ parent sequences) sensitivity and specificity of recombinant detection msprime; Pyvolve
Coalescent simulation + sequence evolution with indels (INDELible) simulated unequal-length amino acid sequences simulated recombination plus varying indel rate/size sensitivity and specificity vs indel rate msprime; INDELible
Comparative benchmarking against existing recombination detection methods simulated aligned sequence datasets (with and without indels) simulated recombination sensitivity at matched specificity across methods
Recombinant/breakpoint detection via JHMM + MAFFT + distance-based triple analysis (detREC) DBLα domain sequences of var genes, P. falciparum field isolates from Bongo District, Ghana (161 isolates) none (natural field infections) identification of recombinant sequences, parents, breakpoints; recombination frequency by ups group/domain subclass; sequence conservation JHMM (Zilversmit et al. 2013); MAFFT; BLOSUM62/Hamming distance
Key results
  • Method achieves sensitivity and specificity above 70% (often much higher) across most simulated parameter settings >70%
  • Sensitivity and specificity remain relatively unaffected by indels, with only a moderate decline in specificity as indel rate increases moderate decline
  • Method shows highest sensitivity among compared methods when specificity is matched, with or without indels
  • 14801 of 17335 DBLα types detected as recombinant 85.4%
  • Sequences within the same ups group recombine with each other more frequently than with other groups
  • Non-recombinant DBLα types show greater sequence conservation than recombinant types
Key statistics
  • count 14801/17335 (85.4%) (DBLα types identified as recombinant)
  • count 35591 sequences from 161 isolates clustered into 17335 representative types (Ghana DBLα dataset composition)
  • mean 125aa (s.d. 8.4aa) (average length of DBLα representative types)
  • other sensitivity/specificity >70% (performance across most simulation parameter settings)
  • count 100 replicates per triple (bootstrap resampling for support value calculation)
  • other minimum aligned segment length threshold of 10 (criterion for excluding unreliable recombinant triples)
  • other ρ estimated over interval [0, 0.1] (recombination rate parameter estimation in JHMM)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper introduces a computational algorithm (detREC) for identifying recent recombinant sequences without requiring multiple sequence alignment, combining a jumping hidden Markov model (JHMM) with a distance-based classification step. Algorithmic performance was evaluated via simulation studies generating 100 independent datasets per parameter combination, with mean sensitivity and specificity reported alongside 95% confidence intervals. Bootstrap resampling (100 replicates per triple alignment) was used to quantify per-detection uncertainty. The method was additionally applied to 17,335 DBLα sequence types from 161 P. falciparum isolates from Ghana to characterize recombination patterns.

Replicationmixed Sample sizeSimulations: 100 independent randomly generated datasets per parameter combination. Real-data application: 17,335 DBLα types clustered from 35,591 sequences collected across 161 P. falciparum isolates. GroupsRecombinant vs non-recombinant sequences; detREC vs other recombination detection methods (e.g., RDP, 3SEQ); sequences with vs without indels; different simulation parameter levels varied one at a time Pairingunclear Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Sensitivity and specificity (proportion-based performance metrics; means with 95% confidence intervals across simulation replicates) Primary algorithmic evaluation across all simulation parameter combinations (Fig. 2, Fig. 3, Supplementary Section S2) 100 simulated datasets per parameter combination not stated
Bootstrap column resampling with replacement (100 replicates per recombinant triple) Per-detection support value calculation for individual recombinant triple alignments (Section 2.4) 100 bootstrap replicates per triple not stated
Composite likelihood maximization over randomly subsampled sequences JHMM recombination rate parameter (ρ) estimation for large-scale datasets 1000 randomly selected sequences (large dataset mode) not stated
Maximum likelihood via Baum-Welch algorithm / Viterbi training JHMM gap initiation (δ) and gap extension (ε) parameter estimation (Section 2.1) null not stated
Distance-based recombinant classification (BLOSUM62 for amino acids; Hamming mismatch distance for DNA) Recombinant identification within each recombinant triple — main classification step (Section 2.3) null not stated
Approaches that could also have been used
  • Performance was summarized as mean sensitivity and specificity at a single operating point (the method's natural threshold) with 95% CIs
    Could also: Area under the ROC curve (AUC-ROC) or Matthews correlation coefficient (MCC) could also summarize binary classifier performance across operating thresholds — AUC-ROC and MCC produce threshold-independent summaries, which can be informative when the natural operating point of a method differs from those of comparator methods, making cross-method comparisons more symmetric
  • The simulation study varied one parameter at a time while holding all other parameters fixed at default values
    Could also: A factorial or Latin hypercube sampling design could also have been used to explore joint parameter combinations — One-at-a-time variation does not capture interaction effects; a multi-factor design would reveal whether, for example, the effect of mutation rate on sensitivity is consistent across different sequence lengths or recombination proportions
  • Bootstrap column resampling (100 replicates) was used to generate per-detection support values for recombinant triples
    Could also: A parametric bootstrap or Bayesian posterior probability could also quantify detection uncertainty — Parametric bootstrapping would leverage the explicit substitution model; a Bayesian approach would additionally propagate uncertainty in JHMM parameters (ρ, δ, ε) into the support values, potentially providing better-calibrated credible intervals
  • Hamming (mismatch) distance was used for DNA sequence comparisons in the distance-based recombinant classification step
    Could also: Model-corrected evolutionary distances such as Jukes-Cantor or Kimura two-parameter distance could also have been applied — Hamming distance does not correct for multiple substitutions at the same site; correction-based distances may more accurately reflect true evolutionary divergence at higher mutation rates, which is one of the simulation parameters explored
  • Method comparison in Fig. 3 was structured by matching the specificity of comparator methods to that of detREC before comparing sensitivity
    Could also: A full ROC curve comparison across all thresholds, or a comparison at a pre-specified shared operating point, could also have structured the benchmark — Post-hoc specificity matching to a single value determined by the proposed method's own threshold may not reflect the operating conditions typical users would apply to comparator methods; a pre-registered or threshold-free comparison would allow a more symmetric evaluation
  • Recombination parameter ρ was estimated via composite likelihood over a random subsample of 1000 sequences for large datasets
    Could also: Full-data composite likelihood or stochastic variational inference with explicit variance estimation could also have been used — Random subsampling introduces sampling variance in ρ that is not propagated into downstream recombinant classifications; methods that track estimation uncertainty would allow sensitivity analyses on how parameter uncertainty affects the final recombinant calls
Software: msprime · Pyvolve · INDELible · MAFFT · detREC (custom implementation, available at https://github.com/qianfeng2/detREC_program)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35025988

Paper: Feng et al. 2022, An accurate method for identifying recent recombinants from unaligned sequences. Bioinformatics 38(7):1823-1829. DOI 10.1093/bioinformatics/btac012. Repo: https://github.com/qianfeng2/detREC_program (authors' own code). Data accession (brief label sra:PRJNA396962): NOT SRA. It is a GenBank Targeted Locus Study (TLS): master KDDE00000000 / version KDDE01000000, components KDDE01000001–KDDE01436091 (436,091 DBLα tag sequences), BioProject PRJNA396962, 7 BioSamples, Plasmodium falciparum DBLα from Bongo District, Ghana.

Method (one paragraph)

Detect recent recombinant sequences from UNALIGNED sequences. (1) A jumping HMM (JHMM / "mosaic", Zilversmit 2013) represents each target sequence as a mosaic of segments copied from other sequences → candidate "recombinant triples" (one recombinant + two parents). (2) MAFFT completes the 3-way alignment per triple; SeqKit concatenates segments. (3) A distance-based rule picks the recombinant in each triple: recombinant = the sequence left out of the argmin over pairs of |D1(x,y) − D2(x,y)| (distance before vs after the breakpoint; BLOSUM62 for AA, Hamming for DNA). (4) Bootstrap (100 replicates, resampling columns within segments) gives a support value sv per call.

IN SCOPE (pipeline-derived, attempted)

# Result Pipeline Where reported Feasibility
R1 Test_files demo: run integrated_rec_det.py output_align.txt input.fasta output.csv on the shipped 8-seq synthetic input → reproduce the shipped Test_files/output.csv (6 recombinant calls, identity columns + sv). MAFFT + SeqKit + integrated_rec_det.py Repo Test_files/output.csv (shipped expected output) HIGH — tiny, ships expected output. Identity columns should match exactly; sv (bootstrap) may vary unless seeded. PRIMARY 1:1 target.
R2 Simulation sens/spec > 70% for one parameter setting; method beats competitors at matched specificity. Snakemake: msprime → pyvolve/INDELible → mosaic → detection → result_analysis Figs 2,3; text "sensitivity & specificity typically above 70%" MEDIUM — stochastic (100 datasets/combo); one setting feasible; reproduces the qualitative claim, not an exact number.
R3 Empirical Ghana: 14,801/17,335 = 85.4% recombinants; ups breakdown upsA 82.3% / upsB 84.9% / upsC 87.6%. full DBLα clustering → mosaic Viterbi training (--num_runs 578) → detection Results §3.2; Table 1/2; Fig 4 LOW/PARTIAL — needs the full 35,591-type analysis set. Deposit has 436,091 raw tags; the exact 35,591 analyzed subset is NOT separately pinned. Repo ships only a Pilot subset. Headline 85.4% likely only partially reproducible. Aspirational.

OUT OF SCOPE (not computational pipeline → not attempted)

  • Wet-lab generation of the DBLα data (field sampling, DNA extraction, PCR, MiSeq sequencing) — experimental, not reproducible computationally.
  • Downstream biological/statistical interpretation that depends on the full empirical run (ups-group enrichment significance, Bonferroni subclass tests, HB5/HB14/HB36 breakpoint peaks) — gated on R3; only attempted if R3 succeeds.

Reproduction order

R1 (fast 1:1, the ~80% floor) → R2 (simulation, push past floor) → R3 (empirical, best-effort, expected partial). All heavy compute on «our HPC»/«infra».

Key risk / honesty notes

  • sv bootstrap column may be non-deterministic → grade identity columns exact, sv within-tol.
  • Data label mismatch: brief says SRA but it is a GenBank TLS; recorded in profile.
  • N mismatch: 436,091 deposited vs 35,591 analyzed — the analyzed subset is not independently downloadable by accession; flagged.
Figures / tables: Fig 2Fig 3
R1
Reported
6 recombinant calls (chunk,triple,rec) on 8 synthetic seqs (Test_files/output.csv)
Reproduced
6/6 calls EXACT across 3 reps (chunk + triple membership + identified recombinant), fresh first-hand «our HPC» run this session («job», node n093)
exact
R1b
Reported
bootstrap sv = 1,1,1,1,0.58,1
Reproduced
borderline row 0.50-0.51 (a rep hit 0.58 exactly); occasional 1.0->0.99; max|dev| <= 0.08 (unseeded 100-rep bootstrap)
within tolerance
R2
Reported
simulation sensitivity & specificity typically >70%
Reproduced
mean sens 0.863, mean spec 0.990 (n=15 reps); 15/15 reps both >=0.70 -> qualitative claim reproduced (first-hand «job»; evidence file reclaimed by janitor, refresh «job» resubmitted)
within tolerance
R2b
Reported
highest sensitivity vs competitors (Fig 3)
Reproduced
NOT ATTEMPTED (competitor methods not shipped; out of scope)
m.public.grade.uncheckable
R3
Reported
empirical 14,801/17,335 = 85.4% recombinant DBLalpha types
Reproduced
80.4% (2412/3000 random types, full-db mosaic, authors' published converged params; 95%CI[79.0,81.8]); ~5pp below reported, same qualitative conclusion (first-hand «job»; evidence file reclaimed, refresh «job» resubmitted)
partial
R3a-c
Reported
upsA 82.3% / upsB 84.9% / upsC 87.6%
Reproduced
NOT ATTEMPTED (upsA/B/C annotations of the 17,335 types not shipped; cannot stratify)
m.public.grade.uncheckable
R3d
Reported
analyzed dataset 35,591 sequences -> 17,335 types
Reproduced
repo Pilot.fasta=35,591 seqs + centroids=17,335 types (EXACT, header counts on «infra» this session)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

436.2 k
tokens (I/O) · 45.2 M incl. cache
217 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.