Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.

PLoS One · 2013
L1 45/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
45/100
Reproducibility score
1.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 6% of all assessed papers rank 1103 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Partial reproduction. The paper's third-party bubble-calling tool (bubbleparse) required four cumulative source patches (build-toolchain fixes plus two real logic bugs, the last a filelist-parsing bug that was completely blocking) but, once fixed, ran to completion on both A. thaliana crosses (Col-0xBur-0: 124,090 ranked bubbles; Col-0xTsu-1: 161,987), reproducing the paper's core SNP/indel/complex bubble-classification taxonomy qualitatively and its clean-SNP-bubble proportion within tolerance (~61% vs paper's 61.99% best case). Total bubble volume and reference-based SNP counts are both substantially below the paper's own reported numbers, but for identifiable, disclosed scope reasons (single depth/k point vs the paper's full sweep; a custom conservative pileup SNP caller used instead of SAMtools) rather than any indication the method itself fails to reproduce. Not attempted: the full depth x k parameter sweep, and the paper's specific 'percent of canonical SNPs recovered' sensitivity metric (no ground-truth SNP set was available). All dataset accessions (SRX000702/703/704, 12 SRR runs) were fully retrieved and integrity-verified.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-31
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can trait-linked single nucleotide polymorphisms be detected and classified directly from next-generation sequencing reads of genetically divergent, non-reference samples by identifying and ranking bubble structures in a multi-coloured de Bruijn graph? The paper tests whether such a reference-free method (Bubbleparse) is sensitive and specific enough to be useful for rapid marker discovery.

Core claims
  • Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module. method
  • The new bubble-detection algorithm identifies more complex bubble structures reference-free than the standard Cortex variant caller, which is restricted to clean bubbles. method
  • Bubbles are classified by type according to the number of paths and the number of colours following each path, enabling classification based on genetic background (e.g. heterozygous vs homozygous in resistant vs susceptible samples). method
  • Increasing graph search depth increases sensitivity for real SNPs but at a substantial cost in specificity, because complex graph structures more often reflect k-length repeats or read errors than true ecotype differences. finding
  • Excessive read depth does not improve SNP detection and increases false positives, because more reads introduce more error, creating false bubbles and more complex bubbles. finding
  • A ranking heuristic based on bubble attributes (coverage, kmer quality score, path length, coverage ratio/total difference) is required and effective for maximising the number of real SNPs recovered from the very large bubble collections. method
  • Bubbleparse is effective on data from unsequenced wild relatives of potato and enabled rapid identification of genes linked to Phytophthora infestans disease resistance in crosses of resistant and susceptible accessions. finding
  • Bubbleparse is released as a software resource: a new bubble-detection module within Cortex plus a ranking/classification tool. resource
Experimental setups
Assay System Perturbation Readout Platform
Reference-free variant detection by coloured de Bruijn graph assembly and bubble detection (Bubbleparse) Arabidopsis thaliana ecotypes Bur-0, Tsu-1 and Ler-1 reads compared against Col-0 reads none (natural ecotype variation); parameter sweep of search depth 0-3, k = 15-31, error cleaning removing nodes with coverage <=2 and tips <100 nt number of bubbles found, percentage of canonical SNPs recovered, percentage of called bubbles that are real SNPs Cortex framework; Illumina short reads (31-40 nt for Bur-0/Tsu-1, 40-80 nt for Ler-1) from the 1001 Genomes Project, ENA SRA experiments SRX000702-SRX000704
Alignment-based SNP calling (reference pipeline for benchmarking) Arabidopsis thaliana Bur-0 and Tsu-1 reads aligned to the Col-0/TAIR9 reference none; minimum alignment coverage thresholds of 10 and 20, requiring 95% of covering reads to carry the same non-reference nucleotide number of homozygous SNPs called (canonical SNP set) BWA for alignment, SAMTools for SNP calling
Sequence similarity search to localise bubble contigs Bubbleparse bubble contigs vs Arabidopsis thaliana TAIR9 reference genome none genomic location of bubbles and predicted SNP positions for comparison with canonical SNPs BLAST
Comparison against an externally curated canonical SNP set at varied sequencing depth Arabidopsis thaliana Ler-1 reads vs Col-0 average kmer coverage varied from 40x up to maximum available 340x; search depths 0-3 percentage of 1001 Genomes canonical SNPs recovered and false positive rate 1001 Genomes Project SNP list generated as described by Schneeberger et al.
Head-to-head variant caller comparison Arabidopsis thaliana Bur-0, Tsu-1 and Ler-1 (50x) reads none; optimal kmer size for each tool, Bubbleparse run across depth parameter values percentage of canonical SNPs called and false positive rate for each tool Cortex_var version 1.0.5.13 vs Bubbleparse
Simulated-read benchmark against a known truth set Escherichia coli reference genome and a simulated mutant genome 100,000 randomly positioned SNPs introduced into the E. coli genome in silico number of true SNPs recovered and number of false positives by Bubbleparse vs BWA/SAMTools SimSeq (Assemblathon project) generating 76 nt simulated Illumina reads at ~20x coverage with the supplied Illumina GAIIx error profile
Runtime/resource benchmarking of graph assembly, bubble identification and contig output Col-0 vs Tsu-1 Arabidopsis thaliana read sets none; maximum sensitivity search settings execution time for assembly, SNP identification and contig output workload-managed compute cluster with nodes of up to 128 Gb RAM
Bubble ranking heuristic evaluation Arabidopsis thaliana Ler-1 50x bubble collection at k = 21, search depth = 1 table sorted by each attribute in turn (bubble index/arbitrary, total Q score, coverage, path length, total difference) proportion of real SNPs in successive sets of 100 ranked bubbles, benchmarked against the 1001 Genomes list
Trait-linked marker discovery (proof-of-principle) Unsequenced wild relatives of potato; crosses of Phytophthora infestans resistant and susceptible accessions none (natural resistant vs susceptible genetic backgrounds encoded as different graph colours) SNPs linked to Phytophthora infestans resistance and identification of disease resistance linked genes
Key results
  • Bubbleparse recovered a maximum of 86.82% of canonical SNPs in Tsu-1 at search depth 3, showing the depth-first bubble search is effective at finding bubbles representing real SNPs. 86.82%
  • Specificity falls steeply as search depth increases: the proportion of Bubbleparse calls that are real SNPs drops from 61.99% (Bur-0, k = 17, depth 0, canonical minimum coverage 10) to 3.53% (Tsu-1, depth 3, k = 31, canonical minimum coverage 20). 61.99% to 3.53%
  • Very large numbers of bubbles are found, maximally 893,627 for Col-0/Bur-0 at depth 3 with k = 19, and 1,226,068 for Ler-1 at depth 2 with k = 21; bubble counts roughly mirror the number of distinct kmers, which peaks at k = 19 or 21. 893,627 and 1,226,068 bubbles
  • Bubbleparse outperformed Cortex_var in sensitivity at optimal kmer size: for Bur-0 and Tsu-1 Cortex_var called around 40% of canonical SNPs versus over 80% for Bubbleparse at higher depths; for Ler-1 Cortex_var called 20% versus 38% (depth 0) and 53% (depth 2) for Bubbleparse. 40% vs >80%; 20% vs 38-53%
  • On simulated E. coli data Bubbleparse recovered more true SNPs than the reference-based pipeline but with far more false positives: 93,446 of 100,000 SNPs with 12,917 false positives, versus 86,275 SNPs with 1 false positive for BWA/SAMTools. 93,446 vs 86,275 true; 12,917 vs 1 false positive
  • Increasing sequencing coverage for Ler-1 gave only modest gains from 40x to 50x and decreased sensitivity thereafter up to 340x, with no reduction in false positives. modest gains 40x-50x, decreases above 50x (up to 340x)
  • Ranking by total Q score alone performed worse than arbitrary ranking: the first 50,000 bubbles contained no real SNPs, after which accuracy rose steadily to a fairly consistent 60%; arbitrary (bubble index) ranking gave a constant proportion of real SNPs throughout the collection. 0% in first 50,000 bubbles, rising to ~60%
  • Maximum sensitivity searches (graph assembly, SNP identification, contig output) completed with a mean time of 29 hours for the Col-0 vs Tsu-1 experiment. mean 29 h
Key statistics
  • count 351,493 and 182,419 SNPs (SAMTools canonical SNP calls for Bur-0 at minimum alignment coverage 10 and 20)
  • count 261,970 and 88,597 SNPs (SAMTools canonical SNP calls for Tsu-1 at minimum alignment coverage 10 and 20)
  • other 86.82% (maximum proportion of canonical SNPs recovered by Bubbleparse (Tsu-1, depth 3))
  • other 61.99% down to 3.53% (range of the proportion of Bubbleparse-identified SNPs present in the canonical set across k and search depth)
  • count 893,627 bubbles (maximum bubbles found for Col-0/Bur-0 at depth 3, k = 19)
  • count 1,226,068 bubbles; 52.66% canonical SNPs found (Ler-1 maximum bubble count at depth 2, k = 21, and maximum canonical SNP recovery)
  • count 93,446 true SNPs with 12,917 false positives (Bubbleparse) vs 86,275 with 1 false positive (BWA/SAMTools) (simulated E. coli mutant carrying 100,000 random SNPs, 76 nt reads at ~20x coverage)
  • other 50x (Bur-0), 38x (Tsu-1), 40x-340x (Ler-1) (approximate read coverage of the Arabidopsis thaliana ecotype read sets used as input)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper describes a computational benchmarking study of a new reference-free SNP/variant-detection algorithm (Bubbleparse) built on de Bruijn graphs. Performance was assessed by comparing Bubbleparse's output against reference-based SNP calls (BWA alignment plus SAMTools calling) and a curated SNP set, and against an existing tool (Cortex_var), using descriptive metrics such as the percentage of canonical SNPs recovered (sensitivity) and the proportion of false positives, evaluated across sweeps of k-mer size, search depth and read coverage, plus one simulated E. coli dataset with a known true SNP set. No formal inferential hypothesis tests, p-values, or measures of statistical uncertainty were reported.

Replicationunclear Sample sizeRead coverage levels are described (e.g. ~50x for Bur-0, 38x for Tsu-1, 40x–340x for Ler-1, ~20x for simulated E. coli), but no formal sample-size/power justification or biological/technical replicate count is given GroupsBubbleparse-called SNPs vs. reference-based (BWA/SAMTools) canonical SNP calls and a curated 1001 Genomes SNP set, across Arabidopsis ecotypes (Bur-0, Tsu-1, Ler-1) and a simulated E. coli dataset; also Bubbleparse vs. Cortex_var caller Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Descriptive sensitivity/precision comparison (percentage of canonical SNPs recovered; percentage of called SNPs matching canonical set) — no named inferential test statistic reported Comparisons of Bubbleparse vs. BWA/SAMTools canonical SNP calls and vs. Cortex_var, across Bur-0, Tsu-1, Ler-1 and simulated E. coli datasets Counts of SNPs/bubbles as stated per comparison (e.g. 351,493 and 182,419 SAMTools SNPs for Bur-0 at coverage 10/20; 100,000 simulated SNPs for E. coli) not stated
Approaches that could also have been used
  • Bubbleparse's SNP detection performance was compared to reference-based calls and to Cortex_var using descriptive sensitivity and false-positive counts, without a formal statistical test of whether differences between methods were significant.
    Could also: A paired comparison such as McNemar's test, or precision-recall/ROC analysis with bootstrap confidence intervals, could also be used — This would let readers assess whether observed differences in detection rate or false-positive rate between methods exceed what might be expected from sampling variability alone
  • Sensitivity and false-positive rates are reported as single point percentages for each parameter combination (k-mer size, search depth, coverage) rather than with an associated measure of variability.
    Could also: Bootstrap resampling of reads, or repeated subsampling of the sequencing data, could also be used to generate confidence intervals around each sensitivity/false-positive estimate — This would convey how much the reported percentages might vary if the underlying read set were resampled, complementing the single-point estimates shown
  • Comparisons across the three ecotypes (Bur-0, Tsu-1, Ler-1) are described qualitatively as showing a 'similar pattern' or 'comparable' rates without a formal statistical comparison.
    Could also: A chi-square or Fisher's exact test on contingency tables of true/false positive counts could also be used to compare detection or false-positive rates between ecotypes or methods — This would provide a quantitative basis for the qualitative similarity/difference statements already made in the text
  • Each ecotype comparison (Bur-0, Tsu-1, Ler-1, E. coli) is based on what appears to be a single sequencing dataset rather than multiple independent replicates.
    Could also: Replicate sequencing runs, or in silico replicate generation via read resampling, could also be used to assess reproducibility of the sensitivity and false-positive estimates — This would allow an estimate of run-to-run variability separate from the biological/genomic differences being studied
  • The ranking heuristic was evaluated by sorting bubbles by each attribute and computing the proportion of real SNPs within successive bins of 100 bubbles.
    Could also: A precision-recall curve or area-under-curve (AUC) summary across the full ranked list could also be used — This would give a single continuous performance summary for each ranking attribute rather than a series of discrete bin-wise proportions
  • Optimal parameter values (k-mer size, search depth, coverage) were selected by visual inspection of plotted sensitivity/false-positive trends across a parameter sweep.
    Could also: A formal grid-search combined with cross-validation on held-out data could also be used to select parameters — This would provide a data-driven, less subjective criterion for parameter choice and an estimate of how performance generalizes beyond the specific datasets examined
Software: BWA · SAMTools · BLAST · Cortex / Cortex_var 1.0.5.13 (Cortex_var, for comparison) · SimSeq

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

bubble_type_taxonomy_qualitative
Reported
Bubbles classified by graph-walking shape into SNP-like, indel-like, and complex/repeat-associated categories (paper's core method)
Reproduced
bubbleparse_31 (patched) classifies all 124,090 (Bur-0) and 161,987 (Tsu-1) bubble matches into S_/i_/y_/z type codes matching this taxonomy
exact
clean_snp_fraction_bur0
Reported
61.99% (paper's highest reported real-SNP proportion in bubble collection, Bur-0 k=17 depth=0 mincov10)
Reproduced
61.4% (S_1_1 bubbles / total = 76200/124090, Col-0xBur-0, k=19, single default depth)
within tolerance
clean_snp_fraction_tsu1
Reported
61.99% (paper's highest reported real-SNP proportion, comparable order for Tsu-1 not separately itemized in extracted text)
Reproduced
59.3% (S_1_1 bubbles / total = 96039/161987, Col-0xTsu-1, k=19, single default depth)
within tolerance
total_bubbles_bur0_vs_paper_max
Reported
893,627 (paper's stated maximum for Col-0/Bur-0, across depth 0-3 x k 15-31 sweep, richest at depth 3 k=19)
Reproduced
124,090 (single run, default depth, k=19 only — sweep not reproduced)
did not match
refbased_snps_bur0_mincov10
Reported
351,493 (paper, SAMtools, Bur-0, mincov10)
Reproduced
194,825 (this project's custom bwa+pileup+snpcall.py, >=0.95 alt-fraction rule, NOT SAMtools)
did not match
refbased_snps_bur0_mincov20
Reported
182,419 (paper, SAMtools, Bur-0, mincov20)
Reproduced
43,032 (custom pileup caller, not SAMtools)
did not match
refbased_snps_tsu1_mincov10
Reported
261,970 (paper, SAMtools, Tsu-1, mincov10)
Reproduced
88,520 (custom pileup caller, not SAMtools)
did not match
refbased_snps_tsu1_mincov20
Reported
88,597 (paper, SAMtools, Tsu-1, mincov20)
Reproduced
4,079 (custom pileup caller, not SAMtools)
did not match
sequencing_coverage
Reported
Bur-0 ~50x, Tsu-1 ~38x (paper)
Reproduced
verified consistent with reported coverage in a prior session's dedicated integrity/coverage-check job; not re-derived this session
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 45/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The core method reproduces: after four documented patches, bubbleparse ran to completion on both crosses (Col-0xBur-0: 124,090 ranked bubbles; Col-0xTsu-1: 161,987) and reproduced the paper's SNP/indel/Y-repeat/complex taxonomy, with the clean SNP-shaped class S_1_1 at 61.4% / 59.3% — within ~0.6 pp of the paper's best-case 61.99%. All large numeric gaps are on our side and disclosed: 124,090 vs the paper's 893,627 reflects a single k=19/default-depth run instead of the full depth 0-3 x k 15-31 sweep, and the reference-SNP shortfalls (e.g. 4,079 vs 88,597 for Tsu-1 mincov20) come from substituting a conservative custom bwa+pileup caller for the paper's SAMtools. On the authors' side there are two real defects that do not touch the reported values: the deposited tool is non-functional as shipped (the fatal 3-column-filelist fscanf bug in look_for_quality_scores_in_fastq() produced zero output and exit(1)), and the ground-truth SNP set underlying the validated 61.99% was never deposited, so that specific metric is not derivable from the shared data. Data availability itself is exemplary — all 12 SRR runs retrieved and integrity-verified — so this is a solid partial reproduction, not a suspect one.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.