Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Optimizing open data to support one health: best practices to ensure interoperability of genomic data from bacterial pathogens.

One Health Outlook · 2020
L1 85/100 3/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

salvaged by watchdog from agreement.json (agent omitted ROOM_RESULT.json)

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-07-29

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This paper poses no experimental hypothesis; it is a Best Practices review arguing that standardized methods, QC thresholds, and metadata curation for submitting bacterial pathogen whole-genome sequencing data to open databases (NCBI Pathogen Detection) are necessary to ensure interoperable, FAIR genomic data supporting One Health surveillance.

Core claims
  • An open-access pathogen surveillance database (NCBI Pathogen Detection) plus contributor Best Practices enables FAIR, interoperable genomic data across human, animal, food, and environmental sources for One Health surveillance. resource
  • Standardized minimum metadata fields submitted to the correct attributes are required for isolates to be properly analyzed, integrated, and labeled in NCBI-PD clusters. method
  • Contributors should only upload sequence and metadata meeting defined QC thresholds to ensure interoperability, accuracy, and usefulness of NCBI-PD. method
  • NCBI-PD computes daily updated phylogenies for clusters of closely related genomes and screens every bacterial genome for AMR, stress response, and virulence genes. method
  • With this guidance, GenomeTrakr is removed from its role as a data broker so individual laboratories worldwide can submit directly to NCBI. resource
  • De novo assembly length is highly variable (e.g., E. coli) or multi-modal (S. enterica, L. monocytogenes), making narrow sequencing-metric guidelines difficult. finding
  • Increasing coverage improves de novo assembly quality (fewer contigs), but the rate of improvement slows once coverage exceeds about 40X. finding
  • Building interoperable systems lets researchers combine public INSDC data with private data across platforms (IRIDA, INNUENDO, PathogenWatch, NextStrain, IDseq, CGE Evergreen, BioNumerics). resource
Experimental setups
Assay System Perturbation Readout Platform
Illumina whole-genome shotgun sequencing with de novo assembly and QC metric analysis Salmonella enterica, Listeria monocytogenes, Escherichia coli, Shigella sp., Campylobacter jejuni, Vibrio parahaemolyticus isolates (NCBI-PD) none (surveillance isolates) average read quality (Q score), average coverage, de novo assembly length (Mbp), number of contigs Illumina (MiSeq, NextSeq, HiSeq, iSeq)
Empirical analysis of assembly length distribution across pathogen databases 51,414 NCBI-PD isolates across six foodborne pathogens none range/distribution of de novo assembly lengths
Coverage vs. assembly quality analysis NCBI-PD foodborne bacterial isolates none number of contigs as a function of sequencing coverage
Key results
  • De novo assembly quality improves with coverage but the improvement rate slows beyond ~40X coverage ~40X coverage inflection
  • Assembly length is highly variable in E. coli and multi-modal in S. enterica and L. monocytogenes
  • GenomeTrakr network has accumulated an open-access archive of genomes from non-human sources ~100K isolates as of July 2020
  • Thirty-two pathogens (31 microbes and one yeast) are under active surveillance and stored at NCBI-PD 32 pathogens
  • Recommended QC thresholds vary by organism (e.g., coverage >=30X for Salmonella, >=40X for E. coli/Shigella/Vibrio, >=20X for Listeria/Campylobacter) Q>=30; coverage 20X-40X
Key statistics
  • count 51,414 isolates (total isolates summarized for assembly length analysis)
  • count 10,000 random isolates per database (selected Jan 6 2020 for S. enterica, L. monocytogenes, E. coli, Shigella, C. jejuni)
  • count 1414 isolates (Vibrio parahaemolyticus isolates selected Jan 9 2020 (smaller database))
  • count ~100 K isolates (GenomeTrakr non-human source genome archive as of July 2020)
  • count 32 pathogens (31 microbes and one yeast) (under active surveillance at NCBI as of March 2020)
  • other >= 30 (average read quality Q score threshold for R1 and R2 across all six pathogens)
  • other ~40X (coverage beyond which de novo assembly improvement slows)
  • other <=300 contigs (de novo assembly contig threshold for Salmonella, Listeria, Campylobacter, Vibrio)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a review and best-practices paper, not a primary experimental study; its quantitative content consists of descriptive analysis of 51,414 whole-genome sequences randomly sampled from NCBI Pathogen Detection databases across six bacterial pathogens to empirically derive QC thresholds for sequence quality metrics. Results were communicated through ranges of assembly-quality metrics and figures showing the relationship between sequencing coverage and de novo assembly quality, with no formal inferential hypothesis tests reported. The paper's primary statistical contribution is the empirical derivation of QC cutoffs (Table 1), presented as practical guidelines for WGS submissions to NCBI.

Replicationunclear Sample size10,000 randomly selected isolates per pathogen for five species (Salmonella enterica, Listeria monocytogenes, E. coli, Shigella sp., Campylobacter jejuni; sampled 6 January 2020); 1,414 isolates for Vibrio parahaemolyticus (full available database, 9 January 2020); total N = 51,414 Groupssix bacterial foodborne pathogens compared descriptively on de novo assembly metrics (total assembly length, number of contigs) and sequencing coverage Pairingna Randomization/blindingnot stated Dispersionrange Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • QC thresholds in Table 1 were described as 'empirically derived' from the sampled isolate distributions, with specific numeric cutoffs chosen based on expert review of the observed data
    Could also: Percentile-based cutoffs (e.g., 5th/95th percentiles of the empirical distributions), k-means clustering on the metric distributions, or mixture-model-based thresholds could also be used to define acceptable ranges — Formal statistical derivation of cutoffs makes the selection criterion explicit, reproducible, and auditable, and quantifies precisely what fraction of the observed data each threshold would exclude — information useful for future threshold revisions as sequencing technology improves
  • The relationship between sequencing coverage and assembly quality (number of contigs) was displayed graphically (Fig. 4) and described qualitatively, noting that 'the rate of improvement slows once coverage exceeds about 40X'
    Could also: A regression model — for example, a log-linear or smoothing spline regression of contig count on coverage — could also formally characterize the coverage–quality curve and locate the inflection point — A fitted curve with confidence bands would make the '~40X' recommendation explicit and quantified, allowing readers to assess uncertainty around that value and supporting future updates as larger or more diverse datasets become available
  • Assembly metric distributions were summarized as ranges (Table 1) without accompanying measures of central tendency or spread
    Could also: Reporting median ± IQR or mean ± SD alongside the observed range for each pathogen would also convey the typical value and distributional spread — Central tendency and dispersion measures communicate where the bulk of isolates fall rather than only the extremes; for small or skewed distributions (e.g., the Vibrio parahaemolyticus subset of 1,414), this distinction is particularly informative for threshold-setting
  • Isolates were sampled randomly at 10,000 per pathogen (or exhaustively for Vibrio) without explicit stratification by submitting laboratory, submission year, or geography
    Could also: Stratified random sampling by submission year, submitting laboratory, or geographic region could also have been applied — Stratification would help ensure that derived thresholds reflect the full diversity of submission practices and sequencing conditions rather than being disproportionately influenced by high-volume contributors, which could affect the generalizability of the recommended cutoffs
  • The six pathogens were treated as independent groups and compared descriptively, without any formal statistical comparison of their metric distributions
    Could also: Non-parametric group comparisons (e.g., Kruskal-Wallis with post-hoc Dunn tests) could also have been used to formally assess whether assembly-length or contig-count distributions differ significantly across pathogens — Formal between-group tests would provide a statistical basis for why species-specific thresholds are warranted rather than a single universal threshold, supporting the rationale for the per-pathogen structure of Table 1
Software: GalaxyTrakr · NCBI Pathogen Detection (NCBI-PD) · BioNumerics (Applied Maths)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33103064

Paper: Timme et al. 2020, "Optimizing open data to support one health: best practices to ensure interoperability of genomic data from bacterial pathogens." One Health Outlook 2:20. DOI 10.1186/s42522-020-00026-3.

This is primarily a best-practices / perspective paper. Most of it (metadata standards, BioProject structure, INSDC ecosystem) is descriptive and NOT a computational pipeline. A minority of the paper presents pipeline-derived quantitative results, which are what we reproduce.

Pipeline(s) named

  • MicroRunQC — QC workflow (Galaxy / GalaxyTrakr; also CLI at github.com/estrain/MicroRunQC). Chains: Trimmomatic -> SKESA assembly -> BWA insert-size -> mlst (PubMLST) -> fastq-scan. Produces per-isolate QC: Contigs, Length, EstCov, N50, MedianInsert, MeanLength/Q R1/R2, MLST Scheme/ST.
  • SKESA v2.2 — de novo assembler; the assembly-stat source for Table 1 / Fig 3 / Fig 4.
  • The Table 1 / Fig 3 / Fig 4 numbers are computed over a random sample of isolates from NCBI Pathogen Detection (snapshot 6–9 Jan 2020), whose assembly metadata (length, #contigs, coverage) are public and queryable.

IN SCOPE (pipeline-derived → attempted)

id result paper location pipeline reproduction route
T1-len Per-species de novo assembly genome-length ranges (e.g. S.enterica ~4.3–5.2 Mbp) Table 1 SKESA via NCBI-PD metadata mean ±3 SD of random 10k/species
T1-contig Per-species max #contigs (e.g. S.enterica ≤300) Table 1 SKESA via NCBI-PD metadata upper percentile of random 10k/species
F3 Genome-length density per species, bars = mean ±3 SD (n=10k, V.para n=1414) Fig 3 SKESA via NCBI-PD reproduce distribution + mean/SD
F4 Mean coverage (SKESA) vs #contigs: contigs decrease as coverage rises, plateau ~40X Fig 4 SKESA via NCBI-PD reproduce scatter + monotone trend
MRQC MicroRunQC produces the 13-col QC report on real isolates; values fall in Table 1 ranges Methods / repo (P16) MicroRunQC run microrunqc.py on real paired Illumina reads

OUT OF SCOPE (descriptive / non-pipeline → not attempted)

  • Table 2 (minimum BioSample metadata fields) — policy/standards, not computed.
  • Fig 1 (INSDC analysis-tool ecosystem), Fig 2 (NCBI-PD browser screenshot), Fig 5 (umbrella-BioProject diagram) — illustrative, not computational.
  • Coverage Q-score thresholds (≥30) and per-species coverage minima — these are recommended cut-offs / policy choices, not values derived by a single rerunnable computation (we still sanity-check them against the metadata distribution).

Data note

Brief lists sra:PRJNA248064 (Public Health England umbrella BioProject, 11 sub-projects). The paper's quantitative results do not come from one BioProject but from a random NCBI Pathogen Detection snapshot; PRJNA248064 is one example data source in the ecosystem. We profile NCBI-PD metadata (the actual statistical substrate of Table 1 / Fig 3 / Fig 4) and use real isolates for the MicroRunQC run.

Figures / tables: TableFig 3Fig 4
T1-sal-len
Reported
4.3-5.2 Mbp
Reproduced
4.35-5.32 Mbp (mu=4.837,sd=0.162,n=10000)
within tolerance
T1-lis-len
Reported
2.7-3.2 Mbp
Reproduced
2.74-3.32 Mbp (mu=3.031,sd=0.097,n=10000)
within tolerance
T1-eco-len
Reported
4.5-5.9 Mbp
Reproduced
4.40-5.94 Mbp (mu=5.173,sd=0.257,n=10000)
within tolerance
T1-shi-len
Reported
4.0-5.0 Mbp
Reproduced
4.05-4.96 Mbp (mu=4.504,sd=0.152,n=10000)
within tolerance
T1-cje-len
Reported
1.5-1.9 Mbp
Reproduced
1.50-2.00 Mbp (mu=1.747,sd=0.084,n=10000)
within tolerance
T1-vpa-len
Reported
4.8-5.5 Mbp
Reproduced
4.62-6.49 Mbp (mu=5.557,sd=0.312,n=1414)
within tolerance
T1-contig-caps
Reported
Sal<=300,Lis<=300,Eco<=500,Shi<=650,Cje<=300,Vpa<=300
Reproduced
98.7-99.8% of isolates within each cap; caps ~= p99
within tolerance
F3-density
Reported
genome-length density, bars=mean+/-3SD (n=10k; Vpara 1414)
Reproduced
density+mean/SD bars reproduced for all 6 species
within tolerance
F4-trend
Reported
contigs decrease as coverage increases, plateau ~40X
Reproduced
PENDING (Track A coverage ladder)
m.public.grade.uncheckable
MRQC-pipeline
Reported
13-col QC report (Contigs,Length,EstCov,N50,...,ST)
Reproduced
PENDING (Track A)
m.public.grade.uncheckable
MRQC-xcheck
Reported
re-assembly matches NCBI-PD reported SKESA v2.2 stats
Reproduced
PENDING (Track A)
m.public.grade.uncheckable

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This best-practices paper's quantitative core (Table 1 ranges, contig caps, Fig 3) reproduces well from open NCBI-PD data: 5/6 genome-length ranges land within ~0.1 Mbp and all six contig caps validate as ~p99 cutoffs, with no fabrication signal. The only real deviation — V. parahaemolyticus (4.8–5.5 → 4.62–6.49) — is snapshot drift on the data side (the set grew ~7×), not an authors' defect, and the exact Jan-2020 sample was never deposited, so the gap is one of data identity/cohort definition rather than derivability. Three Track-A claims (Fig 4, the MicroRunQC report, SKESA cross-check) are still PENDING, keeping this a solid-but-partial reproduction. Overall yellow: the central conclusion holds, deviations are explainable and on the data-availability/our-sampling side.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

322.8 k
tokens (I/O) · 27.7 M incl. cache
72 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.