Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

PHA4GE quality control contextual data tags: standardized annotations for sharing public health sequence datasets with known quality issues to facilitate testin

Microb Genom · 2024
L1 75/100 PQI 82
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to PARTIALLY reproduce, 1:1 on the verifiable parts. This is a standards/specification paper: the GitHub repo ships only a controlled vocabulary of QC tags (CSV/JSON + spreadsheets + a CSV->JSON converter), NOT an analysis pipeline, so there is no shipped data+code+expected-output to run. We reproduced two things. (1) C1 — EXACT: the paper's central real-world-implementation claim (Table 4) that five named GenomeTrakr SRA records carry PHA4GE QC contextual data tags; fetched from NCBI SRA, all five carry quality_control_method/_version/_determination/_issues/_details with values drawn exactly from the spec's GENEPIO picklist (verified against repo commit 2d5e0007). (2) C2 — PARTIAL/consistent: a genuine pipeline-derived check on «our HPC» («job») applying the paper's NAMED tools (minimap2 -> NC_045512.2, samtools 1.19 coverage/depth, FastQC 0.12.1; the 'third-party tool on the paper's data' path the brief endorses) to two contrasting Table-4 samples. SRR21205381 (tagged 'no QC issues') gave 99.9% genome breadth / 24119x depth; SRR20428498 (tagged 'low average genome coverage') gave 61.9% breadth / 2458x — directionally reproducing the recorded determinations. NOT attempted (the hard 20%): the GenomeTrakr pilot percentages (61/10/28/1% of 2255 sequences, Table 3) require the full pilot dataset + per-sequence SSQuAWK/C-WAP determinations + unspecified classification thresholds, none shipped; and the exact SSQuAWK/C-WAP pass/fail logic (thresholds unspecified) — C2 is a consistency check, not a reimplementation. Fabrication scan: no numeric claim is presented as a reproducible computed output tied to shipped data+code, so nothing here is fabrication-prone; the checkable factual claims hold against SRA + repo.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 75
    assessed: 2026-06-15 ⛓ 25662fa6acb8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

There are no standardized attributes or mechanisms for tagging poor-quality or purpose-specific pathogen sequence datasets; the paper proposes that a set of standardized, ontology-based contextual data tags can flag sequence data with known quality issues to improve their discoverability, interpretation, and reuse for public health testing and training.

Core claims
  • PHA4GE developed a set of standardized contextual data tags (five fields plus controlled-vocabulary terms) for annotating pathogen sequence datasets with known quality issues. resource
  • The QC tags are agnostic to organism and sequencing technique, so they can be applied to data from any pathogen using any sequencing platform, as well as to synthetic data. method
  • QC attributes were mapped to and standardized using ontologies (GenEpiO), making them FAIR (Findable, Accessible, Interoperable, Reusable). method
  • The QC tags were piloted, tested, and implemented by the FDA's GenomeTrakr network for SARS-CoV-2 wastewater metagenomic surveillance, becoming part of routine NCBI submission. finding
  • Sharing annotated lower-quality and synthetic datasets enables their reuse for validating, benchmarking, and optimizing algorithms, pipelines, and instruments, and for training personnel. finding
  • The standardized tags, definitions, ontology IDs, and a JSON representation are maintained by PHA4GE in a public GitHub repository with a New Term Request System for community contributions. resource
  • Data deemed insufficient quality for one purpose may still be appropriate and usable for other public health applications. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Metagenomic wastewater sequencing for SARS-CoV-2 surveillance (QC tag implementation/testing) FDA GenomeTrakr laboratory network; wastewater samples none Quality control determination/issues annotated via standardized tags during NCBI/SRA submission NCBI SRA submission form (GenEpiO/PHA4GE QC tags)
Community survey and consultation to identify common QC issues PHA4GE member networks, public health and research community, INSDC representatives none Range and types of common quality control issues identified; feedback on proposed tag list
Ontology mapping and term creation for QC attributes Genomic Epidemiology Ontology (GenEpiO) none Ontology IDs/terms for QC fields and picklist values OBO Foundry / EMBL-EBI OLS
Key results
  • Five standardized QC fields were defined: quality control method name, method version, determination, issues, and details, each with ontology IDs (GENEPIO:0100557–0100561). 5 fields
  • QC determination picklist provides six enumerated values (e.g., sequence passed/failed/flagged QC) and the QC issues field provides nine enumerated values (e.g., low average genome coverage, sequence contaminated). 6 determination + 9 issue values
  • GenomeTrakr scientists reviewed and adopted the tags into routine submission procedures, with many early pandemic datasets flagged as lower quality to allow transparent sharing.
  • QC attributes were made publicly available on GitHub in October 2022 and incorporated into a purpose-built SRA submission form.
Key statistics
  • count five standardized fields ('tags') (Number of standardized QC contextual data fields created by PHA4GE)
  • count six QC determination enum values (Enumerated values for the quality control determination field)
  • count nine QC issues enum values (Enumerated values for the quality control issues field)
  • other CT value of 39 (Example given for the quality control details field (low viral load))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a standards-development and methods paper describing the creation of standardized quality control metadata tags for public health genomic sequence data repositories; it does not present a primary research experiment with inferential statistics. The methodology consists of community needs assessment via survey and consultation, ontology-mapping of identified QC categories, and a descriptive pilot implementation by the FDA GenomeTrakr network for SARS-CoV-2 wastewater metagenomics. No hypothesis tests, effect-size estimates, or p-values are reported.

Replicationunclear GroupsNo comparative groups; the paper describes a consensus-based standard-setting process and a qualitative pilot implementation Pairingna Randomization/blindingna Dispersionnone
Approaches that could also have been used
  • Community data needs were assessed through informal survey distributed via member networks and direct communication, with no formal sampling frame or response-rate reporting
    Could also: A structured Delphi consensus process or formal mixed-methods survey with defined sampling frame, response tracking, and thematic saturation criteria could also be used for standards elicitation — Formal consensus methods (Delphi, nominal group technique) produce documented convergence metrics and make it possible to quantify agreement levels, which can strengthen the reproducibility and perceived legitimacy of the resulting standard
  • Pilot implementation was described narratively for one network (FDA GenomeTrakr SARS-CoV-2 wastewater), without quantitative evaluation of tag adoption, retrieval accuracy, or inter-annotator agreement
    Could also: A quantitative usability evaluation could measure inter-annotator agreement (e.g., Cohen's kappa or Krippendorff's alpha) on whether independent annotators assign the same tags to the same datasets — Reliability coefficients would provide an objective, reproducible measure of tag clarity and could identify which tag terms benefit from additional definitional guidance
  • Tag utility was asserted based on expert consensus and a single-network implementation rather than a controlled comparison
    Could also: A before-after or case-control study comparing retrieval precision and recall of quality-flagged datasets in repositories with and without standardized tags could also be conducted — Empirical retrieval metrics (precision, recall, F1) would quantify the practical discoverability benefit of the tags for downstream users
  • The paper categorizes QC failure reasons into a fixed enumerated list derived from internal review of bacterial and viral datasets
    Could also: Qualitative thematic analysis with formal coding procedures (open coding, axial coding, saturation testing) or latent class analysis on structured survey data could also be used to derive QC categories — Explicit coding methodology with inter-rater reliability reporting would make the category-derivation process transparent and reproducible for future updates or extensions of the tag set
Software: GitHub (repository hosting for tag specifications and JSON representation) · Genomic Epidemiology Ontology (GenEpiO) · EMBL-EBI Ontology Lookup Service (OLS)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
5
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

SRR19851129 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRR20018633 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRR20046849 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRR20428498 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRR21205381 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38860884

Paper: Griffiths et al. 2024, Microb Genom 10(6):001260. "PHA4GE quality control contextual data tags: standardized annotations for sharing public health sequence datasets with known quality issues." PMID 38860884 · PMCID PMC11261899 · DOI 10.1099/mgen.0.001260

Code: https://github.com/pha4ge/contextual-data-qc-tags (redirect target of the brief's contextual_data_QC_tags).

What kind of paper is this?

This is a standards / specification paper, not a computational-analysis paper. Its product is a controlled vocabulary of "quality control contextual data tags" (attribute names, picklist values, GENEPIO ontology IDs) for annotating shared pathogen-genomics datasets that have known quality issues. The GitHub repo ships exactly that specification — QC Contextual Data Tags Specification.csv, QC_Contextual_Data_Tags_Specification.json, two reference spreadsheets, and table_json.py (a CSV→JSON converter). There is no analysis pipeline in the repo and no shipped dataset+expected-output to run.

Reported results, classified

# Reported result Source Pipeline? In scope?
R1 The QC-tag specification itself (fields, picklists, GENEPIO IDs) Table 1, repo CSV/JSON No — a vocabulary, not a computed value out (non_pipeline)
R2 Worked examples 1–8: illustrative tag annotations for hypothetical QC scenarios using ncov-tools/FastQC/Kraken2/Nextclade/samtools-depth/VADR/Quast/PHoeNIx §"Worked examples", Table 2 Illustrative — no numeric values, no specific accession tied to any example out (no_expected_result)
R3 GenomeTrakr pilot: of 2,255 sequences, 61 % no QC issues, 10 % minor, 28 % potential, 1 % significant §"Real-world implementations", Table 3 Yes (GalaxyTrakr SSQuAWK / CFSAN C-WAP) but the full dataset + per-sequence determinations + classification thresholds are not shipped out (hard 20%, docs_insufficient for repro) — see note
R4 Table 4: five named GenomeTrakr SRA records "illustrating PHA4GE QC contextual data tag use" (incl. SRR21205381) Table 4 The claim is that these public records carry the standardized tags IN — directly verifiable from SRA metadata

What we attempt (the 80%)

  1. C1 — Table 4 tag presence (control-plane, exact). Fetch the SRA metadata for all five Table-4 accessions and confirm each carries the PHA4GE QC contextual-data-tag attributes (quality_control_method, _method_version, _determination, _issues, _details) with values drawn from the specification's picklist. This reproduces the paper's central real-world implementation claim 1:1.

  2. C2 — pipeline-derived consistency check («our HPC», third-party tools on the paper's data). The brief explicitly endorses running an existing tool on the paper's own data. The Table-4 determinations for wastewater SARS-CoV-2 amplicon data are coverage-driven (SSQuAWK/C-WAP). We therefore apply the paper's named QC tools — samtools (Table 2; worked example 3 uses samtools depth) and FastQC (Table 2; worked example 4) — to two contrasting Table-4 samples and test whether the computed SARS-CoV-2 genome coverage breadth is consistent with the recorded determination:

    • SRR21205381 → recorded "no quality control issues identified" → expect high breadth.
    • SRR20428498 → recorded "low average genome coverage" (flagged) → expect low breadth. This is a qualitative/directional reproduction, not an attempt to replicate SSQuAWK's exact pass/fail thresholds (those are unspecified).

Explicitly NOT attempted (the hard 20%)

  • R3 GenomeTrakr pilot percentages (61/10/28/1 % of 2,255). Would require the complete pilot dataset and the authors' per-sequence SSQuAWK/C-WAP determinations and their classification thresholds — none of which are shipped. Not reproducible from public artifacts; not attempted.
  • **R1
Figures / tables: Table
C1.1
Reported
SRR21205381 listed in Table 4 as a GenomeTrakr record illustrating PHA4GE QC tag use
Reproduced
SRA record carries quality_control_method=GalaxyTrakr SSQuAWK v4.0.2, determination='no quality control issues identified' (value from spec picklist)
exact
C1.2-C1.5
Reported
SRR19851129, SRR20046849, SRR20428498, SRR20018633 listed in Table 4 as records illustrating PHA4GE QC tag use
Reproduced
All four SRA records carry PHA4GE quality_control_* tags; determinations span 'minor issues' (C-WAP), 'no issues' (C-WAP), 'flagged/low average genome coverage' (SSQuAWK), 'flagged' (SRA human read removal) — all values from the spec GENEPIO picklist
exact
C2.1
Reported
SRR21205381 determination 'no quality control issues identified' (qualitative)
Reproduced
SARS-CoV-2 (NC_045512.2) genome breadth >=1x 99.9% / >=10x 99.78%, meandepth 24119x, FastQC per-base-quality PASS — consistent with 'no QC issues'
partial
C2.2
Reported
SRR20428498 determination 'sequence flagged for potential quality control issues / low average genome coverage' (qualitative)
Reproduced
SARS-CoV-2 genome breadth >=1x 61.87% / >=10x 60.02%, meandepth 2458x — only ~62% of genome covered, consistent with 'low average genome coverage'; stark contrast vs SRR21205381's 99.9%
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

This is a standards/specification paper: the repo ships a controlled QC-tag vocabulary, not an analysis pipeline, so there is no shipped data+code+expected-output to run end-to-end. The central claim reproduced cleanly — all five Table-4 GenomeTrakr SRA records carry the standardized quality_control_* tags drawn from the spec picklist (1:1 against SRA + repo commit 2d5e0007), and a recomputed coverage check (99.9% breadth for the 'no issues' sample vs 61.87% for the 'low coverage' sample) is directionally consistent with the recorded determinations. The only gap is the secondary Table 3 pilot percentages (61/10/28/1% of 2255), which are not derivable because the pilot dataset and per-sequence thresholds were never deposited — a data-availability limitation, not an authors' defect or fabrication. Overall a solid, honest partial reproduction with explainable scope limits.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

133.7 k
tokens (I/O) · 9.9 M incl. cache
20 min
runtime · 0.08 CPU-h
2.6 GB
peak RAM
1
HPC jobs
hummel
machine