PHA4GE quality control contextual data tags: standardized annotations for sharing public health sequence datasets with known quality issues to facilitate testin
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to PARTIALLY reproduce, 1:1 on the verifiable parts. This is a standards/specification paper: the GitHub repo ships only a controlled vocabulary of QC tags (CSV/JSON + spreadsheets + a CSV->JSON converter), NOT an analysis pipeline, so there is no shipped data+code+expected-output to run. We reproduced two things. (1) C1 — EXACT: the paper's central real-world-implementation claim (Table 4) that five named GenomeTrakr SRA records carry PHA4GE QC contextual data tags; fetched from NCBI SRA, all five carry quality_control_method/_version/_determination/_issues/_details with values drawn exactly from the spec's GENEPIO picklist (verified against repo commit 2d5e0007). (2) C2 — PARTIAL/consistent: a genuine pipeline-derived check on «our HPC» («job») applying the paper's NAMED tools (minimap2 -> NC_045512.2, samtools 1.19 coverage/depth, FastQC 0.12.1; the 'third-party tool on the paper's data' path the brief endorses) to two contrasting Table-4 samples. SRR21205381 (tagged 'no QC issues') gave 99.9% genome breadth / 24119x depth; SRR20428498 (tagged 'low average genome coverage') gave 61.9% breadth / 2458x — directionally reproducing the recorded determinations. NOT attempted (the hard 20%): the GenomeTrakr pilot percentages (61/10/28/1% of 2255 sequences, Table 3) require the full pilot dataset + per-sequence SSQuAWK/C-WAP determinations + unspecified classification thresholds, none shipped; and the exact SSQuAWK/C-WAP pass/fail logic (thresholds unspecified) — C2 is a consistency check, not a reimplementation. Fabrication scan: no numeric claim is presented as a reproducible computed output tied to shipped data+code, so nothing here is fabrication-prone; the checkable factual claims hold against SRA + repo.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 75assessed: 2026-06-15 ⛓ 25662fa6acb8
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThere are no standardized attributes or mechanisms for tagging poor-quality or purpose-specific pathogen sequence datasets; the paper proposes that a set of standardized, ontology-based contextual data tags can flag sequence data with known quality issues to improve their discoverability, interpretation, and reuse for public health testing and training.
- ★ PHA4GE developed a set of standardized contextual data tags (five fields plus controlled-vocabulary terms) for annotating pathogen sequence datasets with known quality issues. resource
- ★ The QC tags are agnostic to organism and sequencing technique, so they can be applied to data from any pathogen using any sequencing platform, as well as to synthetic data. method
- ★ QC attributes were mapped to and standardized using ontologies (GenEpiO), making them FAIR (Findable, Accessible, Interoperable, Reusable). method
- ★ The QC tags were piloted, tested, and implemented by the FDA's GenomeTrakr network for SARS-CoV-2 wastewater metagenomic surveillance, becoming part of routine NCBI submission. finding
- ★ Sharing annotated lower-quality and synthetic datasets enables their reuse for validating, benchmarking, and optimizing algorithms, pipelines, and instruments, and for training personnel. finding
- The standardized tags, definitions, ontology IDs, and a JSON representation are maintained by PHA4GE in a public GitHub repository with a New Term Request System for community contributions. resource
- Data deemed insufficient quality for one purpose may still be appropriate and usable for other public health applications. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Metagenomic wastewater sequencing for SARS-CoV-2 surveillance (QC tag implementation/testing) | FDA GenomeTrakr laboratory network; wastewater samples | none | Quality control determination/issues annotated via standardized tags during NCBI/SRA submission | NCBI SRA submission form (GenEpiO/PHA4GE QC tags) |
| Community survey and consultation to identify common QC issues | PHA4GE member networks, public health and research community, INSDC representatives | none | Range and types of common quality control issues identified; feedback on proposed tag list | — |
| Ontology mapping and term creation for QC attributes | Genomic Epidemiology Ontology (GenEpiO) | none | Ontology IDs/terms for QC fields and picklist values | OBO Foundry / EMBL-EBI OLS |
- – Five standardized QC fields were defined: quality control method name, method version, determination, issues, and details, each with ontology IDs (GENEPIO:0100557–0100561). 5 fields
- – QC determination picklist provides six enumerated values (e.g., sequence passed/failed/flagged QC) and the QC issues field provides nine enumerated values (e.g., low average genome coverage, sequence contaminated). 6 determination + 9 issue values
- – GenomeTrakr scientists reviewed and adopted the tags into routine submission procedures, with many early pandemic datasets flagged as lower quality to allow transparent sharing.
- – QC attributes were made publicly available on GitHub in October 2022 and incorporated into a purpose-built SRA submission form.
- count five standardized fields ('tags') (Number of standardized QC contextual data fields created by PHA4GE)
- count six QC determination enum values (Enumerated values for the quality control determination field)
- count nine QC issues enum values (Enumerated values for the quality control issues field)
- other CT value of 39 (Example given for the quality control details field (low viral load))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a standards-development and methods paper describing the creation of standardized quality control metadata tags for public health genomic sequence data repositories; it does not present a primary research experiment with inferential statistics. The methodology consists of community needs assessment via survey and consultation, ontology-mapping of identified QC categories, and a descriptive pilot implementation by the FDA GenomeTrakr network for SARS-CoV-2 wastewater metagenomics. No hypothesis tests, effect-size estimates, or p-values are reported.
-
Community data needs were assessed through informal survey distributed via member networks and direct communication, with no formal sampling frame or response-rate reporting↳ Could also: A structured Delphi consensus process or formal mixed-methods survey with defined sampling frame, response tracking, and thematic saturation criteria could also be used for standards elicitation — Formal consensus methods (Delphi, nominal group technique) produce documented convergence metrics and make it possible to quantify agreement levels, which can strengthen the reproducibility and perceived legitimacy of the resulting standard
-
Pilot implementation was described narratively for one network (FDA GenomeTrakr SARS-CoV-2 wastewater), without quantitative evaluation of tag adoption, retrieval accuracy, or inter-annotator agreement↳ Could also: A quantitative usability evaluation could measure inter-annotator agreement (e.g., Cohen's kappa or Krippendorff's alpha) on whether independent annotators assign the same tags to the same datasets — Reliability coefficients would provide an objective, reproducible measure of tag clarity and could identify which tag terms benefit from additional definitional guidance
-
Tag utility was asserted based on expert consensus and a single-network implementation rather than a controlled comparison↳ Could also: A before-after or case-control study comparing retrieval precision and recall of quality-flagged datasets in repositories with and without standardized tags could also be conducted — Empirical retrieval metrics (precision, recall, F1) would quantify the practical discoverability benefit of the tags for downstream users
-
The paper categorizes QC failure reasons into a fixed enumerated list derived from internal review of bacterial and viral datasets↳ Could also: Qualitative thematic analysis with formal coding procedures (open coding, axial coding, saturation testing) or latent class analysis on structured survey data could also be used to derive QC categories — Explicit coding methodology with inter-rater reliability reporting would make the category-derivation process transparent and reproducible for future updates or extensions of the tag set
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38860884
Paper: Griffiths et al. 2024, Microb Genom 10(6):001260. "PHA4GE quality control contextual data tags: standardized annotations for sharing public health sequence datasets with known quality issues." PMID 38860884 · PMCID PMC11261899 · DOI 10.1099/mgen.0.001260
Code: https://github.com/pha4ge/contextual-data-qc-tags
(redirect target of the brief's contextual_data_QC_tags).
What kind of paper is this?
This is a standards / specification paper, not a computational-analysis
paper. Its product is a controlled vocabulary of "quality control contextual
data tags" (attribute names, picklist values, GENEPIO ontology IDs) for
annotating shared pathogen-genomics datasets that have known quality issues.
The GitHub repo ships exactly that specification — QC Contextual Data Tags Specification.csv, QC_Contextual_Data_Tags_Specification.json, two reference
spreadsheets, and table_json.py (a CSV→JSON converter). There is no
analysis pipeline in the repo and no shipped dataset+expected-output to run.
Reported results, classified
| # | Reported result | Source | Pipeline? | In scope? |
|---|---|---|---|---|
| R1 | The QC-tag specification itself (fields, picklists, GENEPIO IDs) | Table 1, repo CSV/JSON | No — a vocabulary, not a computed value | out (non_pipeline) |
| R2 | Worked examples 1–8: illustrative tag annotations for hypothetical QC scenarios using ncov-tools/FastQC/Kraken2/Nextclade/samtools-depth/VADR/Quast/PHoeNIx | §"Worked examples", Table 2 | Illustrative — no numeric values, no specific accession tied to any example | out (no_expected_result) |
| R3 | GenomeTrakr pilot: of 2,255 sequences, 61 % no QC issues, 10 % minor, 28 % potential, 1 % significant | §"Real-world implementations", Table 3 | Yes (GalaxyTrakr SSQuAWK / CFSAN C-WAP) but the full dataset + per-sequence determinations + classification thresholds are not shipped | out (hard 20%, docs_insufficient for repro) — see note |
| R4 | Table 4: five named GenomeTrakr SRA records "illustrating PHA4GE QC contextual data tag use" (incl. SRR21205381) | Table 4 | The claim is that these public records carry the standardized tags | IN — directly verifiable from SRA metadata |
What we attempt (the 80%)
-
C1 — Table 4 tag presence (control-plane, exact). Fetch the SRA metadata for all five Table-4 accessions and confirm each carries the PHA4GE QC contextual-data-tag attributes (
quality_control_method,_method_version,_determination,_issues,_details) with values drawn from the specification's picklist. This reproduces the paper's central real-world implementation claim 1:1. -
C2 — pipeline-derived consistency check («our HPC», third-party tools on the paper's data). The brief explicitly endorses running an existing tool on the paper's own data. The Table-4 determinations for wastewater SARS-CoV-2 amplicon data are coverage-driven (SSQuAWK/C-WAP). We therefore apply the paper's named QC tools — samtools (Table 2; worked example 3 uses
samtools depth) and FastQC (Table 2; worked example 4) — to two contrasting Table-4 samples and test whether the computed SARS-CoV-2 genome coverage breadth is consistent with the recorded determination:SRR21205381→ recorded "no quality control issues identified" → expect high breadth.SRR20428498→ recorded "low average genome coverage" (flagged) → expect low breadth. This is a qualitative/directional reproduction, not an attempt to replicate SSQuAWK's exact pass/fail thresholds (those are unspecified).
Explicitly NOT attempted (the hard 20%)
- R3 GenomeTrakr pilot percentages (61/10/28/1 % of 2,255). Would require the complete pilot dataset and the authors' per-sequence SSQuAWK/C-WAP determinations and their classification thresholds — none of which are shipped. Not reproducible from public artifacts; not attempted.
- **R1
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a standards/specification paper: the repo ships a controlled QC-tag vocabulary, not an analysis pipeline, so there is no shipped data+code+expected-output to run end-to-end. The central claim reproduced cleanly — all five Table-4 GenomeTrakr SRA records carry the standardized quality_control_* tags drawn from the spec picklist (1:1 against SRA + repo commit 2d5e0007), and a recomputed coverage check (99.9% breadth for the 'no issues' sample vs 61.87% for the 'low coverage' sample) is directionally consistent with the recorded determinations. The only gap is the secondary Table 3 pilot percentages (61/10/28/1% of 2255), which are not derivable because the pilot dataset and per-sequence thresholds were never deposited — a data-availability limitation, not an authors' defect or fabrication. Overall a solid, honest partial reproduction with explainable scope limits.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.