The Dockstore: enhancing a community platform for sharing reproducible and accessible computational protocols.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No authors-side cause for any deviation
- ✓The central claim held under reproduction
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
▸Reproduction agent’s raw note
Dockstore is a NAR Web Server / platform-description paper, not a data-analysis study. It contains NO pipeline-derived computational result: there is no experimental dataset run through a bioinformatic pipeline to yield a reported value. The cited 'data' DOI (zenodo:10.5281/zenodo.4536482) is the Dockstore platform's own software release archive (dockstore-1.10.1.zip, 1.9 MB) -- code, not data. The headline numbers (705 workflows, 240 tools, >25 organizations) are snapshots of a live, continuously-growing production database with no deposited frozen copy, so they are non-reproducible by construction. Verdict: DROP / non_pipeline (well-described paper, but out of reproduction scope). No «our HPC» compute was spent. As control-plane context only, the public Dockstore API was queried once on 2026-06-18 confirming the platform is alive (6093 published workflows, 258 tools, GA4GH TRS v2.0.1 prod) and the GitHub repo is public/Apache-2.0/actively maintained -- documenting WHY the counts cannot be reproduced (they grew ~8.6x) rather than reproducing them. NOT attempted: re-running any pipeline (none exists); reconstructing the 2021 DB snapshot (not deposited).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-18 ⛓ f2965a80061c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper presents enhancements to Dockstore, an open-source platform for publishing, sharing, and finding bioinformatics tools and workflows, aiming to increase the FAIRness and reproducibility of computational analyses across diverse cloud environments.
- ★ Dockstore is a language- and platform-agnostic registry for containerized bioinformatics tools and workflows, distinguishing it from single-language registries like Agora and Galaxy Toolshed. resource
- ★ Dockstore combines Docker containers with workflow descriptor languages to enable reproducible execution across multiple computing environments. method
- ★ Dockstore added 'Launch with' integrations enabling one-click execution of workflows on multiple academic and commercial cloud platforms via the GA4GH TRS API. method
- ★ Dockstore expanded workflow language support to include Nextflow and Galaxy, in addition to existing WDL and CWL support (CWL 1.1, WDL 1.0). resource
- ★ Snapshots, checksums (for descriptors and Docker images), and Zenodo DOI integration make workflow versions immutable and citable, improving reproducibility and security. method
- ★ Dockstore is the leading implementation of the GA4GH Tool Registry Service (TRS) standard, implementing the official 2.0.0 version plus two draft standards. resource
- A GitHub app enables automatic synchronization of new workflow versions from source control without revisiting the Dockstore site. method
- Checker workflows test that a biologically significant workflow runs correctly across new computing environments, explicitly testing scientific reproducibility. method
- – More than twenty-five high-profile organizations share analysis collections through Dockstore in a variety of workflow languages (e.g., Broad GATK/COVID-19 WDL, nf-core Nextflow, IWC Galaxy, Seven Bridges CWL). >25 organizations
- – Dockstore originated from the PCAWG study, running cancer variant-calling workflows reproducibly across 14 cloud and conventional computing infrastructures. 14 infrastructures
- – PCAWG analyzed an internationally federated set of whole cancer genomes totalling roughly 1 PB in size. ~1 PB
- – More than 250 workflow engines have been tracked, many requiring difficult configuration to run outside their home institutions. >250 engines
- – GA4GH TRS became an official GA4GH standard in October 2019; Dockstore implements the official 2.0.0 version and two draft standards. TRS 2.0.0
- count >25 (high-profile organizations sharing analysis collections through Dockstore)
- count >250 (workflow engines tracked to date)
- count 14 (cloud and conventional computing infrastructures used in PCAWG)
- other ~1 PB (total size of whole cancer genomes in PCAWG study)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
-
Platform adoption and usage are described with simple counts and qualitative statements (e.g., '>25 high-profile organizations', '>250 workflow engines tracked').↳ Could also: Longitudinal growth metrics such as registered-tool counts, active-user counts, or API call volumes over time could also be reported, presented as time-series summaries with counts per period. — Quantitative usage trajectories would allow readers to assess the pace and scale of adoption, complementing the qualitative description of organizational uptake.
-
The scope of Launch-with partner integrations is summarized in a descriptive table (Table 2) without quantitative uptake data.↳ Could also: Usage frequency per partner (e.g., number of workflow launches per platform) could also be reported as descriptive counts or proportions. — Counts of launches per partner would give readers a more concrete sense of which integrations are most actively used by the community.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Nothing reproducible exists. This is a Web Server / platform-description paper.
It reports no pipeline-derived computational result; the cited "data" DOI is the
Dockstore software release archive (code, not data); the headline numbers (705
workflows, 240 tools, >25 orgs) are live production-database snapshots with no
deposited frozen copy. Outcome: drop / non_pipeline — see
reproduction/scope.md, reproduction/ROOM_RESULT.json, AUDIT.md.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a NAR Web-Server/platform-description paper (Dockstore) with no pipeline-derived computational result to reproduce: the cited 'data' DOI is a software release archive, and the headline numbers (705 workflows, 240 tools, >25 orgs) are snapshots of a live, ever-growing production database with no deposited freeze. The deviations (705->6093 workflows, 240->258 tools on a 2026 live API read) are fully on the data-availability/scope side — organic platform growth — not an authors' defect or fabrication. The central claim (a live multi-language GA4GH-TRS workflow platform) is confirmed (C4: CWL/WDL/Nextflow/Galaxy still advertised; platform alive and larger). Correct outcome: drop / non_pipeline — solid paper, out of reproduction scope.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.