Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The Dockstore: enhancing a community platform for sharing reproducible and accessible computational protocols.

Nucleic Acids Res · 2021
L1 No computation 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • No authors-side cause for any deviation
  • The central claim held under reproduction
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

Dockstore is a NAR Web Server / platform-description paper, not a data-analysis study. It contains NO pipeline-derived computational result: there is no experimental dataset run through a bioinformatic pipeline to yield a reported value. The cited 'data' DOI (zenodo:10.5281/zenodo.4536482) is the Dockstore platform's own software release archive (dockstore-1.10.1.zip, 1.9 MB) -- code, not data. The headline numbers (705 workflows, 240 tools, >25 organizations) are snapshots of a live, continuously-growing production database with no deposited frozen copy, so they are non-reproducible by construction. Verdict: DROP / non_pipeline (well-described paper, but out of reproduction scope). No «our HPC» compute was spent. As control-plane context only, the public Dockstore API was queried once on 2026-06-18 confirming the platform is alive (6093 published workflows, 258 tools, GA4GH TRS v2.0.1 prod) and the GitHub repo is public/Apache-2.0/actively maintained -- documenting WHY the counts cannot be reproduced (they grew ~8.6x) rather than reproducing them. NOT attempted: re-running any pipeline (none exists); reconstructing the 2021 DB snapshot (not deposited).

💻 Code ↗ 🗄 Data: 10.5281/zenodo.4536482

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-18 ⛓ f2965a80061c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper presents enhancements to Dockstore, an open-source platform for publishing, sharing, and finding bioinformatics tools and workflows, aiming to increase the FAIRness and reproducibility of computational analyses across diverse cloud environments.

Core claims
  • Dockstore is a language- and platform-agnostic registry for containerized bioinformatics tools and workflows, distinguishing it from single-language registries like Agora and Galaxy Toolshed. resource
  • Dockstore combines Docker containers with workflow descriptor languages to enable reproducible execution across multiple computing environments. method
  • Dockstore added 'Launch with' integrations enabling one-click execution of workflows on multiple academic and commercial cloud platforms via the GA4GH TRS API. method
  • Dockstore expanded workflow language support to include Nextflow and Galaxy, in addition to existing WDL and CWL support (CWL 1.1, WDL 1.0). resource
  • Snapshots, checksums (for descriptors and Docker images), and Zenodo DOI integration make workflow versions immutable and citable, improving reproducibility and security. method
  • Dockstore is the leading implementation of the GA4GH Tool Registry Service (TRS) standard, implementing the official 2.0.0 version plus two draft standards. resource
  • A GitHub app enables automatic synchronization of new workflow versions from source control without revisiting the Dockstore site. method
  • Checker workflows test that a biologically significant workflow runs correctly across new computing environments, explicitly testing scientific reproducibility. method
Key results
  • More than twenty-five high-profile organizations share analysis collections through Dockstore in a variety of workflow languages (e.g., Broad GATK/COVID-19 WDL, nf-core Nextflow, IWC Galaxy, Seven Bridges CWL). >25 organizations
  • Dockstore originated from the PCAWG study, running cancer variant-calling workflows reproducibly across 14 cloud and conventional computing infrastructures. 14 infrastructures
  • PCAWG analyzed an internationally federated set of whole cancer genomes totalling roughly 1 PB in size. ~1 PB
  • More than 250 workflow engines have been tracked, many requiring difficult configuration to run outside their home institutions. >250 engines
  • GA4GH TRS became an official GA4GH standard in October 2019; Dockstore implements the official 2.0.0 version and two draft standards. TRS 2.0.0
Key statistics
  • count >25 (high-profile organizations sharing analysis collections through Dockstore)
  • count >250 (workflow engines tracked to date)
  • count 14 (cloud and conventional computing infrastructures used in PCAWG)
  • other ~1 PB (total size of whole cancer genomes in PCAWG study)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Replicationunclear GroupsNo experimental groups compared; this is a software/platform description paper with no hypothesis-driven comparisons Pairingna Randomization/blindingna Dispersionnone
Approaches that could also have been used
  • Platform adoption and usage are described with simple counts and qualitative statements (e.g., '>25 high-profile organizations', '>250 workflow engines tracked').
    Could also: Longitudinal growth metrics such as registered-tool counts, active-user counts, or API call volumes over time could also be reported, presented as time-series summaries with counts per period. — Quantitative usage trajectories would allow readers to assess the pace and scale of adoption, complementing the qualitative description of organizational uptake.
  • The scope of Launch-with partner integrations is summarized in a descriptive table (Table 2) without quantitative uptake data.
    Could also: Usage frequency per partner (e.g., number of workflow launches per platform) could also be reported as descriptive counts or proportions. — Counts of launches per partner would give readers a more concrete sense of which integrations are most actively used by the community.
Software: Not stated — paper describes the Dockstore platform itself, not analysis software

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Nothing reproducible exists. This is a Web Server / platform-description paper. It reports no pipeline-derived computational result; the cited "data" DOI is the Dockstore software release archive (code, not data); the headline numbers (705 workflows, 240 tools, >25 orgs) are live production-database snapshots with no deposited frozen copy. Outcome: drop / non_pipeline — see reproduction/scope.md, reproduction/ROOM_RESULT.json, AUDIT.md.

C1
Reported
705 workflows published
Reproduced
6093 (live API, 2026-06-18; not a reproduction)
partial
C2
Reported
240 tools published
Reproduced
258 (live API, 2026-06-18; not a reproduction)
partial
C4
Reported
CWL/WDL/Nextflow/Galaxy support
Reproduced
confirmed: advertised by live GA4GH TRS service-info
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is a NAR Web-Server/platform-description paper (Dockstore) with no pipeline-derived computational result to reproduce: the cited 'data' DOI is a software release archive, and the headline numbers (705 workflows, 240 tools, >25 orgs) are snapshots of a live, ever-growing production database with no deposited freeze. The deviations (705->6093 workflows, 240->258 tools on a 2026 live API read) are fully on the data-availability/scope side — organic platform growth — not an authors' defect or fabrication. The central claim (a live multi-language GA4GH-TRS workflow platform) is confirmed (C4: CWL/WDL/Nextflow/Galaxy still advertised; platform alive and larger). Correct outcome: drop / non_pipeline — solid paper, out of reproduction scope.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

66 k
tokens (I/O) · 3.4 M incl. cache
6 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.