Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

The Planemo toolkit for developing, deploying, and executing scientific data analyses in Galaxy and beyond.

Genome Res · 2023
L1 No computation 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • No authors-side cause for any deviation
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

Software/methods paper describing the Planemo CLI toolkit for Galaxy. It reports NO pipeline-derived computational result: no analyzed research dataset, no benchmark, no timing/test-pass numbers, no figures of computed values. The single quantitative claim is a usage metric ('downloaded more than 70,000 times from both Anaconda and PyPI'), a monotonic point-in-time counter that is directionally confirmed and now vastly exceeded (PyPI 1,666,288 + bioconda 157,825 = ~1.82M on 2026-06-19) but is NOT bit-reproducible (cumulative download counters only grow; historical totals at the 2022 submission snapshot are not exposed by either index). The brief's data accession (zenodo 4774217) is a CITED REFERENCE to the conda-forge project record (logo.svg + LICENSE), NOT a research-data deposit for this paper -- there is no paper-specific dataset to profile or reproduce. The named code artifacts (planemo, ptdk, iwc, planemo-ci-action) are tools, not analyses. Verdict: drop / non_pipeline -- there is no computational pipeline result to regenerate. NOT attempted: running planemo's tutorials/linting as a 'does the tool execute' demo, because that would not reproduce any REPORTED value. The download claim was verified as supporting evidence and graded partial (consistent, not bit-reproducible).

💻 Code ↗ 🗄 Data: 10.5281/zenodo.4774217

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-19 ⛓ f3e0d3125fc3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Core claims
  • Planemo is a software development kit that streamlines development, testing, deployment, and execution of Galaxy tools, workflows, and training material. resource
  • Planemo encourages and enforces software development best practices such as test-driven development and linting for tool/workflow wrappers. method
  • Planemo also functions as a software development kit for Common Workflow Language (CWL) tools via the same subcommands (tool_init, test, lint) with a --cwl flag. resource
  • Planemo's autoupdate subcommand, combined with Bioconda/conda-forge automation, forms a semiautomated pipeline that updates Galaxy tool and workflow dependencies after upstream software releases. method
  • Galaxy's separation of concerns between simple, reusable tools and higher-level workflows improves usability and allows workflow security to be established via trusted, reviewed component tools. mechanism
  • Planemo's testing and linting functionality has been integrated into CI pipelines of major community tool and workflow repositories (e.g., IUC, IWC). resource
  • More than 8000 tools are available for installation on Galaxy servers. finding
  • Planemo has been downloaded more than 70,000 times from Anaconda and PyPI. finding
Key results
  • Planemo has been downloaded more than 70,000 times from both Anaconda and PyPI, indicating extensive community adoption. 70,000+ downloads
  • More than 8000 tools are available for installation onto Galaxy servers, reflecting the scale of the ecosystem Planemo supports. 8000+ tools
Key statistics
  • count more than 70,000 downloads (Planemo downloads from Anaconda and PyPI)
  • count more than 8000 tools (tools available for installation on any Galaxy server)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Replicationunclear Groupsna Pairingna Randomization/blindingna Dispersionnone
Software: Planemo · Galaxy · cwltool · R

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

downloads-70k
Reported
>70,000 downloads from both Anaconda and PyPI (Results, "Galaxy tool development")
Reproduced
PyPI cumulative 1,666,288 + bioconda 157,825 = ~1.82M as of 2026-06-19 (pepy.tech + api.anaconda.org)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

This is a software/methods paper (Planemo toolkit) with no pipeline-derived computational result — no dataset, benchmark, or computed figure to regenerate, so drop/non_pipeline is correct. Its single quantitative claim, '>70,000 downloads from Anaconda and PyPI', is a monotonic cumulative counter that is directionally confirmed and now far exceeded (~1.82M on 2026-06-19) but not bit-reproducible because neither index exposes the 2022 historical total. The deviation is entirely on the technical/expected side (a growing counter + restricted historical snapshot), not an authors' defect; the cited zenodo record is a citation artifact, not research data. Overall yellow: solid and honest, but no 1:1 reproducible value.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

50.7 k
tokens (I/O) · 1.7 M incl. cache
4 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.