Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

First step toward gene expression data integration: transcriptomic data acquisition with COMMAND>_.

BMC Bioinformatics · 2019
L1 30/100 PQI 65
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
30/100
Reproducibility score
2.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 2% of all assessed papers rank 1148 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Software-tool paper (COMMAND>_, a Django/ExtJS web app for acquiring gene-expression data). NOT well-described as a quantitative pipeline: it reports no pipeline-output values (no probe-mapping counts, no expression numbers, no metrics) -- only software versions, config defaults, and the mapping PARAMETERS. The single data-linked, checkable number is the case study's '273 samples' for the three GEO series it downloads. Reproducing the data-acquisition step against live NCBI GEO (the paper's actual case-study computational step) gives 412 GPL570 samples (54+135+223), NOT 273 -> mismatch. The 273 is citation-traceable (the small-airway-epithelium subset of the cited source study, Yi et al. 2018 / PMID 29616282: 288 SAE samples, 273 after QC), so it is NOT fabricated, but it is mis-attributed as the size of the three GEO experiments -- a reproducer following the paper's own download instructions gets 412. NOT ATTEMPTED (80/20, deliberately): (a) the BLAST+ probe->gene mapping -- fully parameterised but the paper reports NO output value, so there is no 1:1 target to grade (no_expected_result); running it would be demonstrative only and require the authors' exact Google-Drive gene FASTA + Affymetrix probe_tab; (b) deploying/operating the GUI web app -- interactive, no headless reproducibility target (non_pipeline). No «our HPC» SLURM job submitted because no heavy-compute step had a reported numeric target; «our HPC» connectivity was verified («host») and remains available. Honest outcome: partial reproduction with one documented, citation-traceable discrepancy.

💻 Code ↗ 🗄 Data: GSE8545

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 30
    assessed: 2026-06-14 ⛓ ad2d71b2dcd8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Gene expression data integration requires a common first step—data acquisition—which is hampered by heterogeneous data formats; the paper presents COMMAND>_, a tool to simplify acquisition of transcriptomic data from public databases into a coherent local data model.

Core claims
  • COMMAND>_ is a flexible multi-user web application that searches, downloads, parses, re-annotates, and imports gene expression experiments into a coherent data model. resource
  • COMMAND>_ handles both microarray and RNA-seq data and can be extended to other quantitative data platforms. method
  • A two-step filtering of BLAST+ alignments re-annotates microarray probes to genes, retaining high-similarity, uniquely-mapping probes while avoiding cross-hybridization. method
  • Probe-to-gene re-annotation enhances data homogeneity by annotating all platforms against the same genomic background and stores all parameters for full reproducibility. method
  • COMMAND>_ scales with growing experiment numbers by relying on a relational database (PostgreSQL) and the Celery task queue system rather than R/Bioconductor's single-threaded, in-RAM model. method
  • COMMAND>_ has been used to build the COLOMBOS and VESPUCCI gene expression compendia. resource
  • No other existing tool offers the full combined set of functionalities (R-independent, local storage, GUI, GEO/AE/SRA access, internal DB, probe annotation, free-text search) that COMMAND>_ provides. finding
Experimental setups
Assay System Perturbation Readout Platform
Affymetrix microarray gene expression (GeneChip) small airway epithelial samples from COPD patients (human) none (disease vs control observational COPD cohort) probe-level expression intensity from CEL files, re-annotated to genes Affymetrix HGU133Plus2 (GSE8545, GSE20257, GSE11906)
RNA-seq expression quantification (software workflow) user-defined genomic background (FASTA gene sequences) none raw read counts per gene Trimmomatic (trimming) and Kallisto (quantification)
Probe-to-gene re-annotation via sequence alignment microarray platform probe sequences vs genomic background none alignment scores; two-step similarity/uniqueness filtering BLAST+ (short-blastn option)
Key results
  • Demonstrated end-to-end search, download, parse, re-annotate, and export of a COPD small-airway sample collection across three Affymetrix experiments. 273 samples
  • Two-step filtering example correctly retains only the orange and green probes as unique high-similarity mappings, avoiding the false retention/loss seen with single-step filtering.
  • Functional comparison shows COMMAND>_ is the only tool supporting all listed functionalities including GUI, internal DB, annotation, and SRA/GEO/ArrayExpress access.
Key statistics
  • count 273 samples (COPD small airway samples in the case-study collection)
  • count three Affymetrix microarray experiments (GEO experiments GSE8545, GSE20257, GSE11906 used in case study)
  • count 8 by default (default number of concurrent Celery processes for downloading/parsing)
  • other 95% or more similarity threshold; >3% score difference for uniqueness (example thresholds in two-step probe filtering illustration)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a software description paper presenting COMMAND>_, a web application for gene expression data acquisition and management. No formal statistical hypothesis tests are reported; the paper's evaluative content consists of a qualitative feature-comparison table (binary yes/no entries across tools) and a narrative case study demonstrating the tool's workflow on publicly available COPD microarray data. Quantitative performance benchmarking and inferential statistics are absent.

Replicationunclear GroupsNo groups compared inferentially; tool features compared descriptively via a binary feature matrix (8 tools × 10 features) Pairingna Randomization/blindingna Dispersionnone
Approaches that could also have been used
  • Tool comparison was conducted via a binary (yes/no) feature-presence table across 8 tools and 10 features
    Could also: A quantitative benchmark could also be used — e.g., measuring wall-clock time, memory usage, or error rates on a standardised set of experiments across competing tools — Quantitative benchmarking would allow readers to assess practical performance differences in addition to feature coverage, which is particularly informative for scalability claims made in the text
  • Probe-to-gene mapping uses BLAST+ alignment followed by a two-step identity/uniqueness threshold filter, with parameters chosen per-platform by the user
    Could also: Bowtie2 or BWA short-read aligners are also commonly used for probe mapping and can offer faster runtimes; alternatively, manufacturer-supplied annotation files could serve as a reference baseline for comparison — Reporting alignment statistics (e.g., percentage of probes mapped, ambiguity rates) against a manufacturer baseline would allow readers to gauge the empirical effect of re-annotation choices
  • RNA-seq quantification defaults to Kallisto (pseudoalignment-based) with Trimmomatic for trimming
    Could also: STAR + featureCounts, HISAT2 + HTSeq, or Salmon are also widely used alignment-and-quantification pipelines for RNA-seq — Providing a configurable choice among quantification backends, or reporting concordance between Kallisto estimates and a splice-aware aligner on the case-study data, would help users understand how pipeline choice affects downstream counts
  • The case study demonstrates the workflow on 273 samples from 3 Affymetrix experiments but reports no quantitative validation of data quality or re-annotation accuracy
    Could also: Metrics such as median per-sample correlation before and after re-annotation, principal component analysis of batch structure, or comparison of differentially expressed gene lists with the original publication could also be reported — Quantitative quality-control summaries would give readers evidence that the ingested and re-annotated data are concordant with expected biological structure, complementing the workflow narrative
  • Scalability is described qualitatively (scripts 'scale linearly with respect to input size'; default 8 concurrent tasks)
    Could also: Empirical benchmarking — e.g., wall-clock time and peak RAM as a function of experiment count or file size — could also be reported — Empirical scaling curves would make the linearity claim verifiable and help administrators size hardware for their expected workloads
Software: Python/Django Python 3, Django 1.11 · PostgreSQL · Celery · ExtJS 6.2 · BLAST+ · Trimmomatic · Kallisto

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
9
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE11906 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE20257 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE8545 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

Downstream reach in the literature

48 downstream papers · 3 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-30691411

Paper: Moretto M, Sonego P, Villaseñor-Altamirano AB, Engelen K. First step toward gene expression data integration: transcriptomic data acquisition with COMMAND>_. BMC Bioinformatics 2019, article type = Software. DOI 10.1186/s12859-019-2643-6 · PMCID PMC6348648.

Code: https://github.com/marcomoretto/command (public, GPL-3.0, last push 2023-02-07). A multi-user web application (Python 3 / Django 1.11 backend, ExtJS 6.2 JavaScript GUI, Celery task queue, deployed via Docker Compose). Repo composition ~96% JavaScript, ~0.5% Python.

Nature of the paper

This is a software-tool description, not an empirical study. Its "Results" section is a feature/architecture description plus a GUI walkthrough (Figs 1–5) and a tool-comparison table (Table 1). It reports no quantitative pipeline-output values — no probe-mapping counts, no normalized expression values, no performance/runtime metrics, no statistics on imported data.

Complete inventory of numeric content in the paper

value where kind in scope?
273 samples from three Affymetrix experiments (GSE8545, GSE20257, GSE11906) Results → Case study data-acquisition claim (the only data-linked number) YES — checkable against GEO
Python 3 / Django 1.11 / ExtJS 6.2 Implementation software versions (config) no — not a result
Celery concurrency "8 by default" Implementation config default no — not a result
Two-step filter: 95% len / 0 gap / 3 mismatch (sensitivity); 98% len / 0 gap / 1 mismatch (specificity) Implementation input parameters of the probe→gene mapping parameters, not outputs
Fig. 5 worked example: 95/94/96/3% Figure caption hypothetical illustration no — not real data
GPL570 Case study platform identifier context

In scope (attempted)

  1. Data-acquisition step of the case study — reproduce the sample acquisition for the three GEO series the paper's case study downloads, and check the reported "273 samples" against the live GEO deposits. Pipeline = COMMAND>_ "Download Experiment From Public Database" (GEO retrieval). Reproduced on the control plane via NCBI GEO E-utilities / acc.cgi (authoritative metadata; no heavy compute required for a sample count).

Out of scope (not attempted) — and why

  1. Probe-to-gene mapping (BLAST+ two-step filter on GPL570). This is the one genuine heavy-compute pipeline in the paper, and its parameters are fully specified (95/0/3 then 98/0/1). But the paper reports no output value (number of probes/genes mapped) — so there is no 1:1 target to grade against (no_expected_result). Running it would be demonstrative only and would require chasing the exact inputs the authors used (a custom human gene FASTA on Google Drive + the Affymetrix login-walled probe_tab). Per the 80/20 rule this is the optional last ~20% with weak evidential value, so it was not run. Feasible on «our HPC» if a target ever materialises.
  2. Deploying and operating the COMMAND>_ web app (search→download→parse→ preview→import via the ExtJS GUI). It is an interactive, GUI-driven multi- service application with no headless entry point and no expected output to compare — not a reproducible computational result in this study's sense (non_pipeline for the GUI workflow itself).
  3. RNA-seq path (Trimmomatic + Kallisto). Mentioned as a capability; the case study is microarray-only and reports no RNA-seq numbers.

Honesty note

The paper passed screening (code public + data public), but at reproduction time the substance is thin: a single pinnable quantitative claim, which is itself a restatement of the cited source study (ref 31), and it does not match the GEO series sizes a reproducer would obtain by following the paper's own instructions. See AUDIT.md.

Figures / tables: Fig. 4c
C1
Reported
273 samples from three Affymetrix microarray experiments (GSE8545, GSE20257, GSE11906)
Reproduced
412 samples (GSE8545=54, GSE20257=135, GSE11906=223; all GPL570)
did not match
C2
Reported
probe->gene mapping output for GPL570 (BLAST+ two-step filter 95/0/3, 98/0/1) -- NONE reported
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 30/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

A software-tool paper (COMMAND>_) that reports no quantitative pipeline outputs — only software versions, config defaults and mapping parameters. The single data-linked checkable number is the case study's '273 samples from three Affymetrix experiments'; downloading GSE8545+GSE20257+GSE11906 per the paper's own instructions yields 412 (54+135+223, all GPL570) — a mismatch. The agent traced 273 to the small-airway-epithelium subset of the cited source study (Yi et al. 2018, 288 SAE -> 273 after QC), so it is citation-traceable and NOT fabricated, but mis-attributed as the size of the three GEO experiments. The probe->gene mapping (C2) has no reported output value, so it is uncheckable. A benign description error flagged for the human, not a substantive discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

115.6 k
tokens (I/O) · 5.9 M incl. cache
11 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.