Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

First step toward gene expression data integration: transcriptomic data acquisition with COMMAND>_.

BMC Bioinformatics · 2019
L1 30/100 PQI 63
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
30/100
Reproducibility score
2.5 SD below mean
vs. all fields · 1187 studies
🎯 Scores higher than 2% of all assessed papers rank 1159 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Software-tool paper (COMMAND>_, a Django/ExtJS web app for acquiring gene-expression data). NOT well-described as a quantitative pipeline: it reports no pipeline-output values (no probe-mapping counts, no expression numbers, no metrics) -- only software versions, config defaults, and the mapping PARAMETERS. The single data-linked, checkable number is the case study's '273 samples' for the three GEO series it downloads. Reproducing the data-acquisition step against live NCBI GEO (the paper's actual case-study computational step) gives 412 GPL570 samples (54+135+223), NOT 273 -> mismatch. The 273 is citation-traceable (the small-airway-epithelium subset of the cited source study, Yi et al. 2018 / PMID 29616282: 288 SAE samples, 273 after QC), so it is NOT fabricated, but it is mis-attributed as the size of the three GEO experiments -- a reproducer following the paper's own download instructions gets 412. NOT ATTEMPTED (80/20, deliberately): (a) the BLAST+ probe->gene mapping -- fully parameterised but the paper reports NO output value, so there is no 1:1 target to grade (no_expected_result); running it would be demonstrative only and require the authors' exact Google-Drive gene FASTA + Affymetrix probe_tab; (b) deploying/operating the GUI web app -- interactive, no headless reproducibility target (non_pipeline). No «our HPC» SLURM job submitted because no heavy-compute step had a reported numeric target; «our HPC» connectivity was verified («host») and remains available. Honest outcome: partial reproduction with one documented, citation-traceable discrepancy.

💻 Code ↗ 🗄 Data: GSE8545

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 30
    assessed: 2026-06-14 ⛓ ad2d71b2dcd8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Effective integration of transcriptomic (gene expression) data from public repositories is hindered by data-format heterogeneity, and the first, common bottleneck across all integration strategies is raw data acquisition; the paper presents COMMAND>_ as a tool built to address this acquisition step.

Core claims
  • COMMAND>_ is a flexible multi-user web-application that allows users to search and download gene expression experiments, extract relevant information from experiment files, re-annotate microarray platforms, and present data in a coherent data model. resource
  • COMMAND>_ facilitates creation of local datasets of gene expression data from both microarray and RNA-seq experiments and may be a more efficient tool to build integrated gene expression compendia. resource
  • COMMAND>_ supports searching and downloading from NCBI GEO, SRA, and EBI ArrayExpress. method
  • Probe-to-gene mapping is performed via BLAST+ alignment combined with a two-step filtering procedure that mitigates the shortcomings of a single-step filtering approach. method
  • COMMAND>_ has been successfully used to build gene expression compendia such as COLOMBOS and VESPUCCI. finding
  • For RNA-seq platforms, FASTQ files are by default trimmed using Trimmomatic and expression quantified using Kallisto. method
  • Unlike comparable R/Bioconductor packages (GEOquery, ArrayExpress, GEOmetadb, SRAdb, compendiumdb, virtualArray) and Microarray retriever, COMMAND>_ uniquely combines a GUI, local relational database storage, connections to GEO/ArrayExpress/SRA, probe annotation, and free-text search. finding
  • All parameters used for re-annotation are stored within COMMAND>_, making the probe-to-gene mapping procedure completely reproducible. method
Experimental setups
Assay System Perturbation Readout Platform
microarray gene expression profiling small airway samples from COPD patients (human) none (disease cohort, COPD) gene expression measurements across samples Affymetrix HGU133Plus2
probe-to-gene sequence alignment (probe re-annotation) microarray platform probes (HGU133Plus2_Hs_ENSG_probe_tab) none probe-to-gene mapping determined by alignment score and two-step filtering thresholds BLAST+ (short-blastn option)
Key results
  • COMMAND>_ was used to search, download, parse, re-annotate, and export a collection of small airway COPD samples from three GEO microarray experiments (GSE8545, GSE20257, GSE11906). 273 samples, 3 experiments
  • A single-step filtering threshold (e.g., 95% or 96%) either wrongly retains non-unique probe alignments or discards valid ones, whereas the two-step filtering (e.g., 94% then 95%) correctly retains only the expected unique probe-gene alignments.
  • Comparison across tools (Table 1) shows COMMAND>_ is the only one providing GUI, local data storage, GEO/ArrayExpress/SRA connectivity, database storage, annotation, and search functionality together, while other tools are R/Bioconductor packages lacking a GUI and local storage.
Key statistics
  • count 273 (total samples in the COPD small airway case-study collection)
  • count 3 (number of Affymetrix microarray GEO experiments used in the case study (GSE8545, GSE20257, GSE11906))
  • other 95% alignment score threshold (example single-step filtering threshold for probe-to-gene alignment)
  • other 3% score difference threshold (threshold used to judge whether a probe aligns uniquely in the two-step filtering example)
  • count 8 (default number of concurrent processes managed by the Celery task queue)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a software description paper presenting COMMAND>_, a web application for gene expression data acquisition and management. No formal statistical hypothesis tests are reported; the paper's evaluative content consists of a qualitative feature-comparison table (binary yes/no entries across tools) and a narrative case study demonstrating the tool's workflow on publicly available COPD microarray data. Quantitative performance benchmarking and inferential statistics are absent.

Replicationunclear GroupsNo groups compared inferentially; tool features compared descriptively via a binary feature matrix (8 tools × 10 features) Pairingna Randomization/blindingna Dispersionnone
Approaches that could also have been used
  • Tool comparison was conducted via a binary (yes/no) feature-presence table across 8 tools and 10 features
    Could also: A quantitative benchmark could also be used — e.g., measuring wall-clock time, memory usage, or error rates on a standardised set of experiments across competing tools — Quantitative benchmarking would allow readers to assess practical performance differences in addition to feature coverage, which is particularly informative for scalability claims made in the text
  • Probe-to-gene mapping uses BLAST+ alignment followed by a two-step identity/uniqueness threshold filter, with parameters chosen per-platform by the user
    Could also: Bowtie2 or BWA short-read aligners are also commonly used for probe mapping and can offer faster runtimes; alternatively, manufacturer-supplied annotation files could serve as a reference baseline for comparison — Reporting alignment statistics (e.g., percentage of probes mapped, ambiguity rates) against a manufacturer baseline would allow readers to gauge the empirical effect of re-annotation choices
  • RNA-seq quantification defaults to Kallisto (pseudoalignment-based) with Trimmomatic for trimming
    Could also: STAR + featureCounts, HISAT2 + HTSeq, or Salmon are also widely used alignment-and-quantification pipelines for RNA-seq — Providing a configurable choice among quantification backends, or reporting concordance between Kallisto estimates and a splice-aware aligner on the case-study data, would help users understand how pipeline choice affects downstream counts
  • The case study demonstrates the workflow on 273 samples from 3 Affymetrix experiments but reports no quantitative validation of data quality or re-annotation accuracy
    Could also: Metrics such as median per-sample correlation before and after re-annotation, principal component analysis of batch structure, or comparison of differentially expressed gene lists with the original publication could also be reported — Quantitative quality-control summaries would give readers evidence that the ingested and re-annotated data are concordant with expected biological structure, complementing the workflow narrative
  • Scalability is described qualitatively (scripts 'scale linearly with respect to input size'; default 8 concurrent tasks)
    Could also: Empirical benchmarking — e.g., wall-clock time and peak RAM as a function of experiment count or file size — could also be reported — Empirical scaling curves would make the linearity claim verifiable and help administrators size hardware for their expected workloads
Software: Python/Django Python 3, Django 1.11 · PostgreSQL · Celery · ExtJS 6.2 · BLAST+ · Trimmomatic · Kallisto

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
9
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE11906 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE20257 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE8545 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

Downstream reach in the literature

48 downstream papers · 3 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-30691411

Paper: Moretto M, Sonego P, Villaseñor-Altamirano AB, Engelen K. First step toward gene expression data integration: transcriptomic data acquisition with COMMAND>_. BMC Bioinformatics 2019, article type = Software. DOI 10.1186/s12859-019-2643-6 · PMCID PMC6348648.

Code: https://github.com/marcomoretto/command (public, GPL-3.0, last push 2023-02-07). A multi-user web application (Python 3 / Django 1.11 backend, ExtJS 6.2 JavaScript GUI, Celery task queue, deployed via Docker Compose). Repo composition ~96% JavaScript, ~0.5% Python.

Nature of the paper

This is a software-tool description, not an empirical study. Its "Results" section is a feature/architecture description plus a GUI walkthrough (Figs 1–5) and a tool-comparison table (Table 1). It reports no quantitative pipeline-output values — no probe-mapping counts, no normalized expression values, no performance/runtime metrics, no statistics on imported data.

Complete inventory of numeric content in the paper

value where kind in scope?
273 samples from three Affymetrix experiments (GSE8545, GSE20257, GSE11906) Results → Case study data-acquisition claim (the only data-linked number) YES — checkable against GEO
Python 3 / Django 1.11 / ExtJS 6.2 Implementation software versions (config) no — not a result
Celery concurrency "8 by default" Implementation config default no — not a result
Two-step filter: 95% len / 0 gap / 3 mismatch (sensitivity); 98% len / 0 gap / 1 mismatch (specificity) Implementation input parameters of the probe→gene mapping parameters, not outputs
Fig. 5 worked example: 95/94/96/3% Figure caption hypothetical illustration no — not real data
GPL570 Case study platform identifier context

In scope (attempted)

  1. Data-acquisition step of the case study — reproduce the sample acquisition for the three GEO series the paper's case study downloads, and check the reported "273 samples" against the live GEO deposits. Pipeline = COMMAND>_ "Download Experiment From Public Database" (GEO retrieval). Reproduced on the control plane via NCBI GEO E-utilities / acc.cgi (authoritative metadata; no heavy compute required for a sample count).

Out of scope (not attempted) — and why

  1. Probe-to-gene mapping (BLAST+ two-step filter on GPL570). This is the one genuine heavy-compute pipeline in the paper, and its parameters are fully specified (95/0/3 then 98/0/1). But the paper reports no output value (number of probes/genes mapped) — so there is no 1:1 target to grade against (no_expected_result). Running it would be demonstrative only and would require chasing the exact inputs the authors used (a custom human gene FASTA on Google Drive + the Affymetrix login-walled probe_tab). Per the 80/20 rule this is the optional last ~20% with weak evidential value, so it was not run. Feasible on «our HPC» if a target ever materialises.
  2. Deploying and operating the COMMAND>_ web app (search→download→parse→ preview→import via the ExtJS GUI). It is an interactive, GUI-driven multi- service application with no headless entry point and no expected output to compare — not a reproducible computational result in this study's sense (non_pipeline for the GUI workflow itself).
  3. RNA-seq path (Trimmomatic + Kallisto). Mentioned as a capability; the case study is microarray-only and reports no RNA-seq numbers.

Honesty note

The paper passed screening (code public + data public), but at reproduction time the substance is thin: a single pinnable quantitative claim, which is itself a restatement of the cited source study (ref 31), and it does not match the GEO series sizes a reproducer would obtain by following the paper's own instructions. See AUDIT.md.

Figures / tables: Fig. 4c
C1
Reported
273 samples from three Affymetrix microarray experiments (GSE8545, GSE20257, GSE11906)
Reproduced
412 samples (GSE8545=54, GSE20257=135, GSE11906=223; all GPL570)
did not match
C2
Reported
probe->gene mapping output for GPL570 (BLAST+ two-step filter 95/0/3, 98/0/1) -- NONE reported
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 30/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

A software-tool paper (COMMAND>_) that reports no quantitative pipeline outputs — only software versions, config defaults and mapping parameters. The single data-linked checkable number is the case study's '273 samples from three Affymetrix experiments'; downloading GSE8545+GSE20257+GSE11906 per the paper's own instructions yields 412 (54+135+223, all GPL570) — a mismatch. The agent traced 273 to the small-airway-epithelium subset of the cited source study (Yi et al. 2018, 288 SAE -> 273 after QC), so it is citation-traceable and NOT fabricated, but mis-attributed as the size of the three GEO experiments. The probe->gene mapping (C2) has no reported output value, so it is uncheckable. A benign description error flagged for the human, not a substantive discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

115.6 k
tokens (I/O) · 5.9 M incl. cache
11 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.