Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

GREIN: An Interactive Web Platform for Re-analyzing GEO RNA-seq Data.

Sci Rep · 2019
L1 73/100 PQI 93
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
73/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 664 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced paper Table 2 (statistical power analysis of the GSE104193 use-case) end-to-end on «our HPC»: GREP2 pipeline = Salmon 2.1.1 (Ensembl r91 cDNA+ncRNA, k=31, -l A defaults) -> tximport lengthScaledTPM (tx2gene EnsDb.Hsapiens.v86) -> gene counts -> GREIN server.R power code (filter cpm>1 in >=2 samples; edgeR estimateCommonDisp; depth=pseudo.lib.size/1e6; BCOV=sqrt(common.dispersion); RNASeqPower::rnapower alpha=0.01 effect=2). RESULT = strong PARTIAL: all 4 average-sequencing-depths reproduce within 2.5% (42.68/38.93/27.70/29.27 vs 41.72/38.18/27.86/29.52); 3/4 common-BCOV within tolerance; 14/20 Table-2 cells within-tol-or-exact. The 6 partial + 1 mismatch are power columns, driven by (a) Salmon version drift (paper's Salmon version unstated; we used 2.1.1 selective alignment, which post-dates the 2019 study and shifts counts a few % -> dispersion -> power), and (b) an INTERNAL INCONSISTENCY in the paper's own MDA-MB-231 Hypoxia-vs-normoxia row: its printed power 0.73/0.86 is NOT recoverable from its printed depth=27.86 & BCOV=0.24 under its own rnapower call (which yields 0.58/0.74); BCOV0.19 (our reproduced value) reconciles them -> flagged as a likely BCOV typo for human review, not fabrication. An independent cross-check feeding the PAPER's own depth+BCOV into rnapower reproduces 3/4 rows exactly/to-rounding, validating the power formula. The paper is described well enough to reproduce 1:1 and we did. NOT attempted (out of scope / hard 20%): the >6,500-dataset scale claim, all qualitative figures, and DEG gene lists (no pinnable numbers).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-14 ⛓ 6b67e5bd6af0
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Existing resources for reusing GEO RNA-seq data lack comprehensive, user-friendly downstream analytical tools (exploratory analysis, batch-adjusted differential expression, power analysis), so a web platform combining a large library of uniformly processed datasets with an accessible analytical toolbox (GREIN) can remove technical barriers to re-analyzing public RNA-seq data.

Core claims
  • GREIN is a web application providing user-friendly interfaces to manipulate, visualize, and analyze GEO RNA-seq data. resource
  • GREIN's back-end (GREP2 pipeline) provides access to more than 6,500 uniformly processed human, mouse, and rat GEO RNA-seq datasets with over 400,000 samples. resource
  • GREIN uniquely combines interactive exploratory analysis, QC reporting, statistical power analysis, on-the-fly differential expression analysis, and enrichment analysis compared to other RNA-seq resources (Recount2, ARCHS4, Toil, Skymap, Expression Atlas). resource
  • GREIN connects differential expression signatures to iLINCS for enrichment and connectivity analysis against LINCS L1000 signatures. method
  • In re-analysis of GSE104193, MCF10A (non-malignant) cells show a higher number of differentially expressed genes than MDA-MB-231 (TNBC) cells under hypoxia and hypoxia+PP242, suggesting the tumor cell line is better equipped to deal with hypoxia. finding
  • Statistical power analysis shows four replicates per group are needed to achieve ~80% power to detect a 2-fold expression change at α=0.01. finding
  • Enrichment analysis of the hypoxia signature identified an activated HIF-1-alpha transcription factor network common to both cell lines. finding
  • GREIN and GREP2 are released as open-source R package and Docker container for local/offline deployment. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (re-analysis of processed counts) MCF10A (non-malignant breast epithelial) and MDA-MB-231 (TNBC) cell lines, GEO dataset GSE104193 hypoxia (0.5% O2) vs normoxia (21% O2), ± PP242 (mTORC1/2 inhibitor) gene expression counts (total and polysome-bound mRNA fractions) Salmon (read mapping, via GREP2 pipeline)
correlation heatmap / hierarchical clustering MCF10A and MDA-MB-231 cell lines (GSE104193) none (exploratory analysis of full expression profiles / top 500 variable genes) sample-sample correlation, clustering structure
PCA and t-SNE (2D/3D dimensionality reduction) MCF10A and MDA-MB-231 cell lines (GSE104193) cell line, oxygen level, treatment, mRNA fraction as grouping variables sample separation/variance structure
statistical power analysis (power curve, gene detectability) MCF10A and MDA-MB-231 cell lines (GSE104193) hypoxia ± PP242 vs normoxia statistical power to detect 2-fold differential expression at α=0.01 across sample sizes
differential gene expression analysis (GLM with batch/covariate adjustment) MCF10A and MDA-MB-231 cell lines (GSE104193) hypoxia and hypoxia+PP242 vs normoxia control, replicate as covariate differentially expressed genes (FDR, log fold change)
gene list/pathway enrichment analysis DE and NDE&DT gene signatures from MCF10A and MDA-MB-231 hypoxia comparisons none (post-hoc functional analysis) enriched GO terms/pathways DAVID, ToppGene, Enrichr, Reactome (via iLINCS)
connectivity (signature similarity) analysis hypoxia signature vs LINCS L1000 consensus gene knockdown signatures (CGS) none (computational comparison) number/significance of connected knockdown signatures iLINCS
Key results
  • MCF10A shows more differentially expressed genes than MDA-MB-231 in both hypoxia and hypoxia+PP242 comparisons.
  • With 2 samples per group, 2-fold change, α=0.01, statistical power is below 0.55 in all four comparisons. power <0.55
  • Four replicates per group needed to achieve ~80% power to detect a 2-fold expression change. power=0.87 (MCF10A hypoxia vs normoxia, n=4)
  • 3,727 LINCS consensus gene knockdown signatures significantly connected to the uploaded hypoxia signature. n=3,727, p<0.05
  • Top 10 GO categories (response to hypoxia, angiogenesis, oxidation-reduction process, etc.) significantly enriched and common to both cell lines. FDR<0.05
  • HIF-1-alpha transcription factor network identified as activated in both cell lines via ToppGene.
Key statistics
  • count >6,500 datasets (number of uniformly processed GEO RNA-seq datasets in GREIN)
  • count >400,000 samples (number of samples across processed GREIN datasets)
  • count 32 samples (GSE104193 case-study dataset (2 cell lines x 2 O2 levels x treatment x mRNA fraction, 2 replicates))
  • pvalue α = 0.01 (significance threshold used for power analysis of 2-fold expression change)
  • other power = 0.52, 0.74, 0.87 (MCF10A hypoxia vs normoxia at n=2,3,4) (Table 2 statistical power by sample size)
  • other power = 0.28, 0.44, 0.59 (MCF10A hypoxia+PP242 vs normoxia at n=2,3,4) (Table 2 statistical power by sample size)
  • count 3,727 signatures (LINCS consensus gene knockdown signatures significantly connected (pValue<0.05) to hypoxia signature)
  • other power cutoff 0.7, FDR cutoff 0.01 (gene selection criteria for enrichment signature submission)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a software/tool paper presenting GREIN, a web platform for re-analyzing GEO RNA-seq data, with statistical methods illustrated through a re-analysis use case of dataset GSE104193. The reported analysis combines exploratory/unsupervised visualization (correlation heatmaps, hierarchical clustering, PCA, t-SNE), gene-wise statistical power and detectability analysis (using biological coefficient of variation, BCOV, and assumed fold-change/alpha), and differential gene expression analysis fit via a generalized linear model that can adjust for covariates or batch effects, with results ranked and thresholded by false discovery rate (FDR). Downstream gene-set enrichment and L1000 connectivity analyses are summarized by significance thresholds (FDR < 0.05; pValue < 0.05).

Replicationbiological Sample sizeUse-case dataset has 32 samples, two biological replicates per combination of cell line, oxygen level, treatment, and mRNA fraction; power analysis explicitly evaluates 2/3/4 samples per group GroupsHypoxia and hypoxia+PP242 vs normoxia/control, per cell line (MCF10A, MDA-MB-231) Pairingunpaired Randomization/blindingna Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFalse discovery rate (FDR) thresholding/ranking
Statistical tests used
Test Applied to n Assumptions
Differential gene expression analysis via a generalized linear model adjusting for covariates/batch effects (treating 'replicate' as a covariate) Hypoxia and hypoxia+PP242 vs control signatures, for each cell line separately (Fig. 4) two biological replicates per group (32 samples total across conditions) not stated
Gene-wise statistical power / detectability analysis based on BCOV, assumed fold change, and alpha (edgeR-style power model) Power curves and detectability plots; Table 2; Fig. 3 2, 3, and 4 samples per group scenarios; α = 0.01; minimal fold change 2 stated
Hierarchical clustering on Pearson correlation of top 500 most variable genes (median absolute deviation) Exploratory analysis (Fig. 2B) na
Principal component analysis (PCA) and t-SNE dimensionality reduction Exploratory analysis (Fig. 2C, 2D) na
Gene-set / pathway enrichment significance (reported as FDR < 0.05) via DAVID, ToppGene, Enrichr, Reactome Functional characterization of hypoxia signatures (Fig. 5, Supplementary Tables) na
Signature connectivity analysis with LINCS L1000 consensus knockdown signatures (reported as pValue < 0.05) Connectivity analysis of uploaded hypoxia signature na
Approaches that could also have been used
  • Differential expression was fit with a generalized linear model and results were ranked/thresholded by FDR.
    Could also: The specific FDR procedure (e.g., Benjamini-Hochberg) and the underlying count-model package (e.g., edgeR, DESeq2, limma-voom) could also be named explicitly alongside the GLM. — Naming the exact multiplicity method and count framework would make the multiplicity scope and dispersion estimation fully transparent and directly reproducible by readers.
  • The use case relied on two biological replicates per group, which the authors' own power analysis shows yields power below ~0.55.
    Could also: Reporting confidence intervals around fold-change estimates, or framing DE counts with the accompanying power/detectability context (as the paper partly does via NDE&DT genes), is also a common way to convey uncertainty at small n. — Interval estimates would convey the precision of effect sizes directly, complementing the power-based detectability framing already used.
  • Exploratory clustering used Pearson correlation on the top 500 genes by median absolute deviation.
    Could also: Spearman correlation or distance metrics on variance-stabilized counts could also be used for clustering RNA-seq profiles. — Rank-based or variance-stabilized approaches can be less sensitive to a few highly expressed genes and are sometimes preferred for count data; reporting both can show robustness of the cluster structure.
  • Enrichment and connectivity significance were reported as threshold statements (FDR < 0.05; pValue < 0.05).
    Could also: Reporting exact adjusted p-values and effect/enrichment scores in the main text, in addition to thresholds, is also standard. — Exact values let readers gauge the strength and ranking of associations rather than only pass/fail status.
  • Power analysis was based largely on single-gene and gene-wise BCOV estimates with fixed fold-change and alpha inputs.
    Could also: A simulation-based or empirically-resampled power analysis across a range of fold changes and dispersions could also be presented. — Scenario sweeps and simulation can characterize how power varies across the realistic effect-size/dispersion space, complementing the fixed-parameter estimates.
Software: R (GREIN/GREP2 implemented in R; released as R package) · Shiny (web framework) · Salmon (read mapping/quantification) · iLINCS (enrichment/connectivity) · MetaSRA (ontological annotation)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
213
Impact: very high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (4)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE104193 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31110304 (GREIN)

Paper: GREIN: An Interactive Web Platform for Re-analyzing GEO RNA-seq Data. Mahi et al., Sci Rep 2019. PMID 31110304 · PMCID PMC6527554 · DOI 10.1038/s41598-019-43935-8. Code: https://github.com/uc-bd2k/grein (web app). Back-end pipeline = GREP2 (CRAN/uc-bd2k/GREP2), the actual computational pipeline. Data: GSE104193 (Sesé et al. TNBC hypoxia/mTOR RNA-seq) → SRA SRP118788 / PRJNA412005, 32 paired-end samples.

GREP2 back-end pipeline (from Methods + repo R/ source)

  1. GEO/SRA metadata via GEOquery; download SRA → FASTQ (ascp + SRA toolkit).
  2. FastQC + optional Trimmomatic adapter trimming.
  3. Salmon quant against Ensembl release-91 transcriptome (build_index.R: cDNA + ncRNA concatenated, k-mer 31, plain index no decoy; run_salmon.R: salmon quant -l A with default options).
  4. tximport → gene level, countsFromAbundance="lengthScaledTPM", tx2gene from EnsDb.Hsapiens.v86, summarizeToGene (run_tximport.R).
  5. MultiQC report. Downstream analytics (paper Results): power analysis (RNASeqPower + edgeR common/tagwise dispersion), differential expression (edgeR TMM + NB GLM/QL), visualisations.

In scope (pipeline-derived, clearly specified, low-hanging)

  • Table 2 — Statistical power analysis for the GSE104193 use case. Four comparisons (MCF10A & MDA-MB-231, each Hypoxia-vs-normoxia and Hypoxia+PP242-vs-normoxia, Total-mRNA fraction, n=2/group). Reported per row: Average sequencing depth (M), Common BCOV, Power at 2 / 3 / 4 samples. Every value is derived deterministically from the count matrix → edgeR common dispersion → RNASeqPower::rnapower(effect=2, alpha=0.01). This is the cleanest checkable numeric claim in the paper. → Reproduce by running GREP2 (Salmon r91 → tximport lengthScaledTPM) on the 16 Total-fraction GSE104193 samples on «our HPC», then computing the power table with the paper's stated method.

Out of scope / not attempted (with reason)

  • >6,500 already-processed datasets (abstract): infrastructure/scale claim, not a single reproducible pipeline run — non_pipeline for our purposes.
  • Figures (PCA, t-SNE, correlation, heatmaps, MA, detectability plots): qualitative/visual, no pinnable numeric value → low audit value; skipped (the hard/qualitative ~20%).
  • Differential-expression gene lists: paper reports no specific DEG counts/values to pin for the use case (visual only) → no_expected_result for a 1:1 number.
  • Web-app/UI functionality, iLINCS connectivity: interactive, out of scope.

Caveats affecting the 1:1

  • Salmon version is not stated in the paper; GREP2 says "default options". We pin our Salmon version in environment.lock; default-option behaviour changed across Salmon 0.x→1.x (selective alignment), which can shift counts a few %. Depth and common-BCOV are robust summary statistics, so power values should track closely.
  • We download FASTQ from ENA (identical reads) rather than via Aspera+sra-toolkit — content-identical.
  • tx2gene uses EnsDb.Hsapiens.v86 (per GREP2 source) while the index is Ensembl r91; GREP2 itself does this, so we match GREP2 exactly.
Figures / tables: Table
T2_MCF10A_HvN_depth
Reported
41.72
Reproduced
42.68
within tolerance
T2_MCF10A_HvN_bcov
Reported
0.21
Reproduced
0.23
within tolerance
T2_MCF10A_HvN_power234
Reported
0.52/0.74/0.87
Reproduced
0.46/0.68/0.83
partial
T2_MCF10A_HPPvN_depth
Reported
38.18
Reproduced
38.93
within tolerance
T2_MCF10A_HPPvN_bcov
Reported
0.31
Reproduced
0.32
within tolerance
T2_MCF10A_HPPvN_power234
Reported
0.28/0.44/0.59
Reproduced
0.26/0.42/0.57
within tolerance
T2_MDA_HvN_depth
Reported
27.86
Reproduced
27.70
within tolerance
T2_MDA_HvN_bcov
Reported
0.24
Reproduced
0.19
partial
T2_MDA_HvN_power234
Reported
0.38/0.73/0.86
Reproduced
0.51/0.73/0.87
partial
T2_MDA_HPPvN_depth
Reported
29.52
Reproduced
29.27
within tolerance
T2_MDA_HPPvN_bcov
Reported
0.19
Reproduced
0.16
within tolerance
T2_MDA_HPPvN_power234
Reported
0.51/0.73/0.86
Reproduced
0.61/0.82/0.93
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 73/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is an incomplete reproduction, not a discrepancy: the GREP2→Salmon→tximport→GREIN power pipeline was re-implemented 1:1 from the authors' own repos and launched on the cluster, but the 16-sample quant array was only 3/16 done at finalize, so all Table 2 reproduced values (depth/BCOV/power for 4 comparisons) remain TBD. No reported-vs-reproduced comparison exists yet, so severity and core-claim confirmation are unknowable — hence yellow, not green or red. The remaining risk sits on our side (operational cut-off plus self-chosen Salmon 2.0.0 / sample-to-cell-line assignment), with no sign the values are non-derivable or that the authors' reporting is at fault. The agent appropriately refused to assert a grade without numbers, which avoids a fabricated verdict.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

287.8 k
tokens (I/O) · 24.5 M incl. cache
106 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.