Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

ASTRA: a comprehensive resource of stress-induced transcriptional activity in human cell lines.

Nucleic Acids Res · 2026
L1 73/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
73/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 664 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH to reproduce from the paper's own shipped data, with caveats. The GitHub repo is only the Streamlit web-app frontend (no pipeline code), but the Zenodo deposit (10.5281/zenodo.15885686) ships the pipeline OUTPUTS (TPM matrix, edgeR DE table, biomaRt annotations, GEO metadata), so per BRIEF rule P16 we reproduced against those. RESULT = 1:1 on the headline curation counts: 669 samples, 59 studies, 57 cell lines, 46 publications, and the described coding/lncRNA/pseudogene biotype coverage all reproduce EXACTLY from the shipped metadata. The core DE output is validated by internal consistency: recomputing per-gene fold-change directly from the shipped TPM matrix reproduces the SIGN of the shipped edgeR log2FC for ~95-100% of DE-significant genes across 8 GSEs spanning all 4 stress types (Pearson 0.23-0.85; lower Pearson expected because edgeR uses counts+TMM+NB-GLM vs our crude TPM-ratio) - consistent with a genuinely-derived, non-fabricated DE table. DIVERGENCES (honest): (a) the reported '86 differential-expression comparison sets' is NOT recoverable from the deposit - the DE file is keyed only by gse_id with no comparison-set identifier, covers only 45 of 59 studies (14 missing entirely), and yields ~81 stacked sets, while datasets.tsv lists 185 control/treatment contrasts; (b) the per-stress split 20/17/16/33 matches neither study-level (15/9/15/20) nor contrast-level (39/52/27/67) counts; (c) 27 genotypes shipped vs 26 reported. NOT ATTEMPTED (the >20% tail): the full STAR+RSEM+edgeR re-run from raw FASTQ - terabytes of reads, GRCh38 index build, and incompletely-documented per-comparison sample grouping; the internal-consistency check stands in as a lighter validation of the DE step. All raw/large data kept on «infra»; only small derived results on «host».

💻 Code ↗ 🗄 Data: 10.5281/zenodo.15885686

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 73
    assessed: 2026-06-15 ⛓ 007f030898a4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Knowledge of cellular stress response pathways is fragmented; ASTRA was developed to provide a curated, standardized, open-access database of human cell line transcriptomic responses to defined molecular stressors, enabling systematic discovery of conserved and stress-specific coding and noncoding stress-responsive genes and RNA biomarkers.

Core claims
  • ASTRA is a curated, open-access database cataloging standardized bulk RNA-seq transcriptomic responses to stress in human cell lines, organized into four stress categories (oxidative, hypoxia, heat shock, DNA damage). resource
  • A harmonized, standardized computational pipeline (QC, alignment, normalization, differential expression) ensures comparability across samples and studies. method
  • ASTRA profiles both coding (mRNA) and long noncoding RNA/pseudogene transcriptomes to capture the noncoding genome's contribution to the stress response. resource
  • Each dataset pairs stressed and matched control samples annotated with detailed metadata (cell line identity, tissue origin, cell-state, treatment parameters, sequencing protocols). resource
  • The web interface enables gene-level expression queries, differential expression analyses, and visualization of transcriptional dynamics across datasets, conditions, and cell-type groups. method
  • ASTRA supports identification of novel stress-responsive genes and RNA-mediated regulatory mechanisms for functional genomics, systems biology, and translational research. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (oxidative stress, H2O2 50–700 μM) human cell lines H2O2 treatment mRNA and lncRNA expression (TPM) and differential expression
bulk RNA-seq (hypoxia, 1%–3% oxygen) human cell lines low oxygen exposure mRNA and lncRNA expression (TPM) and differential expression
bulk RNA-seq (heat shock, 39°C–45°C) human cell lines increased temperature mRNA and lncRNA expression (TPM) and differential expression
bulk RNA-seq (DNA damage, UV 5–60 J/m2) human cell lines UV radiation mRNA and lncRNA expression (TPM) and differential expression
Key results
  • 59 GEO studies were curated and used to finalize 86 datasets (comparison sets) 59 studies, 86 datasets
  • Datasets distributed as 33 oxidative stress (H2O2), 20 heat shock, 17 UV, and 16 hypoxia studies 33/20/17/16
  • 669 RNA-seq samples retrieved and systematically processed through the standardized pipeline 669 samples
  • Datasets encompass 57 unique cell lines, 26 genotypes, 40 time points, and 27 dosage levels 57 cell lines; 26 genotypes; 40 time points; 27 doses
  • Cell lines grouped into three cell-states: non-diseased reference, cancer, and organoid/3D 3 groups
Key statistics
  • count 59 (number of studies cataloged in ASTRA)
  • count 86 (within-study differential expression comparison sets (datasets))
  • count 669 (individual RNA-seq samples processed)
  • count 33 (oxidative stress (H2O2) studies)
  • count 20 (heat shock response studies)
  • count 17 (UV response studies)
  • count 16 (hypoxia studies)
  • count 57 (unique cell lines cataloged by tissue of origin)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

ASTRA is a database and computational pipeline paper that integrates and reprocesses 669 publicly available human RNA-seq samples from 59 GEO studies into 86 stressed-vs-control comparison sets across four stress categories. Differential expression between stressed and matched control samples was computed with edgeR; transcript abundance was quantified as TPM via RSEM. Pairwise co-expression relationships in the interactive interface are summarized with both Pearson's r and Spearman's ρ. The paper reports gene-level fold-change and average p-value summaries across datasets as a resource tool rather than as a single confirmatory experiment.

Replicationmixed Sample size669 total RNA-seq samples from 59 studies forming 86 treated-vs-control comparison sets; individual dataset sample sizes are not reported per comparison GroupsStressed (treated) human cell line samples vs matched untreated controls across oxidative stress, hypoxia, heat shock, and UV/DNA-damage categories Pairingunclear Randomization/blindingnot stated Dispersionunclear Effect sizesyes Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
edgeR differential expression (negative binomial GLM) Treated (stressed) vs untreated (control) comparisons within each of the 86 within-study comparison sets 669 total samples across 86 comparison sets; per-dataset sample sizes not individually reported in the text not stated
Pearson's r correlation with P-value Pairwise gene co-expression within a stress type (Gene Search / Within Stress-Type tab) not stated
Spearman's ρ correlation with P-value Pairwise gene co-expression within a stress type (Gene Search / Within Stress-Type tab) not stated
Approaches that could also have been used
  • Differential expression was computed with edgeR across all 86 comparison sets
    Could also: DESeq2 (negative binomial GLM with a different shrinkage estimator) or limma-voom (variance-stabilized linear model) could also be applied to RNA-seq count data — DESeq2 and limma-voom are widely used alternatives; running more than one tool and comparing the overlap of DE calls is common practice in large-scale reanalysis pipelines to assess robustness across methods
  • Transcript abundance was normalized and reported as TPM via alignment-based RSEM quantification
    Could also: Pseudo-alignment tools such as kallisto or Salmon could also produce TPM estimates with substantially faster runtimes at comparable accuracy; TMM-normalized counts (edgeR's native scaling) could also serve as the expression unit — Pseudo-alignment approaches are increasingly common in large-scale pipelines; TMM-normalized counts align directly with edgeR's internal model assumptions and can simplify the transition from quantification to DE analysis
  • Cross-study comparability was addressed through a standardized processing pipeline; no explicit statistical batch-correction step is described
    Could also: Methods such as ComBat-seq, RUVSeq, or limma's removeBatchEffect could also be applied to reduce technical variation attributable to study of origin before pooling expression values for cross-dataset visualization — When aggregating data from many independent labs, platforms, and protocols, explicit batch-effect estimation and correction can reduce study-of-origin confounding in cross-dataset expression comparisons
  • Both Pearson's r and Spearman's ρ are reported for co-expression analysis without designating one as primary
    Could also: Designating Spearman's ρ as the primary metric (with Pearson's as supplementary) is also common; mutual information or distance correlation could additionally capture non-linear dependencies — TPM distributions are typically right-skewed and contain zeros; rank-based Spearman's ρ is robust to these features, and making the primary choice explicit aids reproducibility and interpretation for database users
  • The type of p-value used for DEG filtering (raw vs FDR-adjusted) is not explicitly stated in the text
    Could also: Explicitly applying and documenting Benjamini-Hochberg FDR correction within each comparison set, with the correction scope stated, is also standard practice for RNA-seq DE analysis — Clearly specifying whether thresholds apply to raw or adjusted p-values, and the scope of the correction (per-dataset gene family), makes the DEG filtering criteria unambiguous and reproducible for downstream users of the resource
  • Dataset inclusion was determined by manual curation criteria applied to GEO studies
    Could also: Reporting inter-rater reliability (e.g. Cohen's κ) for manual curation decisions and/or explicit automated QC thresholds (e.g. minimum alignment rate, minimum library size) could also be documented alongside the curation criteria — Quantifying the reproducibility of manual inclusion decisions and making automated quality thresholds explicit helps users understand the consistency of the curation process and the boundaries of the resulting resource
Software: edgeR (R/Bioconductor) · RSEM · STAR aligner · FastQC · Cutadapt · biomaRt (R/Bioconductor) Bioconductor v3.22 · Streamlit (Python) · PostgreSQL

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41263110 (ASTRA)

Paper: ASTRA: a comprehensive resource of stress-induced transcriptional activity in human cell lines. Nucleic Acids Res 2026. PMID 41263110 · PMCID PMC12807686 · DOI 10.1093/nar/gkaf1174.

Resource type: a curated transcriptomics database (web app + data deposit), not a wet-lab study. The "results" are database contents and pipeline-derived differential-expression tables.

Pipeline described in Methods

Raw RNA-seq (GEO/ENA) → FastQC → adapter detection (Minion/Reaper) + Cutadapt → STAR (→ GRCh38) → RSEM (TPM quantification) → edgeR (treated vs untreated DE) → biomaRt/Ensembl BioMart annotation (Bioconductor 3.22). Output: per-comparison DE tables (log2FC, p-value) + a TPM expression matrix, served by a Streamlit + PostgreSQL app.

Code artifact (P16 note)

The GitHub repo Mtsif/astra-db is the web-app frontend (Streamlit + Postgres badges, MIT). It ships no read-processing/DE pipeline code. The pipeline itself (STAR/RSEM/edgeR) is described in Methods but not shipped as runnable code. Per BRIEF rule P16, applying the described standard tools to the paper's own data is an equally valid reproduction.

Shipped data (Zenodo 10.5281/zenodo.15885686, CC-BY-4.0, ~1.5 GB)

  • expression_data.tsv (1.32 GB) — TPM expression matrix (long format)
  • differential_expression.tsv (92 MB) — edgeR DE results: log2FC + p-value per gene × comparison (the core pipeline output)
  • sample_metadata.tsv (2.0 MB) — per-sample metadata
  • study_metadata.tsv (130 KB) / datasets.tsv (28 KB) — study index
  • harmonized_dataset.tsv (77 KB) — harmonized metadata
  • gene_annotations.tsv (7.1 MB) — biomaRt gene annotations

IN SCOPE (reproducible from shipped data — what we attempt)

  1. Headline counts (pipeline/curation-derived, stated in abstract & Fig/Table): 669 samples · 59 studies · 86 DE comparison sets · 57 unique cell lines · 46 publications · stress-type split (20 heat shock / 17 UV / 16 hypoxia / 33 oxidative). Verify directly against the shipped metadata + DE deposit.
  2. DE-table internal consistency / reproduction: the deposit ships both the TPM matrix and the edgeR DE table. Recompute fold-changes from the shipped TPM matrix and correlate with the shipped edgeR log2FC across several comparisons. This reproduces the differential-expression output from the paper's own data and detects whether the DE table is faithfully derived (vs fabricated/inconsistent).

OUT OF SCOPE / the hard 20% (not fully attempted — say why)

  • Full STAR+RSEM+edgeR from raw FASTQ for all 669 samples — terabytes of FASTQ, GRCh38 index build, weeks of alignment. Not feasible and explicitly the >20% tail. We may attempt one comparison end-to-end if a small, clearly-mapped study surfaces, but the headline reproduction stands on the shipped deposit (1 & 2).
  • Web-app/Streamlit deployment, browser-compatibility testing — non-pipeline.
  • Wet-lab / stressor-dosage curation choices — manual, not computational.
Figures / tables: Fig.1
C1
Reported
669 samples
Reproduced
669
exact
C2
Reported
59 studies
Reproduced
59
exact
C3
Reported
57 cell lines
Reproduced
57
exact
C4
Reported
46 publications
Reproduced
46
exact
C5
Reported
86 DE comparison sets
Reproduced
81 est / 45 GSEs in DE / 185 contrasts
did not match
C6
Reported
26 genotypes
Reproduced
27
partial
C7
Reported
stress split 20/17/16/33
Reproduced
studyLvl 15/9/15/20; contrastLvl 39/52/27/67
did not match
C8
Reported
coding+lncRNA+pseudogene annotation
Reproduced
pc 20043/lncRNA 17686/pseudo 9481+
exact
C9
Reported
edgeR DE log2FC (pipeline output)
Reproduced
sign-agreement 0.70-1.00 (median ~0.98); Pearson 0.23-0.85 vs TPM-derived log2FC, n=8 GSEs
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 73/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

The headline resource claims reproduce 1:1 from the authors' own Zenodo deposit — 669 samples, 59 studies, 57 cell lines, 46 publications, and the described coding/lncRNA/pseudogene biotype coverage — and the shipped edgeR DE table is internally consistent with the shipped TPM matrix (sign agreement ~95-100% across 8 GSEs), strong evidence it is genuinely derived. The substantive divergences are on the authors'/deposit side: the reported 86 DE comparison sets and their 20/17/16/33 stress split are not derivable (DE table keyed only by gse_id, covers only 45/59 studies; recomputation gives 81/45/185, none=86), and genotypes are 27 vs 26. These are moderate, well-explained under-documentation issues, not fabrication, and the central conclusion of a comprehensive, genuinely-curated resource holds. Overall yellow: solid reproduction with explainable, documentation-driven deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

138.3 k
tokens (I/O) · 5.9 M incl. cache
16 min
runtime · 0.02 CPU-h
0.5 GB
peak RAM
2
HPC jobs
hummel
machine