Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Transcriptomic data meta-analysis reveals common and injury model specific gene expression changes in the regenerating zebrafish heart.

Sci Rep · 2023
L2 90/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. The repo ships the 36 STAR ReadsPerGene.out.tab per-sample count files + design tables + colData, so the downstream DESeq2 analysis is reproducible without re-running STAR on raw FASTQ (the heavy 20%). Implemented the authors' documented pipeline (ComBat_seq batch=Platform / group=Condition -> DESeq2 ~Condition betaPrior=FALSE -> lfcShrink ashr -> DEG padj<0.05 & |log2FC|>1) on «our HPC» SLURM «job». All four Fig 2A DEG counts reproduced BIT-FOR-BIT: Ablation 6858, Resection(Amputation) 5304, Uninjured 3846, Cryoinjury 853. The 148-gene common core: the zebrafish-level intersection of the three injury-model DEG sets is 448 genes; the reported 148 is that set AFTER a 1:1 zebrafish->mouse ortholog conversion + curation (biomaRt, Ensembl-version-sensitive) which I deliberately did NOT chase (80/20). Its upstream inputs (the 3 injury DEG sets) are EXACT, and the repo ships a 148-row core table consistent with the claim, so no fabrication signal. NOT attempted: STAR re-mapping of 36 raw FASTQ, the GO:BP/EnrichmentMap/AutoAnnotate network figures (Fig 4), the Shiny app, and the mouse-ortholog conversion step.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-15 ⛓ 8a4a58335415
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
👤 1 human curator(s) · Level L2 2026-06-15
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Publicly available RNA-seq datasets from three zebrafish heart injury models (ventricular resection, cryoinjury, and genetic cardiomyocyte ablation) can be meta-analyzed at a common regeneration timepoint (7 dpi) to identify a shared core regeneration gene expression signature as well as injury-model-specific signatures.

Core claims
  • Batch correction using sequencing platform as the correcting variable (via Combat-Seq) removes technical variability so that samples cluster by injury condition rather than dataset origin. method
  • The three injury models (resection, cryoinjury, genetic ablation) share a common core set of differentially expressed genes involved in cell proliferation, Wnt signaling, and genes enriched in fibroblasts. finding
  • Resection and genetic ablation show strong injury-specific gene expression signatures, while cryoinjury shows a much weaker injury-specific signature. finding
  • Genetic ablation vs sham yields the greatest number of DEGs and the greatest number of unique enriched GO Biological Process terms among the injury comparisons. finding
  • No biological processes were found to be uniquely enriched in the cryoinjury vs sham comparison. finding
  • The core regeneration gene set is enriched for cell division/proliferation processes and cartilage/bone ossification-associated processes linked to fibroblast genes (e.g. Fn1, Lox, Cthrc1, Prrx1). finding
  • A user-friendly web interface (Shiny dashboard) is provided to browse injury-specific and core regeneration gene expression signatures. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq zebrafish ventricle, cryoinjury vs sham (GSE100892) cryoinjury differential gene expression NextSeq 500
bulk RNA-seq zebrafish heart, cryoinjury vs sham (GSE112452) cryoinjury differential gene expression BGISEQ-500
bulk RNA-seq zebrafish ventricle, resection vs sham (GSE129499, GSE157170) ventricular resection differential gene expression HiSeq X Ten / NovaSeq 6000
bulk RNA-seq zebrafish ventricle, genetic ablation (DTA/NTR) vs uninjured (GSE146859, GSE75894) genetic ablation of cardiomyocytes differential gene expression BGISEQ-500 / Genome Analyzer II
bulk RNA-seq zebrafish ventricle, uninjured vs sham (GSE144831) none (sham vs uninjured control comparison) differential gene expression HiSeq 2000
GO:BP overrepresentation analysis mouse orthologs of zebrafish DEGs none enriched Gene Ontology Biological Process terms clusterProfiler
network clustering of enriched GO terms mouse ortholog GO:BP gene sets none clustered/annotated biological process networks Cytoscape (EnrichmentMap, AutoAnnotate)
cell-type enrichment analysis core regeneration gene set (mouse orthologs) none enriched cell-type populations Enrichr (PanglaoDB)
Key results
  • Ablation vs sham comparison yielded the greatest number of DEGs among injury comparisons. n=6858
  • Resection vs sham yielded the second highest number of DEGs. n=5304
  • Uninjured vs sham comparison unexpectedly yielded more DEGs than cryoinjury vs sham. n=3846 vs n=853
  • Cryoinjury vs sham showed the fewest DEGs of the three injury comparisons. n=853
  • Unique DEGs specific to each injury: ablation highest, resection close second, cryoinjury far lowest (even lower than uninjured vs sham). ablation n=2526; resection n=2448; cryoinjury n=84; uninjured n=654
  • Unique enriched GO:BP terms per injury model: ablation highest, resection second, uninjured lowest; none unique to cryoinjury. ablation 708; resection 532; uninjured 32; cryoinjury 0
  • Core regeneration DEG set (mouse orthologs common across all injury models) predominantly upregulated, with a minority downregulated or showing mixed direction. n=148 total (133 up, 5 down, 10 mixed)
  • Core regeneration cluster includes cell division/proliferation genes (e.g. Aurkb) and cartilage/bone ossification-linked fibroblast genes (Fn1, Lox, Cthrc1, Prrx1).
Key statistics
  • count 36 samples across 7 datasets (total RNA-seq samples used in meta-analysis)
  • count n=6858 DEGs (ablation vs sham DEGs)
  • count n=5304 DEGs (resection vs sham DEGs)
  • count n=3846 DEGs (uninjured vs sham DEGs)
  • count n=853 DEGs (cryoinjury vs sham DEGs)
  • count n=2526 unique DEGs (genetic ablation-specific DEGs)
  • count n=148 common DEGs (mouse orthologs) (core regeneration gene set shared across all injury models (133 up, 5 down, 10 mixed))
  • other |log2FoldChange| > 1 and adjusted p < 0.05 (DEG significance threshold used in DESeq2 analysis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study is a meta-analysis of seven publicly available RNA-seq datasets (36 samples total) from zebrafish hearts at 7 days post-injury across three injury models (ventricular resection, cryoinjury, genetic ablation) and controls. Batch effects attributable to sequencing platform were corrected with Combat-Seq, after which DESeq2 was used for pairwise differential gene expression analysis (each injury model vs. its respective sham or uninjured control) with thresholds of |log2FC| > 1 and adjusted p < 0.05. Downstream functional interpretation relied on GO Biological Process overrepresentation analysis via clusterProfiler, with network clustering visualized in Cytoscape using EnrichmentMap and AutoAnnotate.

Replicationbiological Sample sizeSample sizes per group stated in Table 1: n = 2–4 biological replicates per condition, where each replicate is itself a pool of 1–10 individual hearts or ventricles; no formal power calculation reported GroupsFour pairwise comparisons: resection vs. sham (2 datasets), cryoinjury vs. sham (2 datasets), ablation vs. sham/uninjured (2 datasets), uninjured vs. sham (1 dataset) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionNot explicitly named; DESeq2 default is Benjamini-Hochberg FDR; stated only as 'adjusted p value < 0.05'
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test (negative-binomial GLM) for differential expression Pairwise comparisons: resection vs. sham, ablation vs. sham, cryoinjury vs. sham, uninjured vs. sham n = 2–4 biological replicates per group (each a pool of 1–10 hearts/ventricles) as stated in Table 1 not stated
edgeR likelihood ratio test for differential expression Same pairwise comparisons as DESeq2; used for cross-validation only; DESeq2 results ultimately selected Same as DESeq2 analysis not stated
Hypergeometric / Fisher's exact test (clusterProfiler GO:BP overrepresentation) GO Biological Process enrichment of unique and shared DEG sets, using mouse orthologs as input Gene counts per DEG set (e.g., n = 148 core shared DEGs; up to 708 unique to ablation); background universe not explicitly stated not stated
Principal Components Analysis (PCA) Visualization of batch structure before and after Combat-Seq correction, and after DESeq2 normalization All 36 samples na
Approaches that could also have been used
  • Batch correction was performed using Combat-Seq, treating sequencing platform as the sole batch variable after comparing several candidate variables
    Could also: RUVSeq (Remove Unwanted Variation) or surrogate variable analysis (SVA) could also have been applied, or limma's removeBatchEffect on log-CPM values — RUVSeq and SVA estimate latent confounders from the data itself without requiring a pre-specified batch label, which can be valuable when multiple technical variables are correlated and hard to rank; reporting which variables were tested and how PCA separation changed with each would further document the choice
  • Differential expression results from DESeq2 and edgeR were compared and DESeq2 was selected because it produced a more restrictive (smaller) DEG list
    Could also: A consensus or intersection strategy (retaining only genes called DEG by both tools) could also have been used — An intersection approach explicitly leverages the independent modeling assumptions of each tool to reduce false positives, and the rationale for preferring a smaller list could be made more formally explicit; alternatively, reporting the overlap proportion quantifies concordance rather than leaving it descriptive
  • Each injury model was analyzed via independent pairwise DESeq2 comparisons against its own sham or uninjured control
    Could also: A single multi-group DESeq2 or edgeR model with study-of-origin as a covariate, followed by contrasts, could also have been used — A unified model propagates a common dispersion estimate across all groups and formally accounts for dataset-of-origin as a covariate rather than relying solely on upstream batch correction; this can increase power when group sizes are small (n = 2–4)
  • DEG sets were defined using hard thresholds (|log2FC| > 1, adjusted p < 0.05) and then subjected to overrepresentation analysis
    Could also: Gene Set Enrichment Analysis (GSEA / fgsea) on the full ranked gene list could also have been applied — Rank-based enrichment methods use the entire expression continuum rather than a binary DEG/non-DEG split, which avoids sensitivity to threshold choice and can detect coordinated pathway shifts even when individual genes fall just below cutoffs
  • GO enrichment was performed on mouse orthologs of zebrafish DEGs because zebrafish GO annotations were described as sparse
    Could also: Direct zebrafish GO enrichment using ZFIN or Ensembl zebrafish annotations, or a complementary KEGG/Reactome analysis on zebrafish gene IDs, could also have been applied in parallel — Ortholog conversion introduces uncertainty (many-to-many mappings, genes without orthologs) that is not quantified; reporting how many DEGs were successfully converted and testing sensitivity with direct zebrafish annotations would contextualize how much biological information the conversion step retains
  • Dispersion of gene expression across biological replicates is not reported; results are summarized as DEG counts and volcano plots
    Could also: Reporting normalized count distributions, coefficient of variation across replicates, or MA plots per dataset could also accompany the main results — With n = 2–4 replicates per group (some of which are pools), visualizing per-gene or per-sample dispersion would help readers assess the reliability of DEG calls and the comparability of datasets that differ in pool size (1–10 hearts per sample)
Software: DESeq2 · edgeR · Combat-Seq (sva package) · clusterProfiler · Cytoscape / EnrichmentMap / AutoAnnotate · Enrichr

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
17
Impact: medium
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 81/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE75894 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

  • Fig 2A per-injury-model DEG counts vs Sham: Ablation 6858, Resection 5304, Uninjured 3846, Cryoinjury 853 (DESeq2 on ComBat_seq-corrected shipped STAR counts).
  • Fig 4 common core regeneration signature = 148 mouse-ortholog genes. Result: C1-C4 reproduced EXACT; C5 partial (zebrafish core = 448; ortholog->148 not chased). See ..«path», ..«path», ../AUDIT.md.
Figures / tables: Fig 2AFig 4ATable
C1
Reported
6858
Reproduced
6858
exact
C2
Reported
5304
Reproduced
5304
exact
C3
Reported
3846
Reproduced
3846
exact
C4
Reported
853
Reproduced
853
exact
C5
Reported
148
Reproduced
zf 3-injury intersection = 448; mouse-ortholog conversion to 148 not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4

This is a strong reproduction: four of five pinned claims (Fig 2A DEG counts) match bit-for-bit from the authors' shipped STAR counts and documented ComBat_seq→DESeq2→ashr pipeline. The single open item, the 148-gene mouse-ortholog core (Fig 4), is not a discrepancy — its zebrafish upstream intersection (448) is exact and the repo ships a 148-row table consistent with the claim; only the deterministic, Ensembl-version-sensitive ortholog conversion was deliberately left unrun (our 80/20 choice). No authors'-side or data-availability defect and no fabrication signal; the only caveats sit on our methodology (q3/q4 yellow).

👤 Schlein Lab (curation team) L2 100/100
🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

253.2 k
tokens (I/O) · 22 M incl. cache
62 min
runtime · 0.04 CPU-h
2.5 GB
peak RAM
1
HPC jobs
hummel
machine