Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Synthesis of transcriptomic studies reveals a core response to heat stress in abalone (genus Haliotis).

BMC Genomics · 2025
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH? The analysis repo (roybarkan2020/Abalone-RNAseq-meta-analysis) ships NO runnable code and NO data - only a prose MetaAnalysis_workflow.md, a README, and a PNG; tool versions, per-study sample sheets, the gene-name merge logic and the WGCNA soft-threshold are all under-specified. Re-running the full pipeline (Trimmomatic->Bowtie2/RSEM to H. rufescens GCF_023055435->DESeq2 over ~150 samples) was therefore the heavy, under-specified ~80% and was NOT attempted from raw FASTQ. HOWEVER the paper deposits its pipeline's intermediate outputs in Additional file 1 (Springer ESM xlsx): Table S1 (150-sample metadata), Table S3 (per-gene Up/Down/Non_DEG across the 9 studies), Table S6 (per-sample normalized counts for all 74 genes), Table S4 (GO), Table S8 (STRING clusters). From these the headline computational claims reproduce essentially 1:1: 150 samples/9 studies (exact), the 74-gene core set and its full breakdown 15(>=8)+59(=7) (exact), 15 up-in-all-significant (exact), the 6 named down-in-every-study genes (exact set), 3 STRING clusters / 74 nodes (exact), HSP share >15% (consistent). An INDEPENDENT recomputation of per-study direction from the deposited normalized counts (Table S6) agrees with the reported Up/Down labels (Table S3) 98.1% of the time (523/533) - strong internal consistency, no sign of fabrication in the core result. TWO honest discrepancies flagged for human audit: (C9) the Results TEXT says GO 'Protein folding' has 24 genes but the paper's own Table S4 says 23; (C8) the '50 up in >=6 studies' integer is not cleanly recoverable (49 by DEG label, 53 by normalized mean - definition-dependent). NOT ATTEMPTED: de-novo re-alignment from raw FASTQ (no sample sheets/merge code shipped; deposited counts used instead); STRING 211 edges and g:Profiler/WGCNA re-runs (external-service DB-version dependent); the prepared «our HPC» raw-requant job for PRJNA597237 (reproduction/run.sbatch + deseq.R) was deferred - «our HPC» VPN was down awaiting operator 2FA, and per the 80/20 rule the core claims already reproduce from deposited data. VERDICT: largely 1:1 reproduction of the paper's central meta-analysis result from its deposited intermediate data, with two minor internal-consistency flags.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-15 ⛓ dde355ecbf94
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Across the genus Haliotis, is there a consistent core set of genes that respond to heat stress regardless of species, tissue, or experimental design, identifiable by meta-analysis of publicly available RNA-seq datasets reanalyzed through a standardized pipeline?

Core claims
  • A core set of 74 differentially expressed genes responds to heat stress in at least seven of nine abalone transcriptomic studies, constituting a conserved core heat-stress response across the genus Haliotis. finding
  • The core response is dominated by genes associated with alternative splicing, heat shock proteins (HSPs), the Ubiquitin–Proteasome System (UPS), and other protein folding/processing pathways. mechanism
  • A standardized bioinformatic pipeline (common H. rufescens reference transcriptome, Bowtie2/RSEM/DESeq2) enables cross-study comparison of heterogeneous abalone heat-stress RNA-seq datasets. method
  • The direction of expression change of core genes is remarkably consistent across studies despite differences in species, tissue, and stress intensity. finding
  • HSP genes comprise over 15% of the core response genes. finding
  • The identified core response provides a foundation/resource for conservation, breeding, and aquaculture efforts to enhance thermal tolerance in abalone. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (reanalysis/meta-analysis) H. discus hannai hepatopancreas heat stress (control 20°C → max 30°C, Δ10°C, 24 h) differential gene expression (gene-level read counts) NCBI SRA BioProject PRJNA597237 (Kim et al. 2021)
bulk RNA-seq (reanalysis) H. discus hannai muscle, mantle, gills, blood heat stress (20°C → 30°C, Δ10°C, 1 h) differential gene expression BioProject PRJNA557314 (Kyeong et al. 2020)
bulk RNA-seq (reanalysis) H. discus hannai hemolymph heat stress (17°C → 28°C, Δ11°C, 48 h) differential gene expression CNGBdb CNP0003705 (Zhou Wu et al. 2023)
bulk RNA-seq (reanalysis) H. discus hannai, H. gigantea and their hybrid, gills heat stress (20°C → 30°C, Δ10°C, 2 h) differential gene expression; PCA clustering by treatment and species BioProject PRJNA721743 (Xiao et al. 2021)
bulk RNA-seq (reanalysis) H. discus hannai × H. fulgens, gills heat stress (18°C → 32°C, Δ14°C, 24 h) differential gene expression BioProject PRJNA853707 (Zhang et al. 2022)
bulk RNA-seq (reanalysis, no biological replicates, LFC-based) H. diversicolor hemolymph heat stress (25°C → 31°C, Δ6°C, 96 h) log2 fold change in gene expression (|LFC|>1) BioProject PRJNA481417 (Zhang et al. 2019)
bulk RNA-seq (reanalysis, no biological replicates, LFC-based) H. fulgens pooled gills/mantle/hepatopancreas heat stress (18°C → 32°C, Δ14°C, 12 h) log2 fold change in gene expression BioProject PRJNA453554 (Tripp-Valdez et al. 2019)
bulk RNA-seq (reanalysis) H. laevigata tentacle / H. rufescens × H. corrugata whole organism heat stress (H. laevigata 18.5→21°C, Δ2.5°C, 72 h; hybrid 18→22°C, Δ4°C, 2352 h) differential gene expression BioProjects PRJNA286263 (Shiel et al. 2017) and PRJNA600240 (Tripp-Valdez et al. 2021)
Key results
  • 74 genes were differentially expressed in at least seven of nine studies (core response genes) 74 genes
  • 15 core genes shared among eight studies (padj<0.05) and 59 shared among seven studies 15 + 59 genes
  • 50 of 74 core genes upregulated under heat stress in at least six studies 50 of 74 genes
  • SRSF10, WBP2NL and XBP1 upregulated in all eight studies in which they were identified as DEGs 8 of 8 studies
  • FUS, HNRNPAB, TUBB4B, HNRNPA1, HNRNPLL, HMGB1 downregulated in every study all studies
  • HSPs constituted over 15% of core response genes >15%
  • H. diversicolor study had the fewest HSP DEGs (2 of 11 HSP genes) and lowest number of shared DEGs 2 of 11 HSP genes
  • H. rufescens transcriptome gave the highest average mapping rate as common reference 86%
Key statistics
  • count 74 core response genes (DEGs shared in ≥7 of 9 studies)
  • count 150 samples / 300 FASTQ files (control and heat-stressed samples across nine studies)
  • pvalue adjusted p-value < 0.05 (Benjamini–Hochberg, Wald test) (threshold for differential expression in DESeq2)
  • fold_change |LFC| > 1 (DEG criterion for two studies lacking biological replicates)
  • count 86% (overall highest average mapping rate using H. rufescens reference transcriptome)
  • pvalue padj < 0.01 (significance threshold for GO/KEGG/Reactome term assignment)
  • other >15% (proportion of core response genes that are HSPs)
  • count 15 genes upregulated in all studies where significant (of 50 upregulated core genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This meta-analysis reanalyzed nine publicly available RNA-seq datasets from heat-stress studies on abalone (genus Haliotis) using a standardized bioinformatic pipeline (read QC, trimming, mapping to a single shared reference transcriptome, and count generation). Differential gene expression within each study was assessed using DESeq2's Wald test with Benjamini–Hochberg-adjusted p-values (padj < 0.05), except for two datasets lacking biological replicates where |log2 fold change| > 1 was used as the sole criterion. A 'core response' gene set was defined by a vote-counting approach: genes identified as differentially expressed in at least 7 out of 9 studies. Functional enrichment and co-expression analyses (g:Profiler, STRING, WGCNA) were applied to this core gene set.

Replicationmixed Sample size150 total samples (control and heat-stressed) from 9 studies, yielding 300 paired FASTQ files; per-study sample sizes not explicitly listed in main text (referred to Table S1) GroupsHeat-stressed abalone vs. control (ambient temperature) abalone, across 7 species and 3 hybrids Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini–Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
DESeq2 negative binomial GLM with Wald test Differential gene expression for each of the 7 studies that had biological replicates Varies per study; 150 total samples (control + heat-stressed) across all 9 studies not stated
Log2 fold change threshold (|LFC| > 1, no p-value) Differential gene expression for 2 studies without biological replicates (PRJNA453554 and PRJNA481417) Not stated per study; these 2 studies had no biological replicates na
Principal Component Analysis (PCA) Exploratory analysis of VST-normalized gene count data within individual studies (PRJNA721743, PRJNA557314) to assess effects of species and tissue type Not stated per PCA na
Weighted Gene Co-expression Network Analysis (WGCNA) Co-expression module detection on merged normalized counts from all studies combined, comparing control vs. heat-stressed eigengenes All 150 samples merged not stated
g:Profiler functional enrichment (hypergeometric, BH-corrected) GO, KEGG, and Reactome pathway enrichment of the 74 core response genes 74 core genes as input not stated
STRING protein–protein interaction enrichment (BH-corrected) Functional and physical interaction network of core response genes using H. rufescens proteome 74 core genes as input not stated
Approaches that could also have been used
  • A vote-counting approach was used to define the core gene set: genes had to appear as DEGs in at least 7 of 9 studies.
    Could also: Effect-size-based meta-analysis (e.g., random-effects model on log2 fold changes across studies, or Fisher's combined p-value method) could also synthesize results across studies. — Effect-size meta-analysis weights studies by precision and produces a single pooled effect estimate with a confidence interval, which can be more statistically powerful than vote-counting and explicitly accounts for between-study heterogeneity.
  • All datasets were mapped to a single shared reference transcriptome (H. rufescens) regardless of species, to enable cross-study comparability.
    Could also: Species-specific reference transcriptomes or de novo transcriptome assembly per dataset could also be used, with cross-study comparison done via ortholog mapping (e.g., OrthoFinder or BLASTp reciprocal best hits). — Ortholog-based approaches preserve species-specific annotation accuracy and allow detection of lineage-specific responses that may be missed when reads from divergent species are forced onto a single reference.
  • Two datasets without biological replicates were included using a fold-change-only criterion (|LFC| > 1).
    Could also: These datasets could also be excluded from the primary analysis and treated as a separate validation set, or quasi-likelihood methods (e.g., edgeR's GLM with dispersion borrowed from a related dataset) could provide approximate significance estimates. — Fold-change thresholds alone are sensitive to library depth and baseline expression level; borrowing dispersion from related samples or using these datasets only for directional confirmation preserves the statistical rigor of the main analysis.
  • Variance Stabilizing Transformation (VST) was applied to normalized counts for PCA and visualization.
    Could also: regularized log transformation (rlog, also from DESeq2) or voom-transformed counts (limma) could also be used for this purpose. — rlog is often recommended for small sample sizes as it applies stronger shrinkage for low-count genes, while voom integrates naturally with limma's linear modelling framework if a joint multi-study analysis were desired.
  • WGCNA was applied to all studies merged into a single dataset, treating the combined data as a unified experiment.
    Could also: Consensus WGCNA (multi-set WGCNA) could also be applied, which identifies co-expression modules that are reproducible across individual studies rather than in a single merged matrix. — Consensus WGCNA is specifically designed for multi-dataset settings and explicitly models between-study variation, potentially yielding modules more robust to study-specific confounders such as species and tissue type.
  • Multiplicity correction (BH FDR) was applied independently within each study, with no correction across the vote-counting step itself.
    Could also: A permutation-based approach or resampling procedure could also be used to empirically estimate the expected number of genes appearing in ≥7/9 studies by chance, given the per-study DEG rates. — Empirical null distribution for the vote-counting threshold accounts for dependencies between studies (e.g., shared reference, overlapping biology) and provides a study-level false discovery rate for the core gene list.
Software: FastQC · MultiQC · Trimmomatic · Bowtie2 · RSEM · DESeq2 (R/Bioconductor) · g:Profiler · STRING · NetworkAnalyst · WGCNA (R package) · Cytoscape 3.10.1

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GO:0006457 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0051082 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0140662 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
PRJNA286263 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJNA453554 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJNA481417 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJNA557314 BioProject in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
PRJNA597237 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJNA600240 BioProject in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
PRJNA721743 BioProject in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
PRJNA853707 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Tables Table
C1
Reported
150 samples / 9 studies
Reproduced
150 / 9
exact
C2
Reported
74 core genes (DEG in >=7/9 studies)
Reproduced
74
exact
C3
Reported
15 genes shared among >=8 studies
Reproduced
15 (14 in 8 + 1 in 9)
exact
C4
Reported
59 genes shared among 7 studies
Reproduced
59
exact
C5
Reported
15 genes upregulated in all studies where significant
Reproduced
15
exact
C6
Reported
6 genes down in every study (FUS,HNRNPAB,TUBB4B,HNRNPA1,HNRNPLL,HMGB1)
Reproduced
same 6
exact
C7
Reported
HSPs >15% of core genes
Reproduced
20.3% (broad) / 13.5-14.9% (strict)
within tolerance
C8
Reported
50 genes up in >=6 studies
Reproduced
49 (DEG labels) / 53 (normalized means)
within tolerance
C9
Reported
GO Protein folding GO:0006457 = 24 genes (text)
Reproduced
23 (paper's own Table S4)
did not match
C10
Reported
STRING: 3 clusters / 74 nodes / 211 edges
Reproduced
3 clusters (28/27/19) = 74 nodes; edges not recomputed
partial
C11
Reported
~86% mean cross-species mapping rate
Reproduced
597237=0.91; central per-study ~0.86-0.91
partial
V1
Reported
(internal validation, no paper value)
Reproduced
per-study direction from Table S6 normalized counts vs Table S3 labels = 523/533 = 98.1% agreement
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4

This is largely a 1:1 reproduction of the paper's central meta-analysis result: every headline count (150 samples/9 studies, the 74-gene core set and its 15+59 breakdown, the 15 up-in-all-significant and 6 named down-only genes, 3 STRING clusters/74 nodes) reproduces exactly from the authors' deposited intermediate tables, and an independent recomputation of per-study direction from the normalized counts (Table S6) agrees with the reported labels 523/533 = 98.1% — no fabrication signal. The only deviations are minor and explainable: C9 is the paper's own text-vs-supplement inconsistency (24 in text vs 23 in Table S4), and C8's exact '50' is definition-sensitive (49 by DEG label / 53 by normalized mean). Caveat for scope: the heavy raw-FASTQ pipeline was not re-run, so the reproduction leans on deposited outputs, but the independent direction check corroborates them. Overall: solid, central conclusion fully holds, deviations are trivial and partly on the authors' side.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

178.8 k
tokens (I/O) · 11.6 M incl. cache
20 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.