Synthesis of transcriptomic studies reveals a core response to heat stress in abalone (genus Haliotis).
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH? The analysis repo (roybarkan2020/Abalone-RNAseq-meta-analysis) ships NO runnable code and NO data - only a prose MetaAnalysis_workflow.md, a README, and a PNG; tool versions, per-study sample sheets, the gene-name merge logic and the WGCNA soft-threshold are all under-specified. Re-running the full pipeline (Trimmomatic->Bowtie2/RSEM to H. rufescens GCF_023055435->DESeq2 over ~150 samples) was therefore the heavy, under-specified ~80% and was NOT attempted from raw FASTQ. HOWEVER the paper deposits its pipeline's intermediate outputs in Additional file 1 (Springer ESM xlsx): Table S1 (150-sample metadata), Table S3 (per-gene Up/Down/Non_DEG across the 9 studies), Table S6 (per-sample normalized counts for all 74 genes), Table S4 (GO), Table S8 (STRING clusters). From these the headline computational claims reproduce essentially 1:1: 150 samples/9 studies (exact), the 74-gene core set and its full breakdown 15(>=8)+59(=7) (exact), 15 up-in-all-significant (exact), the 6 named down-in-every-study genes (exact set), 3 STRING clusters / 74 nodes (exact), HSP share >15% (consistent). An INDEPENDENT recomputation of per-study direction from the deposited normalized counts (Table S6) agrees with the reported Up/Down labels (Table S3) 98.1% of the time (523/533) - strong internal consistency, no sign of fabrication in the core result. TWO honest discrepancies flagged for human audit: (C9) the Results TEXT says GO 'Protein folding' has 24 genes but the paper's own Table S4 says 23; (C8) the '50 up in >=6 studies' integer is not cleanly recoverable (49 by DEG label, 53 by normalized mean - definition-dependent). NOT ATTEMPTED: de-novo re-alignment from raw FASTQ (no sample sheets/merge code shipped; deposited counts used instead); STRING 211 edges and g:Profiler/WGCNA re-runs (external-service DB-version dependent); the prepared «our HPC» raw-requant job for PRJNA597237 (reproduction/run.sbatch + deseq.R) was deferred - «our HPC» VPN was down awaiting operator 2FA, and per the 80/20 rule the core claims already reproduce from deposited data. VERDICT: largely 1:1 reproduction of the paper's central meta-analysis result from its deposited intermediate data, with two minor internal-consistency flags.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 80assessed: 2026-06-15 ⛓ dde355ecbf94
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusAcross the genus Haliotis, is there a consistent core set of genes that respond to heat stress regardless of species, tissue, or experimental design, identifiable by meta-analysis of publicly available RNA-seq datasets reanalyzed through a standardized pipeline?
- ★ A core set of 74 differentially expressed genes responds to heat stress in at least seven of nine abalone transcriptomic studies, constituting a conserved core heat-stress response across the genus Haliotis. finding
- ★ The core response is dominated by genes associated with alternative splicing, heat shock proteins (HSPs), the Ubiquitin–Proteasome System (UPS), and other protein folding/processing pathways. mechanism
- ★ A standardized bioinformatic pipeline (common H. rufescens reference transcriptome, Bowtie2/RSEM/DESeq2) enables cross-study comparison of heterogeneous abalone heat-stress RNA-seq datasets. method
- ★ The direction of expression change of core genes is remarkably consistent across studies despite differences in species, tissue, and stress intensity. finding
- HSP genes comprise over 15% of the core response genes. finding
- The identified core response provides a foundation/resource for conservation, breeding, and aquaculture efforts to enhance thermal tolerance in abalone. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (reanalysis/meta-analysis) | H. discus hannai hepatopancreas | heat stress (control 20°C → max 30°C, Δ10°C, 24 h) | differential gene expression (gene-level read counts) | NCBI SRA BioProject PRJNA597237 (Kim et al. 2021) |
| bulk RNA-seq (reanalysis) | H. discus hannai muscle, mantle, gills, blood | heat stress (20°C → 30°C, Δ10°C, 1 h) | differential gene expression | BioProject PRJNA557314 (Kyeong et al. 2020) |
| bulk RNA-seq (reanalysis) | H. discus hannai hemolymph | heat stress (17°C → 28°C, Δ11°C, 48 h) | differential gene expression | CNGBdb CNP0003705 (Zhou Wu et al. 2023) |
| bulk RNA-seq (reanalysis) | H. discus hannai, H. gigantea and their hybrid, gills | heat stress (20°C → 30°C, Δ10°C, 2 h) | differential gene expression; PCA clustering by treatment and species | BioProject PRJNA721743 (Xiao et al. 2021) |
| bulk RNA-seq (reanalysis) | H. discus hannai × H. fulgens, gills | heat stress (18°C → 32°C, Δ14°C, 24 h) | differential gene expression | BioProject PRJNA853707 (Zhang et al. 2022) |
| bulk RNA-seq (reanalysis, no biological replicates, LFC-based) | H. diversicolor hemolymph | heat stress (25°C → 31°C, Δ6°C, 96 h) | log2 fold change in gene expression (|LFC|>1) | BioProject PRJNA481417 (Zhang et al. 2019) |
| bulk RNA-seq (reanalysis, no biological replicates, LFC-based) | H. fulgens pooled gills/mantle/hepatopancreas | heat stress (18°C → 32°C, Δ14°C, 12 h) | log2 fold change in gene expression | BioProject PRJNA453554 (Tripp-Valdez et al. 2019) |
| bulk RNA-seq (reanalysis) | H. laevigata tentacle / H. rufescens × H. corrugata whole organism | heat stress (H. laevigata 18.5→21°C, Δ2.5°C, 72 h; hybrid 18→22°C, Δ4°C, 2352 h) | differential gene expression | BioProjects PRJNA286263 (Shiel et al. 2017) and PRJNA600240 (Tripp-Valdez et al. 2021) |
- – 74 genes were differentially expressed in at least seven of nine studies (core response genes) 74 genes
- – 15 core genes shared among eight studies (padj<0.05) and 59 shared among seven studies 15 + 59 genes
- ▲ 50 of 74 core genes upregulated under heat stress in at least six studies 50 of 74 genes
- ▲ SRSF10, WBP2NL and XBP1 upregulated in all eight studies in which they were identified as DEGs 8 of 8 studies
- ▼ FUS, HNRNPAB, TUBB4B, HNRNPA1, HNRNPLL, HMGB1 downregulated in every study all studies
- – HSPs constituted over 15% of core response genes >15%
- – H. diversicolor study had the fewest HSP DEGs (2 of 11 HSP genes) and lowest number of shared DEGs 2 of 11 HSP genes
- – H. rufescens transcriptome gave the highest average mapping rate as common reference 86%
- count 74 core response genes (DEGs shared in ≥7 of 9 studies)
- count 150 samples / 300 FASTQ files (control and heat-stressed samples across nine studies)
- pvalue adjusted p-value < 0.05 (Benjamini–Hochberg, Wald test) (threshold for differential expression in DESeq2)
- fold_change |LFC| > 1 (DEG criterion for two studies lacking biological replicates)
- count 86% (overall highest average mapping rate using H. rufescens reference transcriptome)
- pvalue padj < 0.01 (significance threshold for GO/KEGG/Reactome term assignment)
- other >15% (proportion of core response genes that are HSPs)
- count 15 genes upregulated in all studies where significant (of 50 upregulated core genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This meta-analysis reanalyzed nine publicly available RNA-seq datasets from heat-stress studies on abalone (genus Haliotis) using a standardized bioinformatic pipeline (read QC, trimming, mapping to a single shared reference transcriptome, and count generation). Differential gene expression within each study was assessed using DESeq2's Wald test with Benjamini–Hochberg-adjusted p-values (padj < 0.05), except for two datasets lacking biological replicates where |log2 fold change| > 1 was used as the sole criterion. A 'core response' gene set was defined by a vote-counting approach: genes identified as differentially expressed in at least 7 out of 9 studies. Functional enrichment and co-expression analyses (g:Profiler, STRING, WGCNA) were applied to this core gene set.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 negative binomial GLM with Wald test | Differential gene expression for each of the 7 studies that had biological replicates | Varies per study; 150 total samples (control + heat-stressed) across all 9 studies | not stated |
| Log2 fold change threshold (|LFC| > 1, no p-value) | Differential gene expression for 2 studies without biological replicates (PRJNA453554 and PRJNA481417) | Not stated per study; these 2 studies had no biological replicates | na |
| Principal Component Analysis (PCA) | Exploratory analysis of VST-normalized gene count data within individual studies (PRJNA721743, PRJNA557314) to assess effects of species and tissue type | Not stated per PCA | na |
| Weighted Gene Co-expression Network Analysis (WGCNA) | Co-expression module detection on merged normalized counts from all studies combined, comparing control vs. heat-stressed eigengenes | All 150 samples merged | not stated |
| g:Profiler functional enrichment (hypergeometric, BH-corrected) | GO, KEGG, and Reactome pathway enrichment of the 74 core response genes | 74 core genes as input | not stated |
| STRING protein–protein interaction enrichment (BH-corrected) | Functional and physical interaction network of core response genes using H. rufescens proteome | 74 core genes as input | not stated |
-
A vote-counting approach was used to define the core gene set: genes had to appear as DEGs in at least 7 of 9 studies.↳ Could also: Effect-size-based meta-analysis (e.g., random-effects model on log2 fold changes across studies, or Fisher's combined p-value method) could also synthesize results across studies. — Effect-size meta-analysis weights studies by precision and produces a single pooled effect estimate with a confidence interval, which can be more statistically powerful than vote-counting and explicitly accounts for between-study heterogeneity.
-
All datasets were mapped to a single shared reference transcriptome (H. rufescens) regardless of species, to enable cross-study comparability.↳ Could also: Species-specific reference transcriptomes or de novo transcriptome assembly per dataset could also be used, with cross-study comparison done via ortholog mapping (e.g., OrthoFinder or BLASTp reciprocal best hits). — Ortholog-based approaches preserve species-specific annotation accuracy and allow detection of lineage-specific responses that may be missed when reads from divergent species are forced onto a single reference.
-
Two datasets without biological replicates were included using a fold-change-only criterion (|LFC| > 1).↳ Could also: These datasets could also be excluded from the primary analysis and treated as a separate validation set, or quasi-likelihood methods (e.g., edgeR's GLM with dispersion borrowed from a related dataset) could provide approximate significance estimates. — Fold-change thresholds alone are sensitive to library depth and baseline expression level; borrowing dispersion from related samples or using these datasets only for directional confirmation preserves the statistical rigor of the main analysis.
-
Variance Stabilizing Transformation (VST) was applied to normalized counts for PCA and visualization.↳ Could also: regularized log transformation (rlog, also from DESeq2) or voom-transformed counts (limma) could also be used for this purpose. — rlog is often recommended for small sample sizes as it applies stronger shrinkage for low-count genes, while voom integrates naturally with limma's linear modelling framework if a joint multi-study analysis were desired.
-
WGCNA was applied to all studies merged into a single dataset, treating the combined data as a unified experiment.↳ Could also: Consensus WGCNA (multi-set WGCNA) could also be applied, which identifies co-expression modules that are reproducible across individual studies rather than in a single merged matrix. — Consensus WGCNA is specifically designed for multi-dataset settings and explicitly models between-study variation, potentially yielding modules more robust to study-specific confounders such as species and tissue type.
-
Multiplicity correction (BH FDR) was applied independently within each study, with no correction across the vote-counting step itself.↳ Could also: A permutation-based approach or resampling procedure could also be used to empirically estimate the expected number of genes appearing in ≥7/9 studies by chance, given the per-study DEG rates. — Empirical null distribution for the vote-counting threshold accounts for dependencies between studies (e.g., shared reference, overlapping biology) and provides a study-level false discovery rate for the core gene list.
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is largely a 1:1 reproduction of the paper's central meta-analysis result: every headline count (150 samples/9 studies, the 74-gene core set and its 15+59 breakdown, the 15 up-in-all-significant and 6 named down-only genes, 3 STRING clusters/74 nodes) reproduces exactly from the authors' deposited intermediate tables, and an independent recomputation of per-study direction from the normalized counts (Table S6) agrees with the reported labels 523/533 = 98.1% — no fabrication signal. The only deviations are minor and explainable: C9 is the paper's own text-vs-supplement inconsistency (24 in text vs 23 in Table S4), and C8's exact '50' is definition-sensitive (49 by DEG label / 53 by normalized mean). Caveat for scope: the heavy raw-FASTQ pipeline was not re-run, so the reproduction leans on deposited outputs, but the independent direction check corroborates them. Overall: solid, central conclusion fully holds, deviations are trivial and partly on the authors' side.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.