Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comparison between short-term stress and long-term adaptive responses reveal common paths to molecular adaptation.

iScience · 2022
L1 59/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
59/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 19% of all assessed papers rank 925 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

INTERIM. Re-run after prior room was requeued for running ZERO compute; compute now RAN. R0 REPRODUCED (C12, within-tol): hardingnj/xpclr built at pinned commit beb3e4a (numpy1.26.4/scipy1.11.4/scikit-allel1.3.13), pytest 2/2 passed, and the tool emits a valid per-window XP-CLR table on its fixture via the hdf5 path the paper describes. Two version/code blockers were found+fixed: scipy>=1.12 removed scipy.integrate.romberg (pinned <1.12); the xpclr VCF loader is broken without --gdistkey (use hdf5 = the paper's stated input). R1 (real partial PRJNA274877 WGS -> fastp -> BWA -> GATK HaplotypeCaller/GenotypeGVCFs -> filtered SNP VCF -> hdf5 -> XP-CLR, highland=9 Yunnan vs lowland=23, restricted to chromosomes carrying STEAP2/STEAP1/CFAP69/ZNF804B) RUNNING. Exact headline numbers (3,277 windows/535 genes/max 33.51) NOT exactly reproducible: 6/38 WGS individuals absent from public deposit, no genotypes/group-labels deposited -> R1 targets qualitative concordance with the reported top candidate genes only.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 0e58c7ff9865
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Short-term stress from UVR exposure, low temperatures, and hypoxia simulating high-altitude extreme environments can induce plastic gene expression responses in cells that facilitate long-term adaptive evolution to high-altitude environments, and short-term plastic responses parallel long-term adaptive genomic evolution.

Core claims
  • Short-term stress and long-term adaptations share common metabolic pathways finding
  • Phenotypic plasticity can promote adaptive evolution finding
  • Metabolic pathways are the most significant signaling pathway activated in both great tit and mouse cells after high-altitude stimuli exposure finding
  • Common paths to molecular adaptation are adopted in mouse and bird cells finding
  • Positively selected genes (PSGs) and differentially expressed genes (DEGs) rarely overlap despite similar functional enrichment, due to high intermodular but not intramodular connectivity of PSGs finding
  • Stress response mechanisms and activated signaling pathways differ between mice and birds after short-term stimulation, indicating species-specific plastic responses finding
  • PSGs in high-altitude passerines are more often classified into lipid metabolism than carbohydrate metabolism finding
  • Genome-wide selective sweep analysis (XP-CLR, XP-EHH) identifies candidate genes under positive selection in high-altitude great tits method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq with WGCNA and DEG analysis great tit embryonic fibroblasts (GEF), Parus major cold (4°C, 1/4h), hypoxia (3% O2, 4/24h), UVR (100 J/m2, 0/1h post-exposure) gene co-expression modules and differentially expressed genes
bulk RNA-seq with WGCNA and DEG analysis mouse embryonic fibroblasts (MEF), Mus musculus cold (4°C, 1/4h), hypoxia (3% O2, 4/24h), UVR (100 J/m2, 0/1h post-exposure) gene co-expression modules and differentially expressed genes
gene set enrichment analysis (GSEA) GEF and MEF cold, hypoxia, UVR KEGG pathway enrichment (e.g., oxidative phosphorylation, glycolysis, metabolic pathways)
transcription factor binding motif overrepresentation analysis GEF and MEF (module gene promoter sequences) none (in silico on stress-responsive modules) overrepresented TF binding motifs (e.g., ZNF384, TBP, SP1, ETV6) CLOVER (Cis-eLement OVERrepresentation)
whole genome resequencing / selective sweep scan (XP-CLR, XP-EHH) wild great tits, 12 high-altitude and 26 low-altitude populations none (natural population comparison) genomic regions/genes under positive selection scikit-allel (hdf5-based)
multiple sequence alignment MAPK1 gene and homologs (cross-species) none conservation of protein kinase active site and functional domains
functional enrichment analysis (GO / KEGG) high- and low-altitude mammal and bird PSGs none enriched biological processes and pathways (e.g., lipid metabolism) g:profiler
Key results
  • Hypoxia produced the strongest gene co-expression response in GEF (blue module) cor=0.69, p=7e-04
  • UVR exposure associated with lightcyan module in GEF cor=-0.61, p=0.004
  • Low temperature produced the weakest co-expression response in GEF (gray module) cor=0.45, p=0.05
  • 11 metabolic genes (A4GALT, ALDOC, AMPD3, ENO2, GBE1, GCLM, GMDS, LDHA, MPI, PFKL, PGAM1) overlapped between GEF and MEF top stress-response modules
  • Four genes with highest XP-CLR selection scores identified in great tits: STEAP2, CFAP69, ZNF804B, STEAP1 XP-CLR score = 33.51
  • Top XP-EHH candidate genes in high-altitude great tits PARD3 XP-EHH=6.93; LOC107202990 XP-EHH=8.18
  • Only four genes overlapped between DEGs and PSGs (COL19A1, SEMA5A, ABCA13, SOX5)
  • PSGs consistently showed higher total connectivity than non-PSGs, but not consistently higher module membership, while DEGs showed higher gene significance than PSGs
Key statistics
  • correlation cor=0.69, p=7e-04 (GEF blue module correlation with hypoxia stress)
  • correlation cor=-0.61, p=0.004 (GEF lightcyan module correlation with UVR exposure)
  • correlation cor=0.45, p=0.05 (GEF gray module correlation with cold exposure)
  • count 26 gene modules (WGCNA modules identified in GEF)
  • count 21 gene modules (WGCNA modules identified in MEF)
  • count 3,277 windows (top 5% XP-CLR), 535 genes (genome-wide selective sweep analysis in great tits)
  • count 2,734 candidate SNPs; 66 genes (top 5%), 60 genes (bottom 5%) (XP-EHH analysis thresholds (1.11 and -1.1) in great tits)
  • fold_change XP-CLR score = 33.51 (highest selective sweep score among candidate PSGs (STEAP1))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combined in vitro RNA-seq of great tit (GEF) and mouse (MEF) embryonic fibroblasts exposed to simulated high-altitude stressors (cold, hypoxia, UVR) with population-genomic selective sweep analyses of 12 highland and 26 lowland great tits to compare short-term plastic transcriptional responses with long-term adaptive signals. WGCNA identified co-expression modules correlated with each treatment condition, and GSEA identified enriched KEGG pathways; differentially expressed genes (DEGs) were then overlapped with positively selected genes (PSGs) from XP-CLR and XP-EHH analyses. Results were reported as mean ± SEM with significance denoted by asterisks (p < 0.05/0.01/0.001), with some exact p-values provided for WGCNA module-trait correlations.

Replicationunclear Sample size12 high-altitude and 26 low-altitude great tits stated for population genomics; biological replicate count for cell-line RNA-seq not stated in available text; total of 1.1 billion reads generated GroupsGEF and MEF cells under cold (1h, 4h), hypoxia (4h, 24h), and UVR (0h, 1h post-exposure) vs. controls; highland vs. lowland great tit populations for selective sweep Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctioncorrection applied but method not specified; text states 'corrected p < 0.05' for DEG pathway enrichment; WGCNA and XP-CLR/XP-EHH p-values appear uncorrected or use empirical top-5% thresholds
Statistical tests used
Test Applied to n Assumptions
WGCNA module-trait Pearson correlation Correlating gene co-expression modules with treatment conditions (cold, hypoxia, UVR) in GEF and MEF cell line samples; biological replicate count not stated in available text not stated
Gene Set Enrichment Analysis (GSEA) Identifying enriched KEGG pathways in GEF and MEF transcriptional responses to stress treatments null not stated
Differential expression analysis (specific test/tool not named in available text) Identifying DEGs in GEF and MEF under cold (1h, 4h), hypoxia (4h, 24h), and UVR (0h, 1h post-exposure) treatments null not stated
XP-CLR (cross-population composite likelihood ratio test) Genome-wide selective sweep analysis comparing highland vs. lowland great tit populations 12 high-altitude + 26 low-altitude great tits not stated
XP-EHH (cross-population extended haplotype homozygosity) Selective sweep analysis comparing highland vs. lowland great tit populations 12 high-altitude + 26 low-altitude great tits not stated
Hypergeometric test Enrichment of KEGG metabolic pathway classes (lipid vs. carbohydrate metabolism) among PSGs and convergent genes in high-altitude passerines null not stated
Unnamed statistical comparison (asterisk-based, p < 0.05/0.01/0.001) Comparing total connectivity (TC), module membership (MM), and gene significance (GS) between PSGs and non-PSGs, and between DEGs and non-DEGs null not stated
CLOVER cis-element overrepresentation (permutation-based motif enrichment) Identifying overrepresented transcription factor binding motifs in 200 bp upstream regions of genes in co-expression modules null not stated
Approaches that could also have been used
  • Dispersion was reported as SEM throughout (Figure legends state 'mean ± SEM')
    Could also: SD or 95% confidence intervals could also be used to summarize spread — For cell-line experiments where biological replicate n is small, SD conveys the actual biological variability in the data rather than precision of the mean estimate; CIs additionally provide inferential context; both are commonly preferred in this setting
  • The DEG identification tool and its internal multiple-testing procedure are not named in the available text
    Could also: Explicit specification of DESeq2 (negative-binomial Wald test + Benjamini-Hochberg FDR), edgeR (quasi-likelihood F-test + FDR), or limma-voom could also be provided — Naming the DEG caller and correction method allows readers to assess model assumptions (e.g., dispersion estimation approach) and to reproduce the analysis; it also makes the built-in multiplicity control transparent
  • DEG counts across treatment conditions were compared between GEF and MEF using individual significance tests per condition (Figure 2G, asterisks)
    Could also: A two-way ANOVA (cell type × treatment) followed by a post-hoc correction (e.g., Tukey HSD or Benjamini-Hochberg) could also be applied — An omnibus test with post-hoc correction would explicitly model the interaction between cell type and treatment and control the family-wise error rate across the multiple simultaneous comparisons
  • WGCNA module-trait correlations were reported with uncorrected p-values across a matrix of multiple modules and multiple treatment conditions
    Could also: Benjamini-Hochberg FDR correction across the full module-by-trait correlation matrix could also be applied, as is common in WGCNA publications — The matrix of correlations (e.g., 26 GEF modules × 6 treatment variables) represents many simultaneous tests; FDR correction would account for this and is frequently reported alongside WGCNA module-trait results to reduce false positives
  • Selective sweep candidates were defined by empirical top-5% thresholds of XP-CLR scores and XP-EHH values
    Could also: Permutation-based or coalescent-simulation-derived significance thresholds could also be used — Empirical top-5% cutoffs depend on the genome-wide score distribution of the specific dataset and may be influenced by population structure; simulation-based thresholds can provide a reference more independent of those distributional properties
  • Pathway enrichment of PSGs used a hypergeometric test while enrichment of DEGs used GSEA, applying different frameworks to the two gene lists
    Could also: Applying a consistent enrichment method (e.g., ORA with the same background gene universe, or GSEA for both) to PSGs and DEGs could also be used — Using the same enrichment framework for both gene lists facilitates direct comparison of short-term (DEG) and long-term (PSG) pathway signals, and removes method-driven differences from the interpretation
Software: WGCNA (R package) · GSEA · CLOVER · scikit-allel (used to generate hdf5 input for XP-CLR) · g:Profiler

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35243257

Paper: Chen X, Ji Y, Cheng Y, Hao Y, Lei X, Song G, Qu Y, Lei F. Comparison between short-term stress and long-term adaptive responses reveal common paths to molecular adaptation. iScience 2022. PMID 35243257 · PMCID PMC8873613 · DOI 10.1016/j.isci.2022.103899.

Code link (brief): https://github.com/hardingnj/xpclr (XP-CLR re-implementation, Python/scikit-allel, MIT). HEAD beb3e4ae010f8aafe408a60374c1c1ffe568d6f0 (2022-06-13). Data link (brief): SRA PRJNA593713.

Important accession correction (established from the full text)

The paper relies on TWO datasets, and the brief's single accession points at the transcriptomic half, not the half the XP-CLR code operates on:

Accession Data type Role in paper What it feeds
PRJNA593713 (brief's accession) bulk RNA-seq, Parus major (GEF) + Mus musculus (MEF) embryonic fibroblasts under cold/hypoxia/UVR short-term stress half WGCNA modules, DESeq2 DEGs, GSEA
PRJNA274877 (NOT in brief; from Qu et al. 2015) whole-genome resequencing, Parus major long-term adaptation half BWA→GATK→VCFtools→XP-CLR/XP-EHH selection scan

The hardingnj/xpclr code therefore operates on PRJNA274877 (WGS genotypes), NOT on PRJNA593713 (RNA-seq). Both are profiled in data/dataset_profile.json.

The paper's Results sentence — "XP-CLR (… designed to run on hdf5 files representing genetic data as generated using scikit-allel)" — is a verbatim match to the hardingnj/xpclr README, confirming the authors used this exact third-party tool (Methods cite the original Reich scripts, but the Results description is hardingnj/xpclr). Per brief rule 2 (P16), applying this third-party tool to the paper's data is a fully valid reproduction.

Pipeline-derived results (candidate reproduction targets)

In scope (computational pipeline outputs)

  1. XP-CLR selective-sweep scan (PRJNA274877 WGS → BWA → GATK → VCFtools MAF<5% → plink phasing → XP-CLR). Reported: 3,277 windows in top-5% of XP-CLR scores harbouring 535 genes; top-4 genes STEAP2, CFAP69, ZNF804B, STEAP1 (max XP-CLR score 33.51). [Results "Candidate genes…"; Fig 3A; Table S8] — PRIMARY (code link).
  2. XP-EHH scan (rehh). Reported: 2,734 candidate SNPs; thresholds 1.11 (66 genes, top-5%) / −1.1 (60 genes, bottom-5%); top genes PARD3 (6.93), LOC107202990 (8.18). [Results; Fig 3B/C; Table S8]
  3. WGCNA modules (PRJNA593713 RNA-seq, β=14): 26 modules in GEF, 21 in MEF; top GEF module-trait corr: hypoxia blue cor=0.69 p=7e-4; UVR lightcyan cor=−0.61 p=0.004; cold gray cor=0.45 p=0.05. [Results; Fig 1B; Table S6]
  4. DESeq2 DEGs (|log2FC|≥1, FDR<0.05) per treatment in GEF/MEF [Fig 2E/2F].
  5. 11 overlapping metabolic-pathway genes between great tit & mouse cells: A4GALT, ALDOC, AMPD3, ENO2, GBE1, GCLM, GMDS, LDHA, MPI, PFKL, PGAM1 [Results; Fig 1C].
  6. 4 DEG∩PSG overlap genes: COL19A1, SEMA5A (GH24); ABCA13 (GU0); SOX5 (GU1) [Table S3].

Out of scope (wet-lab / manual / external — not attempted)

  • siRNA MAPK1 knockdown, apoptosis/cell-cycle flow cytometry, RT-qPCR, western blot (wet-lab; SPSS ANOVA).
  • MAPK1 PhastCons/MEGA phylogeny, Phyre2/Chimera 3D structure, Motif Scan, CLOVER TF motifs (external web tools / curated DBs, not a reproducible pipeline on the deposit).
  • PAML branch-site dN/dS PSGs (relies on multi-species alignments from prior studies Hao et al. 2019 / Tang et al. 2017, not deposited here).

Reproducibility blockers identified up-front (genotype side)

  • Missing samples. Paper uses 38 great tits (12 highland + 26 lowland); the public accession PRJNA274877 contains only 32 runs. The "6 added highland samples" are not in this (or any cited) deposit → exact allele frequencies, hence exact XP-CLR/XP-EHH scores, cannot be reproduced.
  • No deposited genotypes. Only raw FASTQ is public; no VCF/HDF5/SNP
Figures / tables: Fig 3ATableFig 3BFig 3CFig 1BFig 1C
C12
Reported
hardingnj/xpclr (the tool the paper used) executes on scikit-allel hdf5 and emits per-window XP-CLR scores
Reproduced
REPRODUCED: built @beb3e4a; pytest 2/2; xpclr --format hdf5 emits valid 13-col per-window table on fixture (chr3L, 930 SNPs)
within tolerance
C1
Reported
3,277 top-5% XP-CLR windows
Reproduced
PENDING
partial
C2
Reported
535 genes in top-5% windows
Reproduced
PENDING
partial
C3
Reported
top genes STEAP2/CFAP69/ZNF804B/STEAP1, max XP-CLR 33.51
Reproduced
PENDING
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 59/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

117.1 k
tokens (I/O) · 6.3 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.