Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Estimation of genetic parameters and genome-wide association study for carcass traits in native chickens.

Anim Biosci · 2025
L1 71/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough at the data level; partial reproduction. The declared code artifact is fastp (third-party QC tool), NOT the authors' analysis code. The paper's headline results (9,402,908-SNP set, heritabilities h2=0.08-0.50, 311 GWAS SNPs / 73 candidate genes) require a full 565-sample WGS -> BWA 0.7.10 -> GATK 3.5 -> GEMMA/ASReml pipeline (~12 TB raw, commercial ASReml licence, wet-lab carcass phenotypes) -> explicitly OUT OF SCOPE for 'a few clear data points'. REPRODUCED 1:1 (control-plane, ENA filereport): PRJNA942350 = 209 DNBSEQ-T7 PAIRED Gallus-gallus runs, 150 bp reads, mean 21.7 Gbp/run -> confirms paired-end (C2 exact), read length consistent with -l 150 (C3 exact), >10 G/individual (C4 within-tol), species (C5 exact). KEY FLAG (C1, mismatch): paper states 'Illumina HiSeq X Ten' but all 209 PRJNA942350 runs are DNBSEQ-T7 (different manufacturer) - possible methods/data inconsistency; reviewer should check the second BioProject PRJNA1214910. NOT ATTEMPTED: fastp QC run (C6) - sbatch job prepared (reproduction/run.sbatch: download 3 representative runs to «infra», fastp -q 30 -u 30 -l 150) but «our HPC» VPN required SAML+2FA that was not completed before the finalize instruction; also the paper reports no fastp output metrics, so even a successful run has no 1:1 paper value to compare. NOT ATTEMPTED: alignment, variant calling, GWAS, heritability (scope.md).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-15 ⛓ 2e9935d86e9f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper tests whether genome-wide variants and candidate genes (e.g., IGF2BP1) are significantly associated with carcass traits (slaughter weight, eviscerated weight, thigh muscle weight) in Sanhuang native chickens, and estimates the genetic parameters of these traits to facilitate genomic breeding.

Core claims
  • 311 SNPs and 73 candidate genes (e.g., IGF2BP1, BMP3, ACSL5) are significantly associated with carcass traits in Sanhuang chickens finding
  • IGF2BP1 plays a causal role in eviscerated weight and thigh muscle weight, with GIP, SNF8, and PHOSPHO1 located within the same genomic peak on GGA27 mechanism
  • Carcass traits show moderate to high heritability (0.20-0.50) except percentage of slaughter weight (SP, 0.08), and high genetic correlation (0.72-0.93) among SW, EW, and ThW finding
  • 17 of 73 candidate genes (e.g., IGF2BP1, RASGEF1B, BMP3) are differentially expressed between commercial broilers and Sanhuang chickens, with lipid metabolism and immune function altered by selection for carcass traits finding
  • CCND2 is related to SP and significantly higher expressed in commercial broilers than SH chickens finding
  • GWAS using a mixed linear model in GEMMA identifies SNP-trait associations while controlling population structure method
  • A multi-tissue transcriptomic expression profile of 73 candidate genes was constructed as a resource resource
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome resequencing / GWAS 565 Sanhuang (SH) chickens, blood none SNP genotypes associated with carcass traits (SW, EW, ThW, SP, EP, ThP) Illumina HiSeq X Ten; BWA MEM, GATK, GEMMA
Phenotypic carcass measurement 565 SH chickens, 91 days of age none slaughter weight, eviscerated weight, thigh muscle weight and their percentages
Genomic heritability / genetic correlation estimation 565 SH chickens (genomic relationship matrix) none heritability, genetic and phenotypic correlations ASReml v4.1
bulk mRNA sequencing (RNA-seq) SH chicken breast muscle, thigh muscle, lung, liver, heart, fat (n=5); breast muscle of commercial broilers vs SH chickens (n=6) breed comparison (commercial broiler vs SH) differentially expressed genes / expression abundance of candidate genes Illumina PE150; HISAT2, StringTie, DESeq2
RT-PCR (qPCR) breast muscle of commercial broilers and SH chickens (n=6) breed comparison relative expression of candidate genes (IGF2BP1, BMP3, ACSL5, RASGEF1B, MRPL22, ABI3, GIP, CCND2) SYBR Green Pro Taq HS kit; 2^-ΔΔCT method, β-actin control
Key results
  • 311 SNPs and 73 candidate genes significantly associated with carcass traits 311 SNPs, 73 genes
  • Heritability of six traits ranged from 0.08 to 0.50; SP lowest at 0.08 0.08-0.50
  • High genetic correlation among SW, EW, and ThW 0.83-0.95
  • High phenotypic correlation among SW, EW, and ThW 0.72-0.93
  • 17 of 73 candidate genes differentially expressed between broilers and SH chickens; 15 up-regulated in broilers, ABI3 and MRPL22 down-regulated 17/73 genes
  • CCND2 showed increased expression trend in commercial broilers Fold change = 1.41, p = 0.051
  • Peak SNP on GGA4 contributed highest PVE to ThW where BMP3 located 6.85% PVE
  • ACSL5 and RASGEF1B significantly up-regulated in breast muscle of commercial broilers vs SH chickens
Key statistics
  • count 311 SNPs and 73 candidate genes (significant associations with carcass traits)
  • other heritability 0.08-0.50 (SP = 0.08) (genomic heritability of six carcass traits)
  • correlation 0.83-0.95 (genetic correlation among SW, EW, ThW)
  • correlation 0.72-0.93 (phenotypic correlation among SW, EW, ThW)
  • correlation 0.24 (genetic correlation between EW and EP)
  • fold_change Fold change = 1.41, p = 0.051 (CCND2 expression broilers vs SH chickens)
  • mean SW 1,107.97 g; EW 923.12 g; ThW 181.50 g (average carcass weights in SH population)
  • pvalue -log10 p = 6.83 (genome-wide), 5.53 (suggestive) (GWAS significance thresholds (0.05/339733))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study estimated genomic heritability and pairwise genetic/phenotypic correlations for six carcass traits in 565 Sanhuang chickens using univariate and bivariate genomic animal models (REML, ASReml v4.1) with a VanRaden genomic relationship matrix. GWAS was conducted with a mixed linear model in GEMMA, correcting for population structure via top-three PCA components and kinship; significance was assessed by Bonferroni correction through the simpleM effective-test approach. Differential gene expression between commercial broilers and Sanhuang chickens was tested with DESeq2; candidate gene RT-PCR expression differences were compared by Student's t-test; and haplotype-group phenotypic differences were assessed with the Kruskal-Wallis test.

Replicationbiological Sample size565 chickens for GWAS/heritability; n=5 per tissue type for multi-tissue RNA-seq; n=6 per group for broiler vs SH RNA-seq and RT-PCR; no formal a priori power calculation reported GroupsSanhuang chickens stratified by haplotype group (GWAS); commercial broilers vs SH chickens (expression analyses) Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBonferroni via simpleM effective-test count (0.05/339,733 effective tests, -log10 p = 6.83 genome-wide; suggestive threshold 1/339,733, -log10 p = 5.53); DESeq2 internal adjustment method not explicitly specified beyond 'p<0.05'
Statistical tests used
Test Applied to n Assumptions
Mixed linear model (MLM) GWAS via GEMMA, Bonferroni/simpleM significance threshold SNP association with SW, SP, EW, EP, ThW, ThP across ~9.4 M autosomal SNPs 565 SH chickens not stated
Univariate genomic animal model (REML, ASReml v4.1) Heritability estimation for each of the six carcass traits 565 SH chickens not stated
Bivariate genomic animal model (REML, ASReml v4.1) Genetic and phenotypic correlation estimation for all pairwise trait combinations 565 SH chickens not stated
DESeq2 Wald test (fold change >1.5 or <0.67, p<0.05) Differential expression of 73 candidate genes in breast muscle: commercial broilers vs SH chickens; multi-tissue expression profiling in SH chickens n=6 per group (broilers vs SH breast muscle); n=5 for multi-tissue SH not stated
Student's t-test (two-group comparison) RT-PCR quantification of candidate gene expression differences between commercial broilers and SH chickens n=6 per group not stated
Kruskal-Wallis test Phenotypic differences in EW, SW, and ThW among haplotype groups at significant GWAS loci 565 SH chickens not stated
Approaches that could also have been used
  • Candidate gene RT-PCR expression was compared between two groups (n=6 each) using Student's t-test
    Could also: A Mann-Whitney U (Wilcoxon rank-sum) test could also be used for the same two-group comparison — With n=6 per group, the normality assumption underlying the t-test cannot be robustly verified empirically; the non-parametric Mann-Whitney U requires no distributional assumption and is widely applied at small sample sizes as a complement or alternative
  • RT-PCR comparisons were conducted across several candidate genes without a stated multiple-testing correction
    Could also: A Benjamini-Hochberg FDR correction applied across all gene-level RT-PCR tests would also be a standard approach — Testing multiple genes simultaneously at a nominal alpha of 0.05 inflates the expected false-positive count; FDR control quantifies and bounds this inflation and is routinely applied in multi-gene validation panels
  • Heritability and genetic correlations were estimated pairwise via separate univariate and bivariate animal models
    Could also: A single multivariate (six-trait) animal model fitted simultaneously could also be used — Joint multivariate REML estimates all genetic covariances in a single run, potentially yielding more efficient parameter estimates when traits are strongly correlated, and avoids accumulation of estimation error across multiple bivariate analyses
  • GWAS was conducted with a frequentist mixed linear model (GEMMA)
    Could also: Bayesian whole-genome regression methods (e.g., BayesC, BayesB, or BSLMM) could also be applied to the same data — Bayesian approaches estimate marker-specific shrinkage and do not assume uniform effect-size distributions; they are widely used in livestock GWAS as a complement to MLM and can improve fine-mapping resolution when causal variants have heterogeneous effects
  • Heritability estimates are reported as point estimates without accompanying standard errors or confidence intervals
    Could also: Standard errors (available from ASReml output) or parametric bootstrap 95% CIs around h² and rg estimates could also be reported — REML heritability estimates from a single population carry substantial sampling variance, especially for traits with low h² (e.g., SP = 0.08); SE or CI allows readers to judge estimation precision and the reliability of moderate-to-high heritability claims
  • DESeq2 differential expression results are described using 'p<0.05' without explicitly specifying whether this refers to raw or BH-adjusted p-values
    Could also: Explicitly reporting the adjusted p-value (padj from DESeq2's default Benjamini-Hochberg procedure) and using a stated FDR threshold (e.g., padj<0.05 or padj<0.10) would also be a standard presentation — DESeq2 computes both raw and adjusted p-values by default; clarifying which threshold was applied aids reproducibility and lets readers calibrate the expected false-discovery rate among the 17 reported DEGs
Software: GEMMA · ASReml 4.1 · PLINK 1.90 · DESeq2 · SPSS 22.0 · Beagle 5.2 · GATK 3.5 · LDBlockShow 4.3 · HISAT2 2.2.0 · StringTie 2.1.6 · fastp 0.21 · BWA MEM 0.7.10 · VCFtools 0.1.15

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Authors · 4
1Xianghua Zhu 2Houxue Cui 3Nanxi Dong 4Lu liu
Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GCA_000002315.5 GCA in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
PRJNA942350 BioProject in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40211845

Paper: Zhu X, Cui H, Dong N, Liu L. Estimation of genetic parameters and genome-wide association study for carcass traits in native chickens. Anim Biosci 2025. PMID 40211845 · PMCID PMC12229932 · DOI 10.5713/ab.25.0070.

Declared code artifact (per RU brief): https://github.com/OpenGene/fastp (fastp — a third-party FASTQ QC / filtering tool). This is NOT the authors' own analysis code; per brief rule P16, applying this third-party tool to the paper's own data is an equally valid reproduction.

Declared data: SRA BioProject PRJNA942350 (paper also lists PRJNA1214910).

Reported pipeline (from Methods, PMC full text)

Stage Tool reported In scope? Why
Read QC / filtering fastp -q 30 -u 30 -l 150 YES This is the named code artifact; runnable on the deposited data with a few clear data points.
Alignment BWA MEM v0.7.10 no Requires full reference align of 565 WGS samples; not the named tool; heavy.
Variant calling GATK v3.5 no Joint genotyping of all 565 samples needed to reach the reported 9,402,908-SNP set; far beyond "a few clear data points"; not the named tool.
GWAS GEMMA (MLM) no Depends on full variant set above.
Genetic parameters (h², r) ASReml v4.1 no Commercial software (licence-blocked) + needs full pedigree/phenotype + genotype data; phenotypes are wet-lab carcass measurements (out of scope).

What we attempt (80/20)

  1. Dataset-metadata 1:1 check (control-plane, ENA API — no compute): verify the deposited PRJNA942350 data matches the paper's stated data characteristics (platform, read layout, read length, raw volume per individual, species, sample count). Several CLEAR, checkable data points; flags any discrepancy.
  2. fastp QC reproduction on a few representative runs from PRJNA942350, with the exact published parameters -q 30 -u 30 -l 150, on «our HPC»/«infra». Demonstrates the named tool runs on the real data and reports the QC metrics (clean-read retention, Q20/Q30 before/after, GC, duplication, adapter).

What we explicitly do NOT attempt, and why

  • Full WGS alignment → variant calling → 9,402,908-SNP set (565 samples, ~12 TB raw; weeks of compute; not the named tool) — out of feasible scope.
  • Heritability estimates (h² 0.08–0.50, Table 2) and genetic correlations — need ASReml (commercial, licence-blocked) + wet-lab carcass phenotypes + full genotypes.
  • GWAS (311 SNPs, 73 candidate genes) — needs the full variant set + phenotypes.

Honest caveat on the fastp comparison

The paper reports no fastp output metrics (no per-sample clean-read counts, no Q30 table). So our fastp run produces real QC numbers but has no paper value to compare 1:1 — graded partial (tool reproducibly runs on the real data with the published parameters; nothing to match against). The genuinely comparable 1:1 points are the dataset-metadata facts in step 1.

C1
Reported
Illumina HiSeq X Ten platform
Reproduced
DNBSEQ-T7 (209/209 runs in PRJNA942350)
did not match
C2
Reported
paired-end
Reproduced
PAIRED (209/209)
exact
C3
Reported
fastp -l 150 (>=150 bp reads)
Reproduced
150 bp
exact
C4
Reported
over 10 G raw read sequence per individual
Reproduced
mean 21.7 Gbp/run (range 9.6-55.7, n=209)
within tolerance
C5
Reported
Gallus gallus (native Sanhuang chicken)
Reproduced
Gallus gallus (209/209)
exact
C6
Reported
fastp -q 30 -u 30 -l 150
Reproduced
NOT EXECUTED (job script ready, «our HPC» VPN 2FA not completed before finalize)
partial
C7
Reported
565 individuals total (PRJNA1214910 + PRJNA942350)
Reproduced
PRJNA942350 = 209 runs
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

82.8 k
tokens (I/O) · 5.3 M incl. cache
17 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.