Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Extensive transgressive gene expression in testis but not ovary in the homoploid hybrid Italian sparrow.

Mol Ecol · 2022
L1 93/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL (honest). DESCRIBED WELL ENOUGH: yes for the upstream pipeline -- the repo (github.com/Homap/Expression_sparrow @77bca83) ships a clear shell/SLURM pipeline Trimmomatic -> STAR 2.7.2b 2-pass -> HTSeq(cass.gff); downstream DESeq2 DE + transgressive classification are Methods-only (R code not shipped). REPRODUCED 1:1: C3 sample design = EXACT (all 28 PRJNA832330 runs match species/tissue/sex; +13 hybrid superset). C1 mapping rate = WITHIN-TOL: re-ran Trimmomatic+STAR (authors' exact params) on «our HPC» against the same Elgvin-2017 house sparrow assembly (NCBI GCA_001700915.1) for 6 representative libraries -> mean 91.69% total mapped (4/6 >=90%), consistent with the paper's >90%; the two testis libs at 88.9-89.1% are expected given we lacked the cass.gff splice-junction DB and used STAR 2.7.11b. BLOCKED (their-side data gap, not fabrication): C2 (14,734 genes), C4a/C4b (DE), C5/C6 (transgressive), C7 (26x), C8 (GO) ALL depend on the cass.gff annotation, which is not deposited in any public repository and whose host (CEES genome browser) is offline -- confirmed unreachable from both «host» and «our HPC». These downstream numbers are therefore neither confirmed nor refuted. NOT ATTEMPTED: wet-lab steps (out of scope). Compute genuinely ran on «our HPC» («job», 2214065). Dataset PRJNA832330 profiled grade A with file-content QC now passing (downloads decompress + parse clean, STAR input == ENA read_count).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 93
    assessed: 2026-06-21 ⛓ be1bc4470dc3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether gonad gene expression in the stabilized homoploid hybrid Italian sparrow is intermediate to its two parental species (house and Spanish sparrow), reflecting its intermediate genomic composition, or whether break-up of co-evolved cis- and trans-regulatory elements instead produces transgressive expression patterns.

Core claims
  • Italian sparrow testis exhibits extensive transgressive gene expression relative to both parental species finding
  • Italian sparrow ovary gene expression resembles that of the house sparrow parent rather than being transgressive finding
  • Italian sparrow testis transcriptome is far more diverged from parental transcriptomes than the parental transcriptomes are from each other, despite genetic intermediacy finding
  • Genes involved in mitochondrial respiratory chain complexes and protein synthesis are enriched among over-dominantly expressed testis genes, suggesting selection has shaped the hybrid transcriptome mechanism
  • Gene expression inheritance patterns (additive, dominant, under-dominant, over-dominant) were classified following the McManus et al. (2010) framework method
  • Z-linked genes are significantly overrepresented among differentially expressed genes in testis comparisons finding
  • Differential expression between the parental species is strongly asymmetric, with testis more conserved and ovary more divergent finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq testis, wild house/Spanish/Italian sparrow none (interspecies comparison) differential gene expression Illumina HiSeq4000, TruSeq Stranded mRNA library prep
bulk RNA-seq ovary, wild house/Spanish/Italian sparrow none (interspecies comparison) differential gene expression Illumina HiSeq4000, TruSeq Stranded mRNA library prep
bulk RNA-seq testis, captive-bred experimental F1 hybrid (house female x Spanish male) F1 hybrid cross differential gene expression Illumina HiSeq4000
bulk RNA-seq ovary, captive-bred experimental F1 hybrid (house female x Spanish male) F1 hybrid cross differential gene expression Illumina HiSeq4000
Gene Ontology functional enrichment analysis house sparrow protein set / differentially expressed gene sets none enriched GO biological process terms PANNZER, clusterProfiler
protein-protein interaction network analysis differentially expressed gene products (testis) none predicted interaction networks and biological process clusters STRING v11, Cytoscape, ClueGO
RNA integrity quality control gonad RNA samples (testis and ovary) none Transcript Integrity Number (TIN) RSeQC
genomic PCA / ancestry analysis experimental F1 hybrids and whole-genome resequencing data from Passer species none genomic ancestry proportions and mtDNA grouping
Key results
  • 2530 genes (22% of testis genes tested for inheritance) show transgressive expression outside the range of both parent species in Italian sparrow testis 22%
  • 2611 genes (22.71%) in testis showed nonconserved inheritance, with transgressive expression and house-sparrow-dominant as the two largest categories 22.71%
  • Italian sparrow ovary differed from house sparrow in only 22 genes (0.18%) versus 1508 genes (12.63%) differing from Spanish sparrow
  • Italian sparrow testis transcriptome is 26 times as diverged from parental transcriptomes as the parental transcriptomes are from each other 26-fold
  • 3536 genes (30.45%) were differentially expressed in Italian testis vs house sparrow; 3581 genes (30.9%) vs Spanish sparrow ~30%
  • 135 genes (1.16%) were differentially expressed in testis versus 1382 genes (11.65%) in ovary between the parental species
  • 24 Z-linked genes were significantly overrepresented among testis differentially expressed genes between parental species
  • 196 of 3581 differentially expressed genes were Z-linked and overrepresented in the Italian vs Spanish sparrow testis comparison
Key statistics
  • pvalue p < 1.82e-08 (hypergeometric test for Z-linked gene enrichment among testis DE genes, parental comparison)
  • pvalue p = .006 (hypergeometric test for Z-linked gene enrichment among Italian vs Spanish testis DE genes)
  • pvalue p = 1.05e-06 (Chi-squared test for up/down-regulation bias in testis (Italian sparrow vs parents))
  • pvalue p = .001 (Chi-squared test for up/down-regulation bias in ovary (Italian sparrow vs parents))
  • other FST house-Spanish = 0.33; house-Italian = 0.18; Spanish-Italian = 0.25 (genome-wide genetic differentiation between species)
  • count n = 5 house, 5 Spanish, 5 Italian (testis RNA-seq sample sizes)
  • count n = 5 house, 3 Spanish, 5 Italian (ovary RNA-seq sample sizes (Spanish n=3))
  • fold_change shrunken LFC > 0.32 (1.25-fold), padj < .05 (threshold for calling differential expression and nonconserved inheritance)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study compared gonad (testis and ovary) gene expression among wild house sparrows, Spanish sparrows, and their homoploid hybrid, the Italian sparrow (n=3-5 individuals per group per tissue), using RNA-seq read counts. Differential expression between pairs of taxa was tested with DESeq2 (applying a false discovery rate threshold and a shrunken log2 fold-change cutoff), and additional tests (hypergeometric, chi-squared) were used to assess chromosomal enrichment and directional bias among differentially expressed genes. Functional enrichment of Gene Ontology terms among differentially expressed gene sets was assessed with clusterProfiler and ClueGO, each with its own FDR/adjusted-p threshold, and results were reported primarily as gene counts, percentages, log2 fold changes, and exact p-values rather than with traditional descriptive dispersion statistics.

Replicationbiological Sample sizeSample sizes given as counts per group and tissue (testis: house=5, Spanish=5, Italian=5; ovary: house=5, Spanish=3, Italian=5; experimental F1 hybrids: 8 ovary, 5 testis); no power analysis or sample-size justification described. Groupshouse sparrow vs Spanish sparrow vs Italian sparrow (and opportunistically vs experimental F1 hybrids), separately for testis and ovary Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Multiplicity correctionfalse discovery rate (FDR) correction (via DESeq2 padj, and separately for GO enrichment via clusterProfiler and ClueGO adjusted p-values)
Statistical tests used
Test Applied to n Assumptions
DESeq2 differential expression testing (Wald test with empirical Bayes shrinkage of log2 fold changes) pairwise comparisons of gene expression (house vs Spanish, house vs Italian, Spanish vs Italian) separately in testis and ovary testis: house=5, Spanish=5, Italian=5; ovary: house=5, Spanish=3, Italian=5 stated
hypergeometric test (R phyper()) over-representation of Z-linked genes among differentially expressed genes (e.g., testis house-Spanish comparison; testis Italian-Spanish comparison) counts of differentially expressed genes on Z chromosome vs autosomes among all genes tested not stated
chi-squared test proportion of up- vs down-regulated genes in Italian sparrow testis and ovary relative to house and Spanish sparrow not stated
GO term enrichment analysis (clusterProfiler) functional enrichment of Gene Ontology terms among differentially expressed genes relative to a background of all expressed/tested genes, for each pairwise comparison not stated
ClueGO functional term enrichment on STRING protein-protein interaction network biological process enrichment among proteins with significant GO terms in Italian sparrow testis not stated
Approaches that could also have been used
  • Differential expression was tested pairwise per tissue with modest biological sample sizes (n=3-5 per group) and no stated power analysis.
    Could also: Reporting an a priori or post hoc power analysis, or a minimum-detectable-effect-size calculation — This would give readers additional context on how sensitive each comparison was to detect differential expression of a given magnitude, which can be particularly informative when working with modest sample sizes typical of non-model organism RNA-seq studies.
  • Differential expression was assessed using DESeq2's Wald test framework with empirical Bayes shrinkage.
    Could also: edgeR or limma-voom based differential expression testing — These are widely used alternative RNA-seq differential expression tools with different dispersion-estimation and normalization approaches, and comparing results across tools can illustrate the robustness of findings to methodological choice.
  • FDR correction was applied separately within each of the three pairwise comparisons (house-Spanish, house-Italian, Spanish-Italian).
    Could also: A joint correction across the full family of pairwise comparisons (e.g., pooling all tests before applying Benjamini-Hochberg) — Correcting jointly across all comparisons in a study is an alternative that can be more conservative and controls the false discovery rate across the entire set of tests performed, rather than within each comparison independently.
  • Enrichment of Z-linked genes among differentially expressed genes was tested with a hypergeometric test.
    Could also: A permutation-based enrichment test or Fisher's exact test — These provide alternative or complementary ways to assess enrichment significance and can be useful for cross-checking results, especially when gene sets show non-independence such as physical linkage.
  • GO functional enrichment was based on discrete lists of significantly differentially expressed genes (clusterProfiler, ClueGO).
    Could also: Gene set enrichment analysis (GSEA) using continuously ranked log2 fold-change values — GSEA can detect coordinated, sub-threshold shifts in expression across a gene set that would not appear using a fixed significance cutoff, complementing enrichment analyses based on a fixed differentially-expressed gene list.
  • Variability in gene expression across biological replicates was handled internally via DESeq2's per-gene dispersion model rather than reported with descriptive dispersion statistics.
    Could also: Reporting confidence intervals for individual log2 fold-change estimates — Presenting CIs alongside point estimates of fold change can directly convey the precision of individual gene-level estimates, complementing the p-value/FDR-based significance framework.
Software: R 4.0.2 · DESeq2 · STAR 2.7.2b · htseq 0.9.1 · clusterProfiler · STRING 11

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35726533

Paper: Papoli Yazdi et al. 2022, Mol Ecol 31:5575-5590. "Extensive transgressive gene expression in testis but not ovary in the homoploid hybrid Italian sparrow." DOI 10.1111/mec.16572 · PMCID PMC9542029.

Code: https://github.com/Homap/Expression_sparrow (authors' own; P16 N/A). Data: SRA BioProject PRJNA832330 (41 runs deposited; 28 are this paper's design).

Pipeline (as described in Methods + repo)

The repo ships the upstream RNA-seq pipeline as plain shell/SLURM scripts:

Stage Tool (version) Repo script(s) Key params
QC FastQC processing/fastqc*.sh default
Trim/adapter Trimmomatic processing/trimm*.sh PE, TruSeq3-PE-2.fa adapters
Align STAR 2.7.2b, 2-pass mapping/star_aligner/*.sh default params; ref = house sparrow genome (Elgvin et al. 2017)
Count HTSeq-count 0.9.1 counting/htseq*.sh min MAPQ 30; annotation cass.gff
DE DESeq2 (R) NOT in repo median-of-ratios norm; padj<0.05 &
Transgressive classification custom (R) NOT in repo additive/dominant/over-/under-dominant; >1.25-fold deviation from parents = "nonconserved"

So the repo covers fastq → trim → STAR → HTSeq count matrix. The downstream DESeq2 differential expression and the transgressive/inheritance classification are described in Methods but their R code is not shipped. Per BRIEF rule P16, re-implementing those from the described parameters on the paper's own count matrix is an equally valid reproduction — attempted as the harder tier.

IN SCOPE (pipeline-derived; attempt)

# Reported result Pipeline Tier
C1 >90% of reads mapped (STAR) trim→STAR 2-pass 80% floor
C2 14,734 annotated genes (92.52% chromosomal, 7.48% scaffold) HTSeq vs cass.gff annotation 80% floor
C3 Sample design: 15 testis (5 house/5 Spanish/5 Italian) + 13 ovary (5/3/5) = 28 data accounting (PRJNA832330) 80% floor — DONE (control-plane)
C4 DE testis Italian vs house = 3536 genes; vs Spanish = 3581 HTSeq counts → DESeq2 harder ~20%
C5 Transgressive testis Italian = 2530 genes (22% of genes tested for inheritance) DESeq2 + transgressive classification harder ~20%
C6 Transgressive ovary Italian = 4 genes (0.028%) same harder ~20%
C7 Italian testis transcriptome 26× as diverged from parents as parents from each other expression distance hardest
C8 GO enrichment: mitochondrial respiratory chain / protein synthesis in over-dominant testis set topGO/GO hardest, downstream

OUT OF SCOPE (not pipeline-derived → not attempted)

  • RNA extraction, library prep, sequencing (wet-lab).
  • Field/tissue collection, ethics.
  • Reference genome assembly itself (Elgvin et al. 2017; we use it, not rebuild it).

Known dependencies / risks

  • Reference genome + annotation (house_sparrow_genome_assembly-18-11-14.fa, cass.gff) are external (Elgvin et al. 2017) and NOT in the repo or the SRA deposit → must be located on a public host (NCBI/Dryad/figshare/ENA). Possible env_unresolvable/docs_insufficient blocker for C1/C2 if unobtainable.
  • Downstream DESeq2 + transgressive R code not shipped → C4-C6 are re-implementations from the Methods text, so exact-match is not guaranteed (parameter ambiguity in the classification thresholds).
  • Heavy compute (28 PE RNA-seq libraries, 40-66M read pairs each, STAR 2-pass) → «our HPC» SLURM only.
C3
Reported
28 RNA-seq samples: 15 testis (5 house/5 Spanish/5 Italian) + 13 ovary (5 house/3 Spanish/5 Italian)
Reproduced
All 28 present in PRJNA832330 with EXACT species x tissue x sex match; +13 extra captive-bred hybrid runs (superset)
exact
C1
Reported
>90% of reads mapped (STAR 2-pass)
Reproduced
6 representative libs (testis+ovary x house/Spanish/Italian): total mapped 88.89/94.84/90.19/91.33/89.13/95.74%, MEAN 91.69%, 4/6 >=90%; Trimmomatic+STAR 2.7.11b twopassMode vs GCA_001700915.1 (Elgvin 2017 assembly), no cass.gff sjdb available
within tolerance
C2
Reported
14,734 annotated genes (92.52% chr / 7.48% scaffold)
Reproduced
BLOCKED: cass.gff annotation not in any public repo (not NCBI/Dryad/figshare/Zenodo); CEES genome browser offline/unreachable from «host» AND «our HPC»; author-request-only
m.public.grade.error
C4a
Reported
3536 DE genes Italian vs house testis
Reproduced
BLOCKED: needs cass.gff gene counts (unavailable); DESeq2 R code not shipped
m.public.grade.error
C4b
Reported
3581 DE genes Italian vs Spanish testis
Reproduced
BLOCKED: same as C4a
m.public.grade.error
C5
Reported
2530 transgressive genes (22%) Italian testis
Reproduced
BLOCKED: needs cass.gff counts + unshipped transgressive-classification R code
m.public.grade.error
C6
Reported
4 transgressive genes (0.028%) Italian ovary
Reproduced
BLOCKED: same as C5
m.public.grade.error
C7
Reported
Italian testis transcriptome 26x as diverged from parents as parents from each other
Reproduced
NOT ATTEMPTED: downstream of unavailable count matrix
m.public.grade.error
C8
Reported
GO enrichment: mito respiratory chain + protein synthesis (over-dominant testis)
Reproduced
NOT ATTEMPTED: downstream of unavailable DE sets
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

297.9 k
tokens (I/O) · 24.6 M incl. cache
77 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.