Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Accelerated Evolution of Tissue-Specific Genes Mediates Divergence Amidst Gene Flow in European Green Lizards.

Genome Biol Evol · 2021
L1 23/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
23/100
Reproducibility score
2.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 1% of all assessed papers rank 1160 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction with a clean, auditable outcome. The named P16 tool catsequences (commit 59fd7c5) compiled and ran correctly on the paper's own data: it concatenated 2,806 single-copy-ortholog protein MSAs into a 4-taxon supermatrix (1,281,668 aa) and IQ-TREE2 (v3.1.2, JTT+F+I+G4, UFBoot=100) produced an ML species tree -- so the TOOL reproduces. However NONE of the three checkable headline numbers reproduces 1:1: (C1) the tree topology {adriatic,viridis}|{bilineata,podarcis} differs from the reported {bilineata,Adriatic} sister-pair -- but the authors' per-gene MSAs were NOT deposited so the alignments were reconstructed de novo (not like-for-like; difference most plausibly long-branch attraction of the divergent Podarcis outgroup); (C2) the deposited 'Filt' VCF is the raw 4-sample autosomal call set (79.07M records / 65.5M SNPs; 38.9M biallelic-no-missing) and no standard filter reproduces 26,985,663; (C7) the ProteinOrtho table gives 11,313 / 2,806 single-copy orthologs vs reported 1,497 (an autosomal-restricted subset not directly derivable). Root cause across all three: the Zenodo deposit sits one pipeline stage UPSTREAM of the reported figures and lacks the exact intermediate files (filtered VCF, supermatrix MSAs, autosomal gene partition). Fabrication concern LOW -- values are plausible but not independently verifiable from the shipped data, which is itself the auditable finding. Not attempted: C3 D-stat (no population panels in the 4-sample VCF), C4/C5 Ka/Ks (tissue count matrices + Z-chrom variants not deposited), C6 (out of scope).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-24
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The authors hypothesized that divergence among the three lineages of the Lacerta viridis species complex (L. viridis, L. bilineata, and the Adriatic lineage) occurred in the presence of gene flow between multiple lineages and involved tissue-specific gene evolution.

Core claims
  • The Adriatic lineage is a sister taxon to L. bilineata based on mitogenome and autosomal phylogenies finding
  • Gene flow (admixture) occurred between L. viridis and the Adriatic lineage during divergence, in addition to previously known gene flow between L. viridis and L. bilineata finding
  • The Z-chromosome phylogeny is discordant with the autosomal/mitogenome phylogeny, clustering L. viridis with the Adriatic lineage instead finding
  • Genes highly expressed in ovaries show accelerated evolutionary rates (skewed toward fast-evolving) compared to background genes finding
  • Genes strongly co-expressed in the brain network are the most skewed toward faster evolutionary rates among tissue co-expression networks finding
  • Three Z-chromosome genes (CHRDL1, ERCC6L, MTM1) with sex-linked functions show signatures of rapid evolution (omega > 1) finding
  • Whole-genome sequencing of an Adriatic lineage individual and multi-tissue (brain, heart, liver, kidney, ovary) transcriptome sequencing across the three lineages resource
  • Divergence of the L. viridis complex occurred amidst gene flow combined with accelerated evolution of tissue-specific genes, presumably contributing to reproductive isolation mechanism
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome sequencing Adriatic lineage individual (lizard) none genome sequence for phylogenomic/D-statistic analysis
Phylogenetic tree reconstruction (concatenated coding sequences) mitochondrial genes, autosomal genes, and Z-chromosome genes; L. viridis, L. bilineata, Adriatic lineage, with Podarcis muralis as outgroup none tree topology and bootstrap support, divergence time estimates
D-statistics (ABBA-BABA admixture test) autosomal and Z-chromosome gene sequences from L. viridis, L. bilineata, Adriatic, P. muralis none D-statistic and Z-score indicating gene flow between lineages
Bulk RNA-seq / transcriptome sequencing brain, heart, liver, kidney, ovary tissues from L. viridis, L. bilineata, and Adriatic lineage none gene expression levels, tissue-specific differential expression
Differential gene expression / PCA clustering (DESeq2) multi-tissue transcriptome samples none variance-stabilized expression clustering by tissue, log fold change, number of tissue-specific genes DESeq2
GO over-representation analysis tissue-specific gene sets (brain, heart, liver, kidney, ovary) none enriched biological processes per tissue
Molecular evolutionary rate analysis (Ka/Ks, omega) orthologous genes between L. viridis and L. bilineata, stratified by tissue-specific expression none nonsynonymous/synonymous substitution ratio, MWU test comparing tissue-specific vs background genes
Gene co-expression network analysis (weighted topological overlap, wTO) tissue-specific transcriptomes (brain, heart, liver, kidney, ovary) none network connectivity, tissue-specific links, node evolutionary rate (Ka/Ks) mapped onto network
Key results
  • Mitogenome and autosomal trees cluster Adriatic lineage with L. bilineata (bootstrap 64 and 100, respectively)
  • Z-chromosome tree shows L. viridis and Adriatic clustering together instead, with 100 support
  • D-statistics show significant gene flow between L. viridis and Adriatic on both autosomes and Z-chromosome Autosomes D=-0.242 (Z=-79.326); Z-chromosome D=-0.311 (Z=-13.794)
  • Ovaries have by far the highest number of tissue-specific genes, almost twice that of the second-highest tissue (brain) 1066 genes
  • Evolutionary rates (Ka/Ks) between L. viridis and L. bilineata are significantly skewed toward faster evolution for ovary-, liver-, and kidney-specific genes, but not brain or heart Ovaries p<1e-06, q<1e-06; Liver p=1.5e-05, q=6e-04; Kidneys p=0.0141, q=0.0235
  • Three genes on the Z-chromosome (CHRDL1, ERCC6L, MTM1) show signatures of rapid evolution and sex-linked functions omega > 1
  • Brain gene co-expression network has the highest weighted topological overlap and connectivity, and brain co-expressed genes are most skewed toward faster evolutionary rates
  • Autosomal divergence time estimates: L. viridis vs L. bilineata/Adriatic ~4.9 Ma; L. bilineata vs Adriatic ~2.27 Ma; Lacerta vs Podarcis ~33.5 Ma 4.9(±2.7) Ma; 2.27(±1.3) Ma; 33.5(±18.9) Ma
Key statistics
  • other D-statistic autosomes = -0.242, Z-score = -79.326 (gene flow test between L. viridis and Adriatic lineage on autosomes)
  • other D-statistic Z-chromosome = -0.311, Z-score = -13.794 (gene flow test between L. viridis and Adriatic lineage on Z-chromosome)
  • other 4.9(±2.7) Ma (divergence time between L. viridis and L. bilineata/Adriatic clade)
  • other 2.27(±1.3) Ma (divergence time between L. bilineata and Adriatic lineage)
  • pvalue <1e-06 (MWU test), q<1e-06 (BH) (ovary tissue-specific genes vs background evolutionary rate comparison)
  • count 1066 genes (number of ovary-specific expressed genes, highest among tissues tested)
  • count 12,100 genes (genes with significant tissue-specific network integration (co-expression links) across tissues)
  • other sterility of 4% in F1 females and <1% in males (hybridization experiments between L. bilineata and L. viridis showing higher female hybrid sterility (Rykena 1991))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined comparative genomics (whole-genome and multi-tissue transcriptome sequencing across three lizard lineages) with phylogenetic reconstruction (mitochondrial, autosomal, and Z-chromosome gene trees with bootstrap support), a multi-species-coalescent species-tree analysis, and four-population D-statistics with Z-scores to test for gene flow. For gene expression, DESeq2-based differential expression/tissue-specificity calling was paired with one-sided Mann–Whitney U tests comparing evolutionary-rate (Ka/Ks) distributions between tissue-specific and background gene sets across five tissues, with Benjamini–Hochberg correction applied across those five tests. Weighted topological overlap (wTO) co-expression networks were built per tissue, and GO over-representation was used to characterize gene sets. Results were reported via a summary table with p-values and BH q-values, divergence-time point estimates with a ± value, and descriptive/enrichment statistics rather than a single global model.

Replicationunclear Sample sizeSpecific counts are stated for individual analyses (e.g., n = 12 Z-chromosome gene orthologs; per-tissue gene counts in table 2), but no formal power analysis or a priori sample-size justification is described in the provided text GroupsThree lineages (L. viridis, L. bilineata, Adriatic) and five tissues (brain, heart, liver, kidney, ovary) Pairingunpaired Randomization/blindingnot stated Dispersionunclear Exact p-valuesyes Effect sizesyes Multiplicity correctionBenjamini-Hochberg (BH) FDR
Statistical tests used
Test Applied to n Assumptions
Four-population D-statistic (ABBA-BABA admixture test) with Z-scores gene flow testing among L. viridis, L. bilineata, and the Adriatic lineage, using P. muralis as outgroup, separately for autosomes and Z-chromosome (table 1) not stated
Multi-species coalescent species-tree analysis (accounting for incomplete lineage sorting) reconstruction of the unrooted species tree, compared against the autosomal gene tree topology not stated
One-sided (alternative = greater) Mann–Whitney U test comparing evolutionary rates (Ka/Ks between L. viridis and L. bilineata) of tissue-specific genes versus background genes, per tissue (table 2) gene counts per tissue as given in table 2 (e.g., brain: 316 tissue genes vs 9524 background; ovaries: 1066 vs 8774; kidneys: 27 vs 9813) not stated
Benjamini–Hochberg (BH) FDR correction applied to the five per-tissue MWU p-values in table 2 five tests, one per tissue na
DESeq2-based differential expression / tissue-specific gene calling, with variance-stabilizing transformation for PCA sample clustering (fig. 2) and identification of genes with significantly higher expression in one tissue versus others not stated
Gene Ontology (GO) over-representation analysis functional enrichment of tissue-specific gene sets (supplementary table S3) not stated
Bootstrap support values for phylogenetic trees mitochondrial, autosomal, and Z-chromosome gene trees (fig. 1C–E) na
Approaches that could also have been used
  • Evolutionary-rate distributions were compared between tissue-specific and background gene sets using five separate one-sided Mann–Whitney U tests (one per tissue), corrected jointly with Benjamini–Hochberg FDR.
    Could also: A single linear or rank-based model with tissue as a covariate (e.g., a Kruskal-Wallis test followed by pairwise post-hoc contrasts, or a mixed model with tissue as a factor) could also be used to compare rate distributions across all tissues jointly. — Pooling the comparisons into one model can borrow information across tissues and provides a unified framework for the family-wise correction alongside pairwise contrasts.
  • Divergence times between lineages are reported as a point estimate with a ± value (e.g., 4.9(±2.7) Ma) without specifying whether this reflects a standard deviation, standard error, or credible interval.
    Could also: Reporting an explicit 95% highest posterior density (HPD) interval or confidence interval, as commonly output by molecular-clock/coalescent dating software, would also convey the uncertainty range in a directly interpretable probabilistic form. — An explicitly labeled interval removes ambiguity about what the reported spread represents, which can aid comparison with divergence estimates from other studies.
  • Gene flow among the three lineages was assessed using four-population D-statistics (ABBA-BABA) with Z-scores.
    Could also: Complementary approaches such as f-branch statistics, TreeMix, or explicit demographic model fitting (e.g., with dadi or fastsimcoal2) could also be applied. — These methods can help distinguish among alternative admixture scenarios and estimate migration rates or timing, beyond establishing the presence of gene flow.
  • Tissue-specific genes were defined via a DESeq2-based differential expression workflow comparing each tissue against the others.
    Could also: A module-based approach such as WGCNA, correlating co-expression modules with tissue identity, could also be used to characterize tissue specificity. — A module-trait correlation framework can capture coordinated multi-gene expression patterns that a per-gene one-tissue-vs-rest test may not directly summarize.
  • The D-statistic results in table 1 are reported with Z-scores but without an explicit multiple-comparison adjustment across the tested lineage quartets.
    Could also: Block-jackknife-derived standard errors (a standard companion to D-statistics) combined with a multiple-comparison adjustment, if additional quartets were tested, could also be reported. — This would make explicit how the significance threshold accounts for the number of admixture configurations examined.
  • Weighted topological overlap (wTO) co-expression networks were built per tissue to characterize gene-gene relationships.
    Could also: A permutation-based null model or FDR-adjusted threshold for wTO edge significance (as implemented in permutation-based wTO frameworks) could also be used to define network links. — An explicit null-model-based threshold can help control the false discovery rate among the many pairwise gene links being evaluated.
Software: DESeq2

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-33988711

Paper: Kolora et al. 2021, Genome Biol Evol 13(8):evab109. "Accelerated Evolution of Tissue-Specific Genes Mediates Divergence Amidst Gene Flow in European Green Lizards." PMID 33988711 · PMCID PMC8382678 · DOI 10.1093/gbe/evab109.

Linked code (P16, third-party tool): github.com/ChrisCreevey/catsequences — a FASTA-alignment concatenation utility (builds a supermatrix + partition file for IQ-TREE2). NOT the authors' own analysis code; it is one step of the phylogenomics pipeline. Per brief P16, running it (or any pipeline tool) on the paper's own data is an equally valid reproduction.

Data deposited:

  • ENA PRJEB38736 — raw Illumina reads (WGS of 1 Adriatic female; RNA-seq of 5 tissues x 3 lineages).
  • Zenodo 10.5281/zenodo.3900517 (9 files, 4.2 GB): nuclear+mito genome assemblies, 3 transcriptome assemblies, orthology.tar.gz (38 MB), variants_lizards.Filt.Chr1-18.vcf.gz (3.6 GB) = filtered autosomal SNP calls.

Pipeline-derived results (candidate reproduction targets)

# Result (reported) Pipeline Inputs available? In scope?
C1 Concatenated supermatrix + species-tree topology (IQ-TREE2) catsequences → IQ-TREE2 orthology.tar.gz (per-gene alignments) YES — core, exercises the named repo
C2 Autosomal SNP count = 26,985,663 SNP calling (BWA/GATK-like) → filtered VCF deposited VCF (3.6 GB) YES — direct count vs VCF (clean QC)
C3 D-statistic (Adriatic vs L.viridis, autosomes) = -0.242, Z = -79.326 AdmixTools D-stat on SNPs deposited VCF + population labels YES if pop assignment derivable
C4 Ka/Ks: ovary-specific & brain-coexpressed genes evolve faster (MWU p<1e-6) KaKs_Calculator on SCO alignments + tissue sets orthology + tissue assignment STRETCH — needs tissue-gene sets
C5 Z-chrom genes w/ omega>1: CHRDL1, ERCC6L, MTM1 per-gene dN/dS orthology Z-chrom alignments STRETCH
C6 Tissue-specific gene counts (ovary 1066, brain 316, ...) DESeq2 on counts count matrices NOT in Zenodo list OUT — counts not deposited

Out of scope (wet-lab / not deposited): dissection/RNA extraction; divergence-time BPP MCMC (needs full alignment set + calibration, heavy); co-expression wTO/CoDiNA networks (need expression count matrices, not in the Zenodo file list).

Plan: C1 (catsequences→tree) + C2 (SNP count) are the quick floor; C3 (D-stat) and C4/C5 (Ka/Ks) are the harder push. All compute on «our HPC».

BLOCKER (2026-06-18): «infra» per-user quota exceeded

The shared account's «infra» quota is full (write of even a 3-byte file fails with "Disk quota exceeded"; account holds ~14 TB of the operator's real research data in projects/scratch/etc., which must NOT be deleted). Freed ~5 GB of my own regenerable conda package caches; quota accounting is async and had not caught up at last check. Retrying periodically; downloads + compute resume once headroom is available.

C1
Reported
IQ-TREE2 species tree ((L.viridis,(L.bilineata,Adriatic)),Podarcis) from catsequences supermatrix
Reproduced
catsequences (59fd7c5) built 4-taxon supermatrix (1,281,668 aa, 2,806 SCO partitions); IQ-TREE2 v3.1.2 JTT+F+I+G4 UFBoot=100 tree = (adriatic,(bilineata,podarcis),viridis) -> split {adriatic,viridis}|{bilineata,podarcis}
partial
C2
Reported
26,985,663 autosomal bi-allelic SNPs
Reproduced
deposited VCF: 79,072,808 records / 65,511,188 SNPs / 64,507,788 biallelic / 38,936,169 biallelic-no-missing
did not match
C7
Reported
1,497 autosomal single-copy orthologs (viridis vs bilineata)
Reproduced
ProteinOrtho table (90,238 groups): 11,313 single-copy viridis&bilineata; 2,806 single-copy all-4
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 23/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴

The P16 tool catsequences (commit 59fd7c5) compiles and runs correctly on the paper's own data, building the 2,806-SCO / 1,281,668-aa supermatrix and a fully-supported IQ-TREE2 tree — so the tool reproduces. However none of the three checkable headline values reproduces 1:1: the tree topology flips sister pairs (on de-novo-reconstructed alignments, not like-for-like), the SNP count (26,985,663) is not reachable by any standard filter on the deposited 79.07M-record VCF, and the 1,497 autosomal SCO count is an undeposited autosomal subset of the 11,313/2,806 ProteinOrtho counts. The root cause is on the data-availability/authors' side — the open Zenodo deposit sits one pipeline stage upstream and lacks the filtered VCF, supermatrix MSAs and gene→chromosome mapping needed to derive the figures — and the central biological claims (C3–C6) could not be tested at all. Fabrication concern is LOW; the values are plausible but not independently verifiable, which is itself the auditable finding.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

282.3 k
tokens (I/O) · 14 M incl. cache
87 min
runtime · 1.56 CPU-h
1.9 GB
peak RAM
3
HPC jobs
hummel
machine