Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The formation of avian montane diversity across barriers and along elevational gradients.

Nat Commun · 2022
L1 75/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

SCOPED reproduction of pmid-35022441 (Pujolar et al. 2022, Nat Commun 13:268) — one focal pair, Ficedula hyperythra Buru(5) vs Seram(5), as proof-of-pipeline. The full Methods pipeline (Trimmomatic -> BWA-MEM -> BCFtools 1.9 mpileup|call (48 chunks, 34.5M raw records) -> filter DP10-100/QUAL>=20/indel-prox -> vcftools 0.1.14 biallelic-all-individuals = 5,388,792 SNPs -> vcf2sfs) ran end-to-end on the paper's OWN raw reads on «our HPC». S1 EXACT (vcf2sfs on shipped example: 197 ind/11 pops/2265 SNPs/2D 39x29). S2 EXACT-artifact (2D-SFS 11x11 Buru x Seram + 1D per pop; 2D unfolded re-summed = 5,388,792 = n_snps, independently checked; canonical L-shape; this is the fastsimcoal2 input class — no numeric SFS is published to value-match). S4 PARTIAL: Weir&Cockerham WEIGHTED FST = 0.114 (standard estimator) / per-site MEAN = 0.068 vs paper 0.04 — same LOW-differentiation magnitude-class (the paper's qualitative result) but a real ~3x gap on the standard weighted estimator, consistent with unspecified ref/N/filter/version details. C1-N PARTIAL (observed 231 runs/20 sp > reported 226 ind/18 sp; deposit complete). fastsimcoal2 Table 1 (S3) = OUT OF SCOPE. Overall=partial: the pipeline reproduces cleanly and the SFS artifact is exact, but the single hard numeric claim (FST) is only a magnitude-class match. NOT attempted: full 18-species Table 1, de novo assembly, PSMC, fastsimcoal2 demographic inference, and a value-level SFS comparison (none published).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 4ba6db3e3eb9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-26
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether barrier strength and a species' elevational distribution predict population differentiation, historical gene flow, and effective population size changes in montane birds, in order to distinguish among three competing explanations for the origin of montane populations: lowland-to-highland range shifts, mountain-to-mountain colonization with ongoing gene flow, or in situ parapatric speciation along elevational gradients.

Core claims
  • Genetic differentiation is generally greater across the oceanic barrier separating Buru and Seram than across the lowland barrier separating New Guinea mountain ranges. finding
  • Differentiation is greater between montane species' populations than between lowland species' populations in New Guinea, but not in the Moluccas. finding
  • Pleistocene climate oscillations dramatically influenced demographic histories of all species, most pronounced in smaller geographic areas (e.g., Huon). finding
  • There is negligible genetic differentiation between individuals sampled at different elevations along the same slope, arguing against parapatric/sympatric speciation along elevational gradients. finding
  • Even the most divergent montane taxon pairs continue to experience gene flow across barriers, indicating dispersal between montane regions is important for assemblage formation. finding
  • Colonization direction between mountain ranges is predominantly from the larger central range to smaller outlying ranges. finding
  • Assembled 15 de novo genomes and generated whole-genome resequencing data for 18 Indo-Pacific forest bird species across elevational gradients. method
  • Genetic differentiation (F_ST) is significantly positively correlated with a species' altitudinal floor among New Guinean taxa. finding
Experimental setups
Assay System Perturbation Readout Platform
whole-genome de novo sequencing and assembly 18 forest bird species (multiple genera, New Guinea and Moluccas) none genome size, scaffold number, N50, BUSCO completeness
whole-genome resequencing population pairs of 18 bird species from Buru/Seram and New Guinea (Wilhelm, Huon, Scratchley) none SNP variants, kinship coefficients
F_ST calculation (fixation index) population pairs across islands, mountain ranges, and elevations none degree of genetic differentiation between populations
Principal Component Analysis on standardized pairwise F_ST individuals from each population pair none visualization of population structure
STRUCTURE admixture/clustering analysis individuals from each population pair none inferred number of ancestry clusters (K) and individual assignment proportions STRUCTURE
Pairwise Sequentially Markovian Coalescent (PSMC) single individual per species (de novo genome) and resequenced individuals none effective population size (Ne) trajectory over time since ~10 Mya
demographic modeling (fastsimcoal2) population pairs none current Ne estimates, historical gene flow rates fastsimcoal2
correlation analysis New Guinean and Moluccan taxa none relationship between F_ST and species' altitudinal floor
Key results
  • Congeneric species pairs used as validation showed clearly separated F_ST values from 0.08 (Ptiloprora) to 0.20 (Ficedula) 0.08-0.20
  • Five of six Buru/Seram species pairs show high inter-island F_ST (Ceyx lepidus 0.16, Thapsinillas affinis 0.15, Ficedula buruensis 0.13, Pachycephala macrorhyncha 0.09); Ficedula hyperythra lower at 0.04 F_ST 0.04-0.16
  • No significant F_ST difference between individuals sampled at different elevations within Buru/Seram or New Guinea slopes
  • New Guinea Wilhelm-Huon montane species show elevated F_ST (Melipotes fumigatus/ater 0.08, Paramythia montium 0.09, Ifrita kowaldi 0.07) while lowland species show none (Toxorhamphus novaeguineae, Melilestes megarhynchus F_ST=0.00) F_ST 0.00-0.09
  • F_ST significantly positively correlated with altitudinal floor across New Guinean taxa; remained significant excluding lowland taxa r=0.83, p<0.001; r=0.70, p=0.022
  • PSMC-inferred initial ancestral population sizes were mostly ~300,000-400,000 individuals, lowest in Ifrita kowaldi (~200,000) and highest in Toxorhamphus novaeguineae and Peneothello sigillata (~1,000,000) 200,000-1,000,000 individuals
  • Huon populations show demographic decline to a historical low in the Late Pleistocene while Wilhelm populations remained stable, in most species except three (Melanocharis versteri, Toxorhamphus novaeguineae, Melilestes megarhynchus)
  • Two pairs of close relatives (half-siblings) identified via kinship analysis and excluded from downstream demographic analyses kinship coefficients 0.144 and 0.135
Key statistics
  • correlation r=0.83, p<0.001 (F_ST vs altitudinal floor, all New Guinean taxa except Melipotes fumigatus/ater)
  • correlation r=0.70, p=0.022 (F_ST vs altitudinal floor excluding two lowland taxa)
  • fold_change F_ST 0.08-0.20 (congeneric species pairs used as differentiation benchmark)
  • fold_change F_ST 0.04-0.16 (Buru vs Seram population pairs)
  • other kinship coefficient 0.144 (Pachycephala schlegelii individuals A117/A118, indicative of half-siblings)
  • other kinship coefficient 0.135 (Ifrita kowaldi individuals D116/D117, indicative of half-siblings)
  • other BUSCO completeness 66.7%-86.8% (genome assembly completeness across 18 species)
  • count ~200,000-1,000,000 individuals (PSMC-inferred initial ancestral effective population sizes across species)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This comparative population genomics study used whole-genome resequencing data from 18 Indo-Pacific bird species to examine how barrier strength and elevational distribution predict genetic differentiation, gene flow, and effective population size dynamics. Genetic differentiation was quantified via FST, visualised with PCA and STRUCTURE, and correlated with species' altitudinal floors using Pearson correlation. Demographic history was reconstructed per individual using PSMC, and model-based demographic inferences (including gene flow) were conducted with fastsimcoal2.

Replicationbiological Sample sizeSample sizes reported per species and locality (figures/supplementary); no formal power analysis or a priori sample-size calculation described GroupsPopulations across hard oceanic barrier (Buru vs Seram), soft lowland barrier (Mount Wilhelm vs Huon), additional Central Range site (Mount Scratchley), and within-population elevational gradients Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
FST (fixation index) calculation All pairwise population comparisons across 18 species (Buru vs Seram; Mount Wilhelm vs Huon; ± Mount Scratchley; elevational within-population comparisons) not stated
Principal Component Analysis (PCA) on standardised pairwise FST Visualisation of population structure for all species; congeneric species validation na
Bayesian clustering (STRUCTURE) Ancestry component inference for all species; K selection per species not stated
Pairwise kinship coefficient analysis Pre-analysis check for related individuals within populations (e.g. Pachycephala schlegelii, Ifrita kowaldi) not stated
Pearson correlation (r) FST vs altitudinal floor across New Guinean taxa (all taxa: r=0.83, p<0.001; excluding lowland taxa: r=0.70, p=0.022) not explicitly stated (approximately 12–13 New Guinean species/population pairs) not stated
Pairwise Sequentially Markovian Coalescent (PSMC) with bootstrap resampling (100 replicates) Demographic history inference for all 18 species; comparison of population trajectories across sampling localities one individual per species/population for primary PSMC; resequenced individuals for comparative PSMC not stated
Approaches that could also have been used
  • Pearson correlation (r) was used to relate FST to altitudinal floor across species
    Could also: Spearman rank correlation could also be used for this relationship — FST is a bounded statistic (0–1) and may not be normally distributed across species; Spearman's rho makes no distributional assumptions and is robust to outliers, which is relevant when n is small (~12 species)
  • Multiple pairwise FST comparisons were made across 18 species and several population pairs without a stated multiplicity correction
    Could also: A false discovery rate procedure (e.g. Benjamini-Hochberg) or family-wise error rate control (e.g. Bonferroni) could also be applied across the family of pairwise tests — Explicit correction would further quantify the probability of false positives when many comparisons are evaluated simultaneously, which is a common consideration in multi-species comparative genomics
  • PSMC was used for demographic inference, analysing one diploid genome at a time
    Could also: SMC++ or MSMC2 could also be used, as both accommodate multiple haplotypes per population — Multi-haplotype methods can improve resolution of recent demographic history (post ~50 Kya) and allow direct comparison of population-level trajectories, complementing per-individual PSMC inference
  • STRUCTURE was used for ancestry clustering and K selection
    Could also: ADMIXTURE could also be used for the same purpose — ADMIXTURE uses a similar model but employs cross-validation error for K selection and is computationally faster with genome-scale SNP data; results from both tools are commonly compared to assess robustness
  • FST was used as the primary metric of genetic differentiation
    Could also: Standardised metrics such as G'ST or Jost's D could also be calculated alongside FST — FST can be sensitive to within-population heterozygosity levels; standardised alternatives facilitate comparison across loci or species with differing levels of diversity, which is relevant in a multi-species design
  • Within-population elevational comparisons relied on pairwise FST between elevation-stratified subgroups
    Could also: Analysis of Molecular Variance (AMOVA) could also partition genetic variance hierarchically (among elevations within populations vs. among populations) — AMOVA provides a formal test of the proportion of variance attributable to elevation while accounting for the nested structure of the sampling design, and produces a significance test for each hierarchical level
Software: PSMC · STRUCTURE · fastsimcoal2 · BUSCO

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35022441

Paper: Pujolar et al. 2022, Nat Commun 13:268. "The formation of avian montane diversity across barriers and along elevational gradients." DOI 10.1038/s41467-021-27858-5 · PMCID PMC8755808

Tool/repo named for reproduction: https://github.com/shenglin-liu/vcf2sfs (third-party R tool — generates site frequency spectra (SFS) from a VCF + popmap; writes fastsimcoal2 / dadi input. Per BRIEF P16, applying this third-party tool to the paper's own data is a fully valid reproduction.)

Data: resequencing reads sra:PRJNA580350; de novo reference genomes sra:PRJNA637995. No intermediate or derived data is deposited — see below.


The paper's computational pipeline (as described in Methods)

raw PE reads (PRJNA580350)
   │  FastQC + FASTX-Toolkit (QC/trim)
   ▼
BWA v0.7.1  ── map to de novo reference (PRJNA637995, per-species)
   ▼
BCFtools v1.9  mpileup | call               ── variant calling
   ▼
filter: DP<10 or >100 removed; Phred<20 removed; sites near indels removed
   ▼
vcftools v0.1.14  ── retain biallelic SNPs present in ALL individuals
   ▼
vcf2sfs (R, github.com/shenglin-liu/vcf2sfs)  ── build per-population / 2D / 3D SFS
   ▼
┌─────────────────────────────┐     ┌──────────────────────────┐
│ fastsimcoal2 (8 demographic │     │ PSMC (historical Ne over │
│ models; µ=3e-9 [2e-9 Ceyx]; │     │ time; Pleistocene)       │
│ gen time 2 yr) → Table 1    │     │ → Fig. (PSMC trajectories)│
│ Ne, T_div, M, Nm            │     │                          │
└─────────────────────────────┘     └──────────────────────────┘

The fastsimcoal2 model code (.tpl/.est) is only in the paper's Supplementary Methods (docx), not in any code repository.


IN SCOPE (pipeline-derived; we attempt)

# Reported result Pipeline producing it Feasibility
S1 SFS generation works — vcf2sfs turns a VCF+popmap into a (multi-D) SFS vcf2sfs R tool EASY — validate on the repo's shipped example.vcf (2265 SNPs, 197 ind, 11 pops)
S2 Per-species SFS for one focal species pair from the paper's own reads BWA → BCFtools → vcftools → vcf2sfs HEAVY — one species pair (e.g. Ficedula hyperythra Buru/Seram, 10 ind) on «our HPC»
S3 fastsimcoal2 Ne / T_div / M / Nm (Table 1) for that focal pair fastsimcoal2 on the S2 SFS, 8 models, model selection VERY HEAVY / stretch — stochastic; compare to Table 1 ranges
S4 FST between the two populations of the focal species vcftools --weir-fst-pop on the S2 VCF MEDIUM — direct numeric compare to reported FST

OUT OF SCOPE (not pipeline-derived, or infeasible at study cost)

  • Full 18-species × all-comparisons Table 1 (entire ~2.6 TB dataset, 13+ species pairs, each needing its own de novo reference + 8-model fastsimcoal + bootstraps) — not attempted; we attempt a single representative species pair as proof.
  • De novo genome assembly (PRJNA637995) — wet-lab/assembly result, not re-run.
  • Phylogenetics / divergence dating beyond fastsimcoal (external, manual).
  • Field/wet-lab: sampling, sequencing, taxonomy.

Hard blockers known up front (recorded honestly)

  1. No deposited intermediates. VCFs, SFS, and fastsimcoal .tpl/.est/output are NOT deposited (Data-availability lists only the two SRA accessions; Code availability points to Supplementary Methods only). Every reproduced value must be rebuilt from raw reads → cannot be checked against an authors' intermediate.
  2. Reference matching. Reseq reads must be mapped to a per-species de novo reference; the mapping of each reseq species → its assembly in PRJNA637995 is not tabulated and must be inferred.
  3. fastsimcoal2 is stochastic and Table 1 reports wide ranges (e.g. current Ne 80,000–910,000). "Reproduction" of S3 can only mean landing within the reported range/order of magnitude, graded partial at best.
  4. All heavy compute is on «our HPC» (BRIEF rule 1); when the tunnel is do
S1
Reported
vcf2sfs (github.com/shenglin-liu/vcf2sfs) generates an SFS from a VCF+popmap
Reproduced
ran on repo example.vcf (commit a932aa5b): 197 ind, 11 pops, 2265 SNPs, 1D total 2265, 2D-SFS 39x29 total 2265 — internally consistent
exact
S2
Reported
per-population / 2D SFS generated from the final variant calls (no numeric SFS published)
Reproduced
vcf2sfs built 2D-SFS 11x11 (Buru x Seram, unfolded+folded) + per-pop 1D SFS from the 5,388,792-SNP all-individual primary callset; 2D unfolded re-summed = 5,388,792 = n_snps (independently checked this session); 1D Buru 4,102,308 / Seram 4,071,987 segregating; canonical L-shape = fastsimcoal2 input class
exact
S4
Reported
FST = 0.04 (F. hyperythra Buru vs Seram)
Reproduced
Weir&Cockerham WEIGHTED Fst = 0.11420 (standard genome-wide estimator); per-site MEAN Fst = 0.06788; from 5,388,792 all-individual biallelic SNPs (DP10-100), 5 Buru + 5 Seram, ref GCA_040366035. Same low-differentiation magnitude-class but ~2.9x gap on the standard weighted estimator (~1.7x on the mean).
partial
C1-N
Reported
226 individuals / 18 species (PRJNA580350)
Reproduced
231 runs / 231 BioSamples / 20 species observed in ENA; focal F. hyperythra = 10 runs
partial
S3
Reported
fastsimcoal2 demographic parameters Ne / T_div / M / Nm (Table 1)
Reproduced
not attempted (out of scope: downstream of the SFS, stochastic, .tpl/.est only in Supplementary docx)
m.public.grade.out-of-scope

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

115.9 k
tokens (I/O) · 4.9 M incl. cache
24 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.