Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Environmental selection overturns the decay relationship of soil prokaryotic community over geographic distance across grassland biotas.

Elife · 2022
L1 57/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
57/100
Reproducibility score
1.0 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 17% of all assessed papers rank 965 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

P16 third-party-tool reproduction of the eLife distance-decay ecology paper, run on the authors' shipped (already-7500-rarefied) OTU table with R vegan 2.7.1 / geosphere / minpack.lm on «our HPC» (SLURM 2244967, n093). 18/20 claims attempted; counts exact 2, within-tol 3, partial 11, mismatch 2, not-attempted 2. Data descriptors reproduce strongly (samples 258 exact, rarefaction 7500 exact, OTUs 11037 vs 11063 within 0.24%). The paper's CORE THESIS reproduces cleanly: environmental-distance Mantel r (top 0.404) far exceeds geographic Mantel r (top 0.124) - community turnover is environmentally, not geographically, driven. Partial-Mantel Table-1 pattern reproduces qualitatively (all env factors significant positive; subsoil DON r=0.401 near-exact), though several topsoil r's run ~0.05-0.17 higher than reported, likely due to under-specified env-distance construction. DDR U-shape tipping points are close (~1810-1836 vs 1760-1920 km); temperate-topsoil R2 0.119 vs 0.129. Stretch: Sloan neutral model (C19) gives good fits (R2 0.73-0.84) and reproduces the temperate>alpine immigration ordering but ~1.5-2.3x higher absolute m (partial). Two claims flagged for human review: C13 (paper r=-0.544 with p=1.000 is statistically odd; we get r=+0.142, p=0.001 - possible mis-transcription/fabrication) and the C19 magnitude gap. NOT attempted: C18 betaNTI (no phylogeny in repo) and C20 SEM (spec not given). All provisional - a human reviewer signs off in AUDIT.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 68
    assessed: 2026-06-21 ⛓ 7e9e534cce61
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether soil prokaryotic community similarity decreases with geographic distance within and across grassland biotas (distance-decay), and whether turnover rates differ by biota (temperate vs alpine) and soil depth (top- vs subsoil).

Core claims
  • Prokaryotic community similarity follows a significant U-shape relationship over geographic distance up to 4000 km, decreasing within biotas but increasing across biotas after a tipping point of 1760-1920 km finding
  • The U-shape pattern arises because environmental heterogeneity disparities decrease over geographic distance when comparing across biotas mechanism
  • Prokaryotic community similarity still decreases with environmental distance even across biotas, unlike the geographic distance pattern finding
  • Homogeneous environmental selection (a deterministic process) dominates prokaryotic community assembly, accounting for >84% of |βNTI| values mechanism
  • Short-term environmental heterogeneity (e.g., dissolved nutrients) follows the same U-shape pattern spatially as the prokaryotic community finding
  • Environmental selection overturns the classically accepted distance-decay relationship for microbes when comparing across distinct biotas on large scales finding
  • Plant community similarity also exhibits a significant U-shape relationship over geographic distance, paralleling the prokaryotic pattern finding
  • Prokaryotic immigration rates are lower in alpine than temperate biota and higher in topsoil than subsoil finding
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA amplicon/OTU sequencing soil (top- 0-5cm and subsoil 5-20cm) from alpine (Qinghai-Tibet Plateau) and temperate (Inner Mongolia Plateau) grasslands, China none (natural gradient across 258 samples) prokaryotic community composition/similarity (Bray-Curtis-type) vs geographic distance
Plant community survey grassland vegetation, alpine and temperate biotas none plant community similarity vs geographic distance
Soil physicochemical and climate measurement soil top-/subsoil, alpine and temperate grasslands none long-term (MAP, MAT, pH, SOC, TN, TP) and short-term (SWC, AP, DOC, DON, NH4+, NO3-) environmental variables, Bray-Curtis environmental similarity
Mantel and partial Mantel tests soil prokaryotic community, alpine and temperate biotas none correlation between community dissimilarity and geographic/environmental distance
Null model and βNTI (beta-nearest taxon index) analysis soil prokaryotic community, all sites and within/across biotas none relative contribution of deterministic vs stochastic assembly processes
Neutral community model (immigration rate estimation, Hubbell algorithm) soil prokaryotic community, top-/subsoil, alpine and temperate biotas none immigration rate (m)
Structural equation modeling (SEM) soil prokaryotic and plant community, geographic/environmental variables none causal pathways linking geographic distance, environmental variables, plant and prokaryotic community dissimilarity
Key results
  • Significant U-shape relationship for prokaryotic community similarity over geographic distance across all sites, with tipping points at 1760-1920 km R2=0.161 (topsoil), R2=0.114 (subsoil), p<0.001
  • Across biotas, prokaryotic similarity increased significantly with geographic distance after the tipping point slope=0.007, R2=0.199 (topsoil); slope=0.006, R2=0.134 (subsoil)
  • Within alpine and temperate biotas (topsoil), prokaryotic similarity decreased with geographic distance R2=0.034 alpine; R2=0.129 temperate, p<0.001
  • Prokaryotic similarity decreased significantly with environmental distance across biotas long-term turnover -0.161/-0.130 (top/sub); short-term -0.191/-0.175 (top/sub), p<0.001
  • Deterministic homogeneous selection dominated community assembly, more so in temperate than alpine biota >84% overall; 96.63%/94.04% temperate vs 92.91%/87.47% alpine (top/sub)
  • Plant community similarity also showed a U-shape pattern over geographic distance R2=0.071, p<0.001, tipping point 1858 km
  • Prokaryotic immigration rates were lower in alpine than temperate biota and higher in top- than subsoil alpine 0.159±0.008/0.146±0.008 vs temperate 0.261±0.010/0.246±0.009 (top/sub), p<0.01
Key statistics
  • count 11,063 OTUs from 258 samples (total prokaryotic diversity detected across all sites)
  • correlation R2=0.161, p<0.001 (U-shape fit, topsoil prokaryotic similarity vs geographic distance, all sites)
  • correlation R2=0.114, p<0.001 (U-shape fit, subsoil prokaryotic similarity vs geographic distance, all sites)
  • other tipping points 1760-1920 km (geographic distance at which similarity trend reverses)
  • fold_change slope=0.007 (topsoil), slope=0.006 (subsoil) (turnover rate of prokaryotic similarity increase across biotas)
  • pvalue p<0.01 (difference in immigration rates between alpine and temperate biota)
  • correlation DON r=0.398-0.401, p=0.001 (partial Mantel, strongest short-term environmental driver of prokaryotic dissimilarity across biotas)
  • mean immigration rate 0.261±0.010 (temperate topsoil) vs 0.159±0.008 (alpine topsoil) (Hubbell neutral model immigration rate comparison)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This cross-sectional observational study characterized Bray-Curtis-based prokaryotic community similarity across 258 grassland soil samples (topsoil 0–5 cm and subsoil 5–20 cm) spanning up to 4000 km in alpine and temperate Chinese biotas. The main finding—a U-shaped similarity–distance relationship—was quantified with polynomial regression; linear regression described distance-decay within and across biotas. Mantel and partial Mantel tests assessed associations between community dissimilarity and geographic or environmental distance matrices, while β-nearest taxon index (βNTI) null-model analysis and Hubbell-algorithm immigration rates quantified deterministic versus stochastic assembly. A structural equation model (SEM) integrated direct and indirect effects of long-term and short-term environmental variables on community dissimilarity.

Replicationbiological Sample size258 soil samples stated; no power calculation or formal sample-size justification mentioned Groupsalpine vs. temperate biota; topsoil (0–5 cm) vs. subsoil (5–20 cm); within-biota vs. across-biota pairwise site comparisons Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Polynomial (quadratic/binomial) regression U-shaped relationship between prokaryotic community similarity and geographic distance across all sites in topsoil (R²=0.161) and subsoil (R²=0.114) 258 samples (yielding up to ~33,153 pairwise distances); exact pairwise n not stated not stated
Linear regression (ordinary least squares on distance matrices) Distance-decay relationships of prokaryotic community similarity vs. geographic or environmental distance within and across biotas; slopes reported as turnover rates pairwise comparisons from 258 samples; subset n not stated not stated
Mantel test Correlation between prokaryotic composition dissimilarity and geographic distance or altitude within and across biotas (Supplementary file 1) pairwise comparisons; exact n per subset not stated not stated
Partial Mantel test Partial correlation of prokaryotic dissimilarity with individual environmental variables (MAP, MAT, pH, SOC, TN, TP, SWC, AP, DOC, DON, NH4+, NO3−, geographic variables) controlling for remaining predictors, across biotas (Table 1; Supplementary file 2) pairwise comparisons across alpine×temperate sites; exact n not stated not stated
βNTI (β-nearest taxon index) null-model analysis Classification of pairwise community assembly into deterministic (|βNTI|>2) vs. stochastic (|βNTI|<2) processes within and across biotas (Figure 3) 258 samples; permutation count not stated not stated
Null-model permutation test (observed vs. permutated community similarity) Test that observed community similarity exceeds null-model expectation in alpine, temperate, and all sites (Figure 3—figure supplement 1; p<0.001) 258 samples not stated
Immigration rate estimation based on Hubbell neutral theory algorithm Comparison of prokaryotic immigration rate (m) between alpine and temperate biotas and between topsoil and subsoil (Figure 3—figure supplement 2); group differences reported as p<0.01 not stated not stated
Structural equation modeling (SEM) Path analysis of direct and indirect effects of long-term environmental variables, short-term environmental variables, and plant community dissimilarity on prokaryotic community dissimilarity across biotas (Figure 4, Figure 4—figure supplement 1) not stated not stated
Approaches that could also have been used
  • Partial Mantel tests were applied sequentially for each environmental variable to identify drivers of community dissimilarity across biotas
    Could also: Multiple regression on distance matrices (MRM) or distance-based redundancy analysis (db-RDA) with permutation testing could also model community dissimilarity as a function of multiple predictor distance matrices simultaneously — MRM and db-RDA estimate partial effects of all predictors in a single model, which avoids the accumulation of Type I error from sequential partial Mantels and provides an integrated variance-partitioning framework across geography and environment
  • The U-shaped similarity–distance relationship was modeled with a polynomial (quadratic) regression, with tipping points of 1760–1920 km identified descriptively
    Could also: A piecewise (segmented) linear regression or a statistical change-point model could also characterize the transition point — Piecewise regression explicitly estimates the breakpoint location and its confidence interval and formally tests whether the two-segment slopes differ from each other, providing inferential precision around the tipping-point estimate
  • Community assembly was characterized using βNTI, which categorizes pairwise comparisons into deterministic vs. stochastic bins using a fixed ±2 threshold
    Could also: A combined βNTI + Raup-Crick Bray-Curtis (RC_Bray) framework could also be applied to further partition stochastic pairwise comparisons into dispersal limitation, homogenizing dispersal, and ecological drift — The combined framework distinguishes among multiple stochastic mechanisms that |βNTI|<2 groups together, providing a more resolved picture of the relative importance of drift versus dispersal in community assembly
  • Structural equation modeling (SEM) was used to estimate direct and indirect effects of environmental variable sets on prokaryotic community dissimilarity
    Could also: Variation partitioning analysis (VPA) via partial db-RDA could also decompose total explained variation in community composition into fractions uniquely attributable to long-term environment, short-term environment, geography, and their shared components — VPA provides explicit percentage-variance partitions and joint (shared) fractions, which complements SEM path coefficients by showing how much variation each predictor set explains independently versus in common with others
  • Immigration rates were compared between biotas and soil layers with results reported as p<0.01, without identifying the statistical test used for the group comparison
    Could also: A two-sample t-test, Wilcoxon rank-sum test, or a linear mixed model with biota and soil layer as fixed effects could also be stated explicitly for these group comparisons — Naming the test, reporting the test statistic and degrees of freedom, and providing an effect size alongside the p-value allows readers to assess precision, power, and reproducibility of the group difference
  • Spread around immigration-rate means was reported with a ± symbol whose meaning (SD, SEM, or CI) was not labeled in the provided text
    Could also: Explicitly labeling the spread metric as SD (biological variability) or SEM/95% CI (precision of the mean) in both table and figure legends could also be used, and is recommended by many reporting guidelines — SD and SEM convey different things—biological variation versus estimation precision—and distinguishing them helps readers interpret whether the ± value reflects how variable the samples are or how precisely the mean was estimated
Software: not stated

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35073255

Paper: Zhang et al. 2022, eLife 11:e70164. "Environmental selection overturns the decay relationship of soil prokaryotic community over geographic distance across grassland biotas." PMID 35073255 / PMCID PMC8828049 / DOI 10.7554/elife.70164.

Study: 258 soil samples from 129 sites across two grassland biotas — alpine (Qinghai-Tibet Plateau) and temperate (Inner Mongolia). 16S rRNA V4-V5 amplicon sequencing (Illumina MiSeq). Tests distance-decay relationships (DDR) of prokaryotic community similarity over geographic vs environmental distance, and community assembly processes.

Code & data availability

  • Repo: https://github.com/zhangbiao1989/Support-files (commit 254b3d1fb9b81fae868a594f4fc2c9c3583021d2, main, MIT, last push 2022-02-07). ⚠️ The repo ships DATA, not analysis code:
    • OTU-table.xlsx (9.85 MB) — the 97%-OTU count matrix (the bioinformatic output)
    • Environment_variables.xlsx (45 KB) — site env metadata (SOC, TN, DOC, DON, NH4+, coords)
    • Plant composition.xlsx (124 KB) — plant community composition per site
  • Raw reads: SRA PRJNA729210 (16S amplicons; not needed to reproduce the downstream ecology — the OTU table is the upstream pipeline's shipped output).

Because no authors' analysis scripts are shipped, this is a P16 third-party-tool reproduction: apply standard community-ecology tools (R vegan/picante/ape, or python scikit-bio) to the paper's own deposited OTU table + env vars, following the Methods (Bray-Curtis / community similarity, Mantel & partial Mantel, linear DDR regressions). This is explicitly as valid as reproducing with the authors' own code.

IN SCOPE (pipeline/statistics-derived, reproducible from shipped data)

group result claims
Data descriptors sample count, OTU count, rarefaction depth C1, C2, C3
Distance-decay DDR slopes/R2 over geographic distance (overall + per biota, top/subsoil) C4–C7
Partial Mantel (Table 1) community vs each env factor | geography, and vs geography | env C8–C16
Environmental decay community similarity turnover over environmental distance C17

STRETCH (in scope but heavier; need phylo tree / null models / SEM)

group result claims blocker
Community assembly homogeneous-selection % via betaNTI C18 needs OTU phylogenetic tree (not in repo — must build from rep seqs / SRA) + picante null model
Neutral model immigration rate m (Sloan) C19 Sloan/Burns neutral fit
SEM standardized path coefficients C20 structural equation model (lavaan/piecewiseSEM); model spec not fully given

OUT OF SCOPE (not computationally reproducible here)

  • Wet-lab soil chemistry assays that produced the env variables (SOC/TN/DOC/DON/NH4+ measurements) — these are lab measurements, taken as given inputs.
  • Field sampling, DNA extraction, library prep — wet-lab.
  • Upstream raw-read→OTU bioinformatics (USEARCH v8.0.1623 / FLASH / Mothur v1.27, silva.nr_v128) is attemptable from SRA but the OTU table is already shipped; we treat the shipped OTU table as the pipeline output and verify N/depth against it, rather than re-running the full read-processing (would be a separate, larger effort — noted as a possible deeper pass if the downstream results reproduce cleanly).

Tools planned

R 3.5.0-era stack per Methods: vegan (vegdist, mantel, mantel.partial), base lm for DDR slopes, readxl to load the xlsx. Compute on «our HPC» (SLURM); data on «infra».

C1_n_samples
Reported
258 (alpine 128, temperate 130; 129 sites)
Reproduced
258 (topsoil 129 + subsoil 129)
exact
C2_n_otus
Reported
11,063 OTUs
Reproduced
11,037 nonzero union across 258 (11,690 rows)
within tolerance
C3_rarefaction_depth
Reported
7,500 seqs/sample
Reproduced
7,500 (min=median=max)
exact
C4_ddr_overall_topsoil
Reported
U-shape R2=0.161; tipping ~1920 km
Reproduced
quad R2=0.090; tipping 1836 km
partial
C5_ddr_overall_subsoil
Reported
U-shape R2=0.114; tipping ~1760 km
Reproduced
quad R2=0.060; tipping 1811 km
partial
C6_ddr_alpine_topsoil
Reported
R2=0.034, p<0.001
Reproduced
R2=0.0035, p=0.0099 (sign matches)
partial
C7_ddr_temperate_topsoil
Reported
R2=0.129, p<0.001
Reproduced
R2=0.119, p=2e-59
within tolerance
C8_pmantel_topsoil_SOC
Reported
r=0.231, p=0.001
Reproduced
r=0.338, p=0.001
partial
C9_pmantel_topsoil_TN
Reported
r=0.151, p=0.003
Reproduced
r=0.321, p=0.001
did not match
C10_pmantel_topsoil_DOC
Reported
r=0.307, p=0.001
Reproduced
r=0.370, p=0.001
partial
C11_pmantel_topsoil_DON
Reported
r=0.398, p=0.001
Reproduced
r=0.324, p=0.002
partial
C12_pmantel_topsoil_NH4
Reported
r=0.223, p=0.001
Reproduced
r=0.304, p=0.001
partial
C13_pmantel_topsoil_geo
Reported
r=-0.544, p=1.000 (NS)
Reproduced
r=0.142, p=0.001
did not match
C14_pmantel_subsoil_DON
Reported
r=0.401, p=0.001
Reproduced
r=0.401, p=0.001
within tolerance
C15_pmantel_subsoil_DOC
Reported
r=0.206, p=0.001
Reproduced
r=0.251, p=0.005
partial
C16_pmantel_subsoil_SOC
Reported
r=0.120, p=0.010
Reproduced
r=0.202, p=0.012
partial
C17_envdist_turnover
Reported
turnover -0.291(top)/-0.278(sub)
Reproduced
env-Mantel r top=0.404, sub=0.279 (env >> geo)
partial
C18_betaNTI_homog_selection
Reported
>84%; alpine-top 92.91%
Reproduced
NOT ATTEMPTED (no phylo tree / rep seqs in repo)
m.public.grade.not-attempted
C19_immigration_m
Reported
alpine .159/.146; temperate .261/.246
Reproduced
alpine .249/.230; temperate .590/.468 (ordering temperate>alpine holds; R2=0.73-0.84)
partial
C20_sem_geo
Reported
r=0.388(top)/0.320(sub)
Reproduced
NOT ATTEMPTED (SEM spec not given)
m.public.grade.not-attempted

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 57/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

251.1 k
tokens (I/O) · 12.7 M incl. cache
44 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.