Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Estimating biodiversity across the tree of life on Mount Everest's southern flank with environmental DNA.

iScience · 2022
L1 65/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
65/100
Reproducibility score
0.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 27% of all assessed papers rank 843 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PRIMARY result reproduced with fresh end-to-end «our HPC» compute. C1 (the repo-named SingleM v0.13.2 pipeline, run on all 20 WGS runs of PRJNA629845 with the default 14 ribosomal-protein spkgs + the contemporaneous Greengenes-16S spkg, paper filter = drop OTU hits <10 reads): 41 bacterial orders / 11 phyla at >=10 reads, and EXACTLY 9 phyla at >=25 reads -> within-tol of the paper's 40 orders / 9 phyla. The reproduced 41 sits BETWEEN the paper main-text value (40) and the paper's OWN Data S6 table (42) = strong anti-fabrication signal. C6 saturation 97.6% (41/42) reproduces the >=95% claim. C7 reproduces all three named site-dominant orders EXACTLY (Opitutales / Betaproteobacteriales / Sphingomonadales) and matches the paper's per-sample max count for TS13 exactly (26). Deliberately pinned singlem 0.13.2 + 2013 Greengenes packages (modern >=0.15 GTDB build would give materially different counts) = faithful reproduction. NOT reproduced (honest blockers, not fabricated): C2 needs the authors' gone 2014 Kraken1/RefSeq-v76 DB; C4 needs full NCBI nt + heavy multi-step assembly; C5 is their union; C3 is a commercial-GUI pipeline. Dataset PRJNA629845 profiled: 40 runs (20 WGS + 20 amplicon) all present, byte-level QC confirmed (md5-validated downloads), delivers what the paper promises, grade A. All verdicts PROVISIONAL pending human sign-off.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 68
    assessed: 2026-06-22 ⛓ 99ae5c0b6d0f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study asks what the breadth of biodiversity is in the uppermost reaches of the biosphere on Mt. Everest's southern flank, and whether environmental DNA (eDNA) combined with sequencing technologies could be used to establish a baseline biodiversity inventory of high-alpine and aeolian life.

Core claims
  • eDNA from ten high-alpine ponds and streams (4,500-5,500 m) on Mt. Everest's southern flank revealed 187 potential orders from 36 phyla across the Tree of Life. finding
  • Organisms recorded above 4,500 m (an elevational belt comprising <3% of Earth's land surface) represent ~16% of global taxonomic order estimates. finding
  • Metabarcoding (CO1 gene) and whole genome shotgun (WGS) sequencing approaches provide distinct yet complementary information for cataloging biodiversity. method
  • This is the first comprehensive eDNA biodiversity survey across the Tree of Life conducted on Mt. Everest. finding
  • The eDNA inventory generated serves as a baseline genomic resource for future high-Himalayan biomonitoring and retrospective molecular studies of climate-driven change. resource
  • Bacterial community structure differed between sites, with nearby lakes (Kala Pattar Lakes 1 and 2) sharing similar profiles despite being separated by the Khumbu glacier from Kongma La Lake 1 in some order overlap. finding
  • A 310bp CO1 contig from Kala Pattar Lake 1 matched 100% to Fujientomon dicestum (order Protura), a hexapod otherwise known only from China and Japan. finding
  • Asymptotic regression modeling indicates near-saturation detection of taxonomic orders for most sequencing methods used (e.g., ≥95% for bacterial WGS, saturation achieved for microbial eukaryote/cyanobacteria/virus WGS). method
Experimental setups
Assay System Perturbation Readout Platform
Whole genome shotgun (WGS) sequencing eDNA from high-alpine ponds/streams, Mt. Everest Khumbu region none bacterial taxonomic order abundance/diversity (SingleM + Greengenes database)
Whole genome shotgun (WGS) sequencing eDNA from high-alpine ponds/streams, Mt. Everest Khumbu region none eukaryotic microbial, cyanobacterial, and viral taxonomic order abundance (Kraken + RefSeq database)
DNA metabarcoding (mitochondrial CO1 gene) eDNA from high-alpine ponds/streams, Mt. Everest Khumbu region none contig sequence identity/classification to phylum, class, order across Animalia, Chromista, Fungi, Plantae
Whole genome shotgun (WGS) sequencing with reference genome mapping eDNA from high-alpine ponds/streams, Mt. Everest Khumbu region none contig identification/classification to eukaryotic taxonomic order via BLASTn against reference genomes
Satellite imagery comparison Khumbu region lakes (proglacial environments) none lake morphology, presence/age, and glacial silt coloration over time (1962-2016)
Asymptotic regression modeling eDNA sequencing datasets (bacterial, eukaryotic/viral, metabarcoding, WGS-reference) none estimated saturation/proportion of total detectable taxonomic orders
Key results
  • 187 orders from 36 phyla across seven kingdoms detected combining all methods from ten sites and 20 L of water
  • Orders detected above 4,500 m represent ~16% of global taxonomic order estimates despite covering <3% of Earth's land surface ~16%
  • 40 bacterial orders from nine phyla detected via WGS/SingleM; Lake Above EBC had highest order count (27); Kongma La Lake 3 and Lake South of Nuptse lowest (3) 40 orders/9 phyla
  • Bacterial WGS saturation analysis: detected at least 95% of estimated total orders (median 21, 90th percentile 38, total 42) ≥95%
  • 41 orders from 15 phyla of microbial eukaryotes, cyanobacteria, and bacteriophage viruses identified via Kraken/RefSeq; saturation achieved (median 18, 90th percentile 33, total 37) 41 orders/15 phyla
  • 667 CO1 metabarcoding contigs analyzed (>80% identity), yielding 15 orders from nine phyla across four kingdoms; detected ≥83% of estimated total orders (median 8, 90th percentile 16, total 18) 15 orders/9 phyla; ≥83%
  • WGS reference-genome mapping produced 4,889 filtered contigs (>85% identity) identifying 115 potential orders from 21 phyla; Arthropoda comprised 80.1% and Streptophyta 9.1% of listed entities; detected ≥79% of estimated total orders 115 orders/21 phyla; ≥79%
  • Distinct arthropod taxa (Ephemeroptera, Odonata) found only at sites east of the Khumbu Glacier including Nuptse Glacier Mountain Stream
Key statistics
  • count 187 orders from 36 phyla across seven kingdoms (combined result across all sequencing methods and sites)
  • other ~16% of global taxonomic order estimates from <3% of Earth's land surface (elevational belt above 4,500 m diversity representation)
  • count 40 orders from 9 phyla of bacteria (WGS) (bacterial diversity across all sites via SingleM/Greengenes)
  • count 41 orders from 15 phyla (microbial eukaryotes, cyanobacteria, bacteriophage viruses) (Kraken/RefSeq WGS analysis)
  • count 15 orders from 9 phyla (CO1 metabarcoding contigs analyzed after >80% identity filtering)
  • count 115 potential orders from 21 phyla (filtered (>85% identity) contigs from WGS reference-genome mapping (7,070 total generated))
  • other 80.1% of listed entities (proportion of WGS reference-mapped sequences represented by order Arthropoda)
  • fold_change ≥95%, ≥83%, ≥79% saturation detected (bacterial WGS, metabarcoding, WGS-reference methods respectively) (asymptotic regression model estimates of proportion of total detectable orders found)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a descriptive eDNA biodiversity survey of ten ponds/streams on Mount Everest's southern flank, comparing taxonomic composition (bacteria, microbial eukaryotes/cyanobacteria/viruses, and macro-organisms) across sites using several sequencing/bioinformatic pipelines (WGS analyzed with SingleM/Greengenes, WGS analyzed with Kraken/RefSeq, CO1 metabarcoding, and WGS reads mapped to reference genomes). Results are reported mainly as counts of orders/phyla per site, visualized with heatmaps and bar plots, and site/method comparisons are described qualitatively rather than tested with inferential statistics. The one quantitative modeling step reported is an asymptotic (saturation) regression applied to each dataset to estimate what fraction of total detectable taxonomic orders was captured and how many samples would be needed to reach saturation, summarized by median, 90th percentile, and total values.

Replicationunclear Sample sizeTen eDNA samples were collected from ten ponds/streams (4,500-5,500 m); saturation curves reference an 'average number of duplicate samples' but the text does not specify the replicate structure (biological vs. technical) or a formal power/sample-size calculation GroupsTaxonomic composition (order/phylum level) compared descriptively across the 10 sampling sites, and detection compared qualitatively across sequencing/bioinformatic approaches (WGS-SingleM, WGS-Kraken, CO1 metabarcoding, WGS-reference mapping) Pairingna Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Asymptotic regression model (order-accumulation/saturation curve) Figure 2C - bacterial orders from WGS data (SingleM/Greengenes) not stated
Asymptotic regression model (order-accumulation/saturation curve) Figure 3C - microbial eukaryotes, cyanobacteria, and viruses from WGS data (Kraken/RefSeq) not stated
Asymptotic regression model (order-accumulation/saturation curve) Figure 4C - CO1 metabarcoding order detection not stated
Asymptotic regression model (order-accumulation/saturation curve) Figure 5C and Figure S2B - eukaryotic orders from WGS reads mapped to reference genomes, including estimate of reference genomes needed for 90% saturation not stated
Approaches that could also have been used
  • Completeness of taxonomic order detection was estimated using an asymptotic (saturation) regression model applied separately to each sequencing/bioinformatic dataset.
    Could also: Nonparametric richness/rarefaction estimators such as Chao1, ACE, or the iNEXT rarefaction-extrapolation framework — These methods are widely used in eDNA and community-ecology studies to estimate total richness and sampling completeness, and typically provide bootstrap or analytic confidence intervals around the richness estimate, which can complement a single asymptotic curve fit.
  • Saturation results were summarized using the median, 90th percentile, and total value read from the fitted asymptotic curve, without an explicit measure of uncertainty around these values.
    Could also: Bootstrap or jackknife resampling to generate confidence intervals around richness/saturation estimates — Adding resampling-based intervals would convey the uncertainty in the estimated total number of detectable orders, which is useful when comparing saturation across the different sequencing methods used in this study.
  • Differences in bacterial and eukaryotic community composition across the ten sites were described qualitatively (e.g., via heatmaps and shared dominant orders) rather than through a formal statistical test.
    Could also: Multivariate community-comparison methods such as PERMANOVA or ANOSIM on a Bray-Curtis or Jaccard dissimilarity matrix — These approaches are standard in eDNA/metagenomics studies for formally testing whether community composition differs between sites or sample groups, complementing the visual/descriptive comparison presented here.
  • The relative performance of the four sequencing/bioinformatic approaches (WGS-SingleM, WGS-Kraken, CO1 metabarcoding, WGS-reference mapping) in detecting taxonomic orders was compared narratively (e.g., 'distinct yet complementary').
    Could also: A formal statistical comparison of species/order accumulation curves between methods (e.g., permutation-based curve comparison) — This would allow a quantitative statement about whether one method detects significantly more or different orders than another, in addition to the qualitative overlap already described.
  • The study design uses one eDNA sample per site with limited detail on biological versus technical replication.
    Could also: An explicit nested biological/technical replicate design analyzed with mixed-effects models — This would let variance be partitioned between site-level (biological) and sequencing-level (technical) sources, which can strengthen inferences about true site-level diversity differences in future surveys of this kind.
Software: SingleM · Greengenes database · Kraken · RefSeq database · BLASTn (NCBI)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36148432

Paper: Lim et al. 2022, iScience 25:104848. "Estimating biodiversity across the tree of life on Mount Everest's southern flank with environmental DNA." Named repo: https://github.com/wwood/singlem (SingleM — a third-party tool, not the authors' own code; per BRIEF P16 this is equally valid: we run the tool on the paper's data with the described parameters). Data: SRA BioProject PRJNA629845 — 40 runs = 20 WGS (Illumina HiSeq 4000, 2×150 bp, ~25–45 M read-pairs each) + 20 AMPLICON/metabarcoding (MiSeq, CO1).

Reported computational results and pipelines

id reported result pipeline / tool data used scope
C1 40 orders from 9 phyla of bacteria SingleM v0.13.2 singlem pipe (default ribosomal-protein single-copy marker genes + Greengenes 16S), filter: drop taxa repeated between markers + hits <10 reads 20 WGS runs IN SCOPE — PRIMARY (the named repo is SingleM)
C6 bacterial richness saturated: ≥95% of estimated detectable orders accumulation/iNEXT-type modelling on the SingleM order table derived from C1 IN SCOPE (secondary, derives from C1)
C7 per-site bacterial orders: max Lake Above EBC = 27; min Kongma La Lake 3 & Lake S of Nuptse = 3 each per-sample SingleM order counts 20 WGS runs IN SCOPE (secondary, derives from C1)
C2 41 orders from 15 phyla (microbial eukaryotes, cyanobacteria, bacteriophage viruses) Kraken v0.10.5-beta + custom RefSeq v76 k-mer DB (~60k genomes), Jellyfish v1.1.1, threshold 0 20 WGS runs SECONDARY / harder — old Kraken1 + RefSeq v76 DB must be rebuilt; attempt only after C1
C4 115 orders from 21 phyla (reference-mapped eukaryotes) BWA-MEM v0.7.15 → SPAdes v3.10.1 → BLASTn 2.6.0+ vs NCBI nt, >85% id 20 WGS runs + 9 ref organelle genomes SECONDARY / complex — multi-step, NCBI nt dependency
C3 15 orders from 9 phyla (metabarcoding CO1) Geneious Prime v2019.2.3 (commercial GUI) + BBDuk + BLASTn vs NCBI nt 20 AMPLICON runs OUT OF SCOPE — core assembly/filtering done in commercial GUI software (Geneious), not scriptable/reproducible per HARD-RULE; the BLASTn step alone is not the reported pipeline
C5 187 unique orders from 36 phyla across 7 kingdoms (combined) union of C1–C4 after cross-method dedup all DERIVED — only reproducible if C1–C4 all reproduced; we report it as conditional

Decision

  • Primary target = C1 (SingleM bacterial orders/phyla). It is the single cleanest, fully-scriptable, repo-named result and the quick 80% minimum.
  • Then push on C6/C7 (cheap, derive from the same SingleM OTU/order table).
  • Then attempt C2 (Kraken1 + RefSeq v76) and C4 (mapping/assembly/BLAST) as the harder, non-floor results. C5 is conditional on those.
  • C3 dropped from scope: commercial Geneious GUI pipeline is not reproducible.

Out-of-scope (wet-lab / manual / non-pipeline)

DNA extraction, primer design, field sampling, Sanger/manual curation, the Geneious-GUI metabarcoding assembly (C3), and all ecological interpretation.

Reproduction note on faithfulness

SingleM v0.13.2 (2020) shipped a Greengenes-based default metapackage; modern SingleM (≥0.15) uses GTDB and a different package and would give a DIFFERENT order count. Faithful reproduction therefore pins v0.13.2 + its contemporaneous data package. If that exact package is unobtainable we record the discrepancy rather than substituting the GTDB package and calling it a match.

C1
Reported
40 orders from 9 phyla of bacteria (SingleM v0.13.2)
Reproduced
41 orders / 11 phyla at num_hits>=10; exactly 9 phyla (and 20 orders) at num_hits>=25 (combined ribosomal-protein + Greengenes-16S, 20 WGS runs). RE-RUN this session end-to-end on «our HPC» (singlem 0.13.2 + contemporaneous GG16S spkg); identical to prior compute.
within tolerance
C6
Reported
bacterial richness saturated: >=95% of estimated ~42 detectable orders
Reproduced
41 distinct bacterial orders detected = 97.6% of the paper's estimated 42-order ceiling -> reproduces the >=95% saturation conclusion (asymptotic model not re-fit)
within tolerance
C7
Reported
max site Lake Above EBC = 27 (Opitutales-dominant); min Kongma La Lake3 = 3 (Betaproteobacteriales) & Lake S of Nuptse = 3 (Sphingomonadales)
Reproduced
max sample SRR11700437 (ts13) = 26 orders = EXACT match to paper Table S4 TS13 SingleM column; dominant order Opitutales (exact). Kongma La L3 (ts26/27, SRR11700420/418) dominant Betaproteobacteriales (exact). Lake S of Nuptse (ts32/33, SRR11700407/405) dominant Sphingomonadales (exact). All 3 named site-dominant orders reproduce EXACTLY; counts within 1.
within tolerance
C2
Reported
41 orders from 15 phyla (Kraken v0.10.5-beta + custom RefSeq v76 k-mer DB; microbial euks/cyano/phage)
Reproduced
NOT FAITHFULLY REPRODUCIBLE — the exact pipeline needs Kraken1 v0.10.5-beta + the authors' custom RefSeq-v76 (~2014-16, ~60k genomes) k-mer DB, which is not shipped and cannot be reconstructed identically; a modern Kraken2+standard-DB run is an explicitly NON-faithful substitute (different DB + algorithm) that would not match the reported counts. Recorded as honest blocker rather than substituted.
partial
C4
Reported
115 orders from 21 phyla (trim -> BWA-MEM map to refs -> SPAdes -> BLASTn vs NCBI nt)
Reproduced
NOT ATTEMPTED (feasibility-assessed) — multi-step assembly pipeline whose terminal step BLASTns thousands of contigs against the full NCBI nt (~hundreds of GB, version-sensitive); order/phylum counts depend strongly on the nt snapshot date + the >85%/>95% identity & qcov filters. High-cost, low-determinism secondary; documented as out-of-floor.
partial
C5
Reported
187 unique orders from 36 phyla across 7 kingdoms (combined union of C1-C4)
Reproduced
NOT ATTEMPTED — derived = union(C1..C4); only meaningful if C2+C4 also reproduced, which they are not (faithful blockers).
partial
C3
Reported
15 orders from 9 phyla (CO1 metabarcoding)
Reproduced
OUT OF SCOPE — core assembly/filtering done in commercial Geneious Prime v2019.2.3 GUI (not scriptable/reproducible).
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 65/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

The paper's primary bacterial-biodiversity result reproduces near 1:1: running the exact named tool (SingleM v0.13.2 + contemporaneous Greengenes) on the authors' own public SRA data gives 41 orders vs 40 (off by one), saturation 97.6% vs >=95%, and all three named dominant orders match exactly — strong anti-fabrication evidence. The only deviations (phyla 11 vs 9) sit on our preprocessing side (an unparsed marker-dedup sub-filter; exactly 9 at a >=25-read threshold), not the authors' side. Secondary claims (C2/C4/C5 not attempted, C3 out of scope) were not tested, so the overall verdict is solid with small, explainable deviations rather than a perfect 1:1.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

293.3 k
tokens (I/O) · 24.9 M incl. cache
99 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.