Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Insights into the differentiation and adaptation within Circaeasteraceae from Circaeaster agrestis genome sequencing and resequencing.

iScience · 2023
L1 49/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
49/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 7% of all assessed papers rank 1081 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Designated pipeline = OpenGene/fastp (-q 20 -5 -3, v0.20.0) on Illumina WGS reads from BioProject PRJNA877109. RE-CONFIRMED NOT reproducible (re-checked 2026-06-22): PRJNA877109 resolves but contains NO raw reads (elink bioproject->sra=0; biosample SAMN30685011->sra=0; ENA read_run=0; SRA 'PRJNA877109[BioProject]'=PhraseNotFound). Only the genome assembly GCA_047371235.1 (released 2025-02-03) is public. The paper's Data-Availability statement ('resequencing raw reads available from NCBI PRJNA877109') appears to OVERSTATE what is deposited (flag for human audit). The paper reports no fastp output statistic, so there is no fastp value to compare. KEPT GOING with real «our HPC» compute («job», COMPLETED, 28 min): on the deposited assembly (md5-verified) we (1) recomputed descriptive stats independently with seqkit 2.8.2 + pure-python (identical) and (2) ran genome-mode BUSCO 5.8.2 (eukaryota_odb10 + embryophyta_odb10). RESULTS: 15 chromosomes (EXACT vs reported 15), 790.95 Mb (WITHIN-TOL vs reported 791 Mb anchored), GC 34.70% (WITHIN-TOL vs NCBI 34.5%); recomputed gap bases 20,300 == NCBI total-gap-length (internal validation). BUSCO eukaryota_odb10 genome mode = C:98.8% (S:43.5%,D:55.3%) — strongly corroborates the paper's gene-set C:96.04% (S:41.91%,D:54.13%), including the distinctive ~55% duplicated signature, despite different BUSCO version + mode (graded 'partial' for that reason). MISMATCH/caveat: the reported FULL 978.68 Mb / 876-contig / N50 12.3 Mb assembly is NOT in the public deposit — only the ~791 Mb chromosome-anchored portion (218 contigs, contig-N50 14.9 Mb) is public, so the headline 978.68 Mb size is NOT verifiable from public data. NOT ATTEMPTED (no reads/annotation, out of scope): fastp/Trim Galore read QC, 524,766-SNP resequencing, 73,567-gene prediction, Nanopore/Canu/SMARTdenovo assembly, HiCUP Hi-C QC, ~1004 Mb k-mer survey. Two items flagged for human audit: (1) data-availability overstatement, (2) full-assembly size not in public deposit. All grades provisional; a human reviewer decides.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 44
    assessed: 2026-06-16 ⛓ cfac642d30d6
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Using the sister-species pair Circaeaster agrestis (sexual) and Kingdonia uniflora (predominantly asexual) as a comparative system, the study tests whether reproductive mode predictably shapes genome evolution (heterozygosity, transposable element load, linkage disequilibrium, and efficacy of purifying selection) and investigates the genetic basis of adaptation/differentiation within C. agrestis populations.

Core claims
  • C. agrestis and K. uniflora are sister species with contrasting reproductive modes, providing a natural system to test effects of sexual vs asexual reproduction on genome evolution finding
  • Chromosome-level de novo genome assembly of C. agrestis (978.68 Mb, 15 pseudochromosomes) generated via Nanopore and Hi-C resource
  • C. agrestis and K. uniflora have similar genome sizes but C. agrestis encodes many more genes (73,567 vs 43,301) finding
  • C. agrestis experienced two rounds of whole-genome duplication, including one recent event, after filtering tandem duplications finding
  • K. uniflora shows much higher genome heterozygosity, TE load, linkage disequilibrium, and πN/πS ratio than C. agrestis, consistent with predicted genetic consequences of asexual reproduction finding
  • Gene families unique to C. agrestis are enriched for defense response genes, while those unique to K. uniflora are enriched for root system development genes, aligning with each species' reproductive/ecological strategy finding
  • Fst outlier analysis across 25 C. agrestis populations reveals a link between abiotic stress response and genetic variability via divergent selection finding
  • Asexual reproduction in K. uniflora reduces the efficacy of purifying selection, causing faster accumulation of deleterious mutations and TEs relative to sexual C. agrestis mechanism
Experimental setups
Assay System Perturbation Readout Platform
Genome sequencing and de novo assembly (Illumina + Oxford Nanopore + Hi-C) Circaeaster agrestis none genome size, contig N50, pseudochromosome anchoring Oxford Nanopore Technologies; Hi-C
Gene annotation and BUSCO completeness assessment C. agrestis genome none number of protein-coding genes, ncRNAs, BUSCO completeness BUSCO
Kmer-based genome heterozygosity estimation C. agrestis and K. uniflora (Illumina reads) none % intragenomic heterozygosity
Repeat/transposable element annotation (homology-based + de novo) C. agrestis genome (compared to K. uniflora) none % genome as TEs by class (LTR, DNA, LINE, SINE)
Comparative genomics: Ks distribution and collinearity/synteny analysis (OrthoFinder) C. agrestis and K. uniflora genomes none whole-genome duplication events, syntenic blocks OrthoFinder v2.3.12
Orthogroup/gene family clustering and GO enrichment C. agrestis and K. uniflora proteomes none species-specific and shared orthogroups, enriched GO terms OrthoFinder
Whole-genome resequencing and population structure analysis (STRUCTURE, NJ tree, PCoA) 158 C. agrestis individuals from 25 populations none SNP calls, genetic clustering (K groups) Illumina (~20x depth)
Fst outlier detection (BAYESCAN) and gene set enrichment 25 C. agrestis populations, 524,766 SNPs none outlier SNPs, genes under divergent selection, enriched GO terms BAYESCAN
Key results
  • Chromosome-level C. agrestis assembly: 978.68 Mb, 876 contigs, N50 = 12.3 Mb, 791 Mb (80.82%) anchored to 15 pseudochromosomes
  • C. agrestis predicted to have 73,567 protein-coding genes vs 43,301 in K. uniflora despite similar genome size 73,567 vs 43,301 genes
  • K. uniflora genome heterozygosity (0.37%) is up to 20 times higher than C. agrestis (0.0156%) ~20-fold
  • C. agrestis genome contains fewer LTR transposable elements than K. uniflora; TEs comprise 23.31% (228.18 Mb) of C. agrestis genome
  • After filtering tandem duplications, C. agrestis shows two rounds of WGD (one recent) vs one WGD reported in K. uniflora
  • 3,969 orthogroups specific to C. agrestis (enriched for defense response) and 1,322 specific to K. uniflora (enriched for root development) 3,969 vs 1,322 orthogroups
  • πN/πS ratio is lower in C. agrestis (0.37) than in K. uniflora (0.55), indicating stronger purifying selection efficacy in C. agrestis 0.37 vs 0.55
  • BAYESCAN identified 114,697 outlier SNPs (of 524,766 total) located in 3,266 genes, enriched for protein localization/transport functions
Key statistics
  • other 0.0156% (C. agrestis) vs 0.37% (K. uniflora) (genome heterozygosity comparison)
  • count 73,567 protein-coding genes (66,386 functionally annotated) (C. agrestis gene annotation)
  • fold_change 23.31% (228.18 Mb) TEs; 37.76% total repetitive elements (C. agrestis genome repeat content)
  • count 524,766 high-quality SNPs (resequencing of 158 individuals across 25 C. agrestis populations)
  • count 114,697 outlier SNPs in 3,266 genes (2,490 with BLAST hits) (BAYESCAN Fst outlier detection)
  • other Fst ranged from 0.02 to 0.90 (pairwise differentiation between C. agrestis populations)
  • other πN/πS = 0.37 (C. agrestis) vs 0.55 (K. uniflora) (efficacy of purifying selection comparison)
  • pvalue p < 0.05 (gene set enrichment of genes containing outlier SNPs)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper assembled and annotated a chromosome-level genome for Circaeaster agrestis and performed comparative genomic analyses against the sister species Kingdonia uniflora. Population diversity was characterized by whole-genome resequencing of 158 C. agrestis individuals from 25 populations, yielding 524,766 high-quality SNPs analyzed with STRUCTURE, neighbor-joining trees, PCoA, MSMC2 demographic modeling, and BAYESCAN FST-outlier detection. Gene ontology enrichment and πN/πS ratios were used to characterize selection and gene family differentiation between species.

Replicationmixed Sample size158 individuals across 25 populations for population genomics (mean depth ~20×); single individual genome assembled for C. agrestis; K. uniflora genome from prior published study; no power analysis stated GroupsC. agrestis vs K. uniflora (interspecific, n=1 genome each); 7 genetic groups within C. agrestis; East lineage (groups 1–2) vs West lineage (groups 3–7) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
Kmer-based statistical estimation of genome size and heterozygosity Genome survey of C. agrestis (119 Gb Illumina reads); heterozygosity compared descriptively to K. uniflora estimate from prior study 119 Gb Illumina reads for C. agrestis; K. uniflora value from prior published study not stated
BUSCO completeness assessment Gene annotation quality evaluation of C. agrestis genome against Eukaryota universal single-copy orthologs na
STRUCTURE analysis with cross-entropy criterion for K selection (K=2–10 tested, K=7 selected) Population structure inference across 25 C. agrestis populations 158 individuals, 524,766 SNPs not stated
Neighbor-joining (NJ) phylogenetic tree Population-level phylogenetic inference from whole-genome SNP data 158 individuals, 524,766 SNPs not stated
Principal Coordinates Analysis (PCoA) Population structure visualization; first two components reported 158 individuals, 524,766 SNPs not stated
MSMC2 (Multiple Sequentially Markovian Coalescent) effective population size estimation Ne dynamics over time for 7 genetic groups of C. agrestis Subset of 158 resequenced individuals per genetic group; exact per-group n not stated not stated
Linkage disequilibrium decay analysis (r²) over genomic distance LD characterization of C. agrestis up to 200 kb; compared descriptively to K. uniflora from prior study 158 individuals, 524,766 SNPs not stated
πN/πS ratio (π0/π4) comparison Efficacy of purifying selection: C. agrestis (0.37) vs K. uniflora (0.55, from prior published study); descriptive comparison, no formal test stated 158 C. agrestis individuals for C. agrestis estimate; K. uniflora value from Sun et al. not stated
BAYESCAN FST outlier detection Diversifying selection scan across all 25 C. agrestis populations; 114,697 outlier SNPs identified (FST 0.5–0.9) out of 524,766 524,766 SNPs from 158 individuals across 25 populations not stated
Gene Ontology (GO) enrichment analysis (p < 0.05 threshold) Species-specific gene families (C. agrestis and K. uniflora unique orthogroups); genes containing outlier FST SNPs (top 20 GO terms reported) 3,969 C. agrestis-specific and 1,322 K. uniflora-specific orthogroups; 2,490 outlier-associated genes with BLAST hits not stated
Synonymous substitution rate (Ks) distribution analysis with collinearity-based WGD detection (syntenic blocks ≥10 collinear genes) Whole-genome duplication history within and between C. agrestis and K. uniflora; analyses run with and without filtering tandem duplications 62,169 C. agrestis genes and 39,364 K. uniflora genes after filtering not stated
OrthoFinder v2.3.12 orthogroup construction Gene family comparison between C. agrestis and K. uniflora; 31,833 orthogroups identified, 13,271 shared 62,169 C. agrestis genes, 39,364 K. uniflora genes na
Approaches that could also have been used
  • GO enrichment significance was assessed at p < 0.05 across many GO terms without an explicitly stated multiple-testing correction method
    Could also: Benjamini-Hochberg FDR correction (q < 0.05) applied across the full set of GO terms tested, as implemented in tools such as clusterProfiler or topGO — When many GO terms are tested simultaneously, FDR correction reduces the expected proportion of false-positive enrichment signals; it is widely adopted in genomic GO enrichment workflows and facilitates cross-study comparisons
  • BAYESCAN was used for FST-outlier detection to identify loci under divergent selection across 25 populations
    Could also: pcadapt, OutFLANK, or BayPass for selection-scan outlier detection — These alternatives account for population structure in distinct ways (e.g., pcadapt uses a PCA-based confounding correction; OutFLANK fits a neutral FST distribution to control false discovery), providing complementary perspectives; convergence of findings across methods is often used to increase confidence in candidate loci
  • MSMC2 was used to infer Ne history for each of the 7 genetic groups separately from a subset of resequenced individuals
    Could also: SMC++ or PSMC, which leverage multiple haplotypes or model populations jointly — SMC++ can use many more samples per population for improved resolution at recent time scales; comparing results across demographic inference methods is a common practice to assess robustness of inferred Ne trajectories
  • Population structure was inferred with STRUCTURE using the cross-entropy criterion to select K
    Could also: ADMIXTURE with cross-validation error for K selection, or sparse NMF-based approaches such as sNMF — ADMIXTURE uses the same admixture model but scales more efficiently to large SNP panels; reporting convergent results from two methods is a common way to assess stability of the inferred ancestry proportions and optimal K
  • Purifying selection efficacy was compared via the πN/πS point estimate (C. agrestis 0.37 vs K. uniflora 0.55) without a formal statistical test or uncertainty estimate
    Could also: McDonald-Kreitman test, distribution-of-fitness-effects (DFE) inference, or bootstrap confidence intervals around πN/πS — These approaches provide formal tests or uncertainty bounds around the selection-efficacy comparison, allowing assessment of whether the observed ratio difference is consistent with relaxed purifying selection, positive selection, or sampling variance
  • Genome heterozygosity for C. agrestis and K. uniflora was estimated from single-individual assemblies using a kmer-based approach
    Could also: Population-level heterozygosity estimation from multiple resequenced individuals using the site frequency spectrum or mean per-site observed heterozygosity — Multi-individual estimates provide variance around heterozygosity and distinguish between individual- and species-level variation; for C. agrestis, the available 158-individual resequencing dataset could additionally support such an estimate
Software: OrthoFinder 2.3.12 · STRUCTURE · MSMC2 · BAYESCAN · BUSCO

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36895650

Paper: Insights into the differentiation and adaptation within Circaeasteraceae from Circaeaster agrestis genome sequencing and resequencing. iScience 2023. DOI 10.1016/j.isci.2023.106159 · PMCID PMC9988679 · PMID 36895650.

Designated reproduction inputs (from BRIEF):

  • Code: https://github.com/OpenGene/fastp (third-party read QC/trimming tool — P16 valid)
  • Data: sra:PRJNA877109

What the paper's pipeline actually does (Methods, verbatim/near-verbatim)

  • Illumina WGS reads filtered with FASTP v0.20.0 -q 20 -5 -3 (remove adapters + low-quality). "A total of 119 Gb raw Illumina paired-end data were generated."
  • Nanopore long reads (~149×, 149 Gb), corrected with Canu v1.6 -correct; assembler reported by NCBI as SMARTdenovo (Dec-2020).
  • Hi-C (~102×, 102 Gb), QC with HiCUP v0.8.0 (default).
  • Population resequencing: 158 individuals / 25 populations, mean depth ~20×, reads filtered with Trim Galore v0.6.5 (default; not fastp). 524,766 high-quality SNPs.
  • Assembly: full 978.68 Mb / 876 contigs, contig N50 12.3 Mb; 791 Mb (~80.82%) anchored to 15 pseudochromosomes; predicted genome size ~1004 Mb.

In scope vs out of scope

Result Pipeline In scope? Status
fastp QC/trim of Illumina WGS reads OpenGene/fastp designated, BUT see blocker NOT REPRODUCIBLE — data_unavailable
Assembly descriptive stats (length, #seq, N50, GC) seqkit on deposited GCA bonus (verifies a reported pipeline output against the deposited artifact) REPRODUCED (see agreement.json)
Nanopore/Hi-C QC, Canu/SMARTdenovo assembly, HiCUP various out of scope (no reads; heavy de-novo assembly not the designated tool) not attempted
Trim Galore resequencing QC, SNP calling Trim Galore / GATK-like out of scope (no reads) not attempted
Wet-lab / morphology / phylogeny-from-alignment manual/external out of scope not attempted

CRITICAL DATA-AVAILABILITY FINDING (fabrication-relevant)

The paper's Data Availability statement says the genome assembly and "C. agrestis resequencing raw reads are available from NCBI (BioProject PRJNA877109)."

Verified 2026-06-16 via NCBI E-utilities + ENA:

  • BioProject PRJNA877109 resolves (UID 877109, "Circaeaster agrestis isolate:SNJ Genome sequencing", registered 2022-09-06).
  • It links to exactly ONE artifact: the genome assembly GCA_047371235.1 (released 2025-02-03). elink bioproject→sra = 0 runs; elink bioproject→biosample = 0; ENA read_run for the accession = 0 rows; SRA PRJNA877109[BioProject] = PhraseNotFound.
  • The 281 SRA records under organism Circaeaster agrestis belong to unrelated projects (278× PRJNA616258 RAD-Seq 2020; PRJEB38536; PRJEB42299; PRJNA739247) — none to PRJNA877109.

Consequence: the raw Illumina/Nanopore/Hi-C/resequencing reads that fastp (and the rest of the pipeline) operate on are not publicly deposited / not retrievable, so the designated fastp reproduction cannot be run. The data-availability statement appears to overstate what is actually deposited (only the assembly is public, not the reads). The paper also reports no fastp output statistics (no clean-read count / Q20 / Q30 / adapter or duplication rate) in the main text, so even the expected value to compare fastp against is not pinnable. Flagged for human audit.

Compute

All compute on «our HPC» («infra»); data only on «infra» «path». Job: run.sbatch (seqkit + fastp env; downloads GCA_047371235.1, md5-verifies, recomputes assembly stats). «host» holds small results only.

fastp_qc
Reported
fastp v0.20.0 -q 20 -5 -3 on 119 Gb Illumina reads (no output stats reported)
Reproduced
NOT RUN — raw reads not deposited under PRJNA877109 (re-verified 2026-06-22)
did not match
busco_complete
Reported
96.04% complete BUSCOs (S 41.91% + D 54.13%), F 3.63% — BUSCO v3.0.2 on 73,567 predicted genes, eukaryota odb10
Reproduced
98.8% complete (S 43.5% + D 55.3%), F 0.8%, M 0.4% — BUSCO 5.8.2 GENOME mode, eukaryota_odb10, n=255
partial
busco_duplicated
Reported
54.13% complete & duplicated BUSCOs
Reproduced
55.3% (eukaryota_odb10, genome mode)
within tolerance
asm_pseudochromosomes
Reported
15 pseudochromosomes
Reproduced
15 sequences (CM104546.1-CM104560.1)
exact
asm_anchored_length
Reported
791 Mb anchored (~80.82% of assembly)
Reproduced
790,947,575 bp total in deposited assembly
within tolerance
gc_content
Reported
34.5% (NCBI assembly_stats; not in paper main Methods)
Reproduced
34.70% over ACGT (seqkit & python agree)
within tolerance
asm_full_length
Reported
978.68 Mb full assembly
Reproduced
790,947,575 bp (public deposit = anchored set only)
did not match
asm_full_contigs
Reported
876 high-quality contigs (full assembly)
Reproduced
218 contigs / 15 scaffolds (deposited set)
did not match
asm_contig_n50_full
Reported
contig N50 12.3 Mb (full assembly)
Reproduced
contig-N50 14.9 Mb / scaffold-N50 55.3 Mb (deposited set)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 49/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Every value reproducible from the one public artifact (assembly GCA_047371235.1) matched within tolerance — 15 pseudochromosomes (exact), 790.95 Mb ≈ 791 Mb anchored, GC 34.70% ≈ 34.5% — with strong internal checks (md5 verified, recomputed 20,300 gap bases == NCBI). The deviations are on the authors'/data-availability side, not our methodology: the resequencing raw reads claimed at PRJNA877109 are simply not deposited (0 runs), so fastp and the SNP/adaptation results are untestable, and the headline 978.68 Mb full assembly is not in the public deposit (only the ~80.82% anchored portion). Notably the reported full size is internally consistent with the deposited anchored fraction (790.95/0.8082 ≈ 979 Mb), so this reads as a partial-deposit + overstated data-availability issue rather than fabrication. Overall yellow: a clean reproduction of what is available, with explainable, authors-side gaps worth a human flag.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

391 k
tokens (I/O) · 35.7 M incl. cache
136 min
runtime · 5.1 CPU-h
35.5 GB
peak RAM
1
HPC jobs
hummel
machine