Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A near complete genome for goat genetic and genomic research.

Genet Sel Evol · 2021
L1 89/100 PQI 96
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Goat de-novo genome ASSEMBLY paper (Saanen_v1 = GenBank GCA_015443085.1). The full assembly (PacBio CLR + Hi-C, ~2.69 Gb, thousands of CPU-h) is out of scope; the named code repo IsoSeq is only an annotation-evidence track and the paper reports NO IsoSeq-specific numbers to grade. In-scope 80/20: recompute the published assembly-quality statistics from the DEPOSITED assembly on «our HPC»/«infra» using an independent pure-python streamer, cross-checked against NCBI's own assembly_stats.txt and seqkit. Result: 4 headline numbers reproduce 1:1 EXACT (total length 2.696 Gb = paper 2.69 Gb; contig N50 46,208,332 bp = 46.2 Mb; scaffold N50 102,383,509 bp = 102.3 Mb; chromosomal 2.638 Gb = paper 2.637 Gb) and gaps within-tol (NCBI spanned-gaps 169 = paper exactly; my naive N-run count 170, an off-by-one from gap definition). No fabrication signal — every reported assembly metric is independently re-derivable from the public assembly. BUSCO (reported 94.3%) was NOT completed: 5 jobs hit a parser bug + HTTP 404 in compleasm's busco-data downloader (broken upstream in every version), and the comparison would be approximate-only anyway (paper used BUSCO v3.0.2/mammalia_odb9/protein mode vs a modern genome-mode odb10) — left as optional last-20% per brief. NOT attempted: de-novo reassembly, gene-annotation counts (EvidenceModeler integration), IsoSeq transcript counts (no reported target).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 89
    assessed: 2026-06-15 ⛓ 4c3695e94bcb
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can high-depth PacBio long-read sequencing combined with Hi-C scaffolding produce a near-complete, high-quality goat reference genome (Saanen_v1)—including continuous X and the first Y chromosome scaffolds—that improves upon the existing ARS1 assembly for goat genetic and genomic research?

Core claims
  • Saanen_v1 is a high-quality de novo goat genome assembly from a male Saanen buck, including the first goat Y chromosome scaffold. resource
  • Saanen_v1 is more complete than ARS1, with centromeric/telomeric repeats at ends of two-thirds of autosomes, fewer gaps (169 vs 773), and more sequence on chromosomes (2.63 Gb vs 2.58 Gb). finding
  • Eight putative large assembly errors (1 to ~7 Mb each) present in ARS1 were identified and amended in Saanen_v1. finding
  • Saanen_v1 enables reassignment of likely correct positions for 4.4% of SNP probes in the GoatSNP50 chip. finding
  • The substitution rate of the ruminant (goat) Y chromosome was estimated for the first time, allowing estimation of the male-to-female mutation rate ratio. finding
  • A combined approach of high-depth PacBio long-read sequencing (117×) plus Hi-C (118×) scaffolding was used to assemble the genome. method
  • Y chromosome scaffolds were resolved by selecting the Flye assembly, which contained seven Y-linked single-copy genes in conserved order matching the ovine Y chromosome. method
  • Genes were comprehensively annotated combining ab initio, homology-based, and RNA-seq/Iso-Seq-assisted prediction integrated by EvidenceModeler. method
Experimental setups
Assay System Perturbation Readout Platform
PacBio long-read (SMRT) whole-genome sequencing Saanen dairy goat (Capra hircus), male buck, liver tissue DNA none long reads / subreads for de novo assembly PacBio Sequel II, Sequel Binding Kit 1.0, Sequencing Kit 1.0, SMRT Cell 8M
Illumina short-read whole-genome sequencing Saanen dairy goat, same individual, liver DNA none paired-end 150 bp reads for polishing and kmer genome-size estimation Illumina HiSeq X Ten, TruSeq Nano DNA Library Prep Kit
Hi-C chromatin conformation capture sequencing Saanen dairy goat, same individual, blood DNA MboI restriction digestion / cross-linking chromatin interaction matrix for scaffolding to chromosome level Illumina HiSeq X Ten
Short-read RNA-seq goat tissues (publicly available SRA and unpublished Illumina data) none transcript evidence for gene annotation and mapping ratio assessment
PacBio Iso-Seq (long-read RNA sequencing) testicular tissue of Saanen individual and abomasum tissue of three Shanbei white Cashmere goats none full-length transcripts for gene annotation PacBio Iso-Seq
miRNA-seq analysis goat (nine downloaded NCBI SRA miRNA-seq datasets) none known and novel miRNA identification
Whole-genome alignment / comparative assembly assessment Saanen_v1 vs ARS1, CHIR_2.0, Oar_rambouillet_v1.0 goat/ruminant assemblies none structural variations, large assembly errors, SNP probe positions
Multiple sequence alignment for substitution rate estimation cattle, yak, sheep, goat (autosomes and Y chromosome) none substitution rates and male-to-female mutation rate ratio (αm)
Key results
  • Saanen_v1 has far fewer gaps than ARS1 169 vs 773 gaps
  • More assembled sequence anchored to chromosomes in Saanen_v1 than ARS1 2.63 Gb vs 2.58 Gb
  • Eight large assembly errors in ARS1 amended in Saanen_v1 1 to ~7 Mb each
  • Correct positions assigned for a fraction of GoatSNP50 chip SNP probes 4.4% of SNP probes
  • Flye assembly scaffold contained more conserved Y-linked single-copy genes than Wtdbg2 7 vs 5 of 10 Y-linked genes
  • Polished Flye assembly showed higher BUSCO completeness than Wtdbg2 version 94.0% vs 93.1%
  • Final polished de novo assembly length and contiguity 2.69 Gb, contig N50 34.0 Mb
  • Estimated goat genome size from 17-kmer distribution 2.72 Gb
Key statistics
  • count 169 vs. 773 (number of gaps in Saanen_v1 vs ARS1)
  • count 2.63 Gb vs. 2.58 Gb (assembled sequence on chromosomes Saanen_v1 vs ARS1)
  • other 4.4% (GoatSNP50 chip SNP probes given likely correct positions)
  • other 94.0% vs. 93.1% (BUSCO completeness Flye vs Wtdbg2 polished assemblies)
  • other 117× (327.7 Gb) (PacBio SMRT long-read coverage)
  • other 118× (Hi-C data coverage)
  • other contig N50 33.9 Mb vs 35.3 Mb (Flye vs Wtdbg2 contig continuity)
  • count 2.72 Gb (estimated goat genome size from 97 Gb Illumina reads, 17-kmer peak at 71×)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a de novo genome assembly study that is predominantly descriptive and computational rather than inferential. Rather than hypothesis-testing comparisons between experimental groups, the work reports assembly metrics (contig N50, BUSCO completeness, QV, mapping ratios) and uses a single reference individual, with comparisons made against existing assemblies (ARS1, CHIR_2.0, etc.). The one explicitly model-based statistical analysis is phylogenetic estimation of nucleotide substitution rates for autosomes and sex chromosomes, from which a male-to-female mutation rate ratio was derived.

Replicationunclear Sample sizeA single Saanen buck was sequenced for the assembly; additional single individuals were used for QV (one Yunnan black goat plus three downloaded datasets) and other comparisons; no formal sample-size or power calculation was described GroupsSaanen_v1 vs existing assemblies (ARS1, CHIR_2.0, and other species references) Pairingna Randomization/blindingna Dispersionnone Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Phylogenetic substitution-rate estimation (maximum-likelihood under the GTR/REV model with 4 rate categories, PHYLOFIT) substitution rates for autosomes, X and Y chromosomes across cattle, yak, sheep and goat 1.96 Gb aligned autosomal sequence and 645.1 kb of X-degenerate Y sequence across four species stated
Substitution-model selection (jModelTest) choosing the best-fitted model (GTR/REV) for the multiple-sequence alignment stated
k-mer (17-mer) frequency-based genome-size estimation genome size of Saanen_v1 from ~97 Gb Illumina reads one individual; 17-kmer distribution peak at 71× not stated
BUSCO completeness assessment assembly completeness vs mammalia_odb9 (4,104 single-copy orthologues) 4,104 orthologues na
Quality value (QV) estimation via FreeBayes substitution calling per-assembly base accuracy own WGS data plus 3 downloaded individuals (Asian, European, African) na
Approaches that could also have been used
  • Substitution rates and the derived male-to-female mutation ratio (α_m) were reported as single point estimates from PHYLOFIT.
    Could also: Bootstrap resampling of alignment blocks or reporting a likelihood-based confidence/credible interval around the rate and α_m estimates would also be a standard option. — An interval would convey the uncertainty of the estimate, which is helpful given the restricted X-degenerate region (645.1 kb) used for the Y chromosome.
  • Genome size was estimated from a single individual's 17-mer distribution using the peak-depth formula.
    Could also: Fitting a full k-mer mixture model (e.g., GenomeScope) and/or comparing multiple k lengths would also be a common approach. — A model fit can additionally estimate heterozygosity and repeat content and provide a fitted error envelope around the size estimate.
  • Assembly-to-assembly comparisons (continuity, mapping ratios, gene counts) were presented as descriptive metrics for Saanen_v1 versus ARS1 and other references.
    Could also: Reference-free k-mer based metrics such as Merqury QV/completeness could also be reported alongside the alignment-based QV. — Reference-free metrics avoid dependence on a chosen comparison assembly and offer an independent line of evidence for base accuracy and completeness.
  • QV was computed from a small set of individuals (own data plus three downloaded genomes spanning Asian, European and African origins).
    Could also: Summarizing the per-individual QV values with a mean and a measure of spread (range or SD) across the datasets would also be an option. — Showing the spread across individuals would communicate how consistent the accuracy estimate is across genetic backgrounds.
  • Model selection relied on jModelTest choosing GTR/REV with four rate categories.
    Could also: Reporting the information-criterion (AIC/BIC) values for the candidate models, or a sensitivity check under an alternative model, would also be standard practice. — Documenting the selection criteria and a robustness check shows how much the downstream rate estimates depend on the chosen model.
Software: PHYLOFIT (PHAST) · jModelTest · gce (k-mer genome-size estimation) 1.0.2 · BUSCO 3.0.2 · FreeBayes (QV estimation) · FRCBam (CE error evaluation)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
47
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

AY082491 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
AY082500 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
CM001061 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
CM016720 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
CM022046 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
D82963 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
ERR313197 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
ERR313206 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
ERR313211 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
ERR313213 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GCA_000001405.28 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GCA_000003025.6 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GCA_000317765.2 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GCA_001704415.1 GCA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GCA_002263795.2 GCA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GCA_002742125.1 GCA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GCA_002863925.1 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GCA_003121395.1 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM3579832 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSM3579835 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSM3579836 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSM3579838 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSM3579841 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSM3579842 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSM3579844 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSM3579847 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSM3579848 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
LC416985 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
MF448228 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
MF448230 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
MF448232 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
MF448234 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
MF448235 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRP182057 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR11410765 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR1822383 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR1822384 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR1822385 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR1822386 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR1822387 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR1822388 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR1822389 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR1822390 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR1822391 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR5803174 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRR8618141 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

This is a genome-assembly paper (Saanen_v1 goat reference, GenBank GCA_015443085.1_Saanen_v1). The full de-novo assembly (PacBio CLR + Hi-C of a 2.69 Gb genome; Flye/Wtdbg2/MECAT2/Racon/Pilon/3D-DNA/PBjelly) is out of scope (thousands of CPU-hours, the heavy "last 20%++").

In scope (deterministic, gradeable): the published assembly-quality statistics, recomputed from the deposited assembly FASTA:

  • Total length 2.69 Gb (ungapped 2.637 Gb)
  • Contig N50 46.2 Mb
  • Scaffold N50 102.3 Mb
  • Number of gaps 169

Secondary (named pipeline tool, approximate): BUSCO completeness. Paper reports 94.3 % with BUSCO v3.0.2, mammalia_odb9 (4,104 orthologues), protein mode. We run a modern genome-mode completeness check (compleasm / mammalia_odb10) as an audit cross-check — mode + DB differ, so this is an approximate comparison.

Named code repo is PacBio IsoSeq (v3.2.2) used for transcript evidence in annotation, but the paper reports no IsoSeq-specific numbers (CCS/FLNC/isoform counts) in the text or tables — nothing to grade against — so a bare IsoSeq run is not a useful reproduction target and is documented as such rather than run.

Figures / tables: Table
asm_total
Reported
2.69 Gb
Reproduced
2,696,218,280 bp (2.696 Gb)
exact
asm_chromosomal
Reported
2.637 Gb chromosomal (58.3 Mb unplaced)
Reproduced
2,637,907,739 bp assembled-molecule; 58.3 Mb unplaced in 1331 scaffolds
exact
contig_n50
Reported
46.2 Mb
Reproduced
46,208,332 bp (46.21 Mb)
exact
scaffold_n50
Reported
102.3 Mb
Reproduced
102,383,509 bp (102.38 Mb)
exact
n_gaps
Reported
169
Reproduced
169 (NCBI spanned-gaps) / 170 (naive N-run count)
within tolerance
busco
Reported
94.3% (BUSCO v3.0.2, mammalia_odb9, protein mode)
Reproduced
not completed (compleasm/busco-data downloader broken upstream)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This goat genome-assembly paper reproduces exactly: all four published assembly-quality statistics (total length 2.69 Gb, chromosomal 2.637 Gb, contig N50 46.2 Mb, scaffold N50 102.3 Mb) are independently re-derivable from the deposited GenBank assembly and agree to reported precision across a custom Python streamer, NCBI's own stats, and seqkit. The only deviation is a definitional off-by-one in gap counting (naive 170 vs NCBI spanned-gap 169, the latter matching the paper exactly). BUSCO completeness (94.3%) was not finished, but solely due to upstream tool/server breakage and it would have been an approximate cross-check anyway — neither an authors' nor a data-availability defect. No fabrication signal; the central claim of a near-complete reference holds.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

206.8 k
tokens (I/O) · 17.6 M incl. cache
24 min
runtime · 0.04 CPU-h
1.3 GB
peak RAM
5 (1 failed)
HPC jobs
hummel
machine