Chromosome-level genome assembly of Lilford's wall lizard, Podarcis lilfordi (Günther, 1874) from the Balearic Islands (Spain).
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL (BUSCO running). Recomputed reported pipeline-derived assembly QC on deposited GCA_947686815.1 (rPodLil1.2). 7/7 structural/contiguity metrics reproduce EXACTLY: total length 1,460,085,851 bp, 2,148 scaffolds, scaffold N50 89.64 Mb, contig N50 1.48 Mb, 20 chromosomes (18+ZW), 2,084 unplaced + 44 unlocalized, mito 17,251 bp. The only apparent length/count gaps vs paper are precisely explained by NCBI including the 17,251 bp mito that Table 1 excludes - not fabrication. Protein-coding genes 25,676 vs reported 25,663 (within 0.05%). Transcripts: deposit has 34,460 mRNA vs reported 43,578 (deposit carries fewer alt isoforms; flagged). BUSCO vertebrata_odb10 genome completeness (C8/C9) still computing. Full de-novo re-assembly deliberately out of efficient scope.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper aims to produce the first high-quality, chromosome-level genome assembly and annotation of the endangered Balearic lizard Podarcis lilfordi, and to assess whether its genome architecture (size, gene content, repeat content, synteny) is conserved relative to the related species Podarcis muralis despite ~18-20 million years of divergence.
- ★ First high-quality chromosome-level genome assembly and annotation of P. lilfordi, generated via a mixed sequencing strategy (10X linked reads, ONT long reads, Hi-C) plus RNAseq/Iso-Seq resource
- ★ Final assembly (rPodLil1.2) is 1.5 Gb with N50 = 90 Mb, 99% of sequence assigned to candidate chromosomes, and >97% gene completeness finding
- ★ Annotated 25,663 protein-coding genes translating into 38,615 proteins finding
- ★ Comparison to P. muralis reveals substantial similarity in genome size, annotation metrics, repeat content, and strong collinearity despite ~18-20 MYA evolutionary distance finding
- ★ Complete mitogenome of P. lilfordi assembled and annotated alongside the nuclear genome resource
- W sex chromosome contains 14 transposon-derived, short, ab-initio single-copy genes considered artefacts and removed from the final annotation finding
- Given the high quality/contiguity of the YaHS assembly, manual curation required only 17 edits method
- ★ Genome provides a critical resource for eco-evolutionary studies and conservation genomics of an insular, phenotypically diverse species resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| 10X Genomics linked-read sequencing | liver-derived HMW gDNA, single female P. lilfordi | none | base accuracy/polishing input for genome assembly | Illumina NovaSeq 6000, 2x151bp; Chromium Controller |
| Oxford Nanopore long-read sequencing | liver-derived HMW gDNA, single female P. lilfordi | none | genome contiguity and repeat resolution | GridION Mk1, R9.4.1 flow cell, Guppy v4.3.4 |
| Hi-C (Omni-C) sequencing | muscle (heart) tissue, same specimen | none | chromosome-level scaffolding | Dovetail Omni-C kit, Illumina NovaSeq 6000 2x151bp; YaHS scaffolder |
| Short-read RNA-seq | heart, kidney, liver, lungs, tail tissues pooled from multiple individuals | none | transcript evidence for gene annotation | Illumina TruSeq Stranded mRNA kit, NovaSeq 6000 2x150bp |
| Long-read Iso-Seq RNA sequencing | pooled sample of five tissues from five individuals | none | full-length isoform identification for annotation | PacBio Iso-Seq |
| Whole-genome alignment / synteny analysis | P. lilfordi vs P. muralis (PodMur1.0) and Lacerta agilis (rLacAgi1.pri) genome assemblies | none | scaffotyping, collinearity/synteny assessment | nucmer4, Minimap2 (-x asm5), Dot, pafr |
| Contamination screening (BlobTools/Blobtoolkit) | curated genome assembly scaffolds | none | identification and removal of contaminated scaffolds via GC/coverage/taxonomy | BlobTools v1.1, BlobToolKit, NCBI nt database, BUSCO odb10 |
| Mitochondrial genome assembly and annotation | filtered ONT reads from same specimen | none | mitogenome sequence assembly | — |
- ▲ Final chromosome-level assembly spans 1,460,440,873 bp with N50 = 89.54 Mb and scaffold L90 = 17, close to chromosome number (n=19) N50=89.54 Mb
- ▲ 99% of sequence assigned to candidate chromosomal sequences with >97% gene completeness >97%
- – Annotation yielded 25,663 protein-coding genes and 38,615 proteins 25,663 genes / 38,615 proteins
- – Strong collinearity and similarity in genome size, annotation metrics and repeat content between P. lilfordi and P. muralis despite evolutionary distance
- – Initial Flye assembly from ONT reads totaled 1.54 Gb with N50 = 1.36 Mb and consensus quality QV = 27.19 QV=27.19
- ▼ purge_dups removed 4,289 scaffolds accounting for 94,919,135 bp of duplicated/artifactual sequence 94,919,135 bp removed
- – After PCR duplicate removal (46.34%), 130,208,672 Omni-C read pairs were used for YaHS scaffolding, producing an assembly with only 17 manual curation edits needed 46.34% duplicates removed
- ▼ Six contaminated scaffolds removed based on GC cutoff (0.3-0.65), yielding final assembly rPodLil1.2 6 scaffolds removed
- other 1.5 Gb genome size (1,460,440,873 bp) (final chromosome-level assembly span)
- other N50 = 90 Mb (89.54 Mb) (assembly contiguity)
- count 25,663 protein-coding genes; 38,615 proteins (gene annotation output)
- other >97% gene completeness (BUSCO vertebrata_odb10) (assembly/annotation completeness)
- other 99% of sequence assigned to candidate chromosomes (chromosome assignment)
- other ~18-20 MYA divergence from P. muralis (evolutionary distance used for comparison)
- other 50.8 Gb filtered ONT reads, ~34x coverage, read N50 = 31.4 Kb, mean Phred quality 9.8 (long-read sequencing depth/quality)
- count 130,208,672 Hi-C read pairs retained after 46.34% PCR duplicate removal (Hi-C scaffolding input)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genome-resource paper reporting a chromosome-level assembly and annotation of Podarcis lilfordi from a single female specimen, using a combination of 10X Genomics, Oxford Nanopore, Hi-C, and RNAseq data. Results are presented mainly as descriptive assembly, completeness, and annotation metrics (e.g., N50, BUSCO scores, consensus quality QV, coverage ratios, collinearity with a related species) rather than through inferential statistical hypothesis testing across experimental groups.
-
Assembly and annotation quality metrics (N50, BUSCO completeness, consensus QV, coverage ratios) are reported as single summary values.↳ Could also: Reporting bootstrap-based confidence intervals or per-scaffold/per-window distributions (e.g., coverage variability, QV variability across regions) alongside the point estimates — This would additionally convey the uncertainty or heterogeneity underlying the summary metric, which can be useful context for downstream users of the assembly, though point-estimate reporting is standard practice in genome resource papers.
-
Sex-chromosome coverage was compared to autosomal mean coverage as a ratio to infer chromosome identity/ploidy.↳ Could also: A formal statistical comparison (e.g., a distributional test comparing coverage across scaffold classes, or a mixture-model approach) could also be used — This could provide a formal statistical basis (e.g., a p-value or likelihood ratio) for the sex-chromosome assignment in addition to the descriptive ratio, which is a common and sufficient approach in genome assembly papers of this kind.
-
Collinearity and similarity between the P. lilfordi and P. muralis genomes were assessed and visualized descriptively (whole-genome alignments, dot plots).↳ Could also: Quantitative synteny statistics (e.g., percentage of genome in syntenic blocks with associated significance testing, or rearrangement-rate estimates) could also be reported — This would add a quantitative, testable dimension to the comparative genome analysis, complementing the visual/descriptive comparison already provided.
-
Gene annotation combined multiple ab initio predictors, homology evidence, and transcript evidence via EvidenceModeler consensus, without reporting a false discovery rate for annotated gene models.↳ Could also: Reporting an estimated false-positive or false-negative rate for the final gene set (e.g., via manual validation of a random subsample or comparison to a curated reference) could also be included — This would give readers an explicit sense of annotation error rates, complementing the completeness metrics (e.g., BUSCO) already reported.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37137526
Paper: Gomez-Garrido et al. 2023, Chromosome-level genome assembly of Lilford's wall lizard, Podarcis lilfordi. DNA Res. PMID 37137526 · DOI 10.1093/dnares/dsad008.
Assembly: GCA_947686815.1 (rPodLil1.2), chromosome-level. Raw reads: ENA PRJEB50294 (ONT GridION + Illumina 10X linked + Hi-C/Omni-C). Pipeline repo: https://github.com/cnag-aat/assembly_pipeline (third-party CNAG-AAT tool). Sequencing tech (important): NO PacBio HiFi. Base assembly = Flye on ONT nanopore reads, polished (MaSuRCA/NextPolish), purged (purge_dups), scaffolded with linked reads (Tigmint/ARKS/LINKS) + Hi-C (YaHS).
Reproduction strategy (honest 1:1, third-party-tool-on-paper-data)
A full de-novo re-assembly from raw reads (60 Gb ONT + 95 Gb 10X + Hi-C, 1.46 Gb genome, multi-tool Flye→polish→purge→scaffold→Hi-C pipeline) would need hundreds of CPU-hours and exact tool-version pinning across ~10 tools. That is the maximal reproduction; it is not the efficient path to clear data points and risks non-determinism at every stage.
Instead we recompute the reported pipeline-derived QC/contiguity metrics directly on the deposited final assembly (and its deposited annotation). These metrics ARE pipeline outputs the paper reports as its headline results, and they are exactly reproducible by running standard tools (assembly-stats, BUSCO, gffread) on the public deposit. The BRIEF explicitly blesses this ("third-party tool on the paper's data is equally valid").
IN SCOPE (pipeline-derived, attempted)
| # | Result | Reported | Method to reproduce | Tier |
|---|---|---|---|---|
| C1 | Total assembly length | 1,460,085,851 bp (1.46 Gb), Table 1 | sum seq lengths of GCA FASTA | quick |
| C2 | Number of scaffolds | 2,148, §3.1 | count records in FASTA | quick |
| C3 | Scaffold N50 | 89.64 Mb, §3.1/Table S4 | N50 over scaffolds | quick |
| C4 | Contig N50 | 1.48 Mb, §3.1 | split on N-runs, N50 over contigs | quick |
| C5 | Chromosome-level scaffolds | 20 (18+2 sex), §3.1/Fig2 | count chromosome-named seqs / ≥N50-class | quick |
| C6 | Unplaced scaffolds | 2,084; unlocalized 44, §3.1 | classify from assembly_report | quick |
| C7 | Mitogenome size | 17,251 bp, §3.4 | locate mito seq in deposit | quick |
| C8 | BUSCO genome completeness | 98.3% (vertebrata_odb10, BUSCO 5.0.4), §3.1 | run BUSCO genome mode | medium (SLURM) |
| C9 | BUSCO single-copy complete | 96.7%, Fig2 | BUSCO output | medium |
| C10 | GC content (nuclear) | not stated in paper (mito GC 38.73%) | recompute — new value, report only | quick |
| C11 | Protein-coding genes | 25,663, Table 1 | count genes in deposited GFF3 | quick (if GFF avail) |
| C12 | Transcripts | 43,578, Table 1 | count mRNA in GFF3 | quick (if GFF avail) |
| C13 | Repeat content | 565.5 Mb (38.7%), Table 2 | RepeatMasker w/ deposited lib | heavy (stretch) |
OUT OF SCOPE (not pipeline-derivable from the deposit, or wet-lab/manual)
- QV 40 (nuclear) / 44.12 (mito): needs Merqury + raw read k-mer DB — stretch only if time.
- False duplication rate 0.68%: Merqury-derived, same dependency.
- Karyotype 2n=38 / flow-cytometry genome size: wet-lab, out of scope.
- Hi-C contact map / manual curation calls: manual, out of scope.
- Functional annotation % (72%): depends on external DB versions, out of scope.
Datasets to profile (same pass)
- PRJEB50294 (raw reads, 20 runs) — primary.
- GCA_947686815.1 / PRJEB47961 (assembly deposit) — the object we recompute on.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.