Genome-Wide Survey and Development of the First Microsatellite Markers Database (AnCorDB) in Anemone coronaria L.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction, completed end-to-end on «our HPC» compute nodes from the paper's own raw WGS run SRR18072428. Pipeline: download -> Scythe(built from vsbuffalo/scythe)+Sickle Q30 trim (C1) -> KMC k=19 genome size (C2) -> MEGAHIT default assembly (C3, 8.5h/69GB) -> MISA SSR mining with the paper's exact thresholds 1-15 2-8 3-5 4-4 5-4 6-4 (C4,C6,C7,C8,C9). STRONG agreement on SSR COMPOSITION: di/tri motif frequencies (C8,C9) reproduce within ~1-1.5% absolute, and raw contig N50 (2064) is identical. DIVERGENT on absolute magnitudes: our default-MEGAHIT assembly is smaller and more fragmented (filtered N50 2464 vs 6157; 4.89 vs 6.13 Gb), and absolute SSR counts scale down with it (C4 ~70%, C6 ~45% of reported); same MISA thresholds were used so the gap is assembly-driven. Genome-size order confirmed (8.99 vs 7.8 Gb). Honesty signal flagged for the auditor: paper's 'cleaned' volume (91.24 Gb) is only 0.9% below 92.04 Gb raw despite a stated Sickle Q<30 step, whereas a faithful Q30 trim removes ~9%. NOT attempted: C5 (SciRoKo = Windows GUI), C10 genic SSRs (MAKER-P, out of scope), wet-lab validation (lab work). The methods are described well enough to reproduce the third-party-tool pipeline; reproduction = applying documented tools to the authors' own data (no authors' pipeline repo exists).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-19 ⛓ 5b70e13150f1
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-26
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetDue to the lack of species-specific molecular markers for the large, highly heterozygous genome of Anemone coronaria, the study aims to generate the first draft genome sequence and perform a genome-wide survey of microsatellites to develop a public SSR marker resource for breeding, mapping, and fingerprinting applications.
- ★ Generated the first draft genome assembly of A. coronaria by Illumina sequencing a haploid androgenetic plant method
- ★ Identified 401,822 perfect and 188,987 imperfect SSR motifs genome-wide finding
- ★ Developed AnCorDB, the first public online SSR marker database for A. coronaria with integrated Primer3 primer design resource
- ★ Validated 62 SSR markers across 8 commercial cultivars for varietal fingerprinting finding
- About 75% of the assembled genome consists of repetitive elements finding
- Maker-P annotation identified 26,260 genes covering ~0.92% of the estimated genome size finding
- ★ Trinucleotide SSRs are disproportionately enriched in genic regions relative to the whole genome, consistent with negative selection against frameshift mutations mechanism
- ★ A reduced subset of 6 SSR markers reproduces the genetic relationships of the full 62-SSR set and enables cultivar discrimination finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome Illumina sequencing (NovaSeq) and de novo assembly (MEGAHIT) | Haploid androgenetic A. coronaria plant (cv. MISTRAL Magenta) | none | Genome assembly (scaffold number, N50, genome size) | Illumina NovaSeq |
| K-mer based genome size estimation | MISTRAL Magenta genotype sequencing reads | none | Estimated genome size | jelly-bean software |
| Genome-wide SSR mining | Assembled A. coronaria draft genome (in silico) | none | Number, type, and distribution of perfect/imperfect SSR motifs | SciRoKo |
| Repeat masking and transposable element annotation | Assembled draft genome | none | Proportion and classification of repetitive elements/TEs | RepeatMasker v4.1.0; RepeatModeler v1.0.11 |
| Structural and functional gene annotation | Masked assembled draft genome | none | Gene models, IPR domains, GO categories | Maker-P v2.31.08, Augustus v3.3.2, SNAP, InterProScan |
| PCR-based SSR genotyping and varietal fingerprinting | 8 commercial cultivars (6 diploid, 2 tetraploid), Biancheri Creazioni | none | Allele number/length, PIC, genetic similarity, UPGMA/PCoA clustering | — |
| PCR-based SSR genotyping for intra-cultivar variability | 5 plants per cultivar across 8 cultivars, using 6 selected SSRs | none | Fixation index (FIS), observed/expected heterozygosity, intra-cultivar clustering | — |
| Online database development with primer design tool | Genome-wide SSR dataset (AnCorDB) | none | Primer pair design (Tm, GC content, amplicon length) | Primer3 |
- – Draft genome assembly covered ~6.13 Gb across ~2×10^6 scaffolds (N50=6157 bp), ~78.6% of the estimated 7.8 Gb genome
- – 401,822 perfect SSRs (density 65.52 SSR/Mb, including 42,111 compound SSRs) and 188,987 imperfect SSRs identified
- – Dinucleotides were the most abundant SSR class (60.2%), followed by trinucleotides (23.7%) 60.2%
- ▲ About 75% of genome sequences classified as repetitive elements after masking 75%
- – 26,260 genes annotated (AED ≤ 0.4), covering ~56.12 Mb (0.92%) of estimated genome size
- ▲ Trinucleotides were the most common class among genic perfect SSRs (38.2%), versus ~27% triplets in the whole genomic SSR set 38.2%
- – 62 validated SSRs generated 203 alleles across 8 cultivars (mean 3, range 1-8 alleles/locus); PIC ranged 0.13-0.85 (mean 0.52 ± 0.025) PIC mean 0.52
- ▲ Similarity matrix from a 6-SSR subset correlated well with the full 62-SSR matrix, supporting cultivar fingerprinting with fewer markers r=0.92
- count 401,822 perfect SSRs (genome-wide perfect SSR count)
- count 188,987 imperfect SSRs (genome-wide imperfect SSR count)
- other 65.52 SSR/Mb (perfect SSR density in assembled genome)
- other 75% (proportion of genome classified as repetitive elements)
- count 26,260 genes (genes annotated by Maker-P)
- count 203 alleles (alleles generated across 62 SSR loci in 8 cultivars)
- mean PIC mean 0.52 ± 0.025 (range 0.13-0.85) (polymorphism information content of 62 SSRs across 8 cultivars)
- correlation r = 0.92 (correlation between 6-SSR and 62-SSR similarity matrices)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper reports a genome-wide SSR survey of a draft Anemone coronaria genome followed by marker validation and cultivar fingerprinting. The analytical approach is primarily descriptive and tool-development oriented: SSR frequencies were summarised as counts and percentages, 62 validated SSR loci were scored across eight cultivars, and genetic relationships were inferred via UPGMA-based dendrograms with bootstrap support and Principal Coordinate Analysis (PCoA). Intra-cultivar variability was characterised using fixation indices (F_IS), Hardy-Weinberg equilibrium deviation testing, and polymorphism information content (PIC) to guide marker selection.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| UPGMA clustering with bootstrap resampling | Cultivar fingerprinting dendrogram (Figure 7) using 62 SSRs; intra-cultivar variability dendrogram (Figure 8) using 6 SSRs | 8 genotypes (one per cultivar) for Figure 7; 40 plants (5 per cultivar × 8 cultivars) for Figure 8 | not stated |
| Principal Coordinate Analysis (PCoA) | Genetic-relationships scatter plot among cultivars (Figure 7) and extended germplasm panel (Figure 8) | 8 genotypes / 40 plants respectively | not stated |
| Hardy-Weinberg Equilibrium deviation test (specific test not named) | Intra-cultivar variability assessment across 6 SSR loci in 5 plants per cultivar | 5 plants × 8 cultivars = 40 plants; test type not stated | not stated |
| Distance-matrix correlation (Mantel-type; r reported, permutation significance not stated) | Comparison of genetic similarity matrix derived from 6-SSR subset versus 62-SSR full set (r = 0.92) | 8 cultivars | not stated |
| Fixation index (F_IS) calculation | Intra-cultivar variability assessment; values ranged from −0.68 to 0.79 | 5 plants per cultivar × 8 cultivars, 6 SSR loci | na |
| Polymorphism Information Content (PIC) calculation | Marker informativeness assessment across 62 validated SSR loci (range 0.13–0.85; mean 0.52 ± 0.025) | 8 genotypes (one per cultivar) | na |
-
The correlation between the 6-SSR and 62-SSR genetic similarity matrices was reported as r = 0.92 without a formal significance test↳ Could also: A Mantel test with permutation-based significance testing (e.g., 9999 permutations) could also be applied to this matrix correlation — Entries of a genetic distance matrix are not statistically independent, so standard parametric p-values for correlation are not valid; the Mantel test derives an empirical null distribution via permutation and provides a valid p-value and confidence bound for the observed matrix correlation
-
HWE deviation was assessed across multiple SSR loci, with outcomes reported only as 'significant' or 'no significant difference', and no multiplicity correction was mentioned↳ Could also: A Benjamini-Hochberg false-discovery-rate (FDR) adjustment, or Bonferroni correction, across all tested loci could also be applied — Simultaneously testing HWE across 6 (or 62) loci inflates the family-wise type I error rate; FDR control would help identify which deviations are robust across the full set of comparisons
-
UPGMA was used to infer cultivar phylogeny with bootstrap support↳ Could also: Neighbor-Joining (NJ) or model-based clustering (e.g., STRUCTURE / ADMIXTURE) could also have been used — NJ does not assume a molecular clock, making it less sensitive to rate heterogeneity than UPGMA; model-based approaches additionally estimate admixture proportions and probabilistic group membership, which is informative for the highly heterozygous, outcrossing biology of this species
-
PCoA was used to visualise genetic relationships in two dimensions alongside UPGMA↳ Could also: Discriminant Analysis of Principal Components (DAPC) could also be applied — DAPC maximises the ratio of between-group to within-group variance, which can sharpen separation among cultivars and is particularly useful when within-cultivar variability is substantial (as noted for 'Edge')
-
Marker informativeness was characterised solely by PIC (mean 0.52 ± unlabelled dispersion)↳ Could also: Expected heterozygosity (H_E), observed heterozygosity (H_O), allelic richness, and number of effective alleles could also be reported alongside PIC — These complementary statistics are standard in SSR validation studies and together convey information about allele frequency distribution, sampling-bias sensitivity, and cross-study comparability that PIC alone does not fully capture
-
The dispersion around the mean PIC (0.52 ± 0.025) was reported without labelling the measure as SD or SE↳ Could also: Explicitly labelling the dispersion as SD (reflecting marker-to-marker variability) or SE (reflecting uncertainty around the mean), or reporting a 95% CI, would also convey the information — SD and SE differ by a factor of √n and convey different aspects of variability; explicit labelling lets readers judge the spread of informativeness across markers (SD) versus confidence in the average (SE), which matter differently depending on how markers will be selected in practice
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35328546 (AnCorDB, Anemone coronaria microsatellites)
Paper: Martina et al. 2022, IJMS 23(6):3126. PMID 35328546 / PMC8949970. Genome-wide survey + SSR marker database (AnCorDB) for Anemone coronaria.
Pipeline (from Methods §3.1–3.2)
- Sequencing: Illumina NovaSeq 6000, PE 2×150, 270 bp insert, androgenetic haploid plant of cv. MISTRAL® Magenta. → SRA PRJNA808392 (single run SRR18072428, 92.04 Gb raw).
- Read QC/trim: Scythe v0.994 (adapter removal) + Sickle v1.33 (quality trim, Q<30). → "91.24 Gb cleaned reads".
- Genome size: k-mer analysis → ~7.8 Gb estimate.
- Assembly: MEGAHIT (default params). Draft ~4.7×10⁶ scaffolds → after removing <500 bp → 2×10⁶ scaffolds, N50 6157 bp, ~6.13 Gb (~78.6% of est.).
- Repeat masking / gene annotation: RepeatMasker 4.1.0 + custom lib (MITE-Hunter, LTRdigest/LTRharvest, RepeatModeler), Maker-P 2.31.08 + Augustus + SNAP. (annotation only feeds "genic SSR" count)
- SSR mining: SciRoKo (chop scaffolds) + misa.pl (MISA). Thresholds: mono ≥15, di ≥8, tri ≥5, tetra/penta/hexa ≥4; min total length 15 nt; compound max interruption 100 bp, ≤1 mismatch.
- Primer design: Primer3 (built into AnCorDB).
IN SCOPE (pipeline-derived, attempt to reproduce)
- C1 cleaned read volume (91.24 Gb) — trim pipeline. Feasible (I/O-heavy).
- C2 k-mer genome size (~7.8 Gb) — KMC/jellyfish + GenomeScope. Feasible.
- C3 assembly metrics (N50 6157, 6.13 Gb, 2M scaffolds, 78.6%) — MEGAHIT. HIGH compute risk: ~7.8 Gb plant genome assembly from 92 Gb reads needs a very large-memory node; may not fit «infra» std partition. Attempt; may block.
- C4–C7 SSR counts + motif-class breakdown (Table 1) — MISA on the assembly. Depends on C3. Central result.
- C8–C9 di/tri motif frequencies — from same MISA output.
- C10 genic SSR counts (3223/1261) — needs gene annotation (Maker-P); annotation is itself heavy + multi-tool → lower priority / likely not attempted.
OUT OF SCOPE (wet-lab / manual, NOT attempted)
- 150 primers selected → 62 SSRs validated across 8 cultivars; 203 alleles, mean 3/locus (§2.5–2.6). PCR + capillary electrophoresis genotyping.
- AnCorDB website itself.
Code note
Registry "code_url" = github.com/vsbuffalo/scythe = the Scythe trimmer only (step 2). No authors' own pipeline repo exists; reproduction = applying the documented third-party tools (Scythe/Sickle/MEGAHIT/MISA/SciRoKo/Primer3) to the paper's own data with the stated parameters (P16-valid).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.