Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A worldwide map of swine short tandem repeats and their associations with evolutionary and environmental adaptations.

Genet Sel Evol · 2021
L1 63/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
63/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 24% of all assessed papers rank 875 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described-well-enough: YES (repo + Methods specify tools/params). Result is a MIX of exact 1:1 reproduction and an auditable deposit discrepancy. (1) C1 reference STR catalog RE-COMPUTED from scratch with TRF v4.09 on Sscrofa11.1 (Ensembl r95) following 00_TRF_reference_pig11.bash exactly: reproduced 2,806,669 STRs (mono+2-6bp) ~ paper's 2.8M, and 1,716,675 2-6bp STRs ~ paper's 1.71M -- both EXACT. (2) C3 ChIP-seq consensus peaks reproduced WITHIN TOLERANCE from the deposited MACS2 peak files: union-merge 15,442 H3K4me3 / 69,155 H3K27ac vs reported 15,196 / 68,495 (the paper's stricter 'present in all three' filter explains the ~1-2% reduction). (3) C2 pSTR catalog is the one MISMATCH: the deposited catalog (823,667) and the authors' OWN 82w.pSTR_stat.xlsx are 6.3% below the paper-text figure (878,967), consistently across all five motif classes -- an auditable deposit-vs-paper discrepancy flagged 'possible' for human review (NOT confirmed fabrication; the pSTR set cannot be re-genotyped because the per-individual genotype matrix is not deposited). NOT ATTEMPTED (out of scope, justified): 394-WGS STR genotyping and all downstream population-genetics results (Fst/Rst/Nei's D/breed-specific/STR-expansion/GLM) -- they read the undeposited genotype matrix; and MACS2 re-calling from raw reads. No values fabricated.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 40
    assessed: 2026-06-20 ⛓ 513347a1dfd3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whether genome-wide short tandem repeats (STRs) in diverse pig populations show evidence of purifying selection, provide accuracy for breed identification comparable to or better than SNPs, and are associated with signatures of domestication (domestic vs. wild boar differentiation) and environmental adaptation (temperature, altitude).

Core claims
  • Identified 878,967 polymorphic STRs (pSTRs) from 394 deep-sequenced pig/Suidae genomes, the largest pSTR repository in pigs to date resource
  • pSTRs located in coding regions are affected by purifying selection finding
  • Trinucleotide STRs are enriched in CDS, 5'UTR and H3K4me3 regions, suggesting they are important functional components of exons and promoters mechanism
  • pSTRs provide comparable or even greater accuracy than SNPs in determining breed identity of individuals finding
  • A set of pSTRs shows significant population differentiation between domestic pigs and wild boars in both Asia and Europe finding
  • Specific pSTRs are significantly associated with environmental variables (annual temperature, altitude) in Chinese indigenous breeds, including loss-of-function/expanded STRs overlapping AHR, LAS1L and PDK1 finding
  • Some domestication- or environment-associated pSTRs show stronger signals than flanking SNPs within a 100-kb window finding
  • lobSTR-based STR genotyping is reliable, shown by high concordance in replicate samples and against HipSTR calls method
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome sequencing / STR genotyping (lobSTR) 394 pigs and other Sus/Suidae species (domestic breeds, Asian/European wild boars, outgroups) none STR allele calls, genotypes, allele counts lobSTR software, BWA v0.7.17
Tandem Repeat Finder (TRF) genome annotation Sscrofa11.1 reference genome none STR loci locations and counts TRF v4.09
Repeat element annotation (RepeatMasker/RepeatModeler) Sscrofa11.1 reference genome none SINE, LINE, LTR, DNA_TE repeat element locations RepeatMasker v4.0.7, RepeatModeler v1.0.11
Principal component analysis and neighbor-joining phylogenetics 394 pig/Suidae samples none population structure, genetic distance EIGENSOFT/smartPCA v6.1.4, MEGA v7, iTOL
Population differentiation scan (Rst statistic) domestic pigs vs. Asian and European wild boars none Z-transformed Rst values, candidate differentiated regions (CDR)
GLM association with environmental variables 157 Chinese indigenous domestic pigs none association of ~0.3 million STR dosages with annual temperature and altitude R lm() function, WorldClim 2.0 bioclimatic data
ChIP-seq (H3K4me3, H3K27ac) liver samples from 3 pigs none active promoter/enhancer peak locations, STR enrichment in peaks BWA v0.7.17, MACS
Cross-platform STR genotype concordance comparison (lobSTR vs HipSTR) 61 selected pigs; 16 replicate sample pairs from a heterogeneous population none allele/genotype concordance rate, allelic dosage correlation lobSTR, HipSTR
Key results
  • 878,967 polymorphic STRs identified, a >20% increase over the previous largest pig pSTR set (630,906) 878,967 vs. 630,906 (>20% increase)
  • ACAGCC hexanucleotide motif is enriched in SINE elements Odds Ratio = 3.50, P < 2.2 × 10^-16
  • Genotype concordance rate between 16 pairs of replicated samples 97.6% concordance
  • Concordance between lobSTR and HipSTR genotype calls across shared alleles 93.3% concordant (14,415,224/15,447,865 alleles); 0.6% inconsistent
  • Correlation of allelic dosages between HipSTR and lobSTR call sets R2 = 0.91
  • Mean number of alleles per pSTR locus, with dinucleotide STRs showing the highest average allele count and hexanucleotide the lowest mean = 4.6 overall (dinucleotide 6.39, hexanucleotide 3.64)
  • 1.68 million STRs genotyped across 394 samples at average sequencing depth 21.6x average depth (range 5.0x–45.7x)
  • H3K4me3 and H3K27ac ChIP-seq merged peaks retained for enrichment analysis 68,495 H3K27ac peaks; 15,196 H3K4me3 peaks
Key statistics
  • count 878,967 polymorphic STRs (total pSTRs identified across all samples)
  • fold_change >20% increase (878,967 vs. 630,906) (comparison to previous largest pig pSTR study)
  • pvalue P < 2.2 × 10^-16 (chi-square test for ACAGCC motif enrichment in SINE elements)
  • other Odds Ratio = 3.50 (enrichment of ACAGCC motif in SINE elements)
  • correlation R2 = 0.91 (goodness-of-fit between HipSTR and lobSTR allelic dosages)
  • mean 97.6% concordance rate (genotype concordance in replicated low-depth (~7x) sample pairs)
  • count 14,415,224 concordant / 93,186 inconsistent out of 15,447,865 alleles (290,147 STRs) (lobSTR vs HipSTR genotype comparison)
  • mean mean alleles per locus = 4.6, median = 3, range 2–38 (allele count distribution across pSTR loci)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an observational population-genomics study that genotyped short tandem repeats (STRs) from whole-genome sequence data of 394 pig and outgroup individuals and characterized them descriptively (allele counts, motif frequencies) and through exploratory multivariate methods (PCA, neighbor-joining trees). Population differentiation between domestic pigs and wild boars was assessed with the Rst statistic using a top 5‰ empirical outlier threshold, associations between STR dosage and environmental variables (temperature, altitude) were tested with an ordinary linear model corrected for multiplicity by the Bonferroni method, and enrichment of STR types in genomic features was tested with chi-squared tests. Results were reported mainly as summary statistics (means, medians, ranges, percentages), odds ratios, and threshold-based significance calls rather than through a single unified inferential framework.

Replicationmixed Sample sizeSample sizes given as counts of individuals per breed/population (e.g., 394 total samples, 157 Chinese indigenous pigs used for the environmental GLM, 61 pigs for the lobSTR/HipSTR comparison, 16 pairs of replicated low-depth samples for genotyping concordance); no formal power/sample-size calculation stated Groupsdomestic pig breeds/populations vs. Asian and European wild boars; STR dosage regressed against environmental variables (temperature, altitude) Pairingunclear Randomization/blindingnot stated Dispersionrange Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBonferroni correction (for the environmental-association GLM); an empirical top 5‰ threshold on Rst values (for the population-differentiation scan)
Statistical tests used
Test Applied to n Assumptions
Rst (fixation index for STRs), Z-transformed, top 5‰ outlier threshold population differentiation between domestic pigs and Asian/European wild boars not stated
Ordinary (general) linear model, R lm() function association of STR dosage with annual mean temperature and altitude, with top 3 PCs as covariates 157 Chinese indigenous pigs not stated
Chi-squared test (chisq.test in R) fold-enrichment of STR types/motifs (e.g., ACAGCC) in genomic features such as SINE and H3K4me3 regions not stated
Goodness-of-fit correlation (R2) concordance between lobSTR and HipSTR genotype calls 61 pigs / 290,147 shared STRs na
Approaches that could also have been used
  • Population differentiation (Rst) was flagged using a top 5‰ empirical outlier threshold across the genome.
    Could also: a permutation-based empirical p-value or a formal FDR procedure (e.g., Benjamini-Hochberg) applied to the genome-wide Rst distribution — would express significance as a calibrated false-discovery rate rather than a fixed percentile, which can make the stringency of the cutoff more directly comparable across datasets with different numbers of loci
  • Environmental associations (temperature, altitude) with STR dosage were tested with an ordinary linear model including the top 3 PCs as covariates, corrected with Bonferroni.
    Could also: a mixed linear model that includes a genome-wide kinship/relatedness matrix as a random effect (e.g., as implemented in GEMMA or EMMAX), potentially paired with FDR correction — a mixed-model approach can capture cryptic relatedness and fine-scale population structure beyond what a small number of principal components represent, and FDR correction can offer more power than Bonferroni while still controlling the expected proportion of false positives
  • Enrichment of STR motifs/types in genomic features (e.g., SINE elements, H3K4me3 regions) was assessed with chi-squared tests.
    Could also: Fisher's exact test — Fisher's exact test avoids the large-sample approximation used by the chi-squared test and can be preferred when some contingency-table cell counts are small
  • Population structure was summarized primarily through PCA and a distance-based neighbor-joining tree.
    Could also: model-based ancestry estimation such as ADMIXTURE or STRUCTURE — these methods estimate explicit individual ancestry/admixture proportions under a statistical model and can complement PCA and distance-based trees, particularly for populations with mixed ancestry
  • The neighbor-joining tree was constructed from pairwise genetic distances without a stated measure of branch support.
    Could also: bootstrap resampling of loci, or a maximum-likelihood/Bayesian phylogenetic method that reports posterior support values — support values would convey the statistical confidence in specific branching patterns of the tree, which distance-based neighbor joining alone does not provide
  • Agreement between lobSTR and HipSTR genotype calls was summarized with a goodness-of-fit R2.
    Could also: an agreement statistic such as Cohen's kappa (for categorical genotype calls) or a Bland-Altman-style comparison (for allele length/dosage) — these approaches are commonly used to characterize agreement between two measurement methods and can provide complementary information to a correlation-based R2, such as systematic bias between the two callers
Software: Tandem Repeat Finder (TRF) 4.09 · RepeatMasker 4.0.7 · RepeatModeler 1.0.11 · BWA 0.7.17-r1188 · lobSTR · HipSTR · EIGENSOFT/smartPCA 6.1.4 · MEGA 7 · R (lm, chisq.test, ClusterProfiler)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33892623

Paper: Wu Z et al. (2021) A worldwide map of swine short tandem repeats and their associations with evolutionary and environmental adaptations. Genet Sel Evol 53:35. PMID 33892623 · PMCID PMC8063339 · DOI 10.1186/s12711-021-00631-4. Repo: https://github.com/jxlabWzZ/Susrepeats (commit 00caec6, 2019-10-26 — the only commit).

Pipeline overview (from Methods + repo)

  1. Reference STR catalogTandem Repeat Finder (TRF) v4.09 on the Sus scrofa Sscrofa11.1 reference (Ensembl release 95), per chromosome, params 2 7 7 80 10 20 100 -d -h; then length/repeat filtering + bedtools overlap resolution → set of 2–6 bp STR loci. (00_TRF_reference_pig11.bash)
  2. Population STR genotypingbwa mem map 394 WGS samples to Sscrofa11.1, lobSTR v3.0.3 allelotype classify, then VCF filtering (loc-cov 5, loc-log-score 0.8, loc-call-rate 0.6, loc-max-ref-length 80). Cross-validation with HipSTR on 16 replicate pairs / 61 samples. (01_lobSTR_STR_genotype&filter.bash, scripts/HipSTR_16_pairs.sh)
  3. Downstream population genetics on the genotype matrix (per-individual STR copy-number "GB" genotypes): allele freq/heterozygosity (a_gb2af.py), genotype recode (b_gb2gt.py), modal allele (c_Showmax.py), breed-specific alleles (d_Breed_specific.py), STR expansion (e_STRexpansion.py), Fst (f_Fst.py), Rst (g_Rst.py), Nei's D (h_neisD.sh), het sampling (i_Sampling_10_heho.py), motif canonicalisation (j_motif.py).
  4. Functional annotation — pSTR overlap with CDS/UTR and with H3K4me3 / H3K27ac ChIP-seq peaks. ChIP-seq peaks called with MACS2 v2.1.1 callpeak --broad from pig ChIP-seq (ENA PRJEB6906, Villar et al. 2015, 20-mammal enhancer study; pig runs ERR572*). GLM env-association (GLM_info.xlsx).

IN SCOPE (pipeline-derived, attempted)

# Result Pipeline Re-runnable? Status
C1 Total STRs in reference (2.8 M) and 2–6 bp set (1.71 M) TRF v4.09 YES — deterministic on the reference RUN on «our HPC»
C2 pSTR catalog size + per-motif-length breakdown (878,967; di/tri/tet/pen/hex) TRF+lobSTR catalog catalog re-countable from shipped 82w.p.ref.bed DONE (deposit re-count)
C3 Functional-region ChIP-seq peak counts (H3K4me3 / H3K27ac) bowtie2 + MACS2 on PRJEB6906 YES — third-party tool on the paper's data optional add (commands shipped)

OUT OF SCOPE / not attempted (and why)

  • Population STR genotyping from 394 WGS — the per-individual genotype matrix is not shipped in the repo, and re-genotyping 394 whole genomes (multiple TB of FASTQ across 13 BioProjects, lobSTR v3.0.3 which is long-unmaintained python2 software) is far beyond a faithful "few clear data points" pass. The raw WGS accession in the brief (PRJEB6906) is in fact the ChIP-seq project, not the WGS — the WGS sits under PRJNA398176/PRJNA488327/PRJEB1683/… (13 projects). Genotyping NOT attempted.
  • Fst / Rst / Nei's D / breed-specific / STR-expansion / GLM — every one of scripts/{a,c,d,e,f,g,h,i}.* reads the per-individual GB genotype matrix, which is not deposited. They cannot be re-run; the shipped result tables (*.xlsx) are only the outputs. We therefore VERIFY those deposits against the paper text (deposit-vs-paper consistency, fabrication check) rather than re-compute them.
  • Wet-lab / manual — none material; the paper is computational.
C1_total_2.8M
Reported
2.8 million STRs in reference (Sscrofa11.1)
Reproduced
2,806,669 (mono 1,089,994 + 2-6bp 1,716,675)
exact
C1_2to6bp_1.71M
Reported
1.71 million 2-6bp STR reference set
Reproduced
1,716,675
exact
C2_pSTR_total
Reported
878,967 polymorphic STRs
Reproduced
823,667 in deposited catalog (matches authors' own stat sheet; -6.3% vs paper text)
did not match
C2_motif_breakdown
Reported
di 237,296 / tri 105,365 / tet 270,581 / pen 149,351 / hex 116,374
Reproduced
di 226,387 / tri 100,031 / tet 250,466 / pen 140,068 / hex 106,715 (deposit)
did not match
C3_H3K4me3_peaks
Reported
15,196 H3K4me3 merged peaks
Reproduced
15,442 (union-merge of deposited MACS2 peaks)
within tolerance
C3_H3K27ac_peaks
Reported
68,495 H3K27ac merged peaks
Reproduced
69,155 (union-merge of deposited MACS2 peaks)
within tolerance
C4_pSTR_functional_overlap
Reported
6,605 H3K4me3 / 38,999 H3K27ac
Reproduced
8,728 / 69,838 (union consensus; over-counts)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 63/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The deposited STR catalog reproduces the authors' own stat sheet exactly at the motif level (di/tri/tet/pen/hex all match), so there is no fabrication concern — but both the deposit and that sheet (823,668) sit ~6.3% below the paper's printed headline of 878,967 polymorphic STRs, a paper-vs-deposit inconsistency on the authors' side. The reference-catalog totals (2.8M / 1.71M) are still being re-run with TRF v4.09 and remain ungraded, and all downstream evolutionary/environmental associations are unverifiable because the per-individual genotype matrix was never deposited. Severity is moderate (magnitude/direction of the catalog holds), so the central map-of-STRs conclusion stands with limited confirmation rather than being overturned.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

287.7 k
tokens (I/O) · 14.2 M incl. cache
84 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.