Corpus 1,283 assessed · 1,184 scored · 647 reproduced ≥75 · 173 flagged ·∅ 73.9/100
← New search

Mining transcriptomic data to study the origins and evolution of a plant allopolyploid complex.

PeerJ · 2014
L1 54/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
54/100
Reproducibility score
1.1 SD below mean
vs. all fields · 1184 studies
🎯 Scores higher than 15% of all assessed papers rank 1002 of 1184 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the core bioinformatic pipeline (fastq-mcf trim Q30/min50bp -> Bowtie2 default alignment -> MAPQ>=20 filter -> bedtools gene coverage -> bcftools mpileup/call SNP calling, DP>=5 filter) for all 9 accessions in the RU's designated dataset (sra:SRP011928), using the third-party GenoToolBox MultiVcfTool.pl for the multi-sample VCF/hapmap merge step, against a modern RefSeq soybean assembly (GCF_000004515-family NC_016088.1-series) rather than the paper's original Phytozome Glyma v1.0 (unavailable in that exact historical form) - a documented, necessary methodological substitution. Result: 1:1-adjacent but NOT exact reproduction. Raw read counts match Table 1 closely (within ~2-18%) for 6/9 accessions; SRA data for accession 1134 is severely under-deposited (14% of reported reads) and 1403/1820 show unexplained shortfalls (32-49%) not resolvable from the cited archives. Bowtie2 pre-filter alignment counts run 10-36% above/below Table 1's 'Mapped reads' for adequately-represented accessions (best reading of that column: total aligned reads before the separate downstream MAPQ>=20 filter, since MAPQ-filtered counts are 3-4x too low to match Table 1 directly). DP>=5-filtered SNP counts run systematically at ~56-65% of Table 1 for well-represented accessions - consistent direction, plausibly explained by reference-genome/annotation differences, not proportional to data loss. Most importantly, the paper's central biological claim - that the G. dolichocarpa lineage (accessions 1134, 1188, 1393) retains much higher heterozygosity (26-29% in the paper) than the G. syndetika/tomentella-D3 diploid-progenitor lineages (6-13% in the paper) - is QUALITATIVELY REPRODUCED: all 9 reproduced accessions fall in the same high-het (42-44%) or low-het (10-19%) group as the paper, with a consistent ~1.4-1.7x systematic magnitude offset across ALL accessions (both groups), again pointing to a reference-genome/pipeline-version effect rather than a spurious or absent signal. NOT attempted (out of pipeline scope / would require additional tools+data not exercised here): SeparateHomeolog2Sam homeolog-separation, Structure/fineStructure population-structure analysis, PhyML/*BEAST phylogenetics and divergence dating, and full reproduction of the secondary SRP038128 dataset (12 further accessions, profiled at metadata level only). 3 Table-1 accessions (1286, 1854, 1364) could not be located in any cited or searched SRA archive and are recorded as a genuine data-availability gap.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-07
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-07
no human curator yet
Last updated
2026-08-07

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper tests whether three allopolyploid species of the perennial Glycine (G. tomentella T1, G. dolichocarpa/T2, and G. tomentella T5) originated as fixed hybrids combining the genomes of the two extant diploid progenitor species hypothesized from previous crossing and two-gene molecular systematic studies, and asks when and how many times these polyploids arose.

Core claims
  • All three allopolyploid species are fixed hybrids combining the genomes of the two putative diploid parents hypothesized on the basis of previous crossing and molecular work. finding
  • Based on mapping to the soybean reference genome, there appear to be no large regions for which one homoeologous contribution is missing. finding
  • Phylogenetic analyses of 27 selected transcripts using a coalescent approach are consistent with multiple origins for these allopolyploid species. finding
  • The origins of these allopolyploid species occurred within the last several hundred thousand years. finding
  • Transcriptomic (RNA-Seq) data generated for an unrelated light-stress study can be mined with phylogenetic and population-genomic approaches to resolve allopolyploid origins. method
  • Polyploid reads can be assigned to parental subgenomes by preferential mapping against a progenitor reference set built from consensus diploid transcriptomes, enabling homoeologue-resolved SNP datasets and transcript-guided assembly. method
  • Custom Perl tools (MultiVcfTool, SeparateHomeolog2Sam, FastaSeqExtract in the GenoToolBox package) are provided for multi-VCF formatting and homoeologue read separation. resource
  • Alignments whose trees did not place G. max sister to the perennial Glycine species were removed as likely paralogues from the ca. 50 My legume-wide WGD rather than orthologues. method
Experimental setups
Assay System Perturbation Readout Platform
Single-end RNA-Seq (transcriptome sequencing) Pooled leaflet tissue from six individuals per accession of eight perennial Glycine species (G. canescens, G. clandestina, G. dolichocarpa, G. syndetika, G. tomentella D1, D3, T1, T5) plus a synthetic allotetraploid A58; 2-5 accessions per species from the CSIRO Perennial Glycine Germplasm Collection Light-intensity treatment: low light (125 mmol m-2 s-1) vs excess light (800 mmol m-2 s-1) in a growth chamber, 12 h/12 h light/dark, 22 °C/18 °C Raw and processed read counts, reads mapped to the soybean genome, and number of represented G. max reference genes with expression >0 RPKM Illumina GAIIx (88 nt reads) or HiSeq 2000 (100 nt reads); Illumina mRNA-Seq Sample Prep Kit, Dynabeads mRNA DIRECT Kit, Qiagen Plant RNeasy Kit with on-column DNase
Read mapping and SNP calling Perennial Glycine RNA-Seq reads mapped to the Glycine max reference genome version 1.0 (Phytozome) none SNP genotypes (coverage >=5, unique mapping, mapping score >=20) exported to Structure and Hapmap formats Fastq-mcf, Bowtie2 (default parameters), Samtools, MultiVcfTool
Homoeologue read identification and transcript-guided (guided) transcriptome assembly Allopolyploid reads (T1, T2, T5) mapped against progenitor reference sets built from consensus diploid transcriptomes (T1 = D1 + D3; T2 = D3 + D4; T5 = A + D1) none Reads partitioned by preferential mapping to each progenitor subgenome; rebuilt polyploid transcriptomes and homoeologue-resolved SNPs Bowtie2, SeparateHomeolog2Sam, Samtools, Gffread (Cufflinks package)
Population structure analysis (Bayesian clustering) Two SNP datasets (with and without polyploid SNPs separated by homoeologue) from the Glycine accessions none Ancestry/admixture proportions and optimal number of clusters K Structure (admixture model, lambda = 1); three random subsets of 20,000 SNPs, 5 replicates, burn-in 10,000, 10,000 MCMC reps, K = 1-15; K = 6 re-run with 100,000 burn-in; visualized in R
Population structure analysis (chromosome painting / coancestry) The two SNP datasets divided into 20 subsets, each mapping to one soybean reference chromosome none Pairwise coancestry/distance matrix between accessions displayed as a heatmap, plus PCA of the same matrix fineStructure (unlinked model); PCA figure created in R
Phylogenetic reconstruction from concatenated SNP supermatrix 36 OTUs (allopolyploid homoeologues treated as separate OTUs, e.g. D1T1 and D3T1), with G. max accession Williams 82 as outgroup none Maximum likelihood tree topology with bootstrap support, and a phylogenetic network showing reticulation PhyML (GTR model, 100 bootstrap replicates); NeighborNet in SplitsTree (default parameters); FigTree
Gene-based phylogenetic analyses (single-gene trees, networks, dated Bayesian trees) 27 selected transcripts (filtered from 95 alignments) from 24 accessions including two homoeologues per allopolyploid accession; G. max Williams 82 as outgroup none Per-gene ML trees with bootstrap support, NeighborNet networks, and Bayesian divergence-time estimates with the root (G. max vs perennials) scaled to 5 Myr jModelTest2 (BIC model selection), PhyML (1,000 bootstraps), SplitsTree4 NeighborNet, BEAST v2.0 (HKY, 100,000,000 MCMC generations sampled every 1,000)
Coalescent species tree reconstruction 27 selected genes from 24 accessions grouped into 11 OTUs (G. canescens, G. clandestina, G. tomentella D1, G. tomentella D3, G. syndetika/D4, T1-D1, T1-D3, T2-D3, T2-D4, T5-A, T5-D1), with G. max included none Species tree topology under the multispecies coalescent *BEAST
Key results
  • All three polyploid species (T1, T2/G. dolichocarpa, T5) are fixed hybrids combining the two hypothesized parental diploid genomes.
  • No large genomic regions lacking one homoeologous contribution were detected when mapping to the soybean reference genome.
  • Coalescent and gene-tree analyses of 27 transcripts are consistent with multiple independent origins of the allopolyploid species.
  • Estimated origins of the allopolyploids fall within the last several hundred thousand years, based on trees scaled to a 5 Myr G. max/perennial root. within the last several hundred thousand years
  • Of 95 initial alignments screened by exploratory BEAST analysis, 27 genes passed the orthology filter (G. max sister to perennial Glycine); removed alignments showed two long-branch clades suggesting inclusion of paralogues from the ca. 50 My legume WGD. 27 of 95 alignments retained
  • HKY was the best-fitting substitution model for the plurality of the 27 genes by BIC, followed by K80. HKY 40% of genes; K80 26% of genes
  • Structure analysis identified an optimal clustering that was verified at K = 6 with an extended burn-in. K = 6
  • RNA-Seq mapping to the soybean reference recovered roughly 22,500-25,300 expressed reference genes (RPKM > 0) per accession. 22,571-25,278 represented genes per accession
Key statistics
  • other HKY preferred in 40% of genes; K80 in 26% of genes (BIC, jModelTest2) (Substitution model selection across the 27 filtered gene alignments)
  • count 27 genes retained from 95 alignments (Filtering of transcript alignments for gene-based and species-tree analyses)
  • count 36 OTUs in the concatenated SNP supermatrix; 11 OTUs from 24 accessions in the *BEAST species tree (Operational taxonomic units, with allopolyploid homoeologues treated separately)
  • count 202,427,873 raw reads / 187,120,918 processed / 60,712,525 mapped; 23,643 represented genes (Largest accession dataset, Glycine dolichocarpa accession 1134 (13 samples))
  • count 10,401,944 raw reads / 9,604,350 processed / 6,896,983 mapped; 22,802 represented genes (Smallest accession dataset, Glycine tomentella D3 accession 1364 (1 sample))
  • other three subsets of 20,000 SNPs, 5 replicates each, burn-in 10,000, 10,000 MCMC repetitions, K = 1 to K = 15; K = 6 re-run with 100,000 burn-in (Structure population structure run parameters)
  • other 100 bootstrap replicates (concatenated SNP ML tree); 1,000 bootstraps (27-gene ML trees); 10,000,000 MCMC (exploratory BEAST); 100,000,000 MCMC sampled every 1,000 (BEAST v2.0) (Phylogenetic analysis run parameters)
  • other root calibration 5 Myr (G. max vs perennial divergence); Glycine WGD ca. 5-10 Mya; legume-wide WGD ca. 50 Mya; chromosome numbers 2n = 38/40 (diploids), 78/80 (allopolyploids) (Calibration and taxon background used for divergence dating)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a population-genomics and phylogenetic framework rather than classical hypothesis testing: SNPs called from mapped RNA-Seq reads were analyzed with Bayesian clustering (Structure) and coancestry-based clustering (fineStructure/PCA) to assess population structure, and with maximum-likelihood (PhyML) and Bayesian (BEAST, *BEAST) phylogenetic/coalescent methods, plus NeighborNet networks, to reconstruct relationships and divergence times among diploid and allopolyploid Glycine lineages. Support for relationships was conveyed through bootstrap percentages and posterior probabilities rather than through p-values.

Replicationunclear Sample sizeNumber of accessions per species (2–5) and RNA-Seq samples/libraries per accession (1–13, pooled from six individuals each) are listed in Table 1; no formal power analysis is described for the phylogenetic/population-genomic analyses. Groupsdiploid progenitor species vs. their corresponding allopolyploid species (and separated homoeologous subgenomes) within three Glycine 'triads' Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno
Statistical tests used
Test Applied to n Assumptions
Maximum-likelihood phylogenetic reconstruction (PhyML, GTR model), 100 bootstrap replicates concatenated homoeologue SNP supermatrix, 36 OTUs 36 OTUs (SNP supermatrix) not stated
NeighborNet network reconstruction (SplitsTree) same concatenated SNP supermatrix, and the 27-gene alignment set na
Maximum-likelihood phylogenetics (PhyML, HKY/K80 models selected via jModelTest2/BIC), 1,000 bootstrap replicates 27 selected gene alignments 27 genes not stated
Bayesian MCMC phylogenetic analysis (BEAST v2.0, HKY model, 100,000,000 generations) 27 selected gene alignments, 11 OTUs (homoeologues treated as separate OTUs) 27 genes / 24 accessions not stated
Bayesian admixture clustering (Structure, K=1–15, 5 replicates per subset) SNP datasets (with and without homoeologue-separated polyploid SNPs) 3 random subsets of 20,000 SNPs each not stated
Coalescent-based species tree estimation (*BEAST) 27 selected genes, 11 OTUs 27 genes not stated
Approaches that could also have been used
  • Species/homoeologue relationships were inferred using concatenated-SNP ML trees, Bayesian gene trees, and NeighborNet networks.
    Could also: Explicit phylogenetic network or admixture-graph methods (e.g., TreeMix, SNAQ, PhyloNet) — These approaches explicitly model reticulation/hybridization events and can quantify admixture proportions, which could complement the network visualization already used to depict reticulate allopolyploid origins.
  • Population structure was assessed with Structure (Bayesian MCMC admixture model) and fineStructure (coancestry-based clustering/PCA).
    Could also: ADMIXTURE (maximum-likelihood-based clustering) — ADMIXTURE is computationally faster for larger SNP panels and often yields comparable cluster estimates to Structure, offering a useful cross-check on the inferred K and admixture proportions.
  • Nucleotide substitution models were selected using the Bayesian Information Criterion (BIC) in jModelTest2.
    Could also: Akaike Information Criterion (AIC) or corrected AIC (AICc) — Comparing model choice across BIC, AIC, and AICc can show whether the selected model (e.g., HKY) is robust to the criterion used, since these criteria can occasionally favor different models.
  • Divergence times were estimated by scaling the BEAST tree root to a fixed external date (5 Myr).
    Could also: Calibration with explicit prior distributions (e.g., fossil- or secondary-calibration priors) and reporting of highest posterior density (HPD) intervals — Using calibration priors with reported HPD intervals would propagate dating uncertainty into the divergence-time estimates rather than relying on a single fixed calibration point.
  • A multispecies coalescent species tree was estimated from 27 selected genes using *BEAST.
    Could also: Summary coalescent methods such as ASTRAL or SVDquartets — These methods scale efficiently to larger numbers of loci or genome-wide SNP data and can serve as an independent, faster-to-compute complement to the full Bayesian coalescent analysis.
  • Branch support was reported as ML bootstrap percentages and Bayesian posterior probabilities.
    Could also: Gene concordance factors and site concordance factors — Concordance factors quantify the proportion of underlying gene trees or sites actually supporting a branch, which can be informative in groups—like this one—where incomplete lineage sorting and introgression are suspected to affect topology.
Software: Fastq-mcf · Bowtie2 · Samtools · Cufflinks/Gffread · Structure · fineStructure · R (barplot, PCA figure) · PhyML · SplitsTree (NeighborNet) SplitsTree4 · jModelTest2 · BEAST v2.0 · *BEAST · FigTree

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

1820_raw_reads
Reported
71,185,274 raw reads (3 samples, Table 1)
Reproduced
36,597,945 raw reads pooled from 3 SRA runs actually deposited in SRP011928 (single-end)
partial
1820_mapped_reads
Reported
18,625,439 mapped reads (Table 1)
Reproduced
22,404,200 bowtie2-aligned reads (>=1 time, pre-MAPQ-filter; 7,078,497 remain after MAPQ>=20 filtering used downstream for SNP calling)
partial
1820_genes_represented
Reported
22,871 represented genes (RPKM>0 vs Phytozome Glyma1, Table 1)
Reproduced
35,678 genes with >=1x bedtools coverage (raw count, not RPKM-normalized, vs modern RefSeq soybean assembly, 50,961 total annotated genes)
partial
1820_filtered_snps
Reported
641,145 SNPs after DP>=5 filter (Table 2)
Reproduced
365,094 SNPs after DP>=5 filter (bcftools mpileup+call -mv)
partial
1820_pct_heterozygous
Reported
6.6% heterozygous SNPs (Table 2)
Reproduced
10.44% heterozygous SNPs; qualitative high/low-het group assignment MATCHES paper
partial
1134_raw_reads
Reported
202,427,873 raw reads (13 samples, Table 1)
Reproduced
28,389,777 raw reads pooled from 3 SRA runs actually deposited in SRP011928 (single-end)
did not match
1134_mapped_reads
Reported
60,712,525 mapped reads (Table 1)
Reproduced
19,224,805 bowtie2-aligned reads (>=1 time, pre-MAPQ-filter; 4,790,585 remain after MAPQ>=20 filtering used downstream for SNP calling)
did not match
1134_genes_represented
Reported
23,643 represented genes (RPKM>0 vs Phytozome Glyma1, Table 1)
Reproduced
33,610 genes with >=1x bedtools coverage (raw count, not RPKM-normalized, vs modern RefSeq soybean assembly, 50,961 total annotated genes)
partial
1134_filtered_snps
Reported
965,643 SNPs after DP>=5 filter (Table 2)
Reproduced
292,869 SNPs after DP>=5 filter (bcftools mpileup+call -mv)
did not match
1134_pct_heterozygous
Reported
26.4% heterozygous SNPs (Table 2)
Reproduced
44.07% heterozygous SNPs; qualitative high/low-het group assignment MATCHES paper
partial
1188_raw_reads
Reported
19,034,633 raw reads (2 samples, Table 1)
Reproduced
18,305,177 raw reads pooled from 2 SRA runs actually deposited in SRP011928 (single-end)
within tolerance
1188_mapped_reads
Reported
11,960,713 mapped reads (Table 1)
Reproduced
13,225,084 bowtie2-aligned reads (>=1 time, pre-MAPQ-filter; 4,502,234 remain after MAPQ>=20 filtering used downstream for SNP calling)
within tolerance
1188_genes_represented
Reported
22,952 represented genes (RPKM>0 vs Phytozome Glyma1, Table 1)
Reproduced
34,217 genes with >=1x bedtools coverage (raw count, not RPKM-normalized, vs modern RefSeq soybean assembly, 50,961 total annotated genes)
partial
1188_filtered_snps
Reported
423,353 SNPs after DP>=5 filter (Table 2)
Reproduced
262,552 SNPs after DP>=5 filter (bcftools mpileup+call -mv)
partial
1188_pct_heterozygous
Reported
28.9% heterozygous SNPs (Table 2)
Reproduced
42.62% heterozygous SNPs; qualitative high/low-het group assignment MATCHES paper
partial
1393_raw_reads
Reported
21,820,163 raw reads (2 samples, Table 1)
Reproduced
21,578,585 raw reads pooled from 2 SRA runs actually deposited in SRP011928 (single-end)
exact
1393_mapped_reads
Reported
13,643,602 mapped reads (Table 1)
Reproduced
15,570,292 bowtie2-aligned reads (>=1 time, pre-MAPQ-filter; 5,024,939 remain after MAPQ>=20 filtering used downstream for SNP calling)
within tolerance
1393_genes_represented
Reported
23,345 represented genes (RPKM>0 vs Phytozome Glyma1, Table 1)
Reproduced
33,698 genes with >=1x bedtools coverage (raw count, not RPKM-normalized, vs modern RefSeq soybean assembly, 50,961 total annotated genes)
partial
1393_filtered_snps
Reported
367,646 SNPs after DP>=5 filter (Table 2)
Reproduced
239,226 SNPs after DP>=5 filter (bcftools mpileup+call -mv)
partial
1393_pct_heterozygous
Reported
28.8% heterozygous SNPs (Table 2)
Reproduced
42.21% heterozygous SNPs; qualitative high/low-het group assignment MATCHES paper
partial
1300_raw_reads
Reported
25,527,322 raw reads (3 samples, Table 1)
Reproduced
25,282,259 raw reads pooled from 3 SRA runs actually deposited in SRP011928 (single-end)
exact
1300_mapped_reads
Reported
14,092,961 mapped reads (Table 1)
Reproduced
17,384,464 bowtie2-aligned reads (>=1 time, pre-MAPQ-filter; 5,947,553 remain after MAPQ>=20 filtering used downstream for SNP calling)
partial
1300_genes_represented
Reported
24,238 represented genes (RPKM>0 vs Phytozome Glyma1, Table 1)
Reproduced
34,371 genes with >=1x bedtools coverage (raw count, not RPKM-normalized, vs modern RefSeq soybean assembly, 50,961 total annotated genes)
partial
1300_filtered_snps
Reported
477,245 SNPs after DP>=5 filter (Table 2)
Reproduced
280,888 SNPs after DP>=5 filter (bcftools mpileup+call -mv)
partial
1300_pct_heterozygous
Reported
7.8% heterozygous SNPs (Table 2)
Reproduced
11.15% heterozygous SNPs; qualitative high/low-het group assignment MATCHES paper
partial
2073_raw_reads
Reported
12,132,989 raw reads (2 samples, Table 1)
Reproduced
11,931,128 raw reads pooled from 2 SRA runs actually deposited in SRP011928 (single-end)
exact
2073_mapped_reads
Reported
7,087,710 mapped reads (Table 1)
Reproduced
8,159,317 bowtie2-aligned reads (>=1 time, pre-MAPQ-filter; 2,770,422 remain after MAPQ>=20 filtering used downstream for SNP calling)
partial
2073_genes_represented
Reported
24,438 represented genes (RPKM>0 vs Phytozome Glyma1, Table 1)
Reproduced
33,027 genes with >=1x bedtools coverage (raw count, not RPKM-normalized, vs modern RefSeq soybean assembly, 50,961 total annotated genes)
partial
2073_filtered_snps
Reported
282,215 SNPs after DP>=5 filter (Table 2)
Reproduced
158,754 SNPs after DP>=5 filter (bcftools mpileup+call -mv)
partial
2073_pct_heterozygous
Reported
12.6% heterozygous SNPs (Table 2)
Reproduced
19.44% heterozygous SNPs; qualitative high/low-het group assignment DIFFERS FROM paper
partial
2321_raw_reads
Reported
32,796,391 raw reads (2 samples, Table 1)
Reproduced
26,837,309 raw reads pooled from 2 SRA runs actually deposited in SRP011928 (single-end)
partial
2321_mapped_reads
Reported
13,637,368 mapped reads (Table 1)
Reproduced
18,555,846 bowtie2-aligned reads (>=1 time, pre-MAPQ-filter; 5,846,504 remain after MAPQ>=20 filtering used downstream for SNP calling)
partial
2321_genes_represented
Reported
22,571 represented genes (RPKM>0 vs Phytozome Glyma1, Table 1)
Reproduced
34,357 genes with >=1x bedtools coverage (raw count, not RPKM-normalized, vs modern RefSeq soybean assembly, 50,961 total annotated genes)
partial
2321_filtered_snps
Reported
544,101 SNPs after DP>=5 filter (Table 2)
Reproduced
317,646 SNPs after DP>=5 filter (bcftools mpileup+call -mv)
partial
2321_pct_heterozygous
Reported
6.3% heterozygous SNPs (Table 2)
Reproduced
10.75% heterozygous SNPs; qualitative high/low-het group assignment MATCHES paper
partial
1366_raw_reads
Reported
20,631,583 raw reads (2 samples, Table 1)
Reproduced
20,379,626 raw reads pooled from 2 SRA runs actually deposited in SRP011928 (single-end)
exact
1366_mapped_reads
Reported
10,766,169 mapped reads (Table 1)
Reproduced
13,334,397 bowtie2-aligned reads (>=1 time, pre-MAPQ-filter; 4,453,930 remain after MAPQ>=20 filtering used downstream for SNP calling)
partial
1366_genes_represented
Reported
23,364 represented genes (RPKM>0 vs Phytozome Glyma1, Table 1)
Reproduced
33,291 genes with >=1x bedtools coverage (raw count, not RPKM-normalized, vs modern RefSeq soybean assembly, 50,961 total annotated genes)
partial
1366_filtered_snps
Reported
360,327 SNPs after DP>=5 filter (Table 2)
Reproduced
209,173 SNPs after DP>=5 filter (bcftools mpileup+call -mv)
partial
1366_pct_heterozygous
Reported
7.5% heterozygous SNPs (Table 2)
Reproduced
10.69% heterozygous SNPs; qualitative high/low-het group assignment MATCHES paper
partial
1403_raw_reads
Reported
31,631,369 raw reads (3 samples, Table 1)
Reproduced
21,515,507 raw reads pooled from 2 SRA runs actually deposited in SRP011928 (single-end)
partial
1403_mapped_reads
Reported
17,218,424 mapped reads (Table 1)
Reproduced
14,337,219 bowtie2-aligned reads (>=1 time, pre-MAPQ-filter; 4,771,698 remain after MAPQ>=20 filtering used downstream for SNP calling)
partial
1403_genes_represented
Reported
23,352 represented genes (RPKM>0 vs Phytozome Glyma1, Table 1)
Reproduced
32,917 genes with >=1x bedtools coverage (raw count, not RPKM-normalized, vs modern RefSeq soybean assembly, 50,961 total annotated genes)
partial
1403_filtered_snps
Reported
369,661 SNPs after DP>=5 filter (Table 2)
Reproduced
207,934 SNPs after DP>=5 filter (bcftools mpileup+call -mv)
partial
1403_pct_heterozygous
Reported
6.6% heterozygous SNPs (Table 2)
Reproduced
10.55% heterozygous SNPs; qualitative high/low-het group assignment MATCHES paper
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 54/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Descriptive Table 1/2 values reproduce only in part. Raw-read totals match to within ~1% for four accessions (1393, 1300, 2073, 1366), but G1134's reported 202,427,873 reads from 13 samples and G1820's 71,185,274 cannot be reconstructed from what is actually deposited in SRP011928 (28,389,777 and 36,597,945) - a data-completeness defect on the authors'/archive side. All remaining deviations are systematic and explainable by our own tooling: mapping to a modern RefSeq assembly instead of Phytozome Glyma1 and using bedtools coverage instead of RPKM>0 inflates 'represented genes' by ~45%, while bowtie2 + bcftools with a MAPQ>=20 filter gives 35-70% fewer DP>=5 SNPs and a uniform ~1.5x higher heterozygosity. The paper's central qualitative claim - a sharp low-het (6.3-7.8%) vs high-het (26.4-28.9%) split separating diploids from allopolyploid/hybrid accessions - survives for 9 of 10 accessions (10.69-11.15% vs 42.21-44.07%), with accession 2073 the single group-assignment flip, so the conclusion is confirmed in substance but not in any absolute number.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.