Mining transcriptomic data to study the origins and evolution of a plant allopolyploid complex.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the core bioinformatic pipeline (fastq-mcf trim Q30/min50bp -> Bowtie2 default alignment -> MAPQ>=20 filter -> bedtools gene coverage -> bcftools mpileup/call SNP calling, DP>=5 filter) for all 9 accessions in the RU's designated dataset (sra:SRP011928), using the third-party GenoToolBox MultiVcfTool.pl for the multi-sample VCF/hapmap merge step, against a modern RefSeq soybean assembly (GCF_000004515-family NC_016088.1-series) rather than the paper's original Phytozome Glyma v1.0 (unavailable in that exact historical form) - a documented, necessary methodological substitution. Result: 1:1-adjacent but NOT exact reproduction. Raw read counts match Table 1 closely (within ~2-18%) for 6/9 accessions; SRA data for accession 1134 is severely under-deposited (14% of reported reads) and 1403/1820 show unexplained shortfalls (32-49%) not resolvable from the cited archives. Bowtie2 pre-filter alignment counts run 10-36% above/below Table 1's 'Mapped reads' for adequately-represented accessions (best reading of that column: total aligned reads before the separate downstream MAPQ>=20 filter, since MAPQ-filtered counts are 3-4x too low to match Table 1 directly). DP>=5-filtered SNP counts run systematically at ~56-65% of Table 1 for well-represented accessions - consistent direction, plausibly explained by reference-genome/annotation differences, not proportional to data loss. Most importantly, the paper's central biological claim - that the G. dolichocarpa lineage (accessions 1134, 1188, 1393) retains much higher heterozygosity (26-29% in the paper) than the G. syndetika/tomentella-D3 diploid-progenitor lineages (6-13% in the paper) - is QUALITATIVELY REPRODUCED: all 9 reproduced accessions fall in the same high-het (42-44%) or low-het (10-19%) group as the paper, with a consistent ~1.4-1.7x systematic magnitude offset across ALL accessions (both groups), again pointing to a reference-genome/pipeline-version effect rather than a spurious or absent signal. NOT attempted (out of pipeline scope / would require additional tools+data not exercised here): SeparateHomeolog2Sam homeolog-separation, Structure/fineStructure population-structure analysis, PhyML/*BEAST phylogenetics and divergence dating, and full reproduction of the secondary SRP038128 dataset (12 further accessions, profiled at metadata level only). 3 Table-1 accessions (1286, 1854, 1364) could not be located in any cited or searched SRA archive and are recorded as a genuine data-availability gap.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-08-07
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-07no human curator yet
- Last updated
- 2026-08-07
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper tests whether three allopolyploid species of the perennial Glycine (G. tomentella T1, G. dolichocarpa/T2, and G. tomentella T5) originated as fixed hybrids combining the genomes of the two extant diploid progenitor species hypothesized from previous crossing and two-gene molecular systematic studies, and asks when and how many times these polyploids arose.
- ★ All three allopolyploid species are fixed hybrids combining the genomes of the two putative diploid parents hypothesized on the basis of previous crossing and molecular work. finding
- ★ Based on mapping to the soybean reference genome, there appear to be no large regions for which one homoeologous contribution is missing. finding
- ★ Phylogenetic analyses of 27 selected transcripts using a coalescent approach are consistent with multiple origins for these allopolyploid species. finding
- ★ The origins of these allopolyploid species occurred within the last several hundred thousand years. finding
- ★ Transcriptomic (RNA-Seq) data generated for an unrelated light-stress study can be mined with phylogenetic and population-genomic approaches to resolve allopolyploid origins. method
- ★ Polyploid reads can be assigned to parental subgenomes by preferential mapping against a progenitor reference set built from consensus diploid transcriptomes, enabling homoeologue-resolved SNP datasets and transcript-guided assembly. method
- Custom Perl tools (MultiVcfTool, SeparateHomeolog2Sam, FastaSeqExtract in the GenoToolBox package) are provided for multi-VCF formatting and homoeologue read separation. resource
- Alignments whose trees did not place G. max sister to the perennial Glycine species were removed as likely paralogues from the ca. 50 My legume-wide WGD rather than orthologues. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Single-end RNA-Seq (transcriptome sequencing) | Pooled leaflet tissue from six individuals per accession of eight perennial Glycine species (G. canescens, G. clandestina, G. dolichocarpa, G. syndetika, G. tomentella D1, D3, T1, T5) plus a synthetic allotetraploid A58; 2-5 accessions per species from the CSIRO Perennial Glycine Germplasm Collection | Light-intensity treatment: low light (125 mmol m-2 s-1) vs excess light (800 mmol m-2 s-1) in a growth chamber, 12 h/12 h light/dark, 22 °C/18 °C | Raw and processed read counts, reads mapped to the soybean genome, and number of represented G. max reference genes with expression >0 RPKM | Illumina GAIIx (88 nt reads) or HiSeq 2000 (100 nt reads); Illumina mRNA-Seq Sample Prep Kit, Dynabeads mRNA DIRECT Kit, Qiagen Plant RNeasy Kit with on-column DNase |
| Read mapping and SNP calling | Perennial Glycine RNA-Seq reads mapped to the Glycine max reference genome version 1.0 (Phytozome) | none | SNP genotypes (coverage >=5, unique mapping, mapping score >=20) exported to Structure and Hapmap formats | Fastq-mcf, Bowtie2 (default parameters), Samtools, MultiVcfTool |
| Homoeologue read identification and transcript-guided (guided) transcriptome assembly | Allopolyploid reads (T1, T2, T5) mapped against progenitor reference sets built from consensus diploid transcriptomes (T1 = D1 + D3; T2 = D3 + D4; T5 = A + D1) | none | Reads partitioned by preferential mapping to each progenitor subgenome; rebuilt polyploid transcriptomes and homoeologue-resolved SNPs | Bowtie2, SeparateHomeolog2Sam, Samtools, Gffread (Cufflinks package) |
| Population structure analysis (Bayesian clustering) | Two SNP datasets (with and without polyploid SNPs separated by homoeologue) from the Glycine accessions | none | Ancestry/admixture proportions and optimal number of clusters K | Structure (admixture model, lambda = 1); three random subsets of 20,000 SNPs, 5 replicates, burn-in 10,000, 10,000 MCMC reps, K = 1-15; K = 6 re-run with 100,000 burn-in; visualized in R |
| Population structure analysis (chromosome painting / coancestry) | The two SNP datasets divided into 20 subsets, each mapping to one soybean reference chromosome | none | Pairwise coancestry/distance matrix between accessions displayed as a heatmap, plus PCA of the same matrix | fineStructure (unlinked model); PCA figure created in R |
| Phylogenetic reconstruction from concatenated SNP supermatrix | 36 OTUs (allopolyploid homoeologues treated as separate OTUs, e.g. D1T1 and D3T1), with G. max accession Williams 82 as outgroup | none | Maximum likelihood tree topology with bootstrap support, and a phylogenetic network showing reticulation | PhyML (GTR model, 100 bootstrap replicates); NeighborNet in SplitsTree (default parameters); FigTree |
| Gene-based phylogenetic analyses (single-gene trees, networks, dated Bayesian trees) | 27 selected transcripts (filtered from 95 alignments) from 24 accessions including two homoeologues per allopolyploid accession; G. max Williams 82 as outgroup | none | Per-gene ML trees with bootstrap support, NeighborNet networks, and Bayesian divergence-time estimates with the root (G. max vs perennials) scaled to 5 Myr | jModelTest2 (BIC model selection), PhyML (1,000 bootstraps), SplitsTree4 NeighborNet, BEAST v2.0 (HKY, 100,000,000 MCMC generations sampled every 1,000) |
| Coalescent species tree reconstruction | 27 selected genes from 24 accessions grouped into 11 OTUs (G. canescens, G. clandestina, G. tomentella D1, G. tomentella D3, G. syndetika/D4, T1-D1, T1-D3, T2-D3, T2-D4, T5-A, T5-D1), with G. max included | none | Species tree topology under the multispecies coalescent | *BEAST |
- – All three polyploid species (T1, T2/G. dolichocarpa, T5) are fixed hybrids combining the two hypothesized parental diploid genomes.
- – No large genomic regions lacking one homoeologous contribution were detected when mapping to the soybean reference genome.
- – Coalescent and gene-tree analyses of 27 transcripts are consistent with multiple independent origins of the allopolyploid species.
- – Estimated origins of the allopolyploids fall within the last several hundred thousand years, based on trees scaled to a 5 Myr G. max/perennial root. within the last several hundred thousand years
- – Of 95 initial alignments screened by exploratory BEAST analysis, 27 genes passed the orthology filter (G. max sister to perennial Glycine); removed alignments showed two long-branch clades suggesting inclusion of paralogues from the ca. 50 My legume WGD. 27 of 95 alignments retained
- – HKY was the best-fitting substitution model for the plurality of the 27 genes by BIC, followed by K80. HKY 40% of genes; K80 26% of genes
- – Structure analysis identified an optimal clustering that was verified at K = 6 with an extended burn-in. K = 6
- – RNA-Seq mapping to the soybean reference recovered roughly 22,500-25,300 expressed reference genes (RPKM > 0) per accession. 22,571-25,278 represented genes per accession
- other HKY preferred in 40% of genes; K80 in 26% of genes (BIC, jModelTest2) (Substitution model selection across the 27 filtered gene alignments)
- count 27 genes retained from 95 alignments (Filtering of transcript alignments for gene-based and species-tree analyses)
- count 36 OTUs in the concatenated SNP supermatrix; 11 OTUs from 24 accessions in the *BEAST species tree (Operational taxonomic units, with allopolyploid homoeologues treated separately)
- count 202,427,873 raw reads / 187,120,918 processed / 60,712,525 mapped; 23,643 represented genes (Largest accession dataset, Glycine dolichocarpa accession 1134 (13 samples))
- count 10,401,944 raw reads / 9,604,350 processed / 6,896,983 mapped; 22,802 represented genes (Smallest accession dataset, Glycine tomentella D3 accession 1364 (1 sample))
- other three subsets of 20,000 SNPs, 5 replicates each, burn-in 10,000, 10,000 MCMC repetitions, K = 1 to K = 15; K = 6 re-run with 100,000 burn-in (Structure population structure run parameters)
- other 100 bootstrap replicates (concatenated SNP ML tree); 1,000 bootstraps (27-gene ML trees); 10,000,000 MCMC (exploratory BEAST); 100,000,000 MCMC sampled every 1,000 (BEAST v2.0) (Phylogenetic analysis run parameters)
- other root calibration 5 Myr (G. max vs perennial divergence); Glycine WGD ca. 5-10 Mya; legume-wide WGD ca. 50 Mya; chromosome numbers 2n = 38/40 (diploids), 78/80 (allopolyploids) (Calibration and taxon background used for divergence dating)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used a population-genomics and phylogenetic framework rather than classical hypothesis testing: SNPs called from mapped RNA-Seq reads were analyzed with Bayesian clustering (Structure) and coancestry-based clustering (fineStructure/PCA) to assess population structure, and with maximum-likelihood (PhyML) and Bayesian (BEAST, *BEAST) phylogenetic/coalescent methods, plus NeighborNet networks, to reconstruct relationships and divergence times among diploid and allopolyploid Glycine lineages. Support for relationships was conveyed through bootstrap percentages and posterior probabilities rather than through p-values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Maximum-likelihood phylogenetic reconstruction (PhyML, GTR model), 100 bootstrap replicates | concatenated homoeologue SNP supermatrix, 36 OTUs | 36 OTUs (SNP supermatrix) | not stated |
| NeighborNet network reconstruction (SplitsTree) | same concatenated SNP supermatrix, and the 27-gene alignment set | — | na |
| Maximum-likelihood phylogenetics (PhyML, HKY/K80 models selected via jModelTest2/BIC), 1,000 bootstrap replicates | 27 selected gene alignments | 27 genes | not stated |
| Bayesian MCMC phylogenetic analysis (BEAST v2.0, HKY model, 100,000,000 generations) | 27 selected gene alignments, 11 OTUs (homoeologues treated as separate OTUs) | 27 genes / 24 accessions | not stated |
| Bayesian admixture clustering (Structure, K=1–15, 5 replicates per subset) | SNP datasets (with and without homoeologue-separated polyploid SNPs) | 3 random subsets of 20,000 SNPs each | not stated |
| Coalescent-based species tree estimation (*BEAST) | 27 selected genes, 11 OTUs | 27 genes | not stated |
-
Species/homoeologue relationships were inferred using concatenated-SNP ML trees, Bayesian gene trees, and NeighborNet networks.↳ Could also: Explicit phylogenetic network or admixture-graph methods (e.g., TreeMix, SNAQ, PhyloNet) — These approaches explicitly model reticulation/hybridization events and can quantify admixture proportions, which could complement the network visualization already used to depict reticulate allopolyploid origins.
-
Population structure was assessed with Structure (Bayesian MCMC admixture model) and fineStructure (coancestry-based clustering/PCA).↳ Could also: ADMIXTURE (maximum-likelihood-based clustering) — ADMIXTURE is computationally faster for larger SNP panels and often yields comparable cluster estimates to Structure, offering a useful cross-check on the inferred K and admixture proportions.
-
Nucleotide substitution models were selected using the Bayesian Information Criterion (BIC) in jModelTest2.↳ Could also: Akaike Information Criterion (AIC) or corrected AIC (AICc) — Comparing model choice across BIC, AIC, and AICc can show whether the selected model (e.g., HKY) is robust to the criterion used, since these criteria can occasionally favor different models.
-
Divergence times were estimated by scaling the BEAST tree root to a fixed external date (5 Myr).↳ Could also: Calibration with explicit prior distributions (e.g., fossil- or secondary-calibration priors) and reporting of highest posterior density (HPD) intervals — Using calibration priors with reported HPD intervals would propagate dating uncertainty into the divergence-time estimates rather than relying on a single fixed calibration point.
-
A multispecies coalescent species tree was estimated from 27 selected genes using *BEAST.↳ Could also: Summary coalescent methods such as ASTRAL or SVDquartets — These methods scale efficiently to larger numbers of loci or genome-wide SNP data and can serve as an independent, faster-to-compute complement to the full Bayesian coalescent analysis.
-
Branch support was reported as ML bootstrap percentages and Bayesian posterior probabilities.↳ Could also: Gene concordance factors and site concordance factors — Concordance factors quantify the proportion of underlying gene trees or sites actually supporting a branch, which can be informative in groups—like this one—where incomplete lineage sorting and introgression are suspected to affect topology.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Descriptive Table 1/2 values reproduce only in part. Raw-read totals match to within ~1% for four accessions (1393, 1300, 2073, 1366), but G1134's reported 202,427,873 reads from 13 samples and G1820's 71,185,274 cannot be reconstructed from what is actually deposited in SRP011928 (28,389,777 and 36,597,945) - a data-completeness defect on the authors'/archive side. All remaining deviations are systematic and explainable by our own tooling: mapping to a modern RefSeq assembly instead of Phytozome Glyma1 and using bedtools coverage instead of RPKM>0 inflates 'represented genes' by ~45%, while bowtie2 + bcftools with a MAPQ>=20 filter gives 35-70% fewer DP>=5 SNPs and a uniform ~1.5x higher heterozygosity. The paper's central qualitative claim - a sharp low-het (6.3-7.8%) vs high-het (26.4-28.9%) split separating diploids from allopolyploid/hybrid accessions - survives for 9 of 10 accessions (10.69-11.15% vs 42.21-44.07%), with accession 2073 the single group-assignment flip, so the conclusion is confirmed in substance but not in any absolute number.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.