De Novo Assembly and Annotation of the Larval Transcriptome of Two Spadefoot Toads Widely Divergent in Developmental Rate.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED at the pipeline-output level using the authors' deposited data (FigShare 8201825) + the named repo trinotateR. C1 assembly metrics = EXACT 12/12 vs Table 1 (transcripts, genes, N50, total bp, mean & median length, both species) on the deposited Trinity assemblies. C3 annotation counts via trinotateR (read_trinotate/summary_trinotate/split_GO/split_pfam) on sd5_Trinotate.xlsx = 10/12 EXACT + 2 within-tol (Sc Pfam/ORFs off by 2, ~0.001%; GO-total off by 92, 0.004% — spreadsheet round-trip). C2 BUSCO recompute (BUSCO 5.7.1, tetrapoda_odb10) = Complete% within-tol both species (Pc 88.9 vs 89.7; Sc 88.5 vs 86.6) despite odb10(5310) vs the paper's odb9(3950) and v5/metaeuk vs v3/augustus; frag/missing differ (partial). C4 bowtie2 deposited sd1 = paper exactly (83.96%/90.71%). Tally: 22 exact, 4 within-tol, 4 partial, 2 deposited-verified, 0 mismatch. No fabrication indicators — every headline number is regenerable from the shipped data with the described/named tools. NOT attempted: full Trinity re-assembly (non-deterministic; verified at output level) and downstream enrichment. Datasets: SRA 24 PE runs (N exact; reads 101bp vs paper text's 76bp) + FigShare supplement (8/8 complete), both grade A.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 91assessed: 2026-06-21 ⛓ 0677de0b9849
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study investigates the transcriptomic basis underlying the pronounced divergence in larval developmental rate (and associated genome size difference) between the fast-developing Scaphiopus couchii and the slow-developing Pelobates cultripes spadefoot toads at the onset of metamorphosis.
- ★ De novo transcriptome assemblies were generated for larval P. cultripes and S. couchii, providing new genomic resources for spadefoot toads resource
- ★ Both transcriptomes show high completeness (BUSCO) and functional annotation coverage, with S. couchii yielding more annotated transcripts/genes than P. cultripes despite its smaller genome finding
- ★ The two species share the majority (56.5%) of annotated vertebrate proteins, with substantial species-specific gene sets unique to each finding
- ★ GO-slim functional classification (PANTHER) is highly comparable between the two species relative to the X. tropicalis reference, with cell/binding/catalytic/metabolic categories overrepresented in both finding
- ★ Gene pathway enrichment analysis (gProfiler) identifies mostly species-specific enriched terms, with only 16 GO terms shared between species finding
- ★ Genes related to thyroid hormone-regulated development or oxidative stress were not recovered as significantly overrepresented in either species finding
- S. couchii transcriptome shows notable representation of nucleoside triphosphate metabolism-related processes, potentially linked to faster development and higher metabolic rate finding
- A bioinformatics pipeline (Trinity assembly, Trinotate annotation, Kallisto-based expression filtering, OrthoFinder orthology) was used to construct and compare the two larval transcriptomes method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (TruSeq stranded mRNA) | whole-body larval P. cultripes tadpoles (Gosner stage 35) | none | sequence reads for de novo transcriptome assembly | Illumina HiSeq2000, 2x76bp paired-end |
| bulk RNA-seq (TruSeq stranded mRNA) | whole-body larval S. couchii tadpoles (Gosner stage 35) | none | sequence reads for de novo transcriptome assembly | Illumina HiSeq2000, 2x76bp paired-end |
| de novo transcriptome assembly | P. cultripes and S. couchii pooled tadpole RNA-seq reads | none | assembled transcript contigs and 'gene' clusters | Trinity v2.4.0 |
| transcriptome completeness assessment | P. cultripes and S. couchii assembled transcriptomes | none | percent complete/fragmented/missing conserved orthologs | BUSCO v3.0.2, tetrapoda-odb9 |
| homology search (blastx/blastp) | P. cultripes and S. couchii transcripts/CDS | none | protein sequence similarity hits and alignment coverage | blastx/blastp vs SwissProt and X. tropicalis proteome |
| functional annotation (Trinotate pipeline: HMMER/Pfam, SignalP, TmHMM, KEGG, EggNOG, GO) | P. cultripes and S. couchii TransDecoder-predicted CDS | none | domain, signal peptide, transmembrane, pathway and GO annotations | Trinotate v3.0 |
| transcript abundance quantification and expression filtering | P. cultripes and S. couchii transcriptomes | none | TPM-based transcript retention/filtering | Kallisto v0.4.3.1 |
| functional enrichment and orthology analysis (PANTHER, gProfiler, OrthoFinder) | expression-filtered P. cultripes and S. couchii transcriptomes vs X. tropicalis reference | none | overrepresented/enriched GO, KEGG, Reactome terms and orthologous gene groups | PANTHER, gProfiler (g:GOST), OrthoFinder v2.2.3 |
- – Trinity assembled 753,223 transcripts (428,406 'genes') for P. cultripes and 657,280 transcripts (381,135 'genes') for S. couchii
- ▲ S. couchii assembly has a longer N50 than P. cultripes 2,057bp vs 1,496bp
- – BUSCO completeness slightly higher for P. cultripes than S. couchii 89.7% vs 86.6%
- – Bowtie2 read mapping rate higher for S. couchii than P. cultripes 90.71% vs 83.96%
- – Unique vertebrate SwissProt proteins mostly shared between species, with notable species-specific sets 56.5% shared (18,651); 17.2% (5,676) unique to P. cultripes; 26.2% (8,658) unique to S. couchii
- – More significantly enriched functional terms/pathways found in P. cultripes than S. couchii, with limited overlap 54 vs 40 terms, 16 shared
- – Nine Reactome pathways enriched overall; no KEGG pathways enriched in either species 3 (P. cultripes) vs 6 (S. couchii) Reactome pathways
- – Expression-based TPM filtering retained a minority of transcripts, more for S. couchii than P. cultripes 20.8% vs 26.8%
- count 753,223 transcripts | 428,406 genes (P. cultripes Trinity assembly)
- count 657,280 transcripts | 381,135 genes (S. couchii Trinity assembly)
- other N50 1,496bp (transcripts) | 731bp (longest isoform) (P. cultripes assembly quality)
- other N50 2,057bp (transcripts) | 872bp (longest isoform) (S. couchii assembly quality)
- other BUSCO complete 89.7% vs 86.6%; fragmented 7.4% vs 10.5%; missing 2.9% vs 2.9% (BUSCO tetrapoda-odb9 completeness, P. cultripes vs S. couchii)
- fold_change genome size ~3.9 Gbp (P. cultripes) vs ~1.5 Gbp (S. couchii) (Known genome size divergence between species (cited from Zeng et al. 2014))
- pvalue P < 0.05 (adjusted) (Significance threshold for PANTHER overrepresentation and gProfiler enrichment tests)
- count 24,327 (P. cultripes) and 27,309 (S. couchii) unique vertebrate SwissProt proteins identified (Vertebrate-restricted blastp annotation of TransDecoder CDS)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper describes de novo transcriptome assembly and functional annotation for two spadefoot toad species, with statistical analysis limited to gene-set overrepresentation/enrichment testing rather than classical hypothesis testing on continuous outcomes. Results are reported using the PANTHER statistical overrepresentation test (Fisher's exact test with FDR correction) comparing each species' transcriptome to the X. tropicalis reference, and gProfiler's g:GOST overrepresentation/ranked enrichment test (adjusted via the g:SCS algorithm) against GO, KEGG, and Reactome databases. Findings are reported mainly as counts, percentages, and significance thresholds (e.g., adjusted p < 0.05) rather than with conventional dispersion measures or effect sizes.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Fisher's exact test (PANTHER statistical overrepresentation test) with FDR correction | Over/underrepresentation of PANTHER GO-slim terms in each species' transcriptome relative to X. tropicalis reference | 10,666 (P. cultripes) and 11,598 (S. couchii) PANTHER-mapped gene annotations | not stated |
| gProfiler g:GOST overrepresentation/ranked enrichment test with g:SCS multiple-testing correction | Enrichment of GO Biological Process, KEGG, and Reactome terms among expression-ranked gene lists per species | 12,391 of 13,421 queried genes (P. cultripes); 13,315 of 14,811 queried genes (S. couchii) | not stated |
-
Transcript inclusion was based on a fixed expression threshold (>1 TPM) determined by visual inspection of a transcript-count curve.↳ Could also: A model-based or statistically derived expression cutoff (e.g., using a mixture model to separate expressed from background transcripts) could also be used — This would provide a formally justified threshold rather than one selected by visual inspection, which some readers may find useful for reproducibility across datasets
-
Gene set overrepresentation was tested with Fisher's exact test (PANTHER) and a separate ranked-list enrichment test with g:SCS correction (gProfiler), using two different tools and correction methods on overlapping gene lists.↳ Could also: A single unified enrichment framework, such as GSEA (Gene Set Enrichment Analysis) or camera/fgsea in R, could also be applied to both species — Using one consistent method and correction procedure across all enrichment comparisons can make between-species and between-tool results more directly comparable
-
No formal differential expression analysis (e.g., statistical comparison of TPM values) between P. cultripes and S. couchii transcripts was performed; comparisons were largely descriptive (shared vs. unique annotated proteins, enriched term counts).↳ Could also: A count-based differential expression tool such as DESeq2 or edgeR (applied to Kallisto-derived counts) could also be used to formally test for differential transcript abundance between species — This would allow explicit statistical testing (with p-values and fold-change estimates) of which specific genes differ in expression between the two species, complementing the current descriptive functional-annotation comparison
-
Enrichment results were summarized as significant/non-significant term counts and lists (e.g., '54 significantly enriched terms') without accompanying effect-size or magnitude-of-enrichment statistics.↳ Could also: Reporting enrichment score, fold-enrichment, or normalized enrichment score (as in GSEA output) alongside adjusted p-values could also be included — Effect-size-like metrics can help readers gauge the magnitude of overrepresentation, not just its statistical significance
-
Two species' transcriptomes were each compared independently to the X. tropicalis reference for overrepresentation testing, rather than being tested against each other directly in a single combined model.↳ Could also: A direct statistical comparison between the two species' gene lists (e.g., a contingency-table test or ortholog-based differential enrichment analysis) could also be performed — This could directly quantify how enrichment patterns differ between the two species themselves, in addition to each species' comparison with the outgroup reference
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-31217263
Paper: Liedtke et al. 2019, De Novo Assembly and Annotation of the Larval Transcriptome of Two Spadefoot Toads Widely Divergent in Developmental Rate. G3 9(8):2647-2655. DOI 10.1534/g3.119.400389.
Species (NOT Spea multiplicata as the registry hint implied):
- Pelobates cultripes (slow developer; "Pc")
- Scaphiopus couchii (fast developer; "Sc")
Pipeline (Methods): Illumina HiSeq2000 PE RNA-seq → FastQC/MultiQC → Trimmomatic (SLIDINGWINDOW:4:5 LEADING:5 TRAILING:5 MINLEN:25) → Trinity v2.4.0 in-silico normalization (--normalize_max_read_cov 50) + de novo assembly (per species, all samples pooled) → Bowtie2 read-representation check → BUSCO (tetrapoda_odb9) → TransDecoder ORFs → Trinotate v3.0 (blastx/blastp SwissProt, HMMER/Pfam, SignalP 4.1, TmHMM 2.0, eggNOG, KEGG) → trinotateR (the named repo: github.com/cstubben/trinotateR) to summarize the Trinotate report → downstream: blastx vs X. tropicalis, PANTHER, gProfiler, OrthoFinder.
Data available:
- Raw reads: SRA PRJNA490256 / SRP161446 — 24 paired-end runs (12 Pc + 12 Sc), SRR7817202–SRR7817225. OPEN.
- Deposited assemblies: TSA GHBH01000000 (Pc), GHBO01000000 (Sc); also FigShare sd2_transcriptomes.zip (371 MB). OPEN.
- FigShare supplement 10.25387/g3.8201825 (article 8201825): sd1_Bowtie2_results, sd2_transcriptomes.zip, sd3_BUSCO_results, sd4_blastx_xtropicalis.xlsx, sd5_Trinotate.xlsx (111 MB) = the Trinotate report, sd7_gProfiler.xlsx. OPEN.
IN SCOPE (pipeline-derived, will attempt)
| # | Result | Pipeline | Strategy |
|---|---|---|---|
| C1 | Assembly metrics: #transcripts, #genes, N50, total bases, median/avg length (Table 1) | Trinity TrinityStats | Run TrinityStats.pl on the deposited sd2 assemblies → must match Table 1 exactly (same data → verifies deposit delivers reported numbers). |
| C2 | BUSCO completeness 89.7% / 86.6% (tetrapoda_odb9) | BUSCO | Re-run BUSCO on deposited assemblies → within-tol (version-dependent). |
| C3 | Annotation summary counts: blastx/blastp SwissProt hits, Pfam, SignalP, TmHMM, unique GO terms | trinotateR on sd5 Trinotate report | CORE repo reproduction — run trinotateR summary_* on sd5 → match Results-section counts. |
| C4 | Bowtie2 read-representation 83.96% / 90.71% | Bowtie2 | (Stretch) build bowtie2 index from deposited assembly, map raw reads back. Heavy I/O. |
STRETCH (heavy, attempt after floor)
| C5 | Full Trinity de novo assembly from raw reads | Trimmomatic+Trinity | Very heavy (~84M normalized PE reads, large RAM, days). Assembly is non-deterministic across versions, so a 1:1 transcript count is NOT expected — at best within-tol/partial. Lower priority; document honestly. |
OUT OF SCOPE (not attempted, with reason)
- Wet-lab: RNA extraction, library prep, sequencing (not computational).
- PANTHER/gProfiler/OrthoFinder downstream enrichment (multi-tool, web-service dependent; secondary to core assembly+annotation; record reported values only).
- eggNOG/KEGG pathway annotation detail (subsumed in Trinotate report; counts via trinotateR if present in sd5).
Notes / discrepancies to verify on data
- Paper text says 2×76 bp reads; ENA base_count/read_pair = 202 → reads are ~2×101 bp. Verify read length directly on FASTQ.
- ENA read_count = read PAIRS; paper "total raw reads" = mates (2× pairs). Pc: 2×444,132,722 = 888,265,444 ✓ exact. Confirms N concordance.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is essentially a 1:1 reproduction on the authors' own deposited data (FigShare 8201825 assemblies + Trinotate report, SRA PRJNA490256): 12/12 Table-1 assembly metrics EXACT and 10/12 annotation counts EXACT via the named repo trinotateR, with the two off-by-2 values being xlsx→tsv round-trip noise. The only material deviation is BUSCO fragmented/missing, fully explained by using tetrapoda_odb10/v5 vs the paper's odb9/v3 — a technical version difference on our side, not an authors' defect, and Complete% still reproduces within tolerance. No fabrication indicators: every headline number is derivable from the shipped artifacts, so q5/q7/q8 are all green.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
<synthetic>Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.