Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

De Novo Assembly and Annotation of the Larval Transcriptome of Two Spadefoot Toads Widely Divergent in Developmental Rate.

G3 (Bethesda) · 2019
L1 91/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
91/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 82% of all assessed papers rank 197 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED at the pipeline-output level using the authors' deposited data (FigShare 8201825) + the named repo trinotateR. C1 assembly metrics = EXACT 12/12 vs Table 1 (transcripts, genes, N50, total bp, mean & median length, both species) on the deposited Trinity assemblies. C3 annotation counts via trinotateR (read_trinotate/summary_trinotate/split_GO/split_pfam) on sd5_Trinotate.xlsx = 10/12 EXACT + 2 within-tol (Sc Pfam/ORFs off by 2, ~0.001%; GO-total off by 92, 0.004% — spreadsheet round-trip). C2 BUSCO recompute (BUSCO 5.7.1, tetrapoda_odb10) = Complete% within-tol both species (Pc 88.9 vs 89.7; Sc 88.5 vs 86.6) despite odb10(5310) vs the paper's odb9(3950) and v5/metaeuk vs v3/augustus; frag/missing differ (partial). C4 bowtie2 deposited sd1 = paper exactly (83.96%/90.71%). Tally: 22 exact, 4 within-tol, 4 partial, 2 deposited-verified, 0 mismatch. No fabrication indicators — every headline number is regenerable from the shipped data with the described/named tools. NOT attempted: full Trinity re-assembly (non-deterministic; verified at output level) and downstream enrichment. Datasets: SRA 24 PE runs (N exact; reads 101bp vs paper text's 76bp) + FigShare supplement (8/8 complete), both grade A.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 91
    assessed: 2026-06-21 ⛓ 0677de0b9849
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study investigates the transcriptomic basis underlying the pronounced divergence in larval developmental rate (and associated genome size difference) between the fast-developing Scaphiopus couchii and the slow-developing Pelobates cultripes spadefoot toads at the onset of metamorphosis.

Core claims
  • De novo transcriptome assemblies were generated for larval P. cultripes and S. couchii, providing new genomic resources for spadefoot toads resource
  • Both transcriptomes show high completeness (BUSCO) and functional annotation coverage, with S. couchii yielding more annotated transcripts/genes than P. cultripes despite its smaller genome finding
  • The two species share the majority (56.5%) of annotated vertebrate proteins, with substantial species-specific gene sets unique to each finding
  • GO-slim functional classification (PANTHER) is highly comparable between the two species relative to the X. tropicalis reference, with cell/binding/catalytic/metabolic categories overrepresented in both finding
  • Gene pathway enrichment analysis (gProfiler) identifies mostly species-specific enriched terms, with only 16 GO terms shared between species finding
  • Genes related to thyroid hormone-regulated development or oxidative stress were not recovered as significantly overrepresented in either species finding
  • S. couchii transcriptome shows notable representation of nucleoside triphosphate metabolism-related processes, potentially linked to faster development and higher metabolic rate finding
  • A bioinformatics pipeline (Trinity assembly, Trinotate annotation, Kallisto-based expression filtering, OrthoFinder orthology) was used to construct and compare the two larval transcriptomes method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (TruSeq stranded mRNA) whole-body larval P. cultripes tadpoles (Gosner stage 35) none sequence reads for de novo transcriptome assembly Illumina HiSeq2000, 2x76bp paired-end
bulk RNA-seq (TruSeq stranded mRNA) whole-body larval S. couchii tadpoles (Gosner stage 35) none sequence reads for de novo transcriptome assembly Illumina HiSeq2000, 2x76bp paired-end
de novo transcriptome assembly P. cultripes and S. couchii pooled tadpole RNA-seq reads none assembled transcript contigs and 'gene' clusters Trinity v2.4.0
transcriptome completeness assessment P. cultripes and S. couchii assembled transcriptomes none percent complete/fragmented/missing conserved orthologs BUSCO v3.0.2, tetrapoda-odb9
homology search (blastx/blastp) P. cultripes and S. couchii transcripts/CDS none protein sequence similarity hits and alignment coverage blastx/blastp vs SwissProt and X. tropicalis proteome
functional annotation (Trinotate pipeline: HMMER/Pfam, SignalP, TmHMM, KEGG, EggNOG, GO) P. cultripes and S. couchii TransDecoder-predicted CDS none domain, signal peptide, transmembrane, pathway and GO annotations Trinotate v3.0
transcript abundance quantification and expression filtering P. cultripes and S. couchii transcriptomes none TPM-based transcript retention/filtering Kallisto v0.4.3.1
functional enrichment and orthology analysis (PANTHER, gProfiler, OrthoFinder) expression-filtered P. cultripes and S. couchii transcriptomes vs X. tropicalis reference none overrepresented/enriched GO, KEGG, Reactome terms and orthologous gene groups PANTHER, gProfiler (g:GOST), OrthoFinder v2.2.3
Key results
  • Trinity assembled 753,223 transcripts (428,406 'genes') for P. cultripes and 657,280 transcripts (381,135 'genes') for S. couchii
  • S. couchii assembly has a longer N50 than P. cultripes 2,057bp vs 1,496bp
  • BUSCO completeness slightly higher for P. cultripes than S. couchii 89.7% vs 86.6%
  • Bowtie2 read mapping rate higher for S. couchii than P. cultripes 90.71% vs 83.96%
  • Unique vertebrate SwissProt proteins mostly shared between species, with notable species-specific sets 56.5% shared (18,651); 17.2% (5,676) unique to P. cultripes; 26.2% (8,658) unique to S. couchii
  • More significantly enriched functional terms/pathways found in P. cultripes than S. couchii, with limited overlap 54 vs 40 terms, 16 shared
  • Nine Reactome pathways enriched overall; no KEGG pathways enriched in either species 3 (P. cultripes) vs 6 (S. couchii) Reactome pathways
  • Expression-based TPM filtering retained a minority of transcripts, more for S. couchii than P. cultripes 20.8% vs 26.8%
Key statistics
  • count 753,223 transcripts | 428,406 genes (P. cultripes Trinity assembly)
  • count 657,280 transcripts | 381,135 genes (S. couchii Trinity assembly)
  • other N50 1,496bp (transcripts) | 731bp (longest isoform) (P. cultripes assembly quality)
  • other N50 2,057bp (transcripts) | 872bp (longest isoform) (S. couchii assembly quality)
  • other BUSCO complete 89.7% vs 86.6%; fragmented 7.4% vs 10.5%; missing 2.9% vs 2.9% (BUSCO tetrapoda-odb9 completeness, P. cultripes vs S. couchii)
  • fold_change genome size ~3.9 Gbp (P. cultripes) vs ~1.5 Gbp (S. couchii) (Known genome size divergence between species (cited from Zeng et al. 2014))
  • pvalue P < 0.05 (adjusted) (Significance threshold for PANTHER overrepresentation and gProfiler enrichment tests)
  • count 24,327 (P. cultripes) and 27,309 (S. couchii) unique vertebrate SwissProt proteins identified (Vertebrate-restricted blastp annotation of TransDecoder CDS)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper describes de novo transcriptome assembly and functional annotation for two spadefoot toad species, with statistical analysis limited to gene-set overrepresentation/enrichment testing rather than classical hypothesis testing on continuous outcomes. Results are reported using the PANTHER statistical overrepresentation test (Fisher's exact test with FDR correction) comparing each species' transcriptome to the X. tropicalis reference, and gProfiler's g:GOST overrepresentation/ranked enrichment test (adjusted via the g:SCS algorithm) against GO, KEGG, and Reactome databases. Findings are reported mainly as counts, percentages, and significance thresholds (e.g., adjusted p < 0.05) rather than with conventional dispersion measures or effect sizes.

Replicationbiological Sample sizeTwelve individual tadpoles per species (from three egg clutches per species) were sequenced as biological replicates; gene lists for enrichment analysis were based on mean TPM across these replicates GroupsP. cultripes vs. S. couchii transcriptomes, each compared to the X. tropicalis reference transcriptome/proteome Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini-Hochberg-style FDR correction (PANTHER) and the g:SCS algorithm (gProfiler)
Statistical tests used
Test Applied to n Assumptions
Fisher's exact test (PANTHER statistical overrepresentation test) with FDR correction Over/underrepresentation of PANTHER GO-slim terms in each species' transcriptome relative to X. tropicalis reference 10,666 (P. cultripes) and 11,598 (S. couchii) PANTHER-mapped gene annotations not stated
gProfiler g:GOST overrepresentation/ranked enrichment test with g:SCS multiple-testing correction Enrichment of GO Biological Process, KEGG, and Reactome terms among expression-ranked gene lists per species 12,391 of 13,421 queried genes (P. cultripes); 13,315 of 14,811 queried genes (S. couchii) not stated
Approaches that could also have been used
  • Transcript inclusion was based on a fixed expression threshold (>1 TPM) determined by visual inspection of a transcript-count curve.
    Could also: A model-based or statistically derived expression cutoff (e.g., using a mixture model to separate expressed from background transcripts) could also be used — This would provide a formally justified threshold rather than one selected by visual inspection, which some readers may find useful for reproducibility across datasets
  • Gene set overrepresentation was tested with Fisher's exact test (PANTHER) and a separate ranked-list enrichment test with g:SCS correction (gProfiler), using two different tools and correction methods on overlapping gene lists.
    Could also: A single unified enrichment framework, such as GSEA (Gene Set Enrichment Analysis) or camera/fgsea in R, could also be applied to both species — Using one consistent method and correction procedure across all enrichment comparisons can make between-species and between-tool results more directly comparable
  • No formal differential expression analysis (e.g., statistical comparison of TPM values) between P. cultripes and S. couchii transcripts was performed; comparisons were largely descriptive (shared vs. unique annotated proteins, enriched term counts).
    Could also: A count-based differential expression tool such as DESeq2 or edgeR (applied to Kallisto-derived counts) could also be used to formally test for differential transcript abundance between species — This would allow explicit statistical testing (with p-values and fold-change estimates) of which specific genes differ in expression between the two species, complementing the current descriptive functional-annotation comparison
  • Enrichment results were summarized as significant/non-significant term counts and lists (e.g., '54 significantly enriched terms') without accompanying effect-size or magnitude-of-enrichment statistics.
    Could also: Reporting enrichment score, fold-enrichment, or normalized enrichment score (as in GSEA output) alongside adjusted p-values could also be included — Effect-size-like metrics can help readers gauge the magnitude of overrepresentation, not just its statistical significance
  • Two species' transcriptomes were each compared independently to the X. tropicalis reference for overrepresentation testing, rather than being tested against each other directly in a single combined model.
    Could also: A direct statistical comparison between the two species' gene lists (e.g., a contingency-table test or ortholog-based differential enrichment analysis) could also be performed — This could directly quantify how enrichment patterns differ between the two species themselves, in addition to each species' comparison with the outgroup reference
Software: PANTHER classification system (web server) · gProfiler / gprofiler2 (R package) 0.1.3 · Cytoscape with EnrichmentMap add-on 3.7.1 · Trinity 2.4.0 · Kallisto 0.4.3.1 · OrthoFinder 2.2.3

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31217263

Paper: Liedtke et al. 2019, De Novo Assembly and Annotation of the Larval Transcriptome of Two Spadefoot Toads Widely Divergent in Developmental Rate. G3 9(8):2647-2655. DOI 10.1534/g3.119.400389.

Species (NOT Spea multiplicata as the registry hint implied):

  • Pelobates cultripes (slow developer; "Pc")
  • Scaphiopus couchii (fast developer; "Sc")

Pipeline (Methods): Illumina HiSeq2000 PE RNA-seq → FastQC/MultiQC → Trimmomatic (SLIDINGWINDOW:4:5 LEADING:5 TRAILING:5 MINLEN:25) → Trinity v2.4.0 in-silico normalization (--normalize_max_read_cov 50) + de novo assembly (per species, all samples pooled) → Bowtie2 read-representation check → BUSCO (tetrapoda_odb9) → TransDecoder ORFs → Trinotate v3.0 (blastx/blastp SwissProt, HMMER/Pfam, SignalP 4.1, TmHMM 2.0, eggNOG, KEGG) → trinotateR (the named repo: github.com/cstubben/trinotateR) to summarize the Trinotate report → downstream: blastx vs X. tropicalis, PANTHER, gProfiler, OrthoFinder.

Data available:

  • Raw reads: SRA PRJNA490256 / SRP161446 — 24 paired-end runs (12 Pc + 12 Sc), SRR7817202–SRR7817225. OPEN.
  • Deposited assemblies: TSA GHBH01000000 (Pc), GHBO01000000 (Sc); also FigShare sd2_transcriptomes.zip (371 MB). OPEN.
  • FigShare supplement 10.25387/g3.8201825 (article 8201825): sd1_Bowtie2_results, sd2_transcriptomes.zip, sd3_BUSCO_results, sd4_blastx_xtropicalis.xlsx, sd5_Trinotate.xlsx (111 MB) = the Trinotate report, sd7_gProfiler.xlsx. OPEN.

IN SCOPE (pipeline-derived, will attempt)

# Result Pipeline Strategy
C1 Assembly metrics: #transcripts, #genes, N50, total bases, median/avg length (Table 1) Trinity TrinityStats Run TrinityStats.pl on the deposited sd2 assemblies → must match Table 1 exactly (same data → verifies deposit delivers reported numbers).
C2 BUSCO completeness 89.7% / 86.6% (tetrapoda_odb9) BUSCO Re-run BUSCO on deposited assemblies → within-tol (version-dependent).
C3 Annotation summary counts: blastx/blastp SwissProt hits, Pfam, SignalP, TmHMM, unique GO terms trinotateR on sd5 Trinotate report CORE repo reproduction — run trinotateR summary_* on sd5 → match Results-section counts.
C4 Bowtie2 read-representation 83.96% / 90.71% Bowtie2 (Stretch) build bowtie2 index from deposited assembly, map raw reads back. Heavy I/O.

STRETCH (heavy, attempt after floor)

| C5 | Full Trinity de novo assembly from raw reads | Trimmomatic+Trinity | Very heavy (~84M normalized PE reads, large RAM, days). Assembly is non-deterministic across versions, so a 1:1 transcript count is NOT expected — at best within-tol/partial. Lower priority; document honestly. |

OUT OF SCOPE (not attempted, with reason)

  • Wet-lab: RNA extraction, library prep, sequencing (not computational).
  • PANTHER/gProfiler/OrthoFinder downstream enrichment (multi-tool, web-service dependent; secondary to core assembly+annotation; record reported values only).
  • eggNOG/KEGG pathway annotation detail (subsumed in Trinotate report; counts via trinotateR if present in sd5).

Notes / discrepancies to verify on data

  • Paper text says 2×76 bp reads; ENA base_count/read_pair = 202 → reads are ~2×101 bp. Verify read length directly on FASTQ.
  • ENA read_count = read PAIRS; paper "total raw reads" = mates (2× pairs). Pc: 2×444,132,722 = 888,265,444 ✓ exact. Confirms N concordance.
Figures / tables: Table
C1a
Reported
753223 transcripts (Pc)
Reproduced
753223
exact
C1b
Reported
657280 transcripts (Sc)
Reproduced
657280
exact
C1c
Reported
428406 genes (Pc)
Reproduced
428406
exact
C1d
Reported
381135 genes (Sc)
Reproduced
381135
exact
C1e
Reported
N50 1496 bp (Pc)
Reproduced
1496
exact
C1f
Reported
N50 2057 bp (Sc)
Reproduced
2057
exact
C1g
Reported
581.5 Mbp total (Pc)
Reproduced
581464720 bp
exact
C1h
Reported
644.9 Mbp total (Sc)
Reproduced
644907581 bp
exact
C1i
Reported
mean 771.97 bp (Pc)
Reproduced
771.97
exact
C1j
Reported
mean 981.18 bp (Sc)
Reproduced
981.18
exact
C1k
Reported
median 362 bp (Pc)
Reproduced
362
exact
C1l
Reported
median 432 bp (Sc)
Reproduced
432
exact
C3a
Reported
blastx SwissProt 162031 (Pc)
Reproduced
162031
exact
C3b
Reported
blastx SwissProt 204646 (Sc)
Reproduced
204646
exact
C3c
Reported
Pfam 91929 (Pc)
Reproduced
91929
exact
C3d
Reported
Pfam 112740 (Sc)
Reproduced
112738
within tolerance
C3e
Reported
SignalP 10981 (Pc)
Reproduced
10981
exact
C3f
Reported
SignalP 13097 (Sc)
Reproduced
13097
exact
C3g
Reported
TmHMM 25833 (Pc)
Reproduced
25833
exact
C3h
Reported
TmHMM 29991 (Sc)
Reproduced
29991
exact
C3i
Reported
GO unique 18585 (Pc)
Reproduced
18585
exact
C3j
Reported
GO unique 19917 (Sc)
Reproduced
19917
exact
C3k
Reported
TransDecoder ORFs 154906 (Pc)
Reproduced
154906
exact
C3l
Reported
TransDecoder ORFs 175331 (Sc)
Reproduced
175329
within tolerance
C2a
Reported
BUSCO complete 89.7% (Pc, odb9)
Reproduced
88.9% (odb10)
within tolerance
C2b
Reported
BUSCO complete 86.6% (Sc, odb9)
Reproduced
88.5% (odb10)
within tolerance
C2c
Reported
BUSCO fragmented 7.4% (Pc)
Reproduced
4.8% (odb10)
partial
C2d
Reported
BUSCO fragmented 10.5% (Sc)
Reproduced
5.6% (odb10)
partial
C2e
Reported
BUSCO missing 2.9% (Pc)
Reproduced
6.3% (odb10)
partial
C2f
Reported
BUSCO missing 2.9% (Sc)
Reproduced
5.9% (odb10)
partial
C4a
Reported
Bowtie2 representation 83.96% (Pc)
Reproduced
83.96% (deposited sd1)
m.public.grade.deposited-verified
C4b
Reported
Bowtie2 representation 90.71% (Sc)
Reproduced
90.71% (deposited sd1)
m.public.grade.deposited-verified

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 91/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

This is essentially a 1:1 reproduction on the authors' own deposited data (FigShare 8201825 assemblies + Trinotate report, SRA PRJNA490256): 12/12 Table-1 assembly metrics EXACT and 10/12 annotation counts EXACT via the named repo trinotateR, with the two off-by-2 values being xlsx→tsv round-trip noise. The only material deviation is BUSCO fragmented/missing, fully explained by using tetrapoda_odb10/v5 vs the paper's odb9/v3 — a technical version difference on our side, not an authors' defect, and Complete% still reproduces within tolerance. No fabrication indicators: every headline number is derivable from the shipped artifacts, so q5/q7/q8 are all green.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

<synthetic>

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

515.3 k
tokens (I/O) · 69.6 M incl. cache
228 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.