A comparison across non-model animals suggests an optimal sequencing depth for de novo transcriptome assembly.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the core Agalma/Velvet-Oases de novo transcriptome assembly pipeline from Chang et al. 2013 (PMID 23496952) for the paper's primary Mouse-C3 sample (SRR453174, GSE36025) at three sequencing depths (5M, 20M, 70M reads), using containerized seqtk/fastp/velvet/oases on «our HPC» HPC via SLURM. The 70M-depth run is the only one directly gradable against the paper's Table 1: per-transcript composition statistics (mean length, median length, N50, GC%) reproduced within 0.4-2.4% of the published values, but absolute assembly yield (transcript count, total length, Oases loci count) came in systematically 15-21% lower than published, plausibly linked to a documented, honest pipeline-order deviation (subsample-then-filter vs the paper's filter-then-subsample, leaving 62.16M vs the paper's exact 70.00M reads actually assembled) plus unpinned tool versions in the 2026 biocontainers vs the paper's original ~2012 Velvet/Oases builds. The 5M/20M depths qualitatively confirm the paper's Figures 1/3 depth-trend claim (monotonic increase in transcript count, length and N50 with depth) but have no exact numeric target since the paper does not tabulate intermediate depths. The Mouse-C10 cutoff variant, Trinity comparison, and all six invertebrate species datasets were not attempted in this pass (see qc_room) -- the six invertebrate datasets are effectively a drop, since no retrievable SRA/GEO accession exists for them anywhere in the paper, its supplement, or the agalma repo.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhat sequencing depth is adequate for de novo transcriptome assembly in non-model organisms lacking a reference genome? The authors test how read count affects assembly quality by applying an identical assembly strategy across animals from six phyla plus a mouse control.
- ★ Representative de novo transcriptome assemblies are generated with as few as ~20 million reads for single-tissue samples and ~30 million reads for whole animals at the mRNA-coverage level. finding
- ★ Beyond 60 million reads, discovery of new genes is low and sequencing errors of highly-expressed genes are likely to accumulate. finding
- ★ Whole-animal and single-tissue assemblies differ qualitatively: whole-animal samples show rapid gain of transcripts and rapid discovery of conserved genes, while single-tissue samples discover conserved genes more slowly but yield longer transcripts across fewer loci. finding
- ★ Assembly errors (mis-assembled coding sequences, introduced stop codons) become more frequent with more reads, but can be mitigated by more stringent assembly parameters such as a higher k-mer coverage cutoff. finding
- ★ Siphonophores (polymorphic cnidarians) are an exception, producing continuously increasing numbers of short transcripts and loci without reaching an asymptote, and likely require alternate assembly strategies or careful dissection. finding
- ★ A lower coverage cutoff (C3) and Trinity recover more conserved orthologs (KOGs) at fewer reads than a stringent cutoff (C10), suggesting it is better to use a low cutoff and assemble more sequences. finding
- ★ Using computational sub-sampling of a large read library at regular read increments, combined with conserved-ortholog (KOG/CEGMA) recovery, provides a benchmark framework for evaluating de novo assembly saturation without a reference genome. method
- In mouse, whether the translated putative protein falls within the size range of the reference KOG proteins is a reliable predictor of a true full-length protein (no protein designated too short was ever correct). finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-seq de novo transcriptome assembly (Velvet/Oases, k-mers 21-33, coverage cutoffs 3 and 10; and Trinity) | Mus musculus (mouse) heart tissue, ENCODE sample SRR453174 from NCBI SRA | computational sub-sampling of reads (1, 5, 10, 20, 30, 40, 50, 60, 70 million nested subsets) | number of transcripts, mean/median/N50 length, number of Oases loci, loci and transcripts per million reads, transcripts per locus, GC content | 76 bp paired-end reads (ENCODE/NCBI SRA); Velvet/Oases and Trinity assemblers |
| RNA-seq de novo transcriptome assembly (Velvet/Oases) | Chuniphyes multidentata (siphonophore, Cnidaria), whole body (siphosome + nectophore combined) | read-count sub-sampling (increments up to 80 million assembled reads) | assembly size metrics (transcripts, mean/median/N50, loci, GC) and conserved-gene (KOG) recovery | — |
| RNA-seq de novo transcriptome assembly (Velvet/Oases) | Hormiphora californensis (ctenophore), whole body | read-count sub-sampling (up to 57,583,204 assembled reads) | assembly size metrics and KOG recovery | — |
| RNA-seq de novo transcriptome assembly (Velvet/Oases) | Pleuromamma robusta (copepod, Arthropoda), whole body, pooled multiple individuals | read-count sub-sampling (up to 63,867,922 assembled reads) | assembly size metrics and KOG recovery | — |
| RNA-seq de novo transcriptome assembly (Velvet/Oases) | Dosidicus gigas (Humboldt squid, Mollusca), mantle tissue | read-count sub-sampling (up to 56,264,099 assembled reads) | assembly size metrics and KOG recovery | — |
| RNA-seq de novo transcriptome assembly (Velvet/Oases) | Sergestes similis (decapod, Arthropoda), legs | read-count sub-sampling (up to 80 million assembled reads) | assembly size metrics and KOG recovery | — |
| RNA-seq de novo transcriptome assembly (Velvet/Oases) | Harmothoe imbricata (scaleworm, Annelida), scale tissue | read-count sub-sampling (up to 70,340,105 assembled reads) | assembly size metrics and KOG recovery | — |
| Homology search / conserved-ortholog completeness assessment (tblastn against NCBI KOG reference set, blastp of putative proteins against Uniprot/Swissprot reviewed mouse proteins) | Assembled mouse heart transcriptomes and the six invertebrate transcriptomes | none (in silico analysis across read-count subsets); comparison of Oases C3, C10, and Trinity assemblies | number of KOGs with a blast hit, number of translated proteins within expected reference size range, number of proteins 100% identical to canonical mouse protein; detection of erroneous/extraneous coding sequence in single-exon genes | NCBI KOG database (860 clusters), CEGMA 248 single-copy subset, CDH metazoan set (1147 clusters → 202 single-copy), Uniprot/Swissprot |
- – Loci counts approach a plateau with increasing reads in the mouse assemblies, with the greatest increase in loci occurring between 10 and 20 million reads for both C3 and C10; transcript numbers, however, increase steadily so transcripts per locus always increases.
- – For whole-body invertebrate transcriptomes, over 90% of conserved KOGs were detectable at 20 million reads, whereas dissected-tissue transcriptomes detected only 63%–81% of KOGs at 20 million reads. over 90% vs 63-81% of 248 KOGs
- ▼ In whole-body transcriptomes, the number of KOGs whose translated proteins fell within the expected length range decreased beyond 20 million reads, caused mostly by mis-assembly generating stop codons; C. multidentata declined only after 50 million reads.
- – For the mouse C3 assembly at 70 million reads, 186 KOGs were within the expected size range but only 131 were 100% correct; 8 of the 186 differed by a single amino acid. 186 within range, 131 correct, 8 with 1 mismatch
- ▼ At 70 million reads, 3 of 12 presumed single-copy single-exon mouse genes (NAT6, CHMP1B1/DID2, FTSJ) had erroneous alternate coding sequences in the C3 assembly, whereas only NAT6 did in the stricter C10 assembly. 3/12 (C3) vs 1/12 (C10)
- ▲ Four of six marine animals showed modest gains in mean, median, and N50 length with more reads, while P. robusta and H. californensis nearly doubled; most transcript-length increase occurred before 30 million reads. average 20% gain; ~2-fold for P. robusta and H. californensis
- – C. multidentata (siphonophore) produced the highest transcript count and the lowest mean, median, and N50 lengths of all assemblies, with continuous gain of short transcripts and loci and no asymptote. 338,254 transcripts; mean 931 bp, median 421 bp, N50 1,854 bp
- ▼ For all assemblies except mouse, the average GC content of assembled contigs was lower than that of the raw reads. e.g. C. multidentata raw 42.29% vs assembled 31.24%
- count 82,886,668 x76bp paired-end reads (Raw mouse heart reads, ENCODE sample SRR453174)
- other 11.7% of reads removed by filtration, almost 95% of which were due to low quality scores (Mouse heart read filtering; 73,187,048 filtered reads remained)
- count 860 KOG gene clusters across 7 eukaryotes with over 16,000 proteins; 248-gene single-copy CEGMA subset; 1147 CDH clusters reduced to 202 metazoan single-copy KOGs (Conserved-ortholog reference sets used to assess assembly completeness)
- count 186 KOGs within expected size range, 131 correct (Mouse C3 assembly at 70 million reads)
- count over 90% of KOGs detectable at 20 million reads (whole body) vs 63%-81% (dissected tissue) (Conserved gene discovery in invertebrate transcriptomes)
- other Mouse genome GC 42%; conserved gene subset GC 51.24%; raw mouse reads 51.90%; assembled C3 54.08% (GC content comparisons)
- fold_change average 20% increase in mean/median/N50 from fewest to most reads (four animals); nearly 2-fold for P. robusta and H. californensis (Transcript length gains with increasing read count in marine organisms)
- count 12 mouse proteins larger than the longest reference protein and 5 shorter than the shortest reference protein for their KOG (Flexibility of the KOG size-range criterion)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper reports a comparative bioinformatics analysis rather than a hypothesis-testing statistical study: de novo transcriptomes from six invertebrate phyla plus a mouse control were assembled with Velvet/Oases and Trinity from computationally sub-sampled, nested read sets (1 to 70 million reads), and assembly quality (transcript counts, length statistics, locus counts, GC content, and recovery of conserved KOG orthologs) was tracked as descriptive metrics and saturation curves across read depth. Comparisons between organisms, tissue types, and read depths are made by visual/qualitative inspection of these plotted trends rather than by formal significance testing.
-
Assembly-quality metrics were computed once per read depth from nested, cumulative sub-samples of a single original library, without independent replicate sub-samples or dispersion estimates.↳ Could also: Multiple independent random sub-samples (or rarefaction/bootstrap resampling) could also be drawn at each read depth, with metrics reported as mean ± SD or a confidence interval. — This would let variability at each depth be quantified separately from the overall depth trend, since nested cumulative subsets are correlated with one another by construction.
-
The 'optimal' or saturating read depth is identified by visually inspecting where curves of transcript count, locus count, or KOG recovery appear to plateau.↳ Could also: A quantitative saturation/asymptotic model (e.g., a Michaelis-Menten, exponential-plateau, or logistic curve) could also be fit to the depth-vs-metric data, with the fitted asymptote or half-saturation point used as a quantitative depth estimate. — A fitted model would give a reproducible numerical saturation point with an associated uncertainty, complementing visual assessment of curve plateaus.
-
Whole-body versus single-tissue assemblies, and differences among the six taxa, are described qualitatively based on the shape of the plotted metric curves.↳ Could also: If additional replicate assemblies were available, a mixed-effects or ANOVA-type model treating organism and tissue type (whole-body vs. dissected) as factors could also be used to compare metrics such as KOG-recovery rate or N50. — This would formally test whether the whole-body/tissue-type distinction and taxon differences are statistically supported, rather than inferred from visual comparison of curves.
-
Completeness at each read depth is summarized as a raw count or percentage of KOGs detected or found within the expected length range, without an associated confidence interval.↳ Could also: A binomial confidence interval (e.g., Wilson or Clopper-Pearson) on the proportion of KOGs recovered could also be reported at each depth. — This would convey the precision of a recovery percentage, which is particularly relevant when comparing organisms with different total KOG reference-set sizes.
-
GC-content distributions of assembled transcripts versus raw reads are compared by overlaying normalized histograms (Figure 2).↳ Could also: A distributional comparison such as the Kolmogorov-Smirnov test could also be applied to the two GC-content distributions. — This would provide a quantitative statistic for the magnitude/significance of the shift between raw-read and assembled-contig GC distributions, complementing the visual histogram comparison.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The raw input matched the paper exactly (SRR453174: 82,886,668 reads = Table 1), and per-transcript composition statistics reproduced tightly (mean 1,169.8 vs 1,154 bp; median 549 vs 547 bp; N50 2,421 vs 2,364 bp; GC 54.26 vs 54.08%). Absolute yield came in 15-21% low (201,477 vs 254,215 transcripts; 235.68 vs 293.55 Mbp; 65,931 vs 77,411 loci), and the deviation sits on our side: subsampling before filtering left 62,158,312 instead of 70,000,000 reads in the assembly, with unpinned 2026 Velvet/Oases builds as a secondary contributor. The paper's central conclusion — monotonic improvement of assembly metrics with sequencing depth (Figs 1/3) — is fully confirmed by an independent 5M/20M/70M series, so this is a solid, honestly documented partial reproduction rather than a substantive discrepancy. Two genuine authors'-side gaps remain non-fatal but worth logging: the QC filter is not described recoverably, and the six invertebrate Table 1 datasets have no locatable accession anywhere.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.