Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A comparison across non-model animals suggests an optimal sequencing depth for de novo transcriptome assembly.

BMC Genomics · 2013
L1 69/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
69/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 35% of all assessed papers rank 745 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the core Agalma/Velvet-Oases de novo transcriptome assembly pipeline from Chang et al. 2013 (PMID 23496952) for the paper's primary Mouse-C3 sample (SRR453174, GSE36025) at three sequencing depths (5M, 20M, 70M reads), using containerized seqtk/fastp/velvet/oases on «our HPC» HPC via SLURM. The 70M-depth run is the only one directly gradable against the paper's Table 1: per-transcript composition statistics (mean length, median length, N50, GC%) reproduced within 0.4-2.4% of the published values, but absolute assembly yield (transcript count, total length, Oases loci count) came in systematically 15-21% lower than published, plausibly linked to a documented, honest pipeline-order deviation (subsample-then-filter vs the paper's filter-then-subsample, leaving 62.16M vs the paper's exact 70.00M reads actually assembled) plus unpinned tool versions in the 2026 biocontainers vs the paper's original ~2012 Velvet/Oases builds. The 5M/20M depths qualitatively confirm the paper's Figures 1/3 depth-trend claim (monotonic increase in transcript count, length and N50 with depth) but have no exact numeric target since the paper does not tabulate intermediate depths. The Mouse-C10 cutoff variant, Trinity comparison, and all six invertebrate species datasets were not attempted in this pass (see qc_room) -- the six invertebrate datasets are effectively a drop, since no retrievable SRA/GEO accession exists for them anywhere in the paper, its supplement, or the agalma repo.

💻 Code ↗ 🗄 Data: GSE36025

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What sequencing depth is adequate for de novo transcriptome assembly in non-model organisms lacking a reference genome? The authors test how read count affects assembly quality by applying an identical assembly strategy across animals from six phyla plus a mouse control.

Core claims
  • Representative de novo transcriptome assemblies are generated with as few as ~20 million reads for single-tissue samples and ~30 million reads for whole animals at the mRNA-coverage level. finding
  • Beyond 60 million reads, discovery of new genes is low and sequencing errors of highly-expressed genes are likely to accumulate. finding
  • Whole-animal and single-tissue assemblies differ qualitatively: whole-animal samples show rapid gain of transcripts and rapid discovery of conserved genes, while single-tissue samples discover conserved genes more slowly but yield longer transcripts across fewer loci. finding
  • Assembly errors (mis-assembled coding sequences, introduced stop codons) become more frequent with more reads, but can be mitigated by more stringent assembly parameters such as a higher k-mer coverage cutoff. finding
  • Siphonophores (polymorphic cnidarians) are an exception, producing continuously increasing numbers of short transcripts and loci without reaching an asymptote, and likely require alternate assembly strategies or careful dissection. finding
  • A lower coverage cutoff (C3) and Trinity recover more conserved orthologs (KOGs) at fewer reads than a stringent cutoff (C10), suggesting it is better to use a low cutoff and assemble more sequences. finding
  • Using computational sub-sampling of a large read library at regular read increments, combined with conserved-ortholog (KOG/CEGMA) recovery, provides a benchmark framework for evaluating de novo assembly saturation without a reference genome. method
  • In mouse, whether the translated putative protein falls within the size range of the reference KOG proteins is a reliable predictor of a true full-length protein (no protein designated too short was ever correct). finding
Experimental setups
Assay System Perturbation Readout Platform
RNA-seq de novo transcriptome assembly (Velvet/Oases, k-mers 21-33, coverage cutoffs 3 and 10; and Trinity) Mus musculus (mouse) heart tissue, ENCODE sample SRR453174 from NCBI SRA computational sub-sampling of reads (1, 5, 10, 20, 30, 40, 50, 60, 70 million nested subsets) number of transcripts, mean/median/N50 length, number of Oases loci, loci and transcripts per million reads, transcripts per locus, GC content 76 bp paired-end reads (ENCODE/NCBI SRA); Velvet/Oases and Trinity assemblers
RNA-seq de novo transcriptome assembly (Velvet/Oases) Chuniphyes multidentata (siphonophore, Cnidaria), whole body (siphosome + nectophore combined) read-count sub-sampling (increments up to 80 million assembled reads) assembly size metrics (transcripts, mean/median/N50, loci, GC) and conserved-gene (KOG) recovery
RNA-seq de novo transcriptome assembly (Velvet/Oases) Hormiphora californensis (ctenophore), whole body read-count sub-sampling (up to 57,583,204 assembled reads) assembly size metrics and KOG recovery
RNA-seq de novo transcriptome assembly (Velvet/Oases) Pleuromamma robusta (copepod, Arthropoda), whole body, pooled multiple individuals read-count sub-sampling (up to 63,867,922 assembled reads) assembly size metrics and KOG recovery
RNA-seq de novo transcriptome assembly (Velvet/Oases) Dosidicus gigas (Humboldt squid, Mollusca), mantle tissue read-count sub-sampling (up to 56,264,099 assembled reads) assembly size metrics and KOG recovery
RNA-seq de novo transcriptome assembly (Velvet/Oases) Sergestes similis (decapod, Arthropoda), legs read-count sub-sampling (up to 80 million assembled reads) assembly size metrics and KOG recovery
RNA-seq de novo transcriptome assembly (Velvet/Oases) Harmothoe imbricata (scaleworm, Annelida), scale tissue read-count sub-sampling (up to 70,340,105 assembled reads) assembly size metrics and KOG recovery
Homology search / conserved-ortholog completeness assessment (tblastn against NCBI KOG reference set, blastp of putative proteins against Uniprot/Swissprot reviewed mouse proteins) Assembled mouse heart transcriptomes and the six invertebrate transcriptomes none (in silico analysis across read-count subsets); comparison of Oases C3, C10, and Trinity assemblies number of KOGs with a blast hit, number of translated proteins within expected reference size range, number of proteins 100% identical to canonical mouse protein; detection of erroneous/extraneous coding sequence in single-exon genes NCBI KOG database (860 clusters), CEGMA 248 single-copy subset, CDH metazoan set (1147 clusters → 202 single-copy), Uniprot/Swissprot
Key results
  • Loci counts approach a plateau with increasing reads in the mouse assemblies, with the greatest increase in loci occurring between 10 and 20 million reads for both C3 and C10; transcript numbers, however, increase steadily so transcripts per locus always increases.
  • For whole-body invertebrate transcriptomes, over 90% of conserved KOGs were detectable at 20 million reads, whereas dissected-tissue transcriptomes detected only 63%–81% of KOGs at 20 million reads. over 90% vs 63-81% of 248 KOGs
  • In whole-body transcriptomes, the number of KOGs whose translated proteins fell within the expected length range decreased beyond 20 million reads, caused mostly by mis-assembly generating stop codons; C. multidentata declined only after 50 million reads.
  • For the mouse C3 assembly at 70 million reads, 186 KOGs were within the expected size range but only 131 were 100% correct; 8 of the 186 differed by a single amino acid. 186 within range, 131 correct, 8 with 1 mismatch
  • At 70 million reads, 3 of 12 presumed single-copy single-exon mouse genes (NAT6, CHMP1B1/DID2, FTSJ) had erroneous alternate coding sequences in the C3 assembly, whereas only NAT6 did in the stricter C10 assembly. 3/12 (C3) vs 1/12 (C10)
  • Four of six marine animals showed modest gains in mean, median, and N50 length with more reads, while P. robusta and H. californensis nearly doubled; most transcript-length increase occurred before 30 million reads. average 20% gain; ~2-fold for P. robusta and H. californensis
  • C. multidentata (siphonophore) produced the highest transcript count and the lowest mean, median, and N50 lengths of all assemblies, with continuous gain of short transcripts and loci and no asymptote. 338,254 transcripts; mean 931 bp, median 421 bp, N50 1,854 bp
  • For all assemblies except mouse, the average GC content of assembled contigs was lower than that of the raw reads. e.g. C. multidentata raw 42.29% vs assembled 31.24%
Key statistics
  • count 82,886,668 x76bp paired-end reads (Raw mouse heart reads, ENCODE sample SRR453174)
  • other 11.7% of reads removed by filtration, almost 95% of which were due to low quality scores (Mouse heart read filtering; 73,187,048 filtered reads remained)
  • count 860 KOG gene clusters across 7 eukaryotes with over 16,000 proteins; 248-gene single-copy CEGMA subset; 1147 CDH clusters reduced to 202 metazoan single-copy KOGs (Conserved-ortholog reference sets used to assess assembly completeness)
  • count 186 KOGs within expected size range, 131 correct (Mouse C3 assembly at 70 million reads)
  • count over 90% of KOGs detectable at 20 million reads (whole body) vs 63%-81% (dissected tissue) (Conserved gene discovery in invertebrate transcriptomes)
  • other Mouse genome GC 42%; conserved gene subset GC 51.24%; raw mouse reads 51.90%; assembled C3 54.08% (GC content comparisons)
  • fold_change average 20% increase in mean/median/N50 from fewest to most reads (four animals); nearly 2-fold for P. robusta and H. californensis (Transcript length gains with increasing read count in marine organisms)
  • count 12 mouse proteins larger than the longest reference protein and 5 shorter than the shortest reference protein for their KOG (Flexibility of the KOG size-range criterion)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper reports a comparative bioinformatics analysis rather than a hypothesis-testing statistical study: de novo transcriptomes from six invertebrate phyla plus a mouse control were assembled with Velvet/Oases and Trinity from computationally sub-sampled, nested read sets (1 to 70 million reads), and assembly quality (transcript counts, length statistics, locus counts, GC content, and recovery of conserved KOG orthologs) was tracked as descriptive metrics and saturation curves across read depth. Comparisons between organisms, tissue types, and read depths are made by visual/qualitative inspection of these plotted trends rather than by formal significance testing.

Replicationunclear Sample sizeRead-depth sub-samples of 1, 5, 10, 20, 30, 40, 50, 60, and 70 million reads were drawn from each organism's filtered read library in a nested/cumulative fashion (each larger set contains the smaller sets); no statistical power calculation or independent biological replicate count is described. GroupsAssemblies of six invertebrate taxa (whole-body vs. single-tissue samples) and a mouse heart control, compared across increasing sub-sampled read depths and between two assemblers/parameter settings (Velvet/Oases C3, C10; Trinity) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Assembly-quality metrics were computed once per read depth from nested, cumulative sub-samples of a single original library, without independent replicate sub-samples or dispersion estimates.
    Could also: Multiple independent random sub-samples (or rarefaction/bootstrap resampling) could also be drawn at each read depth, with metrics reported as mean ± SD or a confidence interval. — This would let variability at each depth be quantified separately from the overall depth trend, since nested cumulative subsets are correlated with one another by construction.
  • The 'optimal' or saturating read depth is identified by visually inspecting where curves of transcript count, locus count, or KOG recovery appear to plateau.
    Could also: A quantitative saturation/asymptotic model (e.g., a Michaelis-Menten, exponential-plateau, or logistic curve) could also be fit to the depth-vs-metric data, with the fitted asymptote or half-saturation point used as a quantitative depth estimate. — A fitted model would give a reproducible numerical saturation point with an associated uncertainty, complementing visual assessment of curve plateaus.
  • Whole-body versus single-tissue assemblies, and differences among the six taxa, are described qualitatively based on the shape of the plotted metric curves.
    Could also: If additional replicate assemblies were available, a mixed-effects or ANOVA-type model treating organism and tissue type (whole-body vs. dissected) as factors could also be used to compare metrics such as KOG-recovery rate or N50. — This would formally test whether the whole-body/tissue-type distinction and taxon differences are statistically supported, rather than inferred from visual comparison of curves.
  • Completeness at each read depth is summarized as a raw count or percentage of KOGs detected or found within the expected length range, without an associated confidence interval.
    Could also: A binomial confidence interval (e.g., Wilson or Clopper-Pearson) on the proportion of KOGs recovered could also be reported at each depth. — This would convey the precision of a recovery percentage, which is particularly relevant when comparing organisms with different total KOG reference-set sizes.
  • GC-content distributions of assembled transcripts versus raw reads are compared by overlaying normalized histograms (Figure 2).
    Could also: A distributional comparison such as the Kolmogorov-Smirnov test could also be applied to the two GC-content distributions. — This would provide a quantitative statistic for the magnitude/significance of the shift between raw-read and assembled-contig GC distributions, complementing the visual histogram comparison.
Software: Velvet/Oases · Trinity · BLAST (tblastn, blastp)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

raw_reads_srr453174
Reported
82886668
Reproduced
82886668
exact
assembled_reads_70M
Reported
70000000
Reproduced
62158312
did not match
table1_transcripts
Reported
254215
Reproduced
201477
partial
table1_total_length
Reported
293550000
Reproduced
235684135
partial
table1_mean_length
Reported
1154
Reproduced
1169.8
within tolerance
table1_median_length
Reported
547
Reproduced
549
within tolerance
table1_n50
Reported
2364
Reproduced
2421
within tolerance
table1_oases_loci
Reported
77411
Reproduced
65931
partial
table1_gc_pct
Reported
54.08
Reproduced
54.26
within tolerance
depth_trend_qualitative
Reported
Qualitativ (Figs 1/3, nur grafisch): Transkriptzahl, N50 und mittlere Transkriptlaenge steigen monoton mit der Sequenziertiefe
Reproduced
Monoton bestaetigt ueber eigene Tiefen-Serie (SRR453174, gleiche multi-k Velvet/Oases-Pipeline): Transkripte 28.211(5M)->88.638(20M)->201.477(70M); N50 943->1.341->2.421 bp; Mittel 615,6->801,8->1.169,8 bp; Median 404->487->549 bp — alle vier Metriken steigen mit der Tiefe wie in Figs 1/3.
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

The raw input matched the paper exactly (SRR453174: 82,886,668 reads = Table 1), and per-transcript composition statistics reproduced tightly (mean 1,169.8 vs 1,154 bp; median 549 vs 547 bp; N50 2,421 vs 2,364 bp; GC 54.26 vs 54.08%). Absolute yield came in 15-21% low (201,477 vs 254,215 transcripts; 235.68 vs 293.55 Mbp; 65,931 vs 77,411 loci), and the deviation sits on our side: subsampling before filtering left 62,158,312 instead of 70,000,000 reads in the assembly, with unpinned 2026 Velvet/Oases builds as a secondary contributor. The paper's central conclusion — monotonic improvement of assembly metrics with sequencing depth (Figs 1/3) — is fully confirmed by an independent 5M/20M/70M series, so this is a solid, honestly documented partial reproduction rather than a substantive discrepancy. Two genuine authors'-side gaps remain non-fatal but worth logging: the QC filter is not described recoverably, and the six invertebrate Table 1 datasets have no locatable accession anywhere.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.