Corpus 1,284 assessed · 1,185 scored · 647 reproduced ≥75 · 173 flagged ·∅ 73.9/100
← New search

De novo assembly of a transcriptome for Calanus finmarchicus (Crustacea, Copepoda)--the dominant zooplankter of the North Atlantic Ocean.

PLoS One · 2014
59/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
59/100
Reproducibility score
0.8 SD below mean
vs. all fields · 1185 studies
🎯 Scores higher than 19% of all assessed papers rank 930 of 1185 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Raw SRA read counts reproduce essentially exactly (415,469,687 vs 415,469,690 reported, off by 3 reads on one of 6 runs) -- strong validation that the correct dataset was used. The de novo Trinity assembly itself, and everything downstream of it (assembly stats, SwissProt annotation rate, read-mapping-back rates), diverges numerically from the paper because the paper used a bespoke, no-longer-obtainable 2012 tuned Trinity build (2012-03-17-IU_ZIH_TUNED) and FASTX-toolkit trimming, while this reproduction necessarily substitutes modern Trinity 2.5.1 and fastp -- an explicit, documented, unavoidable methodological deviation, not a pipeline failure. The resulting assembly is markedly more fragmented (557,511 vs 206,041 transcripts; 174,932 vs 96,090 genes; N50 606bp vs 1,418bp) yet strikingly similar in GC content (42.82% vs 43%). SwissProt annotation rate is same order of magnitude (35.7% vs 29.8%) but not exact, consistent with a larger, more fragmented gene set and a newer SwissProt release. Read-mapping-back (Table 3) partially reproduces: the longest-isoform/gene-representative alignment rate is close (73.44% vs ~75%, within-tolerance) while the whole-assembly rate is lower than reported (79.55% vs ~89%), plausibly because a much larger, more redundant/fragmented transcript set changes multi-mapping/ambiguous-alignment behavior under bowtie2 end-to-end defaults. NOT attempted: Table 4 (6 separate stage-specific assemblies) and GO/nr-BLAST annotation, both scoped out as disproportionate additional compute for marginal evidentiary gain beyond the already-reproduced core pipeline (assembly -> annotation -> read-mapping). No wet-lab/manual results were in scope. Housekeeping: trinity_out/ intermediate directory (~300GB) was left in place pending explicit operator confirmation to delete, per this session's 'Loeschen immer' autonomy rule, which the operator should resolve.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-08
Rubric version
not recorded
Assessed by
Last updated
2026-08-08

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can deep Illumina RNA-Seq and de novo assembly across six developmental stages produce a sufficiently deep, well-annotated transcriptome for the large-genome copepod Calanus finmarchicus to enable protein discovery and stage-specific gene expression analysis, given the lack of a reference genome and the need for physiological data on this ecologically key North Atlantic zooplankter?

Core claims
  • A de novo transcriptome for Calanus finmarchicus was assembled from six developmental-stage libraries, yielding 206,041 contigs and a reference set of 96,090 unique comps, representing a new molecular resource for this species. resource
  • Assembly coverage was estimated to be at least 65%, as deep or deeper than other crustacean de novo transcriptome studies. finding
  • Expression of many comps was near zero in one or more stages, indicating that 35 to 48% of the transcriptome is 'silent' at any given life stage. finding
  • Relative expression of three lipid-biosynthesis transcripts suggests wax ester biosynthesis in late copepodites but triacylglyceride biosynthesis in adult females. finding
  • Two lipid-biosynthesis transcripts have developmental expression patterns consistent with involvement in the preparatory phase of diapause. mechanism
  • Multiple contigs encoding putative voltage-gated sodium channels were identified, apparently arising from both alternative splicing and gene duplication; this is the first report of multiple NaV1 genes in a protostome. finding
  • Multiplexed sequencing of stage-specific libraries in a single lane, combined with Trinity assembly, is an effective method to capture developmentally differentially expressed genes in a non-model organism with a very large genome. method
  • Targeted transcript discovery via reciprocal BLAST, domain-structure checks, and comparison to extant C. finmarchicus ESTs validates assembly completeness and correctness. method
Experimental setups
Assay System Perturbation Readout Platform
Bulk RNA-Seq (paired-end 100 bp, poly(A)-selected, multiplexed, single lane) Calanus finmarchicus whole individuals; six developmental stages (embryo, NI-II, NV-VI, CI-II, CV, adult CVI female); field-collected (Gulf of Maine) adults/CV and laboratory-reared earlier stages none (developmental stage comparison) sequence reads / transcript abundance per stage Illumina HiSeq 2000; TruSeq RNA sample preparation kit (RS-122-2001), 350 bp insert, random hexamer priming
Total RNA extraction and QC pooled C. finmarchicus individuals per stage (400 eggs, 185 early nauplii, 50 late nauplii, 40 early copepodites, 6 CV, 10 adult females) none RNA concentration (ng/µL) and quality QIAGEN RNeasy Plus Mini Kit (#74134) with Qiashredder (#79654); Agilent 2100 Bioanalyzer / RNA 6000 Nano
Read QC, trimming and filtering raw Illumina reads from all six libraries none over-represented (rRNA) and low-quality read removal; 9 bp random-primer trimming FASTQC v0.10.0; FASTX Toolkit v0.013; blastn; Phred cutoff 20
De novo transcriptome assembly combined reads from all six developmental stages (plus stage-specific and read-subset assemblies) varied sequencing depth (subsets starting at 6 million reads) and stage identity contig number, length distribution, N50/N25/N75, GC content, comps Trinity 2012-03-17-IU_ZIH_TUNED (min_contig_length 300, --edge-thr=0.05, jellyfish, 32 CPU) on NCGAS Mason Linux cluster (Intel Xeon L7555, 512 GB memory)
Read mapping / expression quantification stage-specific and combined read sets against complete assembly and reference transcriptome none overall alignment rate, multi-mapping rate, mapped reads per comp per stage (silent = ≤2 mapped reads) Bowtie v2.0.6 (default, two mismatches)
Functional annotation (blastx and GO assignment) reference transcriptome of 96,090 unique comps none proportion with significant blast hits and GO terms for biological process, molecular function, cellular component; comparison to Drosophila melanogaster precomputed annotation Blast2GO v2.6.4 vs NCBI nr and SwissProt (Feb 2013); E-value ≤10⁻³ blast, ≤10⁻⁶ GO
Targeted homology-based gene discovery (Tera-BLASTP, reciprocal BLAST, domain alignment) complete C. finmarchicus assembly; queries typically D. melanogaster proteins none putative lipid-biosynthesis proteins and voltage-gated sodium channel (NaV1) sequences; conserved functional domains TimeLogic DeCypher server (Mount Desert Island Biological Laboratory); FlyBase and NCBI nr for reciprocal BLAST
Nucleotide-level validation by blastn against existing ESTs assembled C. finmarchicus contigs vs curated C. finmarchicus ESTs at NCBI none identity/similarity of assembled sequences to known ESTs NCBI blastn; <12,000 extant ESTs
Key results
  • Sequencing yielded over 400 million paired-end 100 bp reads across six stage libraries 415,469,690 raw reads; average ~69 million per sample
  • Trinity assembly produced 206,041 contigs totaling 205,480,825 bp, containing 96,090 unique comps of which 73,925 were single-contig average contig 997 bp; N50 1,418 bp; max 23,068 bp; GC 43%
  • Annotation of the reference transcriptome gave significant blast hits for a minority of comps and GO terms for a smaller fraction 40% with blast hits; 11% GO-annotated
  • Estimated transcriptome coverage of the assembly at least 65%; ~60 to 90-fold sequencing coverage; assembled-bp ratio ~150
  • A large fraction of the transcriptome is not expressed (≤2 mapped reads) at any given developmental stage 35 to 48%
  • Mapping against the complete assembly gave high alignment but substantial multi-mapping, while the unique-comp reference reduced both 89% alignment with 44% mapped >1 time (complete) vs 75% alignment with 0.7% mapped >1 time (reference)
  • Lipid-biosynthesis transcript expression shifts between stages, indicating wax ester synthesis in late copepodites versus triacylglyceride synthesis in adult females
  • Multiple putative voltage-gated sodium channel contigs identified, attributed to alternative splicing and gene duplication
Key statistics
  • count 415,469,690 raw reads (total across six libraries) (Illumina HiSeq 2000 sequencing yield, Table 1)
  • count 401,836,653 trimmed, high-quality 91 bp reads assembled (input to Trinity de novo assembly, Table 2)
  • count 206,041 assembled contigs; 96,090 unique comps; 73,925 (77%) single-contig comps (assembly output and reference transcriptome composition)
  • other N50 = 1,418 bp; N25 = 2,748 bp; N75 = 701 bp; average 997 bp; max 23,068 bp; min 301 bp (assembly length statistics, Table 2)
  • other GC content 43% (88,329,861 bp GC of 205,480,825 bp total) (whole-assembly base composition)
  • other 40% of comps with significant blast hits; 11% annotated with GO terms (Blast2GO annotation of reference transcriptome)
  • other 89% overall alignment with 44% of reads mapping >1 time (complete assembly); 75% alignment with 0.7% mapping >1 time (reference transcriptome) (Bowtie read mapping, Table 3)
  • other coverage at least 65%; ~60 to 90-fold sequencing coverage; 35 to 48% of transcriptome 'silent' per life stage (assembly completeness and stage-specific expression estimates)
  • other C-value 6.48 pg, i.e. estimated genome size >6,000 Mb (1 pg = 978 Mb); assumed 7–10% of genome transcribed (C. finmarchicus genome size used for coverage estimation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper describes a de novo transcriptome assembly and annotation project for Calanus finmarchicus rather than a hypothesis-testing study. RNA from six developmental stages was pooled into single samples per stage, sequenced on Illumina, assembled with Trinity, and annotated with Blast2GO using BLAST E-value cutoffs; results are reported as descriptive counts, percentages, and assembly-quality metrics (e.g., N50, read-mapping rates, GC content) rather than through inferential statistical tests with p-values or effect sizes.

Replicationunclear Sample sizeNumber of pooled individuals per developmental-stage sample is stated (e.g., 400 eggs, 180 early nauplii, 50 late nauplii, 40 early copepodites, 6 late copepodites, 10 adult females), but a single RNA sample/library was generated per stage with no described biological or technical replicates, and no power analysis was reported. GroupsSix C. finmarchicus developmental stages: embryo, early nauplius (NI-NII), late nauplius (NV-NVI), early copepodite (CI-CII), late copepodite (CV), adult female (CVI) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
BLAST-based sequence similarity search (blastx/blastn/Tera-BLASTP) with E-value significance cutoff Annotation of the reference transcriptome against NCBI nr and SwissProt databases (Blast2GO), and targeted BLAST searches for lipid biosynthesis and voltage-gated sodium channel transcripts not stated
Approaches that could also have been used
  • A single pooled RNA sample was sequenced per developmental stage, without biological or technical replicates.
    Could also: Collecting multiple biological replicates per stage — Replicates would allow variance estimation and formal differential-expression testing (e.g., DESeq2 or edgeR) with statistical significance measures for stage-specific expression differences, which the current single-sample design does not support.
  • Transcripts with near-zero read counts in a given stage were described as 'silent' based on a simple mapped-read threshold (≤2 reads), without a formal statistical test.
    Could also: A model-based differential expression or expression-presence test (e.g., DESeq2/edgeR Wald or likelihood-ratio test, or NOISeq for no-replicate data) — Such methods provide a statistical framework with associated p-values and effect sizes for calling genes as expressed or not expressed, complementing the descriptive read-count threshold used here.
  • Sequence annotation significance was assessed using fixed BLAST E-value cutoffs (10^-3 and 10^-6).
    Could also: FDR-adjusted significance thresholds across the large number of BLAST comparisons performed — Because annotation involves searching many thousands of sequences against large databases, an FDR-based threshold could additionally quantify the expected proportion of false-positive annotations among those retained.
  • The proportion of annotated sequences in each GO category was compared descriptively to the corresponding proportions in a precomputed Drosophila melanogaster annotation.
    Could also: A formal statistical comparison of proportions (e.g., chi-square or Fisher's exact test) between the two annotation sets — This would let readers assess whether differences in GO category representation between C. finmarchicus and D. melanogaster exceed what might be expected by chance, complementing the direct percentage comparison presented.
Software: FASTQC 0.10.0 · FASTX Toolkit 0.013 · Trinity 2012-03-17-IU_ZIH_TUNED · Bowtie 2.0.6 · Blast2GO 2.6.4 · TimeLogic DeCypher / Tera-BLASTP

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

raw_reads_table1
Reported
415,469,690 total raw reads across 6 developmental-stage SRA runs (Table 1): SRR1150766=59,001,054; SRR1150768=70,036,535; SRR1150770=71,064,356; SRR1150771=71,585,082; SRR1153468=75,871,920; SRR1153469=67,910,746
Reproduced
415,469,687 total (fastp total_reads pre-filter, sum of 6 runs): SRR1150766=59,001,054; SRR1150768=70,036,532 (paper: 70,036,535, diff 3 reads); SRR1150770=71,064,356; SRR1150771=71,585,082; SRR1153468=75,871,920; SRR1153469=67,910,746
exact
assembly_stats_table2
Reported
206,041 total assembled contigs; 96,090 unique comps (genes); N50=1,418 bp; total assembly length=205,480,825 bp; avg contig length=997 bp; GC=43%; Trinity build 2012-03-17-IU_ZIH_TUNED, CPU=32, max_memory=20G, min_contig_length=300
Reproduced
557,511 transcripts; 174,932 Trinity genes; N50(all transcripts)=606 bp (N50 longest-isoform-only=705 bp); total assembled bases=328,748,173 bp; avg contig length=589.67 bp; GC=42.82%; Trinity 2.5.1, CPU=192, max_memory=700G, min_contig_length=300
did not match
read_mapping_table3_whole_assembly
Reported
~89% overall alignment rate, all reads mapped back to whole assembly (Bowtie, default settings, up to 2 mismatches)
Reproduced
79.55% overall alignment rate (bowtie2 --end-to-end default, whole trinity_out.Trinity.fasta as reference, all 6 samples pooled, 194,248,323 read pairs)
partial
read_mapping_table3_longest_isoform
Reported
~75% overall alignment rate, reads mapped back to one representative (longest isoform) transcript per gene
Reproduced
73.44% overall alignment rate (bowtie2 --end-to-end default, trinity_out.Trinity.longest_isoform.fasta as reference, all 6 samples pooled, 194,248,323 read pairs)
within tolerance
swissprot_annotation_table5
Reported
28,616 of 96,090 comps (29.8%) with significant SwissProt BLAST matches, E-value cutoff 1e-3
Reproduced
62,497 of 174,932 genes (35.7%, longest-isoform representative per gene) with significant SwissProt BLAST matches, E-value cutoff 1e-3, current UniProt/SwissProt release (575,503 seqs, not the exact 2013-era release used by the paper)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.