De Novo Transcriptome Meta-Assembly of the Mixotrophic Freshwater Microalga Euglena gracilis.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH; reproduced 1:1 on the deterministic core. The paper's headline deposited final meta-assembly (ENA TSA HBDM01, 91,040 transcripts) reproduces its Table-3 statistics essentially exactly from the public deposit: transcript count, total size (100.1 Mb), N50 (1432), #>1kb (37,294), #>10kb (24) and GC (63.03%) all match exactly; mean length 1099.9 vs 1096 (within-tol). Dataset claims also verified against ENA: 23 runs across 5 experiments, ~2.66 billion reads, 71% from the new experiment E (PRJEB38787) -- all matching. One auditable discrepancy: paper reports 49,922 non-redundant genes but the deposit contains 49,822 unique EvidentialGene loci (delta exactly -100; likely typo). NOT ATTEMPTED / blocked: BUSCO (84.8%), ORF count (62,287), GeneMarkS-T (58,542) are deterministic and feasible but require writing to «our HPC» «infra», whose user quota («user») is fully exhausted account-wide (no byte writable) -- a central janitor must reclaim space; the streamed stat computations above succeeded precisely because they write nothing to disk. The 36-cluster expression analysis (repo clusterization.R = PAM k-medoids on 1-Pearson^2) is readable but its input expression matrix is neither committed to the repo nor deposited, so it is not reproducible from shipped artifacts. The full de-novo Trinity meta-assembly from 2.6B reads is out of scope as a 1:1 target (Trinity is non-deterministic, so re-assembly cannot be bit-identical, plus TB/weeks cost) -- the faithful audit is verifying the deposited consensus, which passed. Also flagged: the brief's data accession sra:PRJEB4713 is a mis-tag (old 454 set, not this paper's data).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 87assessed: 2026-06-18 ⛓ 729070742a3f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a more complete consensus transcriptome of Euglena gracilis be assembled by combining reads from multiple independent studies, and what does cross-condition expression analysis and taxonomic annotation reveal about transcriptional regulation and the mosaic gene ancestry of this secondary green alga?
- ★ A consensus transcriptome assembled by combining reads from five independent studies is the most complete E. gracilis transcriptome released to date, outperforming the two previously available transcriptomes (GEFR01 and GDJR01). resource
- ★ Gene regulation in euglenozoans is not primarily controlled at the transcriptional level, despite emergence of meaningful co-expression gene clusters. finding
- ★ E. gracilis exhibits heavily mixed (mosaic) gene ancestry, and sequence contamination is ruled out as an explanation, indicating evolution through a process involving more than two partners (compatible with a kleptoplastidic phase). finding
- ★ A de novo meta-assembly pipeline combining Trinity multi-parameter assembly with EvidentialGene consensus selection and BLASTN/genome-alignment-based decontamination produces a purified, functionally and taxonomically annotated transcriptome. method
- A functionally annotated co-expression network of E. gracilis genes was inferred from transcript expression across multiple culture conditions. resource
- Remapping reads onto the consensus transcriptome enables comparison of transcript expression across multiple culture conditions simultaneously. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (whole transcriptome, paired-end Illumina) | Euglena gracilis strain SAG 1224-5/25, liquid TMP medium pH 7.0, 25 °C | acetate (60 mM) carbon source vs none, under dark / low PPFD (50 μE) / medium PPFD (200 μE) | transcript abundance / read counts | Illumina HiSeq 2000, Duplex-Specific Nuclease (Evrogen) DSN normalization, Illumina total mRNA kit |
| de novo transcriptome assembly + consensus meta-assembly | E. gracilis reads from five Illumina experiments (5 experiments/23 samples) | none | assembled transcripts / coding sequences / representative isoforms | Trinity v2.4.0, EvidentialGene v2016.07.11 (tr2aacds.pl, evgmrna2tsa2.pl), CD-HIT v4.6.8 |
| transcriptome decontamination (sequence similarity + read filtering) | five per-experiment E. gracilis transcriptomes | none | GC-content distribution, contaminant identification, purified transcripts | BLASTN v2.2.28 vs NCBI nt, Bowtie 2 v2.2.6 |
| transcriptome quality assessment (read representation, reference-free scoring, completeness) | three transcriptomes (consensus, GEFR01, GDJR01) | none | mapping rate, quality scores, BUSCO completeness, gene/ORF counts | Bowtie 2 v2.2.6, Detonate v1.11, TransRate v1.0.3, BUSCO v3.0.1, GeneMarkS-T |
| functional annotation | consensus transcriptome representative isoforms | none | GO and KO term assignments, ortholog annotation | EggNOG-mapper v1 (HMMER), PSI-BLAST v2.2.28 vs Swiss-Prot, TBLASTN v2.2.28 vs NCBI nr |
| taxonomic (LCA) analysis | assembled transcripts of consensus transcriptome | none | last common ancestor taxonomic affiliation per transcript | BLASTX v2.2.28 vs database of 73 eukaryotes + 19,802 prokaryotes, MEGAN-like algorithm |
| tetranucleotide frequency / PCA analysis | taxonomically affiliated transcripts | none | TNF principal components vs GC content and taxonomic group | compseq (EMBOSS), prcomp (STATS v3.4.3 R) |
| expression quantification + co-expression network | consensus transcriptome across multiple culture conditions/studies | varied culture conditions (carbon source, light) | TPM gene-level expression, log2/Z-score normalized, co-expression clusters | RSEM v1.2.31, Bowtie2 v2.2.6, Trinity scripts |
- ▲ Consensus transcriptome is the most complete E. gracilis transcriptome compared to GEFR01 and GDJR01
- – Meaningful co-expressed gene clusters emerge, but transcriptional control is confirmed not to be the primary level of genetic regulation in euglenozoans
- – Heavily mixed gene ancestry observed, with sequence contamination ruled out as the cause
- – Of nine candidate public studies, five Illumina short-read datasets (5 experiments/23 samples) were exploitable for assembly 5 of 9 studies
- – In-house Illumina HiSeq 2000 sequencing yielded on average ca. 235 million reads per sample ~235 million reads/sample
- count five different data sources / 5 experiments / 23 samples (exploitable Illumina whole-transcriptome datasets used for assembly)
- count ~235 million reads per sample (average) (in-house HiSeq 2000 paired-end 2×100 nt sequencing yield)
- count eight studies returned from INSDC + one in-house = nine; five retained (public RNA-Seq data search results)
- other E-value 1×10⁻⁵⁰, identity ≥90% (BLASTN contamination screen against NCBI nt)
- other ≥98% identity exon-sized fragments; CD-HIT 90% amino-acid similarity threshold (isoform clustering criteria in EvidentialGene/CD-HIT)
- other bit-score ≥80 and within 95% of top hit bit-score (LCA computation thresholds for taxonomic affiliation)
- count 73 eukaryotes + 19,802 prokaryotes (from 27,762 genomes) (proteome database for BLASTX taxonomic analysis)
- other <1 TPM (expression threshold below which a 'missing' sequence was deemed invalid)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper describes a bioinformatics pipeline for de novo transcriptome meta-assembly of Euglena gracilis, combining RNA-Seq data from five independent public and in-house experiments (27 total samples). The primary analytical outputs are assembly quality metrics (BUSCO completeness, Detonate, TransRate scores), expression quantification via RSEM (TPM, log2-transformed and Z-score normalised), and principal component analysis (PCA) of tetranucleotide frequencies for taxonomic characterisation. The paper is predominantly a methods and resource study with descriptive comparison across three transcriptomes rather than a classical hypothesis-testing framework. The full methods text is truncated before the co-expression network and enrichment analysis sections, so some downstream statistical approaches cannot be fully characterised here.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Principal component analysis (PCA; R stats::prcomp) | Taxonomic analysis of tetranucleotide frequencies (TNFs) across assembled transcripts | 10 PCAs each on 1000 randomly sampled transcripts | not stated |
| BUSCO completeness scoring (Eukaryota and Protists_ensembl datasets) | Transcriptome completeness evaluation and comparison of three transcriptomes | three transcriptomes (GEFR01, GDJR01, consensus) | na |
| Detonate (reference-free quality score) | Transcriptome quality comparison across three transcriptomes | three transcriptomes | na |
| TransRate (reference-free quality score) | Transcriptome quality comparison across three transcriptomes | three transcriptomes | na |
| RSEM-based expression quantification with log2 transformation and Z-score normalisation | Cross-condition and cross-study expression comparison of assembled transcripts | 23 public samples plus 4 in-house samples (27 total) | not stated |
| BLAST E-value thresholding (BLASTN, BLASTX, TBLASTN, BLASTP, PSI-BLAST, MegaBLAST) with MEGAN-like LCA algorithm | Sequence annotation, decontamination, taxonomic affiliation, and cross-transcriptome similarity | all assembled transcripts; thresholds: E-value 1×10⁻⁵⁰ (decontamination/rRNA), 1×10⁻¹⁹ (isoform detection), 0.001 (annotation); bit-score ≥80 and within 95% of top hit (LCA) | na |
-
Expression quantification used RSEM (alignment-based) with Bowtie2↳ Could also: Pseudo-alignment tools such as Salmon or kallisto could also quantify transcript abundance — Salmon and kallisto are substantially faster and have been shown to produce similarly accurate TPM estimates; they also natively handle multi-mapping reads differently, which can matter for transcriptomes with many highly similar isoforms such as in E. gracilis
-
Cross-study expression values were normalised by log2 transformation followed by per-sample Z-scoring↳ Could also: Quantile normalisation, variance-stabilising transformation (VST from DESeq2), or between-sample normalisation via TMM (edgeR) could also be applied — These methods explicitly model count-based noise and can be more robust to differences in library size and composition across heterogeneous studies; VST in particular is commonly used before clustering and PCA of RNA-Seq data
-
Ten separate PCAs were each computed on 1000 randomly sampled transcripts for tetranucleotide frequency analysis↳ Could also: A single PCA (or UMAP/t-SNE) on all transcripts could also be performed; alternatively, random subsampling results could be aggregated (e.g. consensus PCA) — Using all transcripts in a single dimensionality reduction avoids sampling variability across runs; UMAP or t-SNE can reveal non-linear structure not captured by PCA, which may be informative for taxonomically heterogeneous datasets
-
Transcriptome completeness was assessed with BUSCO using Eukaryota and Protists_ensembl gene sets↳ Could also: rnaQUAST could also provide reference-based and reference-free transcriptome assembly quality metrics including sensitivity and precision against a known gene set — rnaQUAST reports additional metrics (e.g. matched bases, misassembled transcripts, unannotated regions) that complement BUSCO's single-copy orthologue approach and can distinguish completeness from correctness
-
Decontamination relied on mapping reads to contaminant reference genomes and removing mapped reads before reassembly↳ Could also: Tools such as Kraken2 or Centrifuge applied directly to reads, or Blobtools applied to assembled contigs, could also identify and filter contaminant sequences — Read-level classifiers can detect contamination from organisms for which no reference genome is available for subtraction; Blobtools uses GC content, coverage, and BLAST taxonomy jointly, similar in spirit to the approach used here but in an integrated framework
-
Batch effects across the five independent experiments were addressed by a tool not identified in the available text (sentence truncated)↳ Could also: ComBat-seq (for count-level batch correction) or limma::removeBatchEffect (for log-normalised data) are standard alternatives for multi-study RNA-Seq integration — ComBat-seq preserves the count nature of RNA-Seq data and is recommended when downstream analyses require count-based models; limma's approach is appropriate for continuous (log-transformed) expression matrices and allows covariate modelling alongside batch
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34072576
Paper: Cordoba et al. 2021, De Novo Transcriptome Meta-Assembly of the Mixotrophic Freshwater Microalga Euglena gracilis. Genes 12(6):842. PMID 34072576 / PMC8227486 / doi:10.3390/genes12060842.
Repo: https://github.com/microalgues/clustering (HEAD commit 6d390f9a3ec958053532377794bf6b498195ed8a, 2019-11-07)
Deposited assembly (authors' own): ENA/INSDC TSA HBDM01 (HBDM01000001–HBDM01091040), project PRJEB38787, sample SAMEA7015342, first public 2020-08-25.
New raw reads (authors' own): ENA PRJEB38787 (4 runs ERR4227585–88).
Re-used public reads: PRJNA310762 (A), PRJEB10085 (B), PRJNA298469 (C), PRJNA289402 (D).
NOTE: the BRIEF's listed data accession
sra:PRJEB4713is a mis-tag — PRJEB4713 is an old 454 single-end study of E. gracilis + E. mutabilis (10 tiny runs), NOT used by this 2021 Illumina meta-assembly. The real inputs are the five projects above + the HBDM01 deposit.
Pipeline stages described in the paper
| Stage | Tool(s) (paper) | Output | Deterministic? |
|---|---|---|---|
| QC/trim | Trimmomatic v0.32 | clean reads | yes (given params) |
| Per-experiment assembly | Trinity v2.4.0 ×4 param sets ×5 experiments | 20 raw assemblies | NO (Trinity is stochastic / multi-threaded non-deterministic) |
| Consolidation | EvidentialGene tr2aacds.pl v2016.07.11 | non-redundant set | partly |
| Decontamination | Bowtie2 v2.2.6 vs contaminant genomes | filtered set | yes |
| Consensus/redundancy | evgmrna2tsa2.pl + CD-HIT v4.6.8 (90% aa) + BLASTN v2.2.28 | HBDM01 final consensus (91,040 tr / 49,822 genes) | yes |
| Assembly statistics | (length/GC/N50 over final FASTA) | Table 2/3 numbers | yes — deposited as HBDM01 |
| Completeness | BUSCO (eukaryota) | Fig 2 (84.8% complete) | yes (given lineage/version) |
| ORF/protein prediction | EvidentialGene ORFs; GeneMarkS-T | 62,287 ORFs; 58,542 CDS | yes |
| Expression clustering | repo cluster/clusterization.R: PAM/k-medoids on 1−Pearson² distance |
Fig 4/5 (36 clusters of 2,500 var. genes) | yes but input expression matrix NOT shipped |
| Annotation/taxonomy | BLAST/Krona, GO/KEGG | Fig 3, Table 6 | heavy, external DBs |
IN SCOPE — attempted (deterministic, derivable from the deposited final assembly)
- Table 3 (HBDM01 column) assembly statistics, recomputed directly from the deposited HBDM01 FASTA: transcript count, total assembled size, N50, mean length, #>1,000 nt, #>10,000 nt, GC%, and the non-redundant gene count (unique EvidentialGene loci). DONE.
- Dataset N / read-budget claims cross-checked against ENA metadata + the assembly's authoritative run cross-references: 23 samples, 5 experiments, ~2.6 billion reads, ~70% from the new experiment. DONE.
IN SCOPE — not completed (honest blockers)
- BUSCO completeness (84.8%) and ORF/CDS counts (62,287 / 58,542) — deterministic and feasible in principle, but require writing the assembly + running BUSCO/TransDecoder/GeneMarkS-T on «our HPC». BLOCKED: the shared «infra» user quota (account «user») is fully exhausted — not a single byte can be written to /«infra» or /home. A central janitor must reclaim space. The streamed assembly-stat computations above succeeded only because they write nothing to disk.
- Expression clustering (36 clusters / Fig 4–5) — the repo's
clusterization.R(PAM k-medoids on 1−Pearson²) is fully readable, but the input expression matrix (2,500 most-variable genes × 23 samples) is NOT committed to the repo and not deposited; regenerating it requires the full read-mapping + count pipeline. Not reproducible from shipped artifacts.
OUT OF SCOPE (not a clean 1:1 target)
- Full de-novo meta-assembly from the 2.6-billion-read input (20 Trinity runs + EvidentialGene + Bowtie2 decontam). Trinity is non-deterministic, so even a successful re-run would not yield a bit-identical 91,040-transcript s
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.