Corpus 1,283 assessed · 1,184 scored · 647 reproduced ≥75 · 173 flagged ·∅ 73.9/100
← New search

A lack of parasitic reduction in the obligate parasitic green alga Helicosporidium.

PLoS Genet · 2014
30/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
30/100
Reproducibility score
2.5 SD below mean
vs. all fields · 1184 studies
🎯 Scores higher than 2% of all assessed papers rank 1156 of 1184 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Dataset provenance for all three SRA runs (PRJNA188927) was confirmed with high confidence (SRR1019731/SRR1019732 DNA read-pair counts match the paper exactly; SRR1019733 RNA-seq is ~12% short of the paper's reported count and shows read lengths inconsistent with the paper's stated 100bp, suggesting an already-trimmed SRA deposit). The paper's designated code repo (sickle) was used exactly as documented for RNA-seq trimming and reproduced a plausible 98.5% read-retention rate (no exact paper number to compare against). The main computational result - de novo genome assembly with Ray 2.0.0rc8 on raw paired DNA reads at k=21/25/31, matching the paper's documented protocol and hardware spec exactly, with no manual coverage-cutoff override - did NOT reproduce: all three k values gave assemblies 1-2 orders of magnitude smaller and far more fragmented (e.g. 28,860 contigs / 4.7 Mbp at k=21 vs the paper's 11,717 contigs / 13.68 Mbp) than reported, consistently caused by Ray's automatic coverage-cutoff detector landing almost exactly on the histogram peak on a k-mer spectrum that has no distinct genomic peak to begin with. Gene prediction/annotation and all downstream/wet-lab results were not attempted (out of scope: separate heavy third-party pipeline / non-computational). This is reported as a genuine, well-diagnosed mismatch, not a self-imposed time limit or infrastructure failure.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-07
Rubric version
not recorded
Assessed by
Last updated
2026-08-07

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Does the independent transition from a free-living autotrophic green alga to an obligate intracellular animal parasite in Helicosporidium produce the same kind of reductive genome/metabolic evolution seen in apicomplexans such as Plasmodium? The authors sequenced the Helicosporidium genome and transcriptome to test whether its parasitic lifestyle is accompanied by loss of functions associated with host-dependence.

Core claims
  • The Helicosporidium nuclear genome is small and compact (~2.5-fold smaller than Chlorella and Coccomyxa) yet shows almost no evidence of functional/metabolic reduction. finding
  • Gene loss is concentrated in photosynthesis-related pathways (light-harvesting complexes, photosystems I and II, chlorophyll and carotenoid biosynthesis), but photosynthetic reduction is incomplete: the Calvin/carbon fixation pathway is nearly complete except for RuBisCO (rbcL/rbcS) and pyruvate orthophosphate dikinase (ppdK). finding
  • The dominant reductive force is contraction of gene family complexity rather than loss of whole functional categories, and the contracted families relate mostly to genome maintenance and expression (chromosome packing, transcription, translation, post-translational modification, protein turnover), not to host-dependence. finding
  • Transporter gene families are reduced rather than expanded in Helicosporidium, contrary to expectation for a parasite with growing host dependence. finding
  • Chitinase gene families have expanded: 14 GH18 chitinase genes, with 12 copies arising from recent duplication, likely for digesting the chitinous barriers of the insect host and/or remodelling the Helicosporidium cell wall. mechanism
  • The smaller genome size is attributable to greater genome compaction — fewer and smaller introns, smaller exons/intergenic regions, high coding density — rather than to metabolic gene loss. finding
  • Helicosporidium lacks endogenous RNA interference machinery (no Dicer/Argonaute), like Ostreococcus tauri and O. lucimarinus, whereas Chlorella and Coccomyxa have single copies and Chlamydomonas has three paralogs of DC1 and AGO1. finding
  • The draft genome and transcriptome of Helicosporidium sp. ATCC50920 constitute a new genomic resource representing an early stage of the autotroph-to-parasite transition. resource
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome shotgun sequencing and de novo assembly Helicosporidium sp. ATCC50920 (parasite of the black fly Simulium jonesi) none Contig number, assembly size, GC content, N50, sequencing coverage, estimated genome size Illumina
Transcriptome sequencing Helicosporidium sp. ATCC50920 none Transcript contigs; fraction of metabolic-pathway genes detected and mapping rate onto genomic contigs
Gene prediction and genome annotation / structural comparison Helicosporidium genome vs. green algal genomes (Chlorella variabilis NC64A, Coccomyxa subellipsoidea C-169, Ostreococcus tauri, O. lucimarinus, Micromonas spp., Chlamydomonas reinhardtii, Volvox carteri) none Gene number, gene (coding) density, exons/gene, average exon and intron size, chromosome number, GC%
Comparative pathway/metabolic reconstruction (KEGG orthology mapping and GreenCut2 database comparison) Helicosporidium vs. Coccomyxa and Chlorella predicted proteomes none Presence/absence of genes per plastid and metabolic pathway; percentage of GreenCut2 plastid-targeted proteins retained KEGG (ko identifiers); GreenCut2 database
Synteny / gene order conservation analysis Ten largest Helicosporidium contigs vs. Chlorella scaffolds none Percentage of genes arrayed in syntenic clusters
Evolutionary gene network analysis of gene family complexity Helicosporidium, Chlorella and Coccomyxa predicted proteomes none Connected components with lower representation in Helicosporidium, manually curated into functional categories
Glycosyl hydrolase / chitinase gene family identification (PFAM-based annotation) Helicosporidium genome and transcriptome; compared to Chlorella and Coccomyxa none Number of GH18/GH19 chitinase genes and inferred duplication history PFAM
Contamination filtering of assembled contigs (with removal of mitochondrial and plastid genomes) Helicosporidium sequence assembly none Estimated maximum percentage of contaminating sequence and its contig-size distribution
Key results
  • Illumina reads assembled into 11,717 contigs totalling 13,684,556 bp (62.2% GC); after removal of organellar genomes and small contigs, 5,666 contigs ≥500 bp totalling 12,373,820 bp (N50 3,036 bp, 61.7% GC) at 62× average coverage; genome estimated at a maximum of 17±0.5 Mbp, consistent with a prior 13 Mbp CHEF karyotype estimate. 12.4 Mbp assembled; 17±0.5 Mbp estimated
  • 6,035 protein-coding genes predicted, with high coding density (0.487 gene/kb) versus Coccomyxa (0.197) and Chlorella (0.212), but lower than Ostreococcus (0.626 and 0.580); average 2.3 exons/gene (366 bp/exon) and 1.3 introns/gene (168 bp/intron). 0.487 vs 0.197-0.212 gene/kb (~2.3-2.5x higher density)
  • Helicosporidium encodes only 56% of GreenCut2 plastid-targeted proteins, whereas photosynthetic Coccomyxa and Chlorella each encode 96%; losses are non-randomly concentrated on light-harvesting processes. 56% vs 96%
  • Chlorophyll and carotenoid biosynthesis branches are lost while the heme branch of the tetrapyrrole pathway is complete; photosystems I and II and light-harvesting antenna proteins are entirely absent, yet some electron transport proteins, F-type ATPase and cytochrome b6f components, and a nearly complete carbon fixation pathway (lacking rbcL/rbcS and ppdK) are retained.
  • 100 connected components (excluding photosynthesis-related products) showed lower representation in Helicosporidium than in its free-living relatives and were manually curated into 9 functional categories, dominated by chromosome packing, transcription, translation, post-translational modification and protein turnover; transporters were the most reduced functional class. 100 connected components across 9 functional categories
  • 14 glycosyl hydrolase genes, all GH18 chitinases, were identified in both the genome and transcriptome; Chlorella has only two GH18 and one virus-derived GH19 chitinase, and the extra 12 Helicosporidium copies appear to derive from recent duplications. 14 vs 2 GH18 chitinases (12 extra copies)
  • Only 30% of genes in the ten largest Helicosporidium contigs lie in syntenic clusters with Chlorella, with no apparent metabolic relationship among clustered genes. 30%
  • 95.4% of genes assigned to known metabolic pathways were present in both genome and transcriptome and only 3.6% were transcriptome-exclusive; 92.3% (≥1000 bp) and 95.9% (≥1500 bp) of transcriptomic contigs mapped to genomic contigs, indicating the draft genome captures the coding potential well. 95.4% shared; 3.6% transcriptome-only
Key statistics
  • count 11,717 contigs totalling 13,684,556 bp (62.2% GC); filtered set 5,666 contigs, 12,373,820 bp, N50 3,036 bp, 61.7% GC, 62× coverage (Raw and filtered Helicosporidium genome assembly)
  • other 17±0.5 Mbp maximum estimated genome size (vs. 13 Mbp by CHEF karyotype) (Genome size estimate)
  • count 6,035 protein-encoding genes; 2.3 exons/gene (366 bp/exon); 1.3 introns/gene (168 bp/intron) (Gene prediction over 12.4 Mbp assembly)
  • fold_change 2.5-fold smaller genome than Coccomyxa (49 Mbp) and Chlorella (46.2 Mbp) (Genome size comparison with trebouxiophyte relatives)
  • other 56% of GreenCut2 plastid-targeted proteins retained vs 96% in Coccomyxa and Chlorella (Plastid proteome reduction)
  • other Gene density 0.487 gene/kb (Helicosporidium) vs 0.197 (Coccomyxa), 0.212 (Chlorella), 0.626/0.580 (Ostreococcus) (Coding density across green algal genomes)
  • count 14 GH18 chitinase genes in Helicosporidium; 12 from recent duplication; Chlorella has 2 GH18 + 1 GH19 (Chitinase gene family expansion)
  • other 30% of genes in the ten largest contigs syntenic with Chlorella; 95.4% pathway-gene overlap between genome and transcriptome; 92.3%/95.9% transcript contig mapping (≥1000/≥1500 bp) (Synteny and genome/transcriptome concordance)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a comparative genomics study reporting a draft genome assembly and annotation of Helicosporidium, with descriptive comparisons of genome size, gene content, gene density, gene family complexity, and metabolic pathway completeness against related green algal genomes (Chlorella, Coccomyxa, Chlamydomonas, Ostreococcus, Micromonas, Volvox). Results are reported primarily as counts, percentages, and an estimated genome size with an associated range, using genome assembly metrics, synteny/gene-order comparison, GreenCut2 pathway-presence comparison, and an evolutionary gene network analysis to identify gene family expansions/contractions, rather than classical inferential hypothesis testing.

Replicationunclear Sample sizeA genome size estimate is given as 17±0.5 Mbp based on the assembly data, and average coverage (62x) is reported, but no sample size or replicate count for a statistical comparison is described GroupsHelicosporidium genome/gene content vs. other green algal genomes (Chlorella, Coccomyxa, Chlamydomonas, Ostreococcus spp., Micromonas spp., Volvox) Pairingna Randomization/blindingna Dispersionunclear Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Differences in gene family size/complexity between Helicosporidium and its relatives (e.g., the 100 connected components with reduced representation, or the chitinase family expansion) are described qualitatively/by count rather than with a formal statistical test
    Could also: A formal gene family expansion/contraction analysis (e.g., CAFE, or a hypergeometric/Fisher's exact test per family) comparing observed counts against an expected null based on phylogeny — This would let readers see whether a given expansion or contraction (such as the chitinase family) is statistically distinguishable from stochastic gene gain/loss along the tree, complementing the descriptive comparison already presented
  • Pathway completeness is summarized as percentages of genes present/absent per functional category (e.g., 56% vs 96% of GreenCut2 proteins retained)
    Could also: A proportion test (e.g., two-proportion z-test or Fisher's exact test) comparing retention rates between genomes, possibly with a confidence interval on the difference — This would provide a formal measure of whether the difference in retention percentage between Helicosporidium and its photosynthetic relatives exceeds what might be expected by chance, alongside the percentages already reported
  • The estimated genome size is reported as a point estimate with an associated range (17±0.5 Mbp) without specifying whether this reflects a standard deviation, standard error, or another measure
    Could also: Explicitly labeling the dispersion measure (e.g., SD, SE, or a 95% confidence interval) around the genome size estimate — Naming the specific measure would let readers directly interpret the precision of the size estimate and compare it consistently with other genome size estimates in the literature
  • Synteny is reported as a single percentage (30% of genes in syntenic clusters) without a comparison to a null expectation
    Could also: A permutation-based or randomization test comparing observed synteny to a null distribution generated from randomized gene order — This would indicate whether the observed level of synteny conservation is greater than expected by chance given genome size and gene content, adding statistical context to the descriptive percentage
  • Coding density and gene/exon/intron size are compared across genomes as single summary values per species (Table 1) without variance estimates
    Could also: Reporting distributions (e.g., median with interquartile range, or mean with SD) for exon/intron sizes per genome, potentially compared via a non-parametric test such as Mann-Whitney U — This would convey the spread of exon/intron sizes within each genome, not just central tendency, which can be informative when comparing genome compaction across species

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

genome_assembly_ray
Reported
Raw Ray assembly: 11,717 contigs / 13,684,556 bp. Filtered (>=500bp): 5,666 contigs / 12,373,820 bp, N50 3,036 bp, GC 61.7%.
Reproduced
k=21: 28,860 contigs>=100nt / 4,716,834 bp (N50 165); 8 contigs>=500bp / 4,595bp. k=25: 11,094 contigs>=100nt / 2,376,046 bp (N50 225); 164 contigs>=500bp / 95,895bp. k=31: 44 contigs>=100nt / 12,744 bp (N50 370); 5 contigs>=500bp / 4,475bp.
did not match
rnaseq_sickle_trim
Reported
RNA-seq reads filtered with Sickle under default parameters (qualitative description only, no specific retained-read count reported).
Reproduced
72,029,262 / 73,143,800 SRR1019733 reads kept after sickle default-parameter single-end trimming (98.5% retention).
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.