Transcriptome profiling of Giardia intestinalis using strand-specific RNA-seq.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the PMID 23555231 strand-specific RNA-seq pipeline (SeqPrep -> bowtie2 -> samtools -> cufflinks) for all 5 GSE36490 strain datasets (WB, GS, P15, AS175P33, AS175P4). The originally-submitted jobs (from before this session) had all crashed with a segfault due to an apptainer bind-mount path-resolution bug affecting every strain identically; this was root-caused, fixed in all 5 scripts, and all 5 were resubmitted and completed successfully. Mapping rates for all 5 strains are within ~0.5-1.5 percentage points of the paper's reported values (within-tol). The FPKM>=0.5 gene-transcription-fraction claim reproduces closely for GS (exact) and P15 (within-tol), but shows a real ~8-point gap for WB (mismatch, likely a reference-annotation version/size difference: 7444 gene models here vs. presumably fewer in the original study). The AS175 version of this claim could not be directly reproduced because AS175 has no curated genome-wide gene model file (only an ORF-level GFF), forcing de novo cufflinks assembly whose near-100% 'transcribed' rate is a methodological artifact rather than a comparable statistic (graded partial). A QC gap was also found in the input data: only 4 of 20 downloaded fastq.gz files (the WB strain's 2 lanes) have recorded checksums; the other 8 files' integrity rests only on plausible file size and successful downstream alignment, not an explicit checksum match.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study asks what the polyadenylated transcriptome of Giardia intestinalis looks like at single-nucleotide resolution and how gene expression has diverged between genetically distinct assemblages (A, B, E), testing whether transcriptome differences can help explain assemblage-associated host preference and symptoms.
- ★ Most of the G. intestinalis genome is transcribed in in vitro-grown trophozoites, but at vastly different expression levels. finding
- ★ Gene expression divergence between isolates recapitulates the known phylogeny and reveals lineage-specific expression differences. finding
- ★ Strand-specific paired-end RNA-seq of four isolates from three assemblages provides the first deep, comparative gene expression profiling in Giardia and refines/corrects existing genome annotations. resource
- ★ Polyadenylation sites were globally mapped for over 70% of genes, giving the first genome-wide view of 3' UTRs in this parasite, including conserved and unexpectedly long 3' UTRs. finding
- ★ A cluster of 28 open reading frames on chromosome 5 of the WB isolate is not transcribed, likely representing a silenced genomic region. finding
- ★ Allele-specific expression analysis shows a correlation between allele dosage and allele expression in the GS isolate, the first genome-wide evidence of allelic-variant transcription in Giardia. finding
- ★ cis-splicing is very limited: previously reported cis-splicing events were confirmed and global mapping identified only one novel intron. finding
- PolyA tags (≥7 adenines) extracted from the left mate of read-pairs, clustered within 10 nt and requiring ≥4 tags, allow transcript 3'-end mapping from standard RNA-seq data. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Strand-specific paired-end RNA-seq (polyA-selected mRNA) | G. intestinalis trophozoites of four isolates: WB (assemblage AI), AS175 (AII), P15 (E), GS (B), grown in vitro in TYDK | none | Digital gene expression as FPKM (threshold 0.5 FPKM for transcription); transcript structure/annotation refinement | Illumina HiSeq 2000, 2×100-nt paired-end; ScriptSeq mRNA-Seq library prep kit (Epicentre SS10906); cBot cluster generation; Poly(A) Purist MAG (Ambion AM1922); TRIzol RNA extraction |
| Technical and biological variation RNA-seq replication | Same libraries split across two sequencing lanes (technical); AS175 sampled after 4 and 33 in vitro passages (biological) | in vitro passage number (P4 vs P33) | FPKM concordance between lanes; differentially expressed genes by χ² test of fold change versus technical-replicate fold changes | Illumina HiSeq 2000 |
| Reverse transcription quantitative PCR (RT-qPCR) validation | G. intestinalis WB trophozoites (four 10 ml pooled cultures) | none | Expression levels of 49 WB transcripts, triplicate reactions; minus-template and minus-RT controls | Thermo Scientific Maxima SYBR Green/ROX qPCR master mix; RevertAid H Minus First Strand cDNA Synthesis Kit; DNase-I (Fermentas) |
| 3'-Rapid Amplification of cDNA Ends (3' RACE) | G. intestinalis WB total RNA (10 µg) | none | Experimental validation of polyadenylation sites / transcript 3' ends | ExactSTART Eukaryotic mRNA 5'- & 3'-RACE Kit (Epicentre ES80910); APex phosphatase, tobacco acid pyrophosphatase, T4 RNA ligase, MMLV RT; FailSafe PCR Premix E/enzyme mix (Epicentre FSE51100) |
| Computational polyA-tag mapping of transcript 3' ends | RNA-seq read data from the four G. intestinalis isolates mapped to WB, GS, P15 genomes | none | PolyA site clusters (≥4 tags, clustered within 10 nt) assigned to ORFs; most frequent site taken as the 3' end; 3' UTR lengths | cutadapt v.0.9.5; bowtie v.2.0.0-beta6 |
| Whole-genome sequencing | G. intestinalis AS175 isolate (assemblage AII), recently recovered from a human individual | none | Genome sequence/assembly (ENA accession CAHQ00000000) used as reference for AS175 transcriptome | — |
| Oligonucleotide microarray data reanalysis (cross-platform comparison) | G. intestinalis WB; raw array data provided by the A. Hehl lab | none | Probe-level expression assigned to genes (probes fully contained in ORFs; genes <500 bp excluded) compared with RNA-seq FPKM | R limma and multtest; bowtie v.2.0.0-beta6 for probe alignment |
| SAGE data comparison | G. intestinalis curated SAGE tag data from GiardiaDB | none | SAGE tag counts compared with RNA-seq expression estimates | GiardiaDB (v.2.3) |
- – Most of the genome was transcribed in in vitro trophozoites, though transcript levels spanned a very wide range.
- – Gene expression divergence recapitulated the known phylogeny of the isolates, with lineage-specific expression differences uncovered.
- – Polyadenylation sites were mapped for over 70% of genes, revealing conserved and unexpectedly long 3' UTRs. >70% of genes
- – A non-transcribed gene cluster containing 28 open reading frames was found on chromosome 5 of WB. 28 ORFs
- ▲ Allele expression correlated with allele dosage in the GS isolate.
- – Global mapping of cis-splicing confirmed previously reported events and identified only one novel intron. 1 novel intron
- – RNA-seq confirmed many existing gene annotations and was used to refine and correct gene models.
- – Adapter sequences were detected in 45–61% of read-pairs, indicating many sequenced DNA fragments were <200 nt; these mates were joined and mapped as single-end reads. 45–61% of read-pairs
- count 28 open reading frames (Non-transcribed gene cluster on chromosome 5 of the WB isolate)
- other >70% of genes (Fraction of genes for which polyadenylation sites were mapped)
- other ~77% average sequence identity (Comparative genomics between assemblage A and B (background))
- other ~87% average sequence identity (Comparative genomics between assemblage A and E (background))
- count ~5000 genes per genome; haploid genome ~12 Mb across five chromosomes; ~86% protein-coding (Genome features of the three sequenced isolates WB, GS, P15 (background))
- count 49 transcripts (Number of WB transcripts validated by real-time qPCR, each in triplicate)
- other 45 to 61% (Proportion of read-pairs containing library adapter sequence)
- other 0.5 FPKM (Threshold below or equal to which genes were not regarded as expressed)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study profiled polyadenylated transcriptomes of four Giardia intestinalis isolates (three assemblages) using strand-specific, paired-end RNA-seq, with digital gene expression quantified as FPKM via Cufflinks. Differential expression was assessed using a chi-squared test comparing observed fold-changes to the fold-change variation seen between technical replicates (same biological sample split across two sequencing lanes), with biological variation approximated using one isolate (AS175) sampled at two passage numbers. Pearson's correlation coefficient was used to compare expression distributions, and Mantel's test related gene-expression distance matrices to ortholog genetic-distance matrices; multiple-hypothesis correction used the Benjamini-Hochberg procedure. A subset of 49 transcripts was independently checked by triplicate qPCR.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Chi-squared (χ2) test of fold change | Identification of differentially expressed genes, comparing observed fold-changes to fold-change variation from technical replicates | — | not stated |
| Pearson's correlation coefficient (r) | Comparing distributions (e.g., expression values between replicates/samples) | — | not stated |
| Mantel's test (mantel.rtest, ade4 package) | Comparing the matrix of genetic (evolutionary) distances between orthologs to the matrix of gene-expression (Euclidean) distances | — | not stated |
-
Differentially expressed genes were identified using a chi-squared test on fold-changes benchmarked against technical-replicate variation, with only one isolate (AS175) used to approximate biological variability.↳ Could also: A negative-binomial-based RNA-seq DE tool (e.g., DESeq2 or edgeR) using independent biological replicates per condition — Such tools model the mean-variance relationship specific to sequencing count data and can produce replicate-aware dispersion estimates, which becomes especially informative when biological replicates are available for each isolate/condition being compared.
-
Biological variation was estimated from a single isolate (AS175) grown for different numbers of in vitro passages, rather than from independent replicate cultures of each genotype.↳ Could also: Including independent biological replicate cultures for each of the four isolates — Independent replicates per isolate would allow a per-condition variance estimate and support standard replicate-based differential expression frameworks.
-
Pearson's correlation coefficient was used to compare expression distributions between samples/replicates.↳ Could also: Spearman's rank correlation — Spearman's correlation is robust to non-normal or skewed distributions and outliers, which can be common in FPKM-based expression data, without assuming a linear relationship.
-
Mantel's test was used to relate the gene-expression distance matrix to the genetic-distance matrix across orthologs.↳ Could also: A partial Mantel test or distance-based matrix regression controlling for additional covariates — These approaches can account for other structuring variables (e.g., isolate-specific technical factors) when relating expression divergence to phylogenetic distance.
-
Multiple-hypothesis correction was applied using the Benjamini-Hochberg FDR procedure.↳ Could also: Storey's q-value method or a Bonferroni correction — The q-value approach also controls the false discovery rate but can offer different power trade-offs, while Bonferroni provides stricter family-wise error control when the number of comparisons is small.
-
Expression levels of 49 transcripts were validated by triplicate qPCR without a described formal statistical comparison to the RNA-seq FPKM values.↳ Could also: Reporting a correlation coefficient (Pearson or Spearman) or a Bland-Altman-style agreement analysis between qPCR and RNA-seq measurements — This would provide a quantitative measure of concordance between the two platforms in addition to triplicate confirmation.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
What deviates: 7 of 9 claims reproduce cleanly — all five bowtie2 mapping rates land within 0.5–1.5pp of the paper's rounded values, and GS (97.22% vs 97.3%) and P15 (97.05% vs 96.3%) transcription fractions essentially match. The two outliers are WB's transcribed fraction (85.76% vs 93.7%, ~7.9pp) and AS175 (99.72%/99.95% de novo loci vs 98.3% ORFs). Whose side: both sit on the reference-annotation/our-method side, not the authors' — WB's gap is a denominator effect (7444 present-day gene models vs the smaller 2013 ORF catalog), and AS175 had no genome-wide GTF at all, forcing a de novo cufflinks run whose near-100% rate is a methodological artifact. The paper's failure to pin the annotation build is an underspecification, but the values show no sign of being non-derivable or fabricated. Severity: moderate; the core claim that the overwhelming majority of Giardia genes are transcribed in all four isolates holds in every strain. Note also a data-QC gap on our side: only 4 of 20 fastq.gz files carry recorded checksums.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.