RNAseq analysis of the parasitic nematode Strongyloides stercoralis reveals divergent regulation of canonical dauer pathways.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the RNA-seq FPKM-quantification pipeline (SeqPrep -> TopHat2 -> samtools -> Cufflinks de novo) for PMID 23145190 across 18 of 21 E-MTAB-1164 samples (3 align/cufflinks jobs still running at report time, not self-timed-out). Extracted concrete, testable Results-section claims for all 19 canonical dauer-pathway genes directly from the published article, and cross-checked them against both this reproduction's own FPKM values and the paper's own DataS10 supplementary FPKM table (recovered after fixing a Unicode curly-quote matching bug). Of 15 graded claims: 2 exact, 6 within-tolerance, 6 partial, 1 clear mismatch (Ss-tgh-4, where this reproduction's Cufflinks assembly spuriously called a low FPKM in one PP_L1 replicate that is absent from the paper's own data). Also recovered and read TextS1 (a legacy MS Word .doc mislabeled as .pdf, extracted via strings -e l) confirming RNA-isolation methods for all 7 conditions. E-MTAB-1164 raw data delivery is complete (21/21 fastq pairs, ~98% alignment rates); a documented, previously-identified gap remains in primary trim-QC stats for 3/21 samples (P_Female gerbil replicates) due to an unrecoverable original script, though post-alignment flagstat QC exists for 2 of those 3. No p-values or formal statistical tests were recomputed in this reproduction -- claims graded on FPKM direction/magnitude only.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors hypothesized that homologs of the four C. elegans dauer-regulatory pathways (cGMP signaling, insulin/IGF-1-like signaling, dauer TGFβ signaling, and dafachronic acid biosynthesis/DAF-12 nuclear hormone receptor) are present in the parasitic nematode Strongyloides stercoralis, show similar developmental regulation, and are involved in arrest and activation of the infective third-stage larva (L3i). This tests the long-standing "dauer hypothesis" that dauer and L3i are governed by conserved molecular mechanisms.
- ★ S. stercoralis possesses homologs of nearly all C. elegans dauer genes, but with significant differences in protein structure, developmental regulation, and gene family expansion. finding
- ★ Genes encoding cGMP signaling pathway components are coordinately up-regulated in L3i, consistent with a role in L3i regulation. finding
- ★ S. stercoralis has a paucity of genes encoding insulin-like peptide (ILP) ligands relative to C. elegans, and several of these have abundance profiles suggesting involvement in L3i development. finding
- ★ Seven S. stercoralis genes encode homologs of the single C. elegans dauer-regulatory TGFβ ligand Ce-DAF-7, three of which are expressed only in L3i; dauer-like TGFβ signaling is regulated oppositely to C. elegans yet may play a unique role in L3i development. finding
- ★ Putative dafachronic acid (DA) biosynthetic genes are not coordinately regulated during L3i development, unlike in C. elegans dauer formation. finding
- ★ Deep sequencing of the polyadenylated transcriptome across seven developmental stages, combined with genomic-contig alignment and de novo transcript assembly, enables identification and temporal profiling of dauer-pathway homologs previously missing from the small S. stercoralis EST database. method
- The study generates a resource of manually annotated S. stercoralis transcripts, predicted protein sequences, and stage-specific de novo assemblies (ArrayExpress E-MTAB-1164 and E-MTAB-1184). resource
- Understanding the mechanisms governing L3i development may lead to novel chemotherapeutic treatments and environmental control strategies for strongyloidiasis and other parasitic nematode diseases. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Polyadenylated RNAseq (deep sequencing of poly-A transcriptome), 100 bp paired-end | Strongyloides stercoralis PV001 line, seven developmental stages (including post-parasitic L1, post-free-living L1, free-living females, L3i, activated L3+, parasitic females); 21 libraries | none (developmental stage comparison) | Transcript abundance as FPKM (fragments per kilobase of exon per million mapped reads) for coding sequences of manually annotated genes | Illumina HiSeq 2000; TruSeq RNA Sample Preparation Kit; CASAVA v1.8.2; TopHat v1.4.1 with Bowtie v0.12.7 and SAMtools v0.1.18; Cufflinks v2.0.0 |
| De novo transcriptome assembly | S. stercoralis reads from the highest-read sample of each developmental stage | none | Stage-tagged assembled transcripts used as BLAST search space for C. elegans dauer gene homologs | SeqPrep (quality cutoff 35, min merged length 100 bp, no mismatches), FASTX toolkit quality trimmer, Trinity release 2012-04-27 with jellyfish k-mer counting, Geneious v5.5.6 |
| Comparative genomics / BLAST homology search and reverse BLAST validation | S. stercoralis (6 December 2011 draft) and S. ratti genomic contigs; C. elegans protein sequences from WormBase | none | Identification and manual annotation of putative S. stercoralis homologs of C. elegans dauer genes | Geneious (least restrictive parameters), NCBI pBLAST, Integrated Genome Viewer (IGV) v2.0.34 |
| Motif-based search for insulin-like peptide (ILP) ligands in six-frame translations | S. stercoralis and S. ratti genomes plus S. stercoralis de novo assembled transcripts | none | Presence of conserved ILP B peptide motifs (C-11X-C, CPPG-11X-C) and A peptide motifs (C-12X-CC, C-13X-CC, C-14X-CC, CC-3X-C-8X-CC, CC-4X-C-8X-CC, CC-3X-C-8X-C, CC-3X-C-9X-C) | Geneious |
| Protein multiple sequence alignment and neighbor-joining phylogenetic analysis with bootstrapping | S. stercoralis, C. elegans, phylum Nematoda and other Animalia protein sequences (guanylyl cyclases/DAF-11, TGFβ superfamily ligand domains, SMADs, DHS-16-related short-chain dehydrogenases, DAF-9-related cytochrome P450s) | none | Homology assignment / tree topology with 100 bootstrap iterations | Clustal W (BLOSUM matrix) and MUSCLE in Geneious |
| Parasite culture, stage isolation, and total RNA extraction with quality control | S. stercoralis PV001 maintained in prednisolone-treated beagles; free-living stages isolated by migration through agarose into BU buffer | none | Total RNA quantity and RNA integrity number (RIN) | TRIzol reagent (Life Technologies); Bioanalyzer 2100 (Agilent) |
| Experimental infection to obtain activated and parasitic stages | Mongolian gerbils (permissive host) infected with S. stercoralis | in vivo infection (host activation of L3i) | Recovery of activated third-stage larvae (L3+, confirmed by morphological change and resumption of feeding) and parasitic females | — |
| Quantitative PCR library quantification and fragment-size QC | 21 adapter-ligated dsDNA S. stercoralis RNAseq libraries | none | Library molar concentration from a Kapa standard calibration curve; fragment size distribution | Kapa SYBR Fast qPCR Kit for Library Quantification (Kapa Biosystems); Bioanalyzer 2100 High Sensitivity DNA Assay (Agilent) |
- ▲ Genes encoding cGMP signaling pathway components were coordinately up-regulated in L3i
- – S. stercoralis has few genes encoding insulin/IGF-1-like signaling (ILP) ligands compared with C. elegans, several with abundance profiles suggesting involvement in L3i development
- – Seven S. stercoralis genes encode homologs of the single C. elegans dauer TGFβ ligand; three are expressed only in L3i 7 genes; 3 L3i-exclusive
- – Putative dafachronic acid biosynthetic genes were not coordinately regulated during L3i development
- – S. stercoralis dauer-like TGFβ signaling is regulated oppositely to C. elegans, consistent with prior finding that Ss-tgh-1 is transcriptionally regulated opposite to Ce-daf-7
- – Homologs of nearly all C. elegans dauer genes were identified in S. stercoralis, with differences in protein structure, developmental regulation, and gene-family expansion
- – Over 2.3 billion paired-end reads were generated across seven developmental stages, enabling construction of developmental expression profiles over 2.3 billion paired-end reads
- count over 2.3 billion paired-end reads (Total RNAseq reads generated across seven S. stercoralis developmental stages)
- count 21 libraries / 21 samples (Number of sequencing libraries constructed (replicates across the seven developmental stages))
- count seven S. stercoralis genes encoding homologs of the single C. elegans dauer TGFβ ligand; three expressed only in L3i (Expansion of the daf-7-like TGFβ ligand family in S. stercoralis)
- count over 30 genes (C. elegans dauer formation (daf) genes identified by mutant screens)
- count over one billion people (Global burden of parasitic nematode infection)
- count 30–100 million people (Global number of people infected with S. stercoralis)
- mean approximately 170±50 (standard deviation) bp (Fragment size of polyadenylated RNA after fragmentation at 94°C for eight minutes)
- other RNA integrity number (RIN) greater than 8.0 (Quality threshold for total RNA samples used in library construction)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a descriptive RNAseq profiling study rather than a hypothesis-testing study in the classic sense: the authors sequenced polyadenylated RNA from seven S. stercoralis developmental stages (21 libraries), aligned reads with TopHat/Bowtie, assembled transcripts de novo with Trinity, quantified transcript abundance per gene as FPKM using Cufflinks, and identified/verified homologs of C. elegans dauer-pathway genes via BLAST and phylogenetic analysis (Clustal W/MUSCLE alignments, neighbor-joining trees with 100 bootstrap iterations). The provided text is truncated just as the differential-analysis/results section begins ("FPKM values for ent..."), so any formal statistical hypothesis tests, p-values, or multiplicity corrections applied to expression differences are not visible in the excerpt supplied.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Neighbor-joining phylogenetic tree construction with bootstrapping | Resolving homology among S. stercoralis, C. elegans, and other nematode/Animalia protein sequences (e.g., guanylyl cyclases, TGFβ ligands, SMADs, short-chain dehydrogenases, cytochrome P450 proteins) | 100 bootstrap iterations | not stated |
| FPKM transcript abundance estimation (Cufflinks) | Quantifying gene expression across the 21 samples spanning 7 developmental stages | 21 samples (7 stages, apparently ~3 libraries each), though the text is truncated before the differential comparison itself is described | not stated |
-
Transcript abundance was summarized as FPKM per gene per sample using Cufflinks, and the excerpt does not show a described formal statistical test (e.g., a negative-binomial model with hypothesis testing) applied to compare stages.↳ Could also: Count-based differential expression tools such as DESeq2 or edgeR, which model read counts with a negative binomial distribution and provide Wald or likelihood-ratio tests — These approaches directly estimate statistical significance and fold-change confidence for expression differences between developmental stages, which can complement descriptive FPKM profiling with formal uncertainty quantification.
-
Homology and pathway relationships were established using BLAST search plus neighbor-joining phylogenetic trees with 100 bootstrap replicates.↳ Could also: Maximum-likelihood (e.g., RAxML, PhyML) or Bayesian (e.g., MrBayes, BEAST) phylogenetic inference with bootstrap or posterior-probability support — These methods can offer different assumptions about substitution models and branch-length estimation, which some readers use alongside neighbor-joining to cross-check topology support, particularly for more divergent sequences.
-
Transcript abundance was reported as FPKM (fragments per kilobase of exon per million mapped reads).↳ Could also: TPM (transcripts per million) normalization — TPM is sometimes preferred because it is more directly comparable across samples/libraries with differing composition, since it normalizes for library size after length normalization rather than before.
-
RNA quality was screened using a fixed RIN cutoff (>8.0) rather than reporting a distribution or statistical summary of RIN values across all samples.↳ Could also: Reporting the RIN distribution (e.g., mean ± SD or range) for all included samples — Providing the full distribution alongside the cutoff can give readers additional context on sample quality consistency across the seven developmental stages.
-
Library concentration was estimated via qPCR using a calibration curve from three dilutions of Kapa standards.↳ Could also: Including technical replicate qPCR measurements with reported variability (e.g., SD or CV of Ct values) — Reporting measurement variability for the quantification step can help convey the precision of input library concentrations used for pooling.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Data identity is clean - all 21 E-MTAB-1164 FASTQ pairs were retrieved 1:1 - and the paper's own DataS10 FPKM table independently derives every main-text claim, so nothing here points at the authors' data or numbers. The deviations are on our side: an incomplete run (18/21 samples processed; L3i at 2/3 replicates) reversed the Ss-ilp-4 L3i-peak ordering (L3_plus=111.9 > L3i=79.3), and de novo Cufflinks with coordinate-overlap gene attribution produced a spurious Ss-tgh-4 FPKM=5.92 against the paper's own <=0.32 ceiling. A separate, milder issue is on the authors' wording: 'one log/10x' for Ss-ilp-6 is ~4.5x in their own DataS10, and 'exclusively L3i'/'not expressed' understate real low-level background (Ss-tgh-3 PP_L3=0.75-1.29; Ss-tgh-7 P_Female=0.58-1.19) - overstatement in prose, not in data. Overall a solid partial reproduction: the central divergent-dauer-pathway conclusion (Ss-daf-1/daf-4 peaking in L3i/L3+, Ss-ilp-1 down in L3+/P_Female) holds, with no p-values recomputed and all discrepancies explainable.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.