Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genetic analysis of Leishmania donovani tropism using a naturally attenuated cutaneous strain.

PLoS Pathog · 2014
L1 63/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
63/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 24% of all assessed papers rank 875 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Resumed from a prior worker's checkpoint (align+slrnaseq done, variant/CNV/DE pending). Fixed a Java8(GATK3.8)/Java11+(SnpEff5.1) version conflict by installing an isolated openjdk17 for the SnpEff steps only. Completed full pipeline: GATK3.8 joint HaplotypeCaller -> hard-filter -> SnpEff annotate -> classify_variants.py (Table2/3) and somy_cnv.py (Table1); edgeR Table4 was already done. Table3 (5 specific variant positions) reproduces EXACTLY. Table1 CNV: 2/9 loci within-tolerance (chr16,chr23), 5/9 correct-direction-but-off-magnitude, 1/9 genuine mismatch (chr19, no depletion seen), 1/9 (chr22/A2) is the paper's OWN documented non-reproducible locus (manual curation due to ref misassembly) reported as reference only. Table2 SNP/indel counts: correct category structure but ~40-60% of reported magnitude (likely differing variant-calling stringency vs the paper's unspecified original pipeline). Table4 DE: 4/8 highlighted genes match direction+magnitude, 3/8 no signal in my pooled edgeR design vs paper's per-stage contrasts (documented design mismatch), 1/8 direction only. Not attempted: wet-lab/manual curation results (chr22 A2 locus exact values, any non-pipeline validation experiments) - out of scope per brief.

💻 Code ↗ 🗄 Data: GSE48475

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-01
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-01
no human curator yet
Last updated
2026-08-01

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Why do most Leishmania species cause cutaneous infection while others visceralize? The study tests whether closely related L. donovani (MON-37) clinical isolates from Sri Lankan cutaneous (CL-SL) and visceral (VL-SL) leishmaniasis patients differ genetically in ways that determine visceral organ tropism.

Core claims
  • The CL-SL L. donovani isolate is severely attenuated for survival in visceral organs (liver, spleen) of BALB/c mice compared with the VL-SL isolate, while inducing transient footpad swelling that VL-SL does not. finding
  • Gene deletions and CL-SL-specific pseudogenes are not responsible for the difference in disease tropism; no gene deletions were detected in either genome and only 5 genes are unequally affected by frame shift/stop-site change, with no informative homozygous CL-SL-specific pseudogene. finding
  • SNPs and/or gene copy number variations play a major role in the altered pathology/tropism of the CL-SL strain. mechanism
  • Expression of the VL-SL Rag C gene (LdBPK_366140.1/366149, ras-like small GTPase in the mTOR pathway) in the CL-SL isolate significantly increases parasite survival in the spleen, implicating the homozygous R231C substitution in CL-SL attenuation. finding
  • Decreased copy number of the A2 multi-gene family in the CL-SL strain contributes to its attenuation in visceral infection. mechanism
  • Comparing genomes of closely related isolates of the same species that cause different human pathologies is a more effective strategy for defining tropism determinants than cross-species comparison. method
  • The Sri Lankan CL-SL and VL-SL isolates are more closely related to each other than to the L. donovani reference strain BPK282A1 from Nepal, and both are MON-37 as confirmed by 6PGDH isoenzyme gene sequencing. finding
  • Whole-genome and SL-transcriptome datasets for paired cutaneous and visceral L. donovani clinical isolates, including a manually reconstructed A2/A2rel locus that is mis-assembled in the BPK282A1 reference. resource
Experimental setups
Assay System Perturbation Readout Platform
Experimental visceral infection with parasite burden quantification (limiting dilution of spleen homogenates; Leishman Donovan Units for liver) and spleen weight BALB/c mice (5 mice/group) Intravenous tail vein injection of 5×10^7 stationary phase promastigotes of CL-SL or VL-SL L. donovani Spleen weight (splenomegaly), liver parasite burden in LDU, spleen parasite numbers at 4 weeks post-infection
Experimental cutaneous infection with lesion measurement BALB/c mice (5 mice/group) Subcutaneous rear footpad injection of 5×10^6 stationary phase promastigotes of CL-SL or VL-SL Footpad swelling/lesion development monitored over 11 weeks; parasite numbers
Transgenic complementation followed by in vivo spleen infection CL-SL L. donovani promastigotes transfected with VL-SL genes; BALB/c mice Overexpression/transfection of 7 cloned VL-SL genes (LdBPK171200.1 steroid dehydrogenase, LdBPK220120 phosphoinositide phosphatase, LdBPK252290 DnaJ, LdBPK320370 hypothetical, LdBPK322160 rab-2a, LdBPK341900 NAD-dependent deacetylase, LdBPK366149 Rag C) into CL-SL Spleen infection level 4 weeks after intravenous infection
Whole genome sequencing (gDNA-seq) with alignment to reference, CNV/SNP/indel and chromosome somy analysis CL-SL and VL-SL L. donovani isolates; reference strain BPK282A1 (Nepal) none Read coverage in 10 kb tiles (VL/CL ratios), gene copy number, SNPs/indels and their coding effects, chromosome read depth scaled to 2 for disomic chromosomes Illumina GAIIx, >200× coverage
Spliced-leader (SL) cDNA sequencing / transcriptome profiling L. donovani CL-SL and VL-SL axenic promastigotes, axenic amastigotes, and macrophage-derived amastigotes (infected macrophages) Life-cycle stage differentiation (promastigote vs axenic amastigote vs intracellular amastigote) cDNA coverage per chromosome normalized to the entire genome; median chromosome transcript levels
RT-PCR Transfected CL-SL L. donovani cells Transfection with VL-SL genes Transcription/expression of the corresponding VL-SL transgenes
Sanger sequencing of isoenzyme gene (6-phosphogluconate dehydrogenase, 6PGDH) CL-SL and VL-SL L. donovani clinical isolates none Sequence-based strain typing confirming both isolates are L. donovani MON-37
Sanger sequencing verification of candidate SNPs and manual assembly/read alignment of the A2–A2rel multi-gene family region CL-SL and VL-SL L. donovani genomic DNA none Confirmation of CL-SL-specific non-synonymous SNPs; relative A2 gene copy number between strains
Key results
  • VL-SL infection produced splenomegaly and high liver and spleen parasite burdens, whereas CL-SL caused negligible/very little detectable visceral infection
  • In cutaneous infection both isolates showed low virulence, but CL-SL induced transient footpad swelling while VL-SL did not; the swelling was not accompanied by a significant increase in parasite number and was attributed to a stronger inflammatory response
  • Of 7 VL-SL genes transfected into CL-SL, only Rag C significantly increased spleen infection levels 5- to 40-fold across separate experiments
  • The CL-SL Rag C gene carries a homozygous non-conservative R231C substitution at a residue highly conserved across Leishmania and trypanosomatids
  • The A2/A2rel repeat cluster (LdBPK_220670.1, chromosome 22) is present at higher copy number in VL-SL than CL-SL gene VL/CL ratio 1.83
  • Nine regions of copy number variation were identified; higher in VL-SL: LdBPK_111220.1 ABC transporter, LdbpK_161030.1–161110.1 hypothetical cluster, LdbpK_200120.1 phosphoglycerate kinase B, chromosome 27 rRNA locus; higher in CL-SL: chromosome 23 MRPA/YIP1 region, chromosome 1 eIF4a, chromosome 19 glycerol uptake proteins, chromosome 29 hypotheticals tile VL/CL ratios 0.46–2.61; gene ratios 0.35–2.94
  • No gene deletions were detected in either genome, and only 5 genes were unequally affected by frame shift or stop-site change; the sole homozygous CL-SL-specific change (LdBPK_311390.1 stop lost) adds only 3 amino acids
  • Chromosome somy: most chromosomes disomic in both isolates, chromosome 31 tetrasomic and chromosome 23 trisomic in both; chromosomes 13 and 20 trisomic in VL-SL but diploid in CL-SL; chromosome-level transcript levels largely did not track somy (e.g. tetrasomic chromosome 31 had average mRNA levels)
Key statistics
  • fold_change 5 to 40 fold increase in spleen infection (CL-SL transfected with VL-SL Rag C vs control, separate experiments, statistically significant)
  • fold_change 1.83 (gene VL/CL ratio) (A2 and A2rel repeat cluster copy number, chromosome 22)
  • fold_change 2.94 (gene VL/CL ratio); tile ratio 2.61 (LdBPK_111220.1 ABC transporter-like protein, chromosome 11)
  • count 117 non-synonymous (70 homozygous + 47 heterozygous) and 92 synonymous (60 homozygous + 32 heterozygous) SNP changes specific to CL-SL; 15 homozygously mutated genes have putative functions (CL-SL-specific coding SNPs vs VL-SL and reference)
  • count over 70 pseudogenes from frame shifts or stop codon variations (Common to CL-SL and VL-SL but functional in reference L. donovani BPK282A1)
  • other more than 80% of differences common to VL-SL and CL-SL; about 20% of variants heterozygous; about 20% of SNPs in coding regions (Variant comparison against reference strain BPK282A1)
  • count 38,243 homozygous and 6,132 heterozygous variants common to VL & CL; 1,977/1,908 unique to VL; 2,251/2,587 unique to CL (Table 2 total SNP/indel counts vs reference genome)
  • other >200× coverage; mRNA level for chromosome 9 about 40% higher in VL-SL; 5 mice/group; 5×10^7 promastigotes IV and 5×10^6 subcutaneous (Sequencing depth, chromosome 9 transcript difference, and infection dose/group size)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper compares two Leishmania donovani clinical isolates (CL-SL and VL-SL) primarily through descriptive experimental biology: mouse infection experiments (5 mice/group) reporting organ parasite burdens and footpad swelling as mean plus standard error, a gene-complementation experiment (RagC transfection) described as producing a 'statistically significant' increase in spleen infection without naming the specific test, and genomic/transcriptomic comparisons (SNP/indel counts, copy-number ratios, chromosome somy, RNA-seq coverage) presented largely as descriptive counts and ratios rather than through named inferential statistical tests.

Replicationbiological Sample size5 mice per group stated for tail-vein (visceral) and footpad (cutaneous) infection experiments; the RagC complementation result is described as reproduced across 'separate experiments' without a stated number of replicates GroupsCL-SL vs VL-SL clinical isolates for visceral/cutaneous infection outcomes; CL-SL transfected with individual VL-SL candidate genes (including Rag C) vs other transfected genes Pairingunclear Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesyes Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
not specified (result described only as 'statistically significant') Fig. 2B, spleen infection levels in CL-SL parasites transfected with VL-SL Rag C versus other transfected/control genes 5 mice per group (as stated for infection experiments); described as replicated in 'separate experiments' not stated
Approaches that could also have been used
  • Differences between isolates and transfected lines were summarized as mean plus standard error and described as 'statistically significant' without naming the specific test used or reporting exact p-values.
    Could also: Explicitly naming the test applied (e.g., Student's t-test or Mann-Whitney U) and reporting exact p-values alongside SEM or a 95% confidence interval. — This would let readers independently judge the strength of evidence and the precision of the estimate, since 'significant' alone does not convey effect magnitude or uncertainty.
  • Seven candidate VL-SL genes were transfected individually into CL-SL and each tested for an effect on spleen infection levels.
    Could also: A multiple-comparison correction such as Bonferroni or Benjamini-Hochberg FDR applied across the panel of tested genes. — Testing several candidate genes within the same experimental round is a setting where a family-wise error rate or false discovery rate correction is commonly used to guard against chance findings among the several comparisons.
  • Mouse infection group sizes were small (5 mice/group).
    Could also: A nonparametric or exact test (e.g., Mann-Whitney U, permutation test) suited to small-sample comparisons. — Such methods do not depend on distributional (e.g., normality) assumptions that can be difficult to verify with small group sizes.
  • Chromosome-level transcript abundance was compared between the two isolates (Fig. 3B) using normalized RNA-seq/SL-sequencing coverage, presented descriptively rather than through a formal differential-expression test.
    Could also: A count-based differential expression framework such as DESeq2 or edgeR with model-based significance testing and FDR control. — These tools model the variance structure of sequencing count data and provide formal per-feature significance and multiple-testing-adjusted estimates across many genes or chromosomes at once.
  • Randomization of mice to groups and blinding of outcome assessment (parasite counts, footpad measurements) are not described.
    Could also: Explicitly stating randomized group allocation and blinded outcome assessment, in line with animal-study reporting frameworks such as ARRIVE. — Documenting these procedures helps readers assess the risk of observer bias in the infection and measurement outcomes.
  • SNP/indel and gene copy-number differences between isolates were reported as counts and ratios (Tables 1-3) without a formal enrichment or comparative statistical test.
    Could also: A statistical enrichment test such as Fisher's exact test to compare the distribution of variant categories (e.g., frame shift, stop gained) between isolates. — This would provide a formal probability estimate of whether the observed differences in variant-class distribution between isolates exceed what would be expected by chance.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

table1_gene_cnv
Reported
9 CNV loci (chr1,11,16,19,20,22,23,27,29) w/ VL/CL tile+gene coverage ratios from paper Table 1 (e.g. chr16 repeat cluster ~1.77-2.00, chr23 MRPA region ~0.38-0.46, chr19 depletion 0.55, chr22/A2 locus 1.83 [author-curated, paper states NOT programmatically reproducible due to reference misassembly])
Reproduced
somy_cnv.py (10kb tiles + per-gene BPK282A1 coverage, WGS bwa alignments). chr16: 1.70-2.13 (matches, n=10 genes); chr23: 0.38-0.42 (matches, n=5 genes); chr1/11/20/27/29: correct direction, ratio 20-40% off reported magnitude; chr19: reproduced ratio 0.93-1.03 (NO depletion) vs reported 0.49-0.55 (MISMATCH); chr22: 1.77 vs 1.83 reported but paper itself flags this locus as manually curated/not reproducible -> reported as reference only, not scored
partial
table2_snp_indel_counts
Reported
9 effect categories x 6 columns (Common/Unique-VL/Unique-CL x Homo/Hete) from GATK+SnpEff on paper's own calls; e.g. Non-Syn unique-VL=142 (70+72), unique-CL=117 (70+47)
Reproduced
GATK 3.8 joint HaplotypeCaller (hard-filtered, no VQSR truth set available) + SnpEff 5.1 custom BPK282A1 DB, classify_variants.py. Non-Syn unique-VL=85 (64+21), unique-CL=74 (68+6). Category ORDER/structure matches (Intergenic>>NonSyn>Syn>indels) but magnitudes are ~40-60% of reported, esp. heterozygous unique calls undercounted
partial
table3_specific_variants
Reported
5 exact chr:pos calls with CL/VL genotype+effect (chr30:579030, chr31:602800, chr31:1000447, chr32:9103, chr32:174033)
Reproduced
same 5 positions extracted from annotated.vcf: all genotype calls (Homo/Hete/wt) and effects MATCH exactly
exact
table4_diffexp_direction
Reported
8 genes >4-fold VL/CL DE (5 up-in-VL, 3 up-in-CL) with per-stage (Am/Ax/Pro) fold-changes from paper Table 4
Reproduced
edgeR GLM ~stage+condition (pooled, not per-stage as paper does) on SL-RNAseq counts. 4/8 genes correct direction+comparable magnitude, 1/8 correct direction but far-off magnitude, 3/8 no significant signal in pooled model
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 63/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Both datasets (SRP026479 WGS, GSE48475 SL-RNAseq) are fully public and were used 1:1, so this is a genuine reproduction rather than a proxy. Table 3 reproduces exactly (5/5 genotypes and effects) and the two headline CNV loci hold (chr16 1.70-2.13 vs ~1.77-2.00; chr23 0.38-0.42 vs 0.38-0.46), which supports the paper's central tropism thesis. The shortfalls are mostly on our side or in the paper's under-specification: Table 2 counts reach only ~40-60% of reported (Non-Syn unique-VL 85 vs 142) because the original calling/filtering pipeline is never described, and 3/8 Table 4 genes lose signal under our pooled edgeR design instead of the paper's per-stage contrasts. The one substantive unexplained discrepancy is chr19, where the reported 0.49-0.55 depletion is simply not present in the shared reads (0.93-1.03); chr22/A2 is excluded because the paper itself declares it manually curated and non-reproducible. Overall: solid but partial — yellow, not red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.