Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Metabolic profiling during malaria reveals the role of the aryl hydrocarbon receptor in regulating kidney injury.

Elife · 2020
L1 62/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation was attributed to the published material
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
62/100
Reproducibility score
0.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 23% of all assessed papers rank 891 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (re-queued room; prior attempt lost the C3 result to a «infra» reclaim before pull-back and had a silently-broken stepB join). The paper's pipeline-derived result is the liver RNA-seq (GEO GSE150268). The registry 'Code' link (github.com/tidyverse/ggplot2) is a text-mining false-positive (generic plot library, not authors' code), so per BRIEF P16 we reproduced with a standard third-party quantifier on the paper's own data. THREE claims: (C1/C2, grade partial) heme-metabolism gene induction of Fig 1-fig suppl 2A computed from the deposited FPKM matrix -- 15/22 panel genes up incl. Hmox1 +4.24 log2FC (~19x) and the paper-highlighted Hrg1/Slc48a1 (+0.99) and Ferroportin/Slc40a1 (+1.03); 'partial' only because the figure prints no exact numbers (pattern/direction match, no disagreement). (C3, grade within-tol) the central reproducibility test: salmon 2.1.1 (GENCODE vM25) re-quantification of all 24 raw runs reproduces the deposited FPKM matrix at per-sample median Spearman 0.957 with 24/24 samples uniquely re-identifying their own deposited column. NOTE for auditor: the deposited FPKM column names (day0_0...) differ from the GEO sample titles, so a naive name-join matches 0/24 (this silently happened in the prior run) -- pairing here was recovered by exact group-count matching + optimal assignment, and is non-circular (each sample's global-best match is its assigned column). NOT ATTEMPTED (out of scope): all metabolomics (LC-MS) and kidney-injury/histology/flow/qRT-PCR panels -- wet-lab, not pipeline-derived, no public processed data; and the authors' exact FPKM pipeline (aligner/quantifier/normalisation unspecified in Methods -> correlation, not byte-identical recompute). No data-fabrication signal; the only integrity issue is the mis-resolved code link. See AUDIT.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ 518c1914f4f8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether host metabolic reprogramming during malaria—specifically the production of aryl hydrocarbon receptor (AHR) ligands from heme and tryptophan metabolism—affects infection outcome, hypothesizing that AHR signaling protects against tissue damage (e.g., kidney injury) and improves survival during malaria.

Core claims
  • Plasma AHR ligands derived from heme metabolism (biliverdin, bilirubin) and tryptophan/kynurenine pathway metabolites increase during acute Pc malaria in mice and acute cerebral malaria in humans finding
  • Ahr-/- mice are more susceptible to Pc malaria, developing higher parasitemia, worse anemia, and 100% mortality finding
  • Loss of AHR signaling leads to high plasma heme and acute kidney injury during malaria finding
  • AHR-dependent protection during malaria requires AHR expression in Tek-expressing radioresistant cells mechanism
  • AHR signaling limits tissue damage and promotes survival during malaria finding
  • Untargeted plasma metabolomics reveals extensive, infection-stage-specific metabolic reprogramming during Pc malaria (370 of 587 metabolites significantly altered) resource
  • Acute Pc malaria is characterized by stable plasma heme despite hemolysis, alongside increased heme breakdown products (biliverdin, bilirubin), consistent with hepatic Hmox1/Blvra induction finding
  • Mouse Pc malaria recapitulates key metabolic changes (77 shared altered metabolites) observed in human pediatric cerebral malaria finding
Experimental setups
Assay System Perturbation Readout Platform
untargeted plasma metabolomics C57BL/6 mice, blood/plasma Pc infection (daily 0-25 DPI) vs mock infection scaled intensity of 587 identified metabolites
targeted mass spectrometry (tryptophan/kynurenine pathway) C57BL/6 mice, plasma Pc infection, 9 DPI concentration of tryptophan, L-kynurenine, 3-hydroxy-DL-kynurenine, quinolinic acid
bilirubin quantification C57BL/6 mice, plasma Pc infection, 9 DPI plasma bilirubin concentration
peripheral blood immune cell and cytokine profiling (flow cytometry/cytokine assay) C57BL/6 mice, peripheral blood Pc infection, 0-25 DPI NK cells, neutrophils, B cells, γδ T cells, IL-12p70, IFNγ, TNF, IL-10
gene expression analysis C57BL/6 mice, liver tissue Pc infection expression of heme metabolism genes (Hmox1, Blvra) and kynurenine pathway genes (Ido1 and downstream genes)
clinical plasma metabolomics (published dataset comparison) pediatric P. falciparum cerebral malaria patients, plasma natural infection, acute vs 30 days post-treatment convalescent scaled intensity of metabolites including heme and kynurenine pathway compounds
parasitemia, pathology and survival monitoring littermate Ahr+/+, Ahr+/-, Ahr-/- mice (female and male) Ahr genetic knockout + Pc infection parasitemia, parasite density, RBC density, body weight, temperature, survival over 15 days
plasma biochemistry C57BL/6 mice, blood/plasma Pc infection ALT (liver damage), RBC density (anemia)
Key results
  • 370 of 587 detected plasma metabolites were significantly altered (≥2-fold, p<0.05 FDR) during Pc malaria, mostly during acute infection; 66% increased ≥2-fold
  • Biliverdin, Z,Z-bilirubin, and E,E-bilirubin increased in scaled intensity during acute/late infection while heme scaled intensity remained stable despite hemolysis
  • 77 metabolites were significantly altered with similar magnitude and direction in both Pc-infected mice and human pediatric cerebral malaria patients p=1.572e-14
  • Tryptophan decreased while kynurenine, kynurenate, and quinolinate increased in plasma during acute infection in mice and humans
  • Ahr-/- mice developed higher parasitemia, more severe anemia, greater weight loss and temperature decrease, and all died between days 9-12, unlike Ahr+/+ and Ahr+/- mice which mostly survived
  • Targeted quantification confirmed elevated plasma bilirubin at 9 DPI in infected mice
  • Targeted quantification confirmed increased L-kynurenine, 3-hydroxy-DL-kynurenine, and quinolinic acid, and decreased tryptophan, at 9 DPI
Key statistics
  • count 587 metabolites identified (untargeted plasma metabolomics in Pc-infected mice)
  • count 370 metabolites significantly altered (≥2-fold change, p<0.05 t-test with FDR correction) (plasma metabolome during Pc malaria vs uninfected)
  • other 66% increased in scaled intensity (direction of the 370 significantly altered metabolites)
  • pvalue p=1.572e-14 (similarity of fold-change magnitude/direction for 77 metabolites shared between Pc mice and human cerebral malaria patients)
  • count n=10, 8, 11 (Ahr+/+, Ahr+/-, Ahr-/- female mice in survival/parasitemia experiment)
  • count n=11 patients per condition (pediatric cerebral malaria patient metabolomics comparison)
  • other all Ahr-/- mice died between days 9 and 12 post-infection (survival outcome of Ahr-/- mice vs Ahr+/+/Ahr+/-)
  • count n=12-13 mice (bilirubin), n=5-6 mice (kynurenine pathway metabolites) (targeted AHR ligand quantification at 9 DPI)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper used a longitudinal mouse model of Plasmodium chabaudi malaria combined with untargeted plasma metabolomics, targeted mass spectrometry, gene expression, and immune-cell profiling across a 25-day time series. Time-point comparisons of infected versus uninfected mice relied primarily on two-way ANOVA with FDR correction; human clinical metabolomics data were compared with Wilcoxon tests; and survival across Ahr genotypes was assessed by log-rank test. Results were reported as mean ± SEM with asterisk-based significance thresholds, with one exact p-value provided for the mouse–human metabolite correlation.

Replicationbiological Sample sizeExplicitly stated per figure: n=5 infected mice per day, n=2 uninfected per day, n=5 at DPI 0 for time series; n=10/8/11 per genotype for survival/disease experiment (combined from three independent experiments); n=11 patients per condition for clinical data; n=12–13 and n=5–6 for targeted metabolite validations GroupsInfected vs. uninfected mice across time; Ahr+/+, Ahr+/−, and Ahr−/− littermates during infection; acute vs. convalescent pediatric cerebral malaria patients; mouse vs. human fold changes Pairingunpaired Randomization/blindingnot stated DispersionSEM Effect sizesno Confidence intervalsno Multiplicity correctionFDR correction (specific algorithm, e.g., Benjamini-Hochberg, not stated)
Statistical tests used
Test Applied to n Assumptions
Two-way ANOVA with FDR correction Infected vs. uninfected comparisons across the 25-day time series for metabolites, immune cells, cytokines, and liver gene expression (Figures 1A–C, E, F, H; 2B; Figure 1—supplement 2A; Figure 2—supplement 1B) n=5 mice on DPI 0, n=5 infected mice per day, n=2 uninfected mice per day not stated
t-test with FDR correction Initial filter to identify significantly altered metabolites from the untargeted metabolomics dataset (Figure 1D selection criteria) n=5 mice on DPI 0, n=5 infected per day, n=2 uninfected per day not stated
Wilcoxon test Human pediatric cerebral malaria patient comparisons: heme metabolites, bilirubin isomers, kynurenine pathway metabolites, arachidonate, stachydrine, BHBA (Figures 1J, 2C, 2D, 2E; Figure 1—supplement 2B) n=11 patients per condition; n=12–13 mice (bilirubin); n=5–6 mice (kynurenine metabolites) not stated
Log-rank test Survival comparison across Ahr+/+, Ahr+/−, and Ahr−/− mice during Pc infection (Figure 3F) n=10, 8, and 11 for Ahr+/+, Ahr+/−, Ahr−/−, respectively not stated
Two-way ANOVA with FDR correction Genotype comparisons (Ahr+/− vs. Ahr−/−) for parasitemia, parasite density, RBCs, body weight, and temperature over time (Figures 3A–E) n=10, 8, and 11 for Ahr+/+, Ahr+/−, Ahr−/−, respectively; data combined from three independent experiments not stated
Linear model (regression) Correlation of fold-change direction and magnitude between mouse Pc malaria and human cerebral malaria for 77 shared metabolites (Figure 1I) n=77 metabolites; n=11 patients per condition for human data not stated
Approaches that could also have been used
  • Dispersion was reported as SEM throughout
    Could also: Standard deviation (SD) or 95% confidence intervals could also summarize spread — With small group sizes (n=2–5 per timepoint for uninfected mice), SEM scales with 1/√n and can appear narrow; SD conveys the actual biological variability in the sample, and a 95% CI explicitly bounds the uncertainty around the mean — both are commonly preferred when n is small
  • Separate two-way ANOVAs with FDR correction were run independently for each metabolite across the 25-day time series
    Could also: Multivariate approaches such as PCA, PERMANOVA, or linear mixed models on the full metabolite matrix could also characterize global metabolic trajectories — Univariate per-metabolite testing treats each of 370+ metabolites independently; multivariate methods account for correlations among metabolites, reduce dimensionality, and provide a single test of overall metabolic shift across time without inflating the false-discovery burden across hundreds of comparisons
  • The FDR correction method was described only as 'FDR correction' without naming the specific procedure
    Could also: Explicitly naming the algorithm (e.g., Benjamini-Hochberg, Storey q-value) could also be reported — Different FDR procedures have different assumptions and yield different adjusted p-values, particularly at small n; naming the method allows readers to assess its appropriateness and reproduce the analysis
  • Mouse–human metabolite fold-change correspondence was assessed with a linear model (Figure 1I)
    Could also: Pearson or Spearman correlation coefficients with 95% CI could also quantify and report this relationship — A linear model p-value indicates that the slope differs from zero, but an explicit correlation coefficient (r or ρ) would provide an interpretable effect size describing the strength of the mouse–human concordance, aiding comparison with similar cross-species studies
  • Survival differences across three Ahr genotypes were assessed with a log-rank test
    Could also: A Cox proportional hazards model could also be used to compare genotype groups — The log-rank test tests equality of survival curves; a Cox model additionally quantifies hazard ratios with confidence intervals, providing an effect size and enabling adjustment for covariates such as sex when combining experiments
  • The time-series metabolomics and gene-expression experiments were performed once, with between-mouse variability as the only replication
    Could also: An independent biological replicate experiment could also be included for the time-series arms — A single experimental run means that batch effects or run-specific variation cannot be distinguished from biological signal; a replicate run would allow assessment of inter-experiment reproducibility and increase confidence in the identified metabolic trajectories
Software: Metabolon (untargeted metabolomics platform, implied by 'scaled intensity' reporting convention)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
12
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

RRID:AB_11125547 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2239227 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2270597 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2722659 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_476692 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_880536 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33021470

Paper: Lissner et al. 2020, eLife 9:e60165. "Metabolic profiling during malaria reveals the role of the aryl hydrocarbon receptor in regulating kidney injury." PMID 33021470 · PMCID PMC7538157 · DOI 10.7554/eLife.60165

Data / code resolution

  • Code link in registry = github.com/tidyverse/ggplot2 → this is a generic plotting library, NOT the authors' analysis code. It is a text-mining false-positive (the paper cites ggplot2 as a figure tool). Treated as no own repo. Per BRIEF rule P16, reproduction via a standard third-party tool on the paper's own data is equally valid → that is what we do.
  • Data = GEO GSE150268 ("Liver transcriptomics during malaria"): RNA-seq, liver, female C57BL/6-Crl mice, Plasmodium chabaudi AJ infection. 24 PE samples (SRR11768314–11768337; SRA SRX8321587–8321610). Timecourse: 0 dpi uninfected (n=4), 7/8/9/10/11 dpi infected (n=16 total), plus 8 & 10 dpi uninfected controls (n=4). Public, fully obtainable. GEO also ships a processed FPKM matrix GSE150268_liver_fpkm.txt.gz = the authors' pipeline output.

What the paper reports from these RNA-seq data

  • The liver RNA-seq underlies Figure 1—figure supplement 2A: "Expression of heme metabolism genes in livers of Pc-infected mice."
  • Methods give library prep + sequencer only (TruSeq RNA v2, HiSeq 4000, 75bp PE). No aligner, quantifier, normalization, or DE statistics are stated. → exact authors'-pipeline reproduction is not specifiable (docs_insufficient for the pipeline itself); the deposited FPKM matrix is the only pinnable output.

In scope (pipeline-derived, attempted)

  1. Re-quantification of the deposited expression matrix from raw reads (headline 1:1). Tool: salmon (GENCODE vM25 mouse, decoy-free) → gene-level TPM, per-sample correlated (Spearman + log-Pearson) against the deposited FPKM matrix. Tests whether the deposited processed data is reproducible from the raw reads with a standard third-party pipeline. Job B (SLURM 2176832).
  2. Biological pattern of Fig 1—fig suppl 2A — heme-metabolism genes (Hmox1, Slc48a1/Hrg1, Slc40a1/Ferroportin, Hp, Hpx, Cd163, Alas2 …) induced in infected vs uninfected liver, computed directly from the deposited FPKM matrix. Job A (SLURM 2176830).

Out of scope (not pipeline / not attempted, with reason)

  • All metabolomics (LC-MS plasma/kidney metabolite profiling) — wet-lab instrument data, not a public pipeline; no deposited processed matrix linked.
  • Kidney-injury / histology / flow-cytometry / qRT-PCR panels (incl. Fig 4 suppl 2D Ahr+/- vs Ahr-/-) — wet-lab measurements, not derived from GSE150268.
  • Exact authors' FPKM pipeline (aligner/quantifier/version) — not specified in Methods; we reproduce the matrix with a standard tool and report rank/scale agreement, not a byte-identical recompute.

Grading intent

  • Claim 1 graded by per-sample correlation (within-tol if median Spearman high).
  • Claim 2 graded on direction/pattern agreement with the figure (figure prints no exact numbers → at best partial, never exact). Human reviewer decides.
Figures / tables: Fig 1figure supplementfig suppl
C1
Reported
heme metabolism genes induced during infection (Fig 1-fig suppl 2A, qualitative panel, no printed values)
Reproduced
15/22 heme-panel genes up in infected liver; direction/pattern reproduced from deposited FPKM matrix
partial
C2
Reported
Hmox1/heme-oxygenase response upregulated in infected liver (no printed number)
Reproduced
Hmox1 mean FPKM 35.5 -> 670.9, log2FC +4.24 (~19-fold) infected vs uninfected
partial
C3
Reported
deposited GSE150268_liver_fpkm.txt (authors' processed matrix, 24346 genes x 24 samples) reproducible from raw reads
Reproduced
salmon 2.1.1 (GENCODE vM25) re-quant of all 24 raw runs; all 24 matched; per-sample median Spearman 0.957 (min 0.942, max 0.964), median log2-Pearson 0.963; 24/24 each sample's globally-best match is its correct deposited column (non-circular)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 62/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Within the reproducible scope this is a clean result: the heme-metabolism gene induction of Fig 1—fig suppl 2A is fully derivable from the authors' own deposited FPKM matrix (15/22 panel genes up, Hmox1 35.5→670.9 FPKM, +4.24 log2FC), with no fabrication signal — q5 green, severity negligible. The explainable limitations are on method-reporting/availability, not result integrity: Methods omit the aligner/quantifier/DE (underspecified pipeline), the registry code link is a ggplot2 false-positive, and the target figure is qualitative so only pattern can be compared. The paper's actual central conclusion (AhR regulating malaria kidney injury) rests on out-of-scope metabolomics/wet-lab data and was untested, so core-claim confirmation is limited rather than full; the C3 matrix-reproducibility correlation was still pending at finalization.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

328.2 k
tokens (I/O) · 20.5 M incl. cache
97 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.