Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

FoxO transcription factors are required for hepatic HDL cholesterol clearance

· 2018
PubMed 29408809 ↗ pmid-29408809
L1 49/100 3/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Total score +9
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
49/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 7% of all assessed papers rank 1081 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Predominantly a wet-lab mouse paper; only 2 pipeline-derived results in scope, both attempted on «our HPC». The microarray it 'queried' is not deposited with this paper but reused from Haeusler 2014 (GSE60527), which is open, complete (12 samples, balanced 2x2x3) and delivers what is promised. Reanalysis (GEOquery+limma) REPRODUCES the qualitative direction of the headline finding - Scarb1 and Lipc trend DOWN in L-FoxO1,3,4 livers - and the HOMER motif claim (137/346 FoxO motifs near Scarb1/Lipc, 'dozens' confirmed). However it is a PARTIAL, not clean 1:1: (a) the Scarb1/Lipc microarray changes are not statistically significant on their own (the paper's significance comes from qPCR, which is wet-lab/out-of-scope); (b) two genes the paper lists as unchanged - Abcg5 and Abcg8 - are in fact significantly reduced in the pooled genotype model (a real discrepancy, but analytical/power-related, not fabrication). NOT attempted (out of scope): qPCR, HDL-C/lipid assays, lipoprotein kinetics, selective CE uptake, SR-BI re-expression rescue, Westerns, histology, all mouse phenotyping. No fabrication indicated.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 49
    assessed: 2026-06-18 ⛓ 650fb67616e8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether the insulin-repressible FoxO transcription factors in the liver mediate insulin's effect on HDL cholesterol, hypothesizing that hepatic FoxOs are required for HDL-C clearance and cholesterol homeostasis.

Core claims
  • Mice with liver-specific triple FoxO knockout (L-FoxO1,3,4) have increased HDL cholesterol. finding
  • Hepatic FoxOs are required for cholesterol homeostasis and HDL-mediated reverse cholesterol transport to the liver. mechanism
  • Increased HDL-C in L-FoxO1,3,4 mice is associated with decreased expression of the HDL-C clearance factors SR-BI and hepatic lipase and defective hepatic selective uptake of HDL cholesteryl ester. mechanism
  • The HDL-C phenotype can be rescued by re-expression of SR-BI. finding
  • L-FoxO1,3,4 mice have reduced hepatic glucose production. finding
Experimental setups
Assay System Perturbation Readout Platform
plasma HDL cholesterol measurement liver-specific triple FoxO knockout (L-FoxO1,3,4) mice liver-specific FoxO1,3,4 knockout HDL cholesterol levels
gene/protein expression analysis L-FoxO1,3,4 mouse liver liver-specific FoxO1,3,4 knockout expression of SR-BI and hepatic lipase
HDL cholesteryl ester selective uptake assay L-FoxO1,3,4 mouse liver liver-specific FoxO1,3,4 knockout hepatic selective uptake of HDL cholesteryl ester
SR-BI re-expression rescue experiment L-FoxO1,3,4 mice SR-BI re-expression HDL-C phenotype rescue
Key results
  • L-FoxO1,3,4 mice have increased HDL cholesterol
  • Decreased expression of SR-BI and hepatic lipase in L-FoxO1,3,4 liver
  • Defective hepatic selective uptake of HDL cholesteryl ester
  • Re-expression of SR-BI rescued the HDL-C phenotype
  • L-FoxO1,3,4 mice have reduced hepatic glucose production

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper reports a mouse genetic study comparing liver-specific triple FoxO knockout (L-FoxO1,3,4) mice with controls to examine effects on HDL-C metabolism, SR-BI and hepatic lipase expression, and selective hepatic cholesteryl ester uptake. A rescue experiment involving SR-BI re-expression is also described. Specific statistical tests, sample sizes, and reporting conventions are not detailed in the text provided (abstract and metadata only; full methods and results sections were not included in the supplied text).

Replicationbiological Groupsliver-specific triple FoxO1,3,4 knockout mice vs. control mice; SR-BI rescue subgroup Pairingunpaired Randomization/blindingnot stated Dispersionunclear
Statistical tests used
Test Applied to n Assumptions
not stated not stated — full methods/results text not supplied not stated
Approaches that could also have been used
  • The study compares a single knockout genotype against controls across multiple metabolic endpoints (HDL-C, SR-BI expression, hepatic lipase, cholesteryl ester uptake), which typically involves multiple separate tests.
    Could also: A single mixed-model ANOVA or linear mixed model with genotype as a fixed effect and animal as a random effect, followed by a single multiplicity correction (e.g., Benjamini-Hochberg FDR) across all endpoints. — Controlling the family-wise error rate or FDR across multiple correlated outcomes reduces the probability of false positives and makes the inference across endpoints more coherent.
  • The rescue experiment (SR-BI re-expression) introduces a three-group comparison (control, knockout, knockout + rescue), which is a common design in genetic mouse studies.
    Could also: One-way ANOVA with a post-hoc test (e.g., Tukey HSD or Dunnett's test against the control group) applied to the three groups jointly. — A single ANOVA followed by a planned post-hoc comparison controls the familywise error rate across the three pairwise comparisons more explicitly than conducting separate t-tests.
  • Mouse metabolic phenotyping studies of this type routinely measure outcomes (e.g., plasma cholesterol) that may not be normally distributed, particularly at small sample sizes typical of genetic mouse work.
    Could also: Non-parametric alternatives such as the Mann-Whitney U (for two-group comparisons) or Kruskal-Wallis test with Dunn's post-hoc (for three groups) could also be applied. — Non-parametric tests do not assume normality and can be more appropriate when n per group is small (e.g., ≤8), as is common in transgenic mouse studies; they provide an independent check on parametric results.
  • Gene expression outcomes (SR-BI, hepatic lipase mRNA/protein) are likely compared between genotypes, a context where specialized tools are sometimes used.
    Could also: For mRNA quantification by RT-qPCR, a ΔΔCt-based analysis with appropriate reference gene normalization and reporting of efficiency-corrected fold-changes with 95% CIs would also be standard. — Reporting fold-changes with confidence intervals alongside p-values conveys both the magnitude and precision of expression differences, aiding biological interpretation.
  • Results in studies of this type are frequently summarized with SEM to represent group means.
    Could also: SD or 95% CIs around the mean could also be used to summarize spread. — SD describes the variability of individual animals (biologically interpretable), while 95% CIs convey estimation uncertainty; both are often preferred over SEM for small-n animal studies because SEM can visually understate biological variability.
  • The study uses a complete gene knockout to establish that FoxOs are required for hepatic HDL clearance.
    Could also: A dose-response or graded knockdown design (e.g., heterozygous knockouts or inducible partial suppression) could also characterize the relationship between FoxO activity and HDL-C. — Graded designs can reveal whether the phenotype is linearly related to FoxO dosage or threshold-dependent, providing additional mechanistic resolution beyond the binary knockout comparison.
Software: not stated

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 29408809

Paper: Lee SX, Heine M, Schlein C, … Haeusler RA. FoxO transcription factors are required for hepatic HDL cholesterol clearance. J Clin Invest. 2018;128(4):1615-1626. DOI 10.1172/JCI94230 · PMCID PMC5873864.

NOTE: the user/operator («email», Christian Schlein) is a co-author.

Nature of the study

Predominantly a wet-lab mouse-physiology study: liver-specific triple FoxO knockout (L-FoxO1,3,4), plasma HDL-C, lipoprotein kinetics, selective CE uptake, SR-BI re-expression rescue, qPCR, Western blots. The bulk of the figures are wet-lab and out of scope for computational reproduction.

In-scope pipeline-derived (computational) results

R1 — Microarray query (CORE, clearly reproducible)

The authors "queried microarrays from livers of L-FoxO1,3,4 mice" for HDL-metabolism genes. Reported result:

  • Scarb1 (SR-BI) and Lipc (hepatic lipase): reduced in L-FoxO1,3,4 livers.
  • No change in HDL synthesis / biliary-excretion genes: Abca1, Apoa1, Apoa2, Lcat, Apoe, Abcg5, Abcg8.

The microarray is not deposited with this paper; it is the L-FoxO1,3,4 liver microarray from the prior Haeusler-lab paper (Haeusler et al., Nat Commun 2014; Integrated control of hepatic lipogenesis vs glucose production requires FoxO TFs), deposited as GEO GSE60527 (GPL6096, Affymetrix Mouse Exon 1.0 ST; 12 samples = WT vs TKO × fast 22h / refed 4h × 3). 1:1 reproduction = re-derive genotype effect (KO−WT) for the 9 named genes from GSE60527. Pipeline: GEOquery → limma differential expression.

R2 — HOMER motif analysis (harder; attempted)

"motif analysis using HOMER … identified dozens of FoxO consensus sequences in and near Scarb1 and Lipc" using "the entire mouse gene sequences of Scarb1 and Lipc (including introns) plus 50 kb upstream of each TSS" (Supplemental Fig. 8A–C). Pipeline: extract mm10 regions → HOMER FoxO motif scan → count consensus matches.

Out of scope (wet-lab / manual / not pipeline-derived)

qPCR (Fig 2A,B), HDL-C and lipid measurements, lipoprotein turnover/kinetics, selective cholesteryl-ester uptake assays, adenoviral SR-BI re-expression rescue, Western blots, histology, all mouse phenotyping. Not attempted.

Figures / tables: Fig 8A
C1
Reported
Scarb1 (SR-BI) reduced in L-FoxO1,3,4 liver microarray
Reproduced
down in KO, logFC(KO-WT)=-0.176 log2 (P=0.155, ns); strongest in refed (-0.324, P=0.067)
partial
C2
Reported
Lipc (hepatic lipase) reduced in microarray
Reproduced
logFC=-0.134 (ns); fasted -0.485 (P=0.11, down), refed +0.217 (up)
partial
C3
Reported
no difference in Abca1, Apoa1, Apoa2, Lcat, Apoe, Abcg5, Abcg8
Reproduced
Apoa1/Apoa2/Apoe/Lcat unchanged (match); Abcg5 -0.96 (adj.P=0.016) and Abcg8 -1.31 (adj.P=0.008) significantly DOWN; Abca1 +0.32 (P=0.004) up
did not match
C4
Reported
dozens of FoxO consensus sequences in/near Scarb1 and Lipc (HOMER)
Reproduced
Scarb1=137, Lipc=346 FoxO motif occurrences (gene body + 50kb upstream, mm10)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 49/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🔴3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Total score +9

Only two pipeline-derived results were in scope and the reused microarray (GSE60527) is fully public and complete, so data identity is clean. The headline direction reproduces (Scarb1/Lipc trend down; 137/346 FoxO motifs confirm 'dozens'), but neither reaches microarray significance — the paper's significance rests on out-of-scope qPCR — and two genes called 'unchanged' (Abcg5 adj.P=0.016, Abcg8 adj.P=0.008) are in fact significantly reduced. The discrepancies sit on the analysis/statistics side and are plausibly driven by our self-chosen pooled model and low power (n=3/group), not by the authors or by fabrication; all values are derivable from the shared data. Overall a solid, direction-consistent partial reproduction with explainable deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

213.4 k
tokens (I/O) · 23.9 M incl. cache
66 min
runtime · 0.07 CPU-h
7.3 GB
peak RAM
4
HPC jobs
hummel
machine