Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genetic Dissection of Tissue-Specific Apolipoprotein E Function for Hypercholesterolemia and Diet-Induced Obesity

PLoS ONE · 2015
L1 No computation 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Total score +7
✓ What held up
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

DROP (non_pipeline). PLoS ONE 2015 wet-lab mouse genetics/physiology paper (tissue-specific conditional Apoe knockouts; diet studies) by Wagner, Bartelt, Schlein, Heeren. Methods are entirely bench assays (Cre/loxP genetics, FPLC, qPCR/TaqMan, Western blot, ELISA, ALT, Tyloxapol VLDL secretion, 125I-TC turnover, H&E+ImageJ, Student's t-test) with NO bioinformatic/computational pipeline, NO sequencing/microarray/proteomics, NO code repository (own or third-party, so P16 does not apply), and NO public data accession. Data availability = 'all data within the paper and SI'; the SI is only two summary TIF figures (S1 food intake, S2 OGTT), both verified to resolve as valid TIFFs. Nothing pipeline-derived exists to reproduce, so no «our HPC» compute was run. Profiled the only deposited data (2 SI figures, grade D, delivers_promised=partial) and recorded 7 representative reported wet-lab claims with paper locations for audit. NOT attempted: re-deriving any reported number, because no raw/machine-readable data or code was deposited (a data-availability limitation, not evidence of fabrication). Honest, evidence-backed drop; no result forced or fabricated.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-18 ⛓ df07d296cf51
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study investigates whether the metabolic phenotypes seen in globally apoE-deficient mice (hypercholesterolemia and altered diet-induced obesity/adiposity) are attributable specifically to hepatocyte-derived or adipocyte-derived apoE, using a novel Cre-loxP conditional tissue-specific Apoe knockout mouse model.

Core claims
  • Hepatocyte apoE is required for normal VLDL production and protects against diet-induced dyslipidemia finding
  • Adipocyte-specific apoE deletion does not reproduce the lean/insulin-sensitive adipose phenotype seen in global Apoe-/- mice finding
  • A novel conditional (Cre-loxP) tissue-specific Apoe knockout mouse model was generated (Apoe ΔHep and Apoe ΔAT) resource
  • Apoe ΔHep mice show increased body/liver weight, hepatic steatosis, elevated ALT and inflammation markers on high-fat diet finding
  • Apoe ΔAT mice show no detectable metabolic phenotype on either HFD or WTD finding
  • Circulating plasma apoE is predominantly derived from hepatic production, not adipose tissue finding
  • VLDL organ-specific tissue uptake is disturbed in Apoe ΔHep mice while plasma VLDL clearance rate (half-life) is unaltered finding
  • ApoE produced by cell types other than hepatocytes or adipocytes likely explains the lean and insulin-sensitive phenotype of global Apoe-/- mice mechanism
Experimental setups
Assay System Perturbation Readout Platform
real-time RT-PCR (TaqMan) liver and white adipose tissue, mouse Cre-loxP tissue-specific Apoe deletion Apoe mRNA expression (ΔΔCt, normalized to Tbp) TaqMan Gene Expression Assay, Applied Biosystems
Western blot liver lysates, mouse Apoe ΔHep knockout (Cre+ vs Cre-) total apoE protein amount (34 kDa) NuPAGE Bis-Tris gels, Invitrogen; apoE antibody, Santa Cruz
ELISA plasma, mouse diet (chow/HFD/WTD) x genotype apoE plasma levels
FPLC lipoprotein profiling pooled plasma, mouse diet x genotype cholesterol and apoE distribution across lipoprotein fractions AKTA FPLC with S6-superose sizing columns, GE Healthcare
VLDL production assay (Tyloxapol injection) plasma, mouse diet (chow/WTD) x Apoe ΔHep genotype VLDL triglyceride and cholesterol secretion rate
125I-VLDL turnover/organ distribution assay plasma and organs (liver, heart, fat pads, kidney, spleen, muscle), mouse injection of 125I-labeled apoE-free VLDL tracer, Apoe ΔHep genotype plasma decay (half-life) and organ-specific radioactivity uptake 125I-tyramine cellubiose labeling
biochemical liver lipid quantification liver homogenate, mouse diet x genotype triglyceride and cholesterol content normalized to protein commercial kits, Roche
histology (H&E staining) liver and epididymal WAT, mouse diet x genotype steatosis grade and adipocyte size distribution ImageJ
Key results
  • Apoe ΔHep mice showed higher body weight on HFD compared to WT controls about 10%
  • Liver mass increased in Apoe ΔHep mice on HFD
  • Hepatic cholesterol elevated in Apoe ΔHep mice on both HFD and WTD; hepatic triglycerides unaltered by genotype
  • Plasma ALT activity higher in Apoe ΔHep mice than WT on HFD 4-fold
  • Plasma cholesterol increased in Apoe ΔHep mice on WTD, with increased TRL/remnant cholesterol by FPLC
  • VLDL triglyceride secretion (post-Tyloxapol) was lower in Apoe ΔHep mice on both chow and WTD; VLDL cholesterol unchanged
  • Plasma half-life of 125I-VLDL was similar between genotypes, but organ-specific VLDL uptake/delivery differed
  • Apoe ΔAT mice showed no genotype-specific differences in body weight, tissue mass, adipocyte size, or oral glucose tolerance on HFD or WTD
Key statistics
  • pvalue p<0.05 (significance threshold used for all Student's t-test comparisons)
  • fold_change 4-fold (ALT plasma activity higher in Apoe ΔHep vs WT controls on HFD)
  • other about 10% (higher body weight of Apoe ΔHep mice vs WT controls on HFD)
  • count n≥8 (sample size for hepatic Apoe expression comparison in Apoe ΔHep mice (Fig 1B))
  • count at least 10 generations (backcrossing of floxed Apoe mice onto C57BL6/J background)
  • other 16 weeks (duration of dietary feeding (chow/HFD/WTD) before tissue harvest)
  • other 0.5 g kg-1 body weight (intravenous Tyloxapol dose used for VLDL production assay)
  • other 1 g/kg (oral glucose dose used for oral glucose tolerance test)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used a conditional Cre-loxP mouse model to compare tissue-specific apoE-deficient mice (Apoe^ΔHep or Apoe^ΔAT) against floxed littermate controls across three dietary regimens (chow, HFD, WTD) for 16 weeks. All pairwise group comparisons were performed with unpaired Student's t-tests, with results expressed as mean ± SEM. Statistical significance was defined by p < 0.05, indicated by asterisks in figures; no multiplicity correction was stated despite numerous outcomes tested across conditions and genotypes.

Replicationbiological Sample sizeGroup sizes stated per figure legend as n≥3 to n≥9; no a priori power calculation reported GroupsApoe^ΔHep or Apoe^ΔAT vs. floxed littermate controls (WT), each across three dietary conditions (chow, HFD, WTD) Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Student's t-test (two-sample, unpaired) Apoe mRNA expression in liver (Fig 1B) and WAT (Fig 1C) n≥8 (liver); n≥6 (WAT) not stated
Student's t-test (two-sample, unpaired) ApoE plasma levels across diets (Fig 1E) n≥3 to n≥8 per group per dietary condition not stated
Student's t-test (two-sample, unpaired) Body weight curves and tissue weights (liver, BAT, IngWAT, EpiWAT; Fig 2A–J) n≥3 to n≥9 per group per diet not stated
Student's t-test (two-sample, unpaired) Hepatic TG and cholesterol, plasma ALT, OGTT, liver gene expression (Tnf, Cd68, Cxcl-2, Cxcl-10; Fig 3) n≥4 to n≥8 per group per diet not stated
Student's t-test (two-sample, unpaired) Plasma total cholesterol and triglycerides across diets (Fig 4A,B) n≥3 to n≥8 per group per diet not stated
Student's t-test (two-sample, unpaired) VLDL production (tyloxapol assay) and 125I-VLDL organ distribution (Fig 5A–D, 5H) n≥5 to n≥8 per group not stated
Approaches that could also have been used
  • Multiple independent Student's t-tests were applied across many outcomes and dietary conditions without a stated multiplicity correction
    Could also: A two-way ANOVA (genotype × diet) with a post-hoc correction (e.g., Tukey HSD or Bonferroni) could also have been applied to outcomes measured across all three diets — Two-way ANOVA directly models the genotype-by-diet interaction — a central question of the study — and an omnibus test with post-hoc correction controls the family-wise error rate across the simultaneous comparisons, making the inferential structure explicit
  • FPLC lipoprotein profiles (Figs 4C–H, 5E–F) were generated from pooled plasma samples per group and no between-group statistics were applied to these profiles
    Could also: Individual-sample FPLC profiling followed by area-under-curve or fraction-specific quantification could also support formal statistical comparison between genotypes — Pooling samples precludes estimation of within-group variability and formal hypothesis testing on fraction-specific cholesterol levels; individual profiles would allow the differences visible in the traces to be accompanied by confidence intervals or p-values
  • Dispersion was reported as SEM throughout all figures
    Could also: Standard deviation (SD) or 95% confidence intervals could also be used to represent spread — SEM reflects precision of the mean estimate and narrows mechanically with increasing n; SD describes the biological variability of individual animals and is often preferred for that purpose, while 95% CIs simultaneously convey both precision and inferential content — particularly informative for small group sizes (n≥3 to n≥9)
  • Statistical significance was reported as binary asterisk notation (p<0.05) without exact p-values
    Could also: Exact p-values (e.g., p=0.03, p=0.18) could also be reported for each comparison — Exact p-values allow readers to evaluate borderline results, support downstream meta-analyses, and distinguish a p=0.049 from a p=0.001 finding — a distinction obscured by the single-threshold asterisk system
  • The Student's t-test (a parametric test) was used for all comparisons, including groups with the smallest stated sample sizes (n≥3 to n≥5)
    Could also: A non-parametric alternative such as the Mann-Whitney U test could also be applied, particularly for the smallest group sizes where distributional assumptions are difficult to verify — With very small n the normality assumption underlying the t-test cannot be robustly assessed; non-parametric rank-based tests make fewer distributional assumptions, though they are generally less powerful when the normality assumption is met
  • No a priori sample size justification or power calculation was reported
    Could also: A power calculation based on a pilot estimate or published values for a primary outcome (e.g., plasma cholesterol difference) could also have been reported — Reporting a power calculation or the targeted effect size helps readers interpret null results — such as the absence of phenotype in Apoe^ΔAT mice — as having been adequately powered to detect a biologically relevant difference, rather than leaving the power of those negative findings uncertain
Software: not stated

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope analysis — PMID 26695075

Title: Genetic Dissection of Tissue-Specific Apolipoprotein E Function for Hypercholesterolemia and Diet-Induced Obesity Authors: Wagner T, Bartelt A, Schlein C, Heeren J. Venue: PLoS ONE 2015 Dec 22;10(12):e0145102 · DOI 10.1371/journal.pone.0145102 · PMCID PMC4687855 Note: the operator («email») is a co-author of this paper.

What kind of study is this?

A wet-lab mouse genetics / metabolic physiology study. The authors generated a novel conditional (floxed) Apoe allele and used Cre/loxP to delete apoE tissue-specifically in hepatocytes (Apoe^ΔHep) or adipocytes (Apoe^ΔAT), then fed the mice experimental diets (chow / high-fat diet HFD / Western-type diet WTD) and phenotyped them.

Methods inventory (from Materials & Methods, PMC4687855)

Technique Output Class
Cre/loxP, genotyping PCR, Southern blot allele validation wet-lab
Body weight curves, food intake physiology wet-lab
Oral glucose tolerance test (AccuChek) glucose curves wet-lab
Plasma cholesterol / triacylglyceride kits concentrations wet-lab
ELISA (apoE) plasma apoE wet-lab
ALT (Cobas Mira) liver enzyme wet-lab
FPLC (S6 Superose, ÄKTA) lipoprotein sizing lipoprotein profiles wet-lab
Tyloxapol VLDL secretion; ¹²⁵I-TC turnover secretion/clearance rates wet-lab
Hepatic lipid extraction + Lowry tissue lipids wet-lab
Real-time RT-PCR (TaqMan, ΔΔCt) targeted gene expression wet-lab
Western blot (apoE) protein wet-lab
H&E histology + ImageJ adipocyte sizing morphometry manual image analysis
Student's t-test (p<0.05) significance basic stats

In scope vs out of scope (pipeline-derived results only)

Bioinformatic / computational-pipeline results in this paper: NONE.

  • No high-throughput sequencing (RNA-seq/WGS/WES/ATAC), no microarray, no proteomics MS, no genome-wide assay, no alignment/quantification/DE pipeline.
  • Gene expression is targeted qPCR (a handful of genes, ΔΔCt) — bench measurement, not a pipeline.
  • The only "computational" step is ImageJ adipocyte-area measurement on H&E micrographs (manual ROI quantification). The source micrographs are not deposited, so even this is not reproducible from shipped data.
  • Statistics are pairwise Student's t-tests — not a pipeline.

Every reported number is a wet-lab measurement on mice. There is no pipeline-derived computational result to reproduce.

Code & data availability

  • Code: none. No repository, no scripts (own or third-party). The brief's P16 third-party-tool path does not apply — there is no computational analysis to re-run on the paper's data.
  • Data availability statement (verbatim): "All relevant data are within the paper and its Supporting Information files."
  • Public accession: none (no GEO/SRA/ENA/ArrayExpress/PRIDE/Zenodo).
  • Supporting Information: two TIF figure images only —
    • S1 Fig "Effect of tissue-specific Apoe deletion on food intake" (.s001, 231 634 B)
    • S2 Fig "Effect of adipose tissue-specific deletion of apoE on oral glucose tolerance" (.s002, 377 770 B) No raw per-animal data tables, no spreadsheets — only summary figures.

Decision

DROP — drop_reason = non_pipeline (secondary: no_code, no_data_accession). This is a text-mineable biomedical paper with no computational pipeline, no code artifact, and no deposited machine-readable dataset; only summary figures of wet-lab measurements are available. No «our HPC» compute is warranted. Per the brief, drops are valid outcomes — recorded honestly with full evidence rather than forcing a result.

Figures / tables: Fig 2BFig 2GFig 3IFig 1EFig 5AFig 4A

No individual results have been recorded for this entry yet.

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 56/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Total score +7

PMID 26695075 is an all-wet-lab mouse genetics/physiology paper (PLoS ONE 2015, tissue-specific conditional Apoe knockouts) with no computational pipeline, no code, and no deposited machine-readable data — only two summary SI figures — so it was correctly dropped as non_pipeline. All reported endpoints are bench measurements that cannot be placed against any output (q1/q2 red on availability, not an authors' defect). No re-derivation was possible and no fabrication signal exists; overall yellow reflects a sound study with nothing computationally reproducible.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

56.9 k
tokens (I/O) · 2.5 M incl. cache
5 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.