Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
About

Does It Reproduce?

We automatically re-run the analyses behind computational papers and check whether — and how far — their results reproduce from the deposited data and code. Every record is readable on its own, transparently justified, and version-sealed.

How it works

  1. 1

    An agent reads each paper, fetches its data and code, and re-runs the analysis autonomously on the brainbox compute.

  2. 2

    Every reported claim is graded against the re-run; the grades roll up into one 0–100 reproducibility score.

  3. 3

    Each record carries a level: L1 (AI) → L2 (a scientist plausibility-checks the grades) → L3/L4 (a scientist re-runs it with documented methods and results).

  4. 4

    Every verdict is version-sealed and contestable — a failed reproduction is a finding, never an accusation.

≥ 75 · reproduced 50–74 · partial < 50 · discrepant neutral · nothing to recompute

Progress

How far we are toward all reproducible literature. The middle stages are extrapolated from a broad random sample.

All literature (target) 39M
With code/pipeline (est.) 757k
Open full text (est.) 624k
Assessed 1k
1,272 of an estimated 624k reproducible papers assessed (0.2038%).

As of 2026-06-18 · counted live, every figure auditable.

How much code the literature discloses

The “with code/pipeline” stage above is an estimate. This curve is counted.

Papers that name a code repository

145,204 since 2000
2010
first year

Counted is every paper naming a GitHub repository in its methods section.

0 8k 16k 24k 32k 2000 2005 2010 2015 2020 2025

Extracted from PMC full text archive · 2026 is still in progress (lighter bar).

How this is counted: the base set is the open PMC full-text archive — every paper of a given year that can be read in full. Within it we count those naming a GitHub repository in their methods section (Europe PMC section search). Two limits to the figure: a mention in the methods may be the authors’ own repository or a third-party tool they used, and we cannot separate the two — so as a measure of “discloses its own code” the first curve is an upper bound. And whatever sits behind a paywall is not included at all: we can only check what is open.

Not an accusation — an invitation

A failed reproduction is not an integrity finding and not peer review. It measures what can be recomputed from public data today — failures often come from missing files, undocumented steps or access restrictions, not from the authors. We deliberately show no league table of “bad” papers, and we invite authors to have their work re-measured.

Who is behind it

Schlein Lab, University Medical Center Hamburg-Eppendorf (UKE). Reproductions run autonomously on the brainbox compute (brainarbeit.com); scientists do the plausibility checks.

Questions, corrections, collaborations: support@doesitreproduce.com

Imprint · Privacy