Does It Reproduce?
We automatically re-run the analyses behind computational papers and check whether — and how far — their results reproduce from the deposited data and code. Every record is readable on its own, transparently justified, and version-sealed.
How it works
-
1
An agent reads each paper, fetches its data and code, and re-runs the analysis autonomously on the brainbox compute.
-
2
Every reported claim is graded against the re-run; the grades roll up into one 0–100 reproducibility score.
-
3
Each record carries a level: L1 (AI) → L2 (a scientist plausibility-checks the grades) → L3/L4 (a scientist re-runs it with documented methods and results).
-
4
Every verdict is version-sealed and contestable — a failed reproduction is a finding, never an accusation.
Progress
How far we are toward all reproducible literature. The middle stages are extrapolated from a broad random sample.
As of 2026-06-18 · counted live, every figure auditable.
How much code the literature discloses
The “with code/pipeline” stage above is an estimate. This curve is counted.
Papers that name a code repository
Counted is every paper naming a GitHub repository in its methods section.
Extracted from PMC full text archive · 2026 is still in progress (lighter bar).
How this is counted: the base set is the open PMC full-text archive — every paper of a given year that can be read in full. Within it we count those naming a GitHub repository in their methods section (Europe PMC section search). Two limits to the figure: a mention in the methods may be the authors’ own repository or a third-party tool they used, and we cannot separate the two — so as a measure of “discloses its own code” the first curve is an upper bound. And whatever sits behind a paywall is not included at all: we can only check what is open.
Not an accusation — an invitation
A failed reproduction is not an integrity finding and not peer review. It measures what can be recomputed from public data today — failures often come from missing files, undocumented steps or access restrictions, not from the authors. We deliberately show no league table of “bad” papers, and we invite authors to have their work re-measured.
Who is behind it
Schlein Lab, University Medical Center Hamburg-Eppendorf (UKE). Reproductions run autonomously on the brainbox compute (brainarbeit.com); scientists do the plausibility checks.
Questions, corrections, collaborations: support@doesitreproduce.com