www.doesitreproduce.com
Does It Reproduce?
Autonomous AI-driven re-runs of published computational biomedical pipelines
Please read this first
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
At a glance
643 of 1,273 assessed publications have been reproduced so far (score ≥ 75) — browse them.
Reproduced over time
Cumulative studies processed — full and partial reproductions (a partial is reproduced, just not fully). Dashed line = the fully-reproduced subset.
1273 processed total · 24/day average over 53 days
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
Agents read each paper, fetch its data, rerun the analysis on the brainbox compute, and grade what reproduces — end to end, without a human in the loop.
The pipeline, live
Reproductions run continuously. Here is what just finished — and the papers readers have asked for next.
- ✓ A comprehensive resource of genomic, epigenomic and tran...
- ✓ RNA-sequence analysis of primary alveolar macrophages af...
- ✓ Elucidation of the molecular responses to waterlogging i...
- ✓ Ancient gene duplicates in Gossypium (cotton) exhibit ne...
- ✓ A consensus approach to vertebrate de novo transcriptome...
- ✓ Bayesian transcriptome assembly.
- ✓ Revised annotations, sex-biased expression, and lineage-...
- ✓ TP53 engagement with the genome occurs in distinct local...
- ✓ Polymorphism identification and improved genome annotati...
- ✓ WikiPathways App for Cytoscape: Making biological pathwa...
- nothing yet
How it works
Each publication is read on two independent axes. The Level (L1–L4) says how rigorously it was examined; the 0–100 reproducibility score is the outcome of the attempt. Both are shown together (e.g. L3 · 82). A score is never a statement about the authors.
Level — how it was curated
- L1Fully automated, AI-curated re-run of the bioinformatic pipeline
- L2Human scientist plausibility check of the L1 AI curation
- L3Scientist-curated partial re-run of the bioinformatic pipeline
- L4Scientist-curated full re-run of the bioinformatic pipeline
Score — the reproducibility verdict
A 0–100 reproducibility score from hands-on re-running of the paper's computational claims — the share that held up, colour-banded:
- No computationPure wet-lab / in-vivo paper — nothing computational to recompute. Neutral: neither pass nor fail.
- InconclusiveThe reproduction could not be completed or scored (e.g. data unavailable, pipeline error). No conclusion is drawn.
- ⚑A separate, human-reviewed flag — independent of the score and never assigned automatically.
Every score stores an auditable, ACMG-style worksheet: each item that counted toward it, the exact part of the reproduction it came from, its points, and any decisive gate. Open any record to inspect it line by line.