Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Exploring microproteins from various model organisms using the mip-mining database.

BMC Genomics · 2023
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. The pipeline-derived in-scope claim is Case study 1 (budding yeast): three microprotein-encoding genes' limma Log2FC under acetic-acid stress from GEO GSE52160. I ran the authors' own pipeline (GlancerZ/Mipmining DEG_function.R limma block) on the full GSE52160 RNA-seq (11 paired MiSeq runs: sra-tools 2.10.9 -> fastp -> HISAT2 2.2.1 vs Ensembl R64-1-1 r110 -> StringTie 2.1.4 -e -B -> ballgown FPKM -> FPKM->TPM -> normalizeBetweenArrays -> log2 -> topTable coef=2). The paper omits the case/control split; of four plausible groupings I tested, the 6h-acetic-acid vs 0h-baseline contrast reproduces all three values essentially exactly: PMP2 +1.94->+1.9555, ATP15 -2.92->-2.9142, SDH6 -3.4->-3.4339 (all |diff|<=0.034, all adj.P<1e-3). No fabrication: every reported number is re-derivable from the cited data. NOT attempted (out of scope): the database-scale curation tables (Table 1: 336/360 datasets, 8626/8677 samples, 9-species stress tallies = manual GEO curation + MySQL website), GO/KEGG enrichment figures (no pinned numbers), and the UniProt <=100aa microprotein catalogue (only labels genes; the 3 Log2FC are gene-level). One inference made: the exact grouping (reported transparently, all 4 groupings shipped).

💻 Code ↗ 🗄 Data: GSE52160

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-14 ⛓ 60eb11bcf53c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Transcriptome data can be systematically mined to reveal microprotein functions in environmental stress responses and disease, a use case overlooked by existing microprotein databases; this is exemplified by testing whether ATP15 contributes to acetic acid stress tolerance in budding yeast.

Core claims
  • Mip-mining is a database of 336 curated RNA-seq datasets from 8626 samples across nine species, built specifically to explore microprotein functions under stress and disease conditions resource
  • Existing microprotein databases are limited in species coverage, lack personalized/customized analysis tools, largely overlook transcriptomic data, and do not focus on environmental stress or disease finding
  • Overexpression of ATP15 severely inhibits yeast growth under acetic acid stress, confirming its role in acetic acid stress tolerance finding
  • Deletion of PMP2 is known to decrease acetic acid resistance, validating Mip-mining's ability to identify microproteins with known stress-related roles finding
  • Mip-mining provides three major functions: browsing/searching primary data, differential expression identification with functional enrichment analysis, and result visualization/download method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (database-wide reprocessing) nine species including E. coli, S. cerevisiae, A. thaliana, O. sativa, C. elegans, D. rerio, D. melanogaster, M. musculus, H. sapiens various (chemical, physical, biological stress, disease) differential gene/microprotein expression (FPKM, Log2FC) Hisat2, StringTie, ballgown, limma
transcriptome differential expression analysis (Mip-mining case study) S. cerevisiae, dataset GSE52160 acetic acid stress Log2FC of microprotein-encoding genes (PMP2, ATP15, SDH6)
growth curve assay S. cerevisiae BY4741 strain (plasmid pJFE3) ATP15 overexpression (high-copy plasmid) vs empty vector control growth/biomass (OD600), lag phase duration under stress-free, low pH (2.3), and acetic acid (4.2 g/L) conditions
transcriptome differential expression analysis (Mip-mining case study) A. thaliana, dataset GSE116004 heat stress (37°C vs normal temperature) global differential gene transcription
Key results
  • ATP15 overexpression caused an approximately 24-hour longer lag phase under acetic acid stress compared to control ~24 h longer lag phase
  • ATP15 overexpression also reduced biomass under non-stress and low pH (2.3) conditions
  • ATP15 transcription changed in GSE52160 acetic acid stress analysis Log2FC = -2.92
  • PMP2 gene showed significant transcriptional change under acetic acid stress, consistent with known role in acetic acid resistance upon deletion Log2FC = 1.94
  • SDH6 gene showed significant transcriptional change under acetic acid stress Log2FC = -3.4
  • Mip-mining database compiles 336 RNA-seq datasets from 8626 samples across nine organisms
Key statistics
  • fold_change Log2FC = -2.92 (ATP15 expression change, GSE52160 acetic acid stress dataset)
  • fold_change Log2FC = 1.94 (PMP2 expression change, GSE52160 acetic acid stress dataset)
  • fold_change Log2FC = -3.4 (SDH6 expression change, GSE52160 acetic acid stress dataset)
  • count 336 datasets (total curated RNA-seq datasets and samples in Mip-mining across nine species)
  • other ~24 h longer lag phase (ATP15-overexpressing yeast strain growth delay under acetic acid stress)
  • other initial OD600 = 0.03 (starting inoculation density for yeast growth assay)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Mip-mining is a database paper describing an RNA-seq collection and analysis pipeline for microprotein mining across nine model organisms (336 datasets, 8626 samples). The primary statistical framework uses the limma R package (moderated t-statistics with empirical Bayes shrinkage) applied to FPKM expression matrices generated via a Hisat2–StringTie–ballgown pipeline to identify differentially expressed microprotein-encoding genes. Results are reported as Log2FC with raw and adjusted p-values and supplemented by GO, KEGG, and GSEA enrichment analyses. A yeast case study experimentally validates candidate microprotein ATP15 by comparing growth curves between an overexpression strain and a control under three conditions, with no formal statistical test reported for those comparisons.

Replicationunclear Sample size8626 total samples across 336 datasets from 9 species stated at database level; per-experiment biological replicate count not stated for any case study GroupsStress or disease condition vs. control across five stress categories and nine species; yeast case: ATP15 overexpression strain vs. empty-plasmid control under three growth conditions Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionAdjustment method not explicitly named; paper reports 'Adj.P.val' from limma output (limma's default is Benjamini-Hochberg FDR)
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test (empirical Bayes) Differential expression of microprotein-encoding genes across all 336 RNA-seq datasets and case studies not stated
GSEA (Gene Set Enrichment Analysis) Functional enrichment of differentially expressed microprotein gene sets not stated
GO/KEGG overrepresentation analysis Pathway and term enrichment for differentially expressed microproteins not stated
Principal component analysis (PCA) Sample distribution visualization and quality control na
No formal test stated Yeast growth curve comparison (Fig. 6): ATP15 overexpression vs. empty-plasmid control under stress-free, low-pH, and acetic acid conditions na
Approaches that could also have been used
  • Differential expression was computed from FPKM values using limma's standard linear-model pipeline
    Could also: DESeq2 or edgeR could also be applied, operating on raw integer read counts with negative binomial dispersion modeling; limma-voom is a closely related option that transforms counts to log-CPM with precision weights before applying the same linear-model framework — DESeq2, edgeR, and limma-voom are purpose-built for count-based RNA-seq data and are among the most extensively benchmarked tools for this purpose; for heterogeneous multi-study collections like this database, count-based methods or voom may provide more consistent variance stabilization across datasets
  • The multiple-testing adjustment is reported only as 'Adj.P.val' without naming the specific procedure
    Could also: Explicitly stating the correction method (e.g., Benjamini-Hochberg FDR) is also standard practice and is what most reporting guidelines recommend — Naming the procedure allows readers to assess false discovery rate control and to reproduce results exactly; BH FDR is limma's default and the most common choice for transcriptome-wide testing
  • Yeast growth curves comparing ATP15 overexpression vs. control across three conditions (Fig. 6) were presented visually without a stated formal statistical test
    Could also: A two-way ANOVA (strain × condition) with post-hoc pairwise comparisons, or area-under-the-growth-curve analysis followed by a t-test or Mann-Whitney U, could also be used — Formal tests quantify uncertainty around observed differences and provide p-values that complement visual inspection of growth curves, particularly when biological replicates are available
  • No dispersion measure (SD, SEM, or CI) is reported for the yeast growth experiment
    Could also: Reporting SD or SEM error bars alongside mean OD600 at each time point would also convey biological variability across replicates — Dispersion measures allow readers to assess reproducibility and gauge the magnitude of differences relative to natural variation within each condition
  • FPKM was used as the expression unit for the ballgown-to-limma pipeline
    Could also: TPM (Transcripts Per Million) is also widely used and avoids a known cross-sample comparability issue where FPKM values do not sum to the same total across samples — TPM values sum to a constant within each sample, making cross-sample comparisons more straightforward; several downstream tools and databases also assume TPM input
  • Sample size and statistical power are not described for the experimental case study validating ATP15
    Could also: Reporting the number of biological replicates per group, and optionally an effect-size estimate with confidence interval, would also be informative — Knowing n per group helps readers assess the sensitivity of the experiment to detect the observed growth differences and evaluate the generalizability of the findings
Software: R 3.6.1 · limma 3.42.0 · ballgown 2.18.0 · clusterProfiler 3.14.0 · enrichplot 1.6.0 · FactoMineR 2.4 · factoextra 1.0.7 · ggplot2 3.3.3 · ggrepel 0.9.1 · Hisat2 2.2.1 · StringTie 2.1.4 · FastQC 0.11.9 · fastp 0.19.5 · multiQC 1.9 · Sratoolkit (Prefetch / fasterq-dump) 2.10.9 · MySQL 5.5

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37919660

Paper: Zhao et al. 2023, BMC Genomics. "Exploring microproteins from various model organisms using the mip-mining database." PMID 37919660 / PMC10623795 / doi:10.1186/s12864-023-09735-1.

Real analysis code (not the BRIEF's sra-tools link): https://github.com/GlancerZ/Mipmining (MIT, default branch main). The repo ships the actual published DEG pipeline script/DEG_function.R plus one worked example (GSE21341_result/). The BRIEF's "code" field (ncbi/sra-tools) is just the first tool in the workflow; the authors' own filtering code is the repo above. Per BRIEF rule P16, running the authors' / pipeline tools on the paper's own data is a valid reproduction.

Pipeline (Methods + DEG_function.R)

sra-tools prefetch/fasterq-dump v2.10.9 -> fastp v0.19.5 / FastQC -> HISAT2 v2.2.1 align to species reference -> StringTie v2.1.4 (-e -B) -> ballgown v2.18.0 FPKM -> FPKM->TPM -> limma normalizeBetweenArrays -> log2(x+1) -> limma lmFit/eBayes topTable(coef=2) -> Log2FC per gene -> subset to microproteins (<=100 aa, UniProt). Enrichment (clusterProfiler/GO/KEGG) is downstream decoration, not a pinnable number.

IN SCOPE (pipeline-derived, pinnable)

Case Study 1 (budding yeast), GEO GSE52160 ("Tolerance to acetic acid is improved by mutations of the TATA-binding protein gene", S. cerevisiae RNA-seq, 4 biological samples AT-00/AT-06/AC-00/AC-06, 11 paired-end MiSeq runs SRR1026832..SRR1026842, BioProject PRJNA227050 / SRP032768).

Reported result (Results, Case study 1 / Table 2): three significantly altered microprotein-encoding genes under acetic acid stress, with limma Log2FC:

  • PMP2 Log2FC = +1.94
  • ATP15 Log2FC = -2.92
  • SDH6 Log2FC = -3.4 These three Log2FC values are the reproduction target. The Log2FC equals the gene-level limma topTable value (microprotein subsetting does not alter it), so gene_results.csv suffices.

OUT OF SCOPE / NOT ATTEMPTED

  • Database-scale tallies (Table 1: 336/360 datasets, 8626/8677 samples, 9 species, stress-type counts). These are manual GEO curation + a MySQL website, not a re-runnable pipeline -> non_pipeline, not attempted.
  • The web database itself (weilab.sjtu.edu.cn/mipmining), MySQL/Tomcat backend.
  • GO/KEGG enrichment figures for the case study (qualitative, no pinned numbers).
  • Reference microprotein catalogue construction from UniProt (<=100 aa) — only needed to label genes as microproteins; the 3 target Log2FC are gene-level.

Known reproduction ambiguity (the hard 20%)

GSE52160 = 2 strains (AT tolerant / AC control) x 2 timepoints (00h/06h), single biological replicate each (lane-split into 2-3 SRR runs). The paper does not state the exact case/control grouping fed to limma. We compute the 3 genes' Log2FC under the plausible groupings (strain AT-vs-AC; timepoint 06-vs-00) and report which matches, rather than guessing one.

Figures / tables: Table
MP_PMP2
Reported
1.94
Reproduced
1.9555
within tolerance
MP_ATP15
Reported
-2.92
Reproduced
-2.9142
within tolerance
MP_SDH6
Reported
-3.4
Reproduced
-3.4339
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

All three pipeline-derived claims (PMP2 +1.94→+1.9555, ATP15 -2.92→-2.9142, SDH6 -3.4→-3.4339) reproduce essentially exactly from the public GEO data GSE52160 via the authors' own GlancerZ/Mipmining limma pipeline, with residuals ≤0.034 log2 units consistent with reference/tool-version drift. The only friction is on the authors' side as an underspecification: the paper omits the case/control split, which the curator had to infer — but the inferred 6h-vs-0h contrast matched all three numbers and no alternative grouping did, confirming it. No fabrication concern; sign, magnitude, rank and significance (adj.P<1e-3) all hold. Database-scale tallies and enrichment figures were out of scope (manual curation, no pinned numbers).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

129.5 k
tokens (I/O) · 8.2 M incl. cache
22 min
runtime · 0.79 CPU-h
1.9 GB
peak RAM
1 (1 failed)
HPC jobs
hummel
machine