Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The archives are half-empty: an assessment of the availability of microbial community sequencing data.

Commun Biol · 2020
L1 95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough; 1:1 reproduction. Jurburg et al. is a meta-research/text-mining paper; its code+data are openly deposited on Zenodo (10.5281/zenodo.3953307 = github a-h-b/Data_availability_study @8db57c1, MD5-verified exact). The in-scope part is stage 5: the R statistical analysis on the shipped mined data. Re-running the authors' exact data-prep logic on a «our HPC» compute node (R 3.6.3; paper used 3.6.1) reproduced every headline number to the decimal: 2,015 / 2,656 study counts, 75.9% INSDC deposition, 7.2% not-public (n=146), the V3-V4 fate split 34% reusable (n=216) / 40.3% not-available (n=256) / 25.5% partial (n=162), and the error rates 11.8% one-fasta (n=52), 12% mislabeled amplicon (n=53), 16.8% mislabeled paired-end (n=74). 10/12 claims exact, 1 within-tol (C8 off by 1 study -- and the paper's own fate categories sum to 634 vs the 635 universe, a 1-study slack in the paper itself), 1 partial (C1 = upstream 26,927-PDF Web-of-Science corpus count, which is NOT in the deposit; the deposit ships 21,572 unique mined DOIs). NOT attempted (out of scope): the GROBID 0.5.4 text-mining of the copyrighted PDF corpus (data not shippable -- only its output is) and the read-level FastQC/cutadapt QC of the V3-V4 subset (would require re-downloading fastqs from SRA; its outputs are shipped in raw_data.txt/Metadata.RDS). No fabrication signal: every published number is derivable from the shipped data+code.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.3953307

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 95
    assessed: 2026-06-18 ⛓ 5e67b4866f5b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper tests whether 16S rRNA gene amplicon sequencing data archived in public genetic repositories (SRA, ENA, DRA) is actually available and reusable for future meta-analyses and synthesis efforts, despite reported deposition.

Core claims
  • More than half of surveyed amplicon sequencing studies were affected by lack of data deposition, improper file formatting, or inconsistent labeling that impede reuse. finding
  • Only 34% of the V3-V4 subset studies contained fully reusable datasets, while 40.3% contained data that was not available or not reusable and 25.5% were partially available. finding
  • 7.2% of studies listed accession numbers correctly but had not made their sequence data public at the time of publication. finding
  • Major contributors to data loss are loss due to data location, errors in data deposition, errors in data formatting, and errors in data labeling. finding
  • Deposition to non-INSDC alternative databases (Qiita, MG-RAST, figshare) risks long-term information loss because they are not designed for long-term archiving of amplicon data. mechanism
  • A custom pattern-based text extraction algorithm combined with manual curation was used to screen publications and identify 16S rRNA amplicon sequencing studies with INSDC accession numbers. method
  • The study provides concrete recommendations (Table 1) for improving data archiving practices. resource
  • Legacy non-demultiplexed single-file deposits and missing quality scores render sequence data unusable on INSDC platforms. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Literature text-mining / bibliometric survey 17 microbiology journals, 26,927 articles (Jan 2015–Mar 2019) none Identification of 16S rRNA amplicon sequencing studies and INSDC accession numbers Custom pattern-based text extraction algorithm
Manual curation/inspection of articles 150 randomly selected articles mentioning 16S rRNA without detected accession numbers none Verification of missed accession numbers and data deposition status
Repository data availability assessment INSDC databases (SRA, ENA, DRA) for 2015 V3-V4 subset studies none Findability/accessibility of accession numbers and public vs private data status SRA, ENA (ENA), DRA
Sequence file format inspection 441 V3-V4 subset studies with public data (45,440 samples) none Demultiplexing status, single-file deposits, quality scores, primer presence Illumina / 454 pyrosequencing platforms
Metadata/labeling inspection V3-V4 subset studies with accession numbers none Correctness of 'Amplicon' labeling and paired-end read file labeling
Key results
  • Of 2015 studies with accession numbers, 7.2% (n=146) had not made sequence data public despite correct accession numbers 7.2% (n=146)
  • In the V3-V4 subset, 40.3% (n=256) contained data not available or not reusable 40.3% (n=256)
  • In the V3-V4 subset, only 34% (n=216) contained fully reusable datasets 34% (n=216)
  • 25.5% (n=162) of V3-V4 subset studies contained partially available datasets 25.5% (n=162)
  • Of 2,656 studies employing 16S amplicon sequencing, 75.9% deposited data to an INSDC database 75.9%
  • Proportion of V3-V4 studies claiming INSDC deposition rose from 33/56 (2015) to 172/214 (2018) 33/56 to 172/214
  • 11.8% (n=52) of public V3-V4 studies uploaded a single sequence file despite multiple samples; this decreased from 24.5% to 9.5% over time 24.5% to 9.5%
  • 12% (n=53) of V3-V4 studies incorrectly labeled sequences using terms other than 'Amplicon'; 16.8% (n=74) mislabeled paired-end files 12% (n=53); 16.8% (n=74)
Key statistics
  • count 26,927 articles screened; 2015 16S studies identified; 145,203 samples (Initial literature survey of 17 journals)
  • count 635 studies in V3-V4 subset (515-806 region) (Focused reusability subset)
  • pvalue χ2 = 6.6, p = 0.01 (Increase in INSDC deposition over time (V3-V4))
  • pvalue χ2 = 14.04, p < 0.001 (Decrease in deposition to alternative databases over time)
  • pvalue χ2 = 16.92, p < 0.001 (Decline in single-sequence-file deposits over time)
  • pvalue χ2 = 9.18, p < 0.001 (Increase in incorrect accession numbers from 1.3% (2015) to 5.3% (2019))
  • other 2.2% (n=45) listed incorrect accession numbers; 2.5% (n=51) did not make metadata public (Data deposition errors)
  • count 441 V3-V4 studies with public data representing 45,440 samples (Format inspection subset)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper is a cross-sectional audit of 26,927 publications across 17 microbiology journals (January 2015–March 2019), combining automated text-parsing with manual curation to identify 2,015 studies reporting INSDC accession numbers for 16S rRNA amplicon sequencing, with a focused sub-analysis of 635 studies targeting the V3–V4 region. Data availability, formatting, and labeling were categorized descriptively as proportions. Temporal trends in those proportions across years were evaluated with chi-squared tests for trend, with results reported as percentages and chi-squared statistics accompanied by p-values.

Replicationunclear Sample size26,927 total publications screened; 2,015 with INSDC accession numbers; V3–V4 subset n=635; n=441 with publicly accessible INSDC data; 150 articles randomly sampled to validate the parsing algorithm GroupsStudies grouped by year (2015–2019) for temporal trend analyses; studies categorized by data availability/format/labeling status Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Chi-squared test for trend in proportions Proportion of studies depositing data to INSDC databases over time (2015–2019); χ²=6.6, p=0.01 V3–V4 subset, n=635 studies not stated
Chi-squared test for trend in proportions Proportion of studies depositing to alternative (non-INSDC) databases over time; χ²=14.04, p<0.001 V3–V4 subset, n=635 studies not stated
Chi-squared test for trend in proportions Proportion of studies without any publicly available data over time; χ²<0.28, p=0.6 V3–V4 subset, n=635 studies not stated
Chi-squared test for trend in proportions Proportion of studies with incorrect/placeholder accession numbers over time; χ²=9.18, p<0.001 Studies with reported accession numbers, n=2015 not stated
Chi-squared test for trend in proportions Proportion of studies with data not yet made public; χ²=3.9, p=0.05; and with non-public metadata; χ²=14.83, p<0.001 Studies with reported accession numbers, n=2015 not stated
Chi-squared test for trend in proportions Proportion of studies using single-file deposition (χ²=16.92, p<0.001), Illumina platforms (χ²=10.96, p<0.001), 454 pyrosequencing (χ²=10.46, p=0.001), primer presence (χ²=2.33, p=0.13), amplicon mislabeling (χ²=1.41, p=0.24), and paired-end labeling errors (χ²=2.09, p=0.15) over time V3–V4 studies with public INSDC data, n=441 not stated
Approaches that could also have been used
  • Temporal trends in proportions across years were assessed with chi-squared tests
    Could also: The Cochran-Armitage test for trend could also be used — The Cochran-Armitage test is designed specifically for testing a monotonic linear trend in binomial proportions across ordered categories (years), yielding a single directional statistic that makes the ordinal nature of time explicit, whereas the general chi-squared test treats year categories as unordered
  • Approximately 12 chi-squared tests were conducted across multiple outcome proportions and overlapping subsets without a multiplicity adjustment
    Could also: A Benjamini-Hochberg false discovery rate (FDR) correction or a Bonferroni adjustment could also be applied to the family of trend tests — Explicitly controlling the FDR or family-wise error rate when testing many related hypotheses is a common practice that makes the probability of spurious findings transparent; its absence or presence is a choice worth noting in an audit study with many parallel comparisons
  • Key proportions (e.g., 40.3% not available or not reusable; 18% not depositing data at all) were reported as point estimates without uncertainty bounds
    Could also: Wilson score or Clopper-Pearson confidence intervals for each proportion could also be reported — Confidence intervals convey the precision of proportion estimates, which is especially informative for the inferred 18% figure derived from a 150-article random subsample, where sampling variability is non-negligible
  • The proportion of studies that performed sequencing but provided no data link was estimated by manual review of a convenience sample of 150 articles flagged by the algorithm
    Could also: A pre-specified stratified random sampling scheme with a formal sample-size calculation (targeting a given margin of error) could also be used — A sample-size-justified design would make the precision target of the extrapolated estimate explicit and reproducible, and would support a formal confidence interval around the 18% figure
  • Each data-quality dimension (deposition, formatting, labeling) was analyzed with separate trend tests
    Could also: Logistic regression with year as a continuous predictor could also be used for each outcome — Treating year as continuous in a logistic model yields an odds ratio per year as an interpretable effect size, allows adjustment for covariates (e.g., journal, sequencing platform), and is consistent with modeling year as an ordinal rather than categorical variable
  • Study-level data-quality categories (reusable / partially usable / not available) were reported as counts and percentages without uncertainty quantification
    Could also: A multinomial confidence interval (e.g., Sison-Glaz or bootstrap-based) for the joint distribution across the three categories could also be reported — Simultaneous confidence intervals for a multinomial partition make the joint uncertainty of the three-category breakdown explicit, which is useful when the categories are mutually exclusive and collectively exhaustive and readers wish to compare proportions across them
Software: custom pattern-based text extraction algorithm (bespoke, no named package cited)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32859925

Paper: Jurburg, Konzack, Eisenhauer, Heintz-Buschart (2020). The archives are half-empty: an assessment of the availability of microbial community sequencing data. Commun Biol 3:474. DOI 10.1038/s42003-020-01204-9. PMCID PMC7455719.

Type: meta-research / bibliometric text-mining study. The authors mined the literature (16S rRNA amplicon-sequencing studies, 17 journals, 2015–Mar 2019) for INSDC accession numbers, then checked whether the deposited data was actually present and reusable.

Artifacts

  • Code + data: Zenodo 10.5281/zenodo.3953307 (v1.0.1) = GitHub a-h-b/Data_availability_study @ commit 8db57c1. One 5.6 MB zip, MD5 c6518e1eb4f0246cfd9d3c9173ee8869 (verified — exact match).
  • The code_url in the registry pointed at kermitt2/grobid; that is only the upstream dependency (GROBID v0.5.4, used to turn PDFs into TEI XML). The authors' own analysis code + the mined data live in the Zenodo deposit. Per brief rule P16, applying/here re-running the authors' shipped pipeline on the shipped data is fully valid.

Pipeline stages and scope decision

Stage Tool In scope? Why
1. Literature corpus (26,927 WoS articles → PDFs) Web of Science OUT the raw PDFs are copyrighted and are NOT in the deposit; cannot be re-fetched at scale
2. PDF → TEI XML → accession/primer mining GROBID 0.5.4 + custom python (pdf2seqs.ipynb) OUT needs the stage-1 PDF corpus; only the output of this stage (mining_output.txt, the RDS tables) is shipped
3. Run-level metadata mining NCBI Entrez Direct OUT (output shipped) network-/time-dependent; outputs captured in shipped RDS
4. Read-level QC of V3–V4 subset (1000 reads/fastq) FastQC 0.11.3 + cutadapt 1.18 OUT (not attempted) would require re-downloading fastqs from SRA; its outputs are in raw_data.txt/Metadata.RDS. A possible future extension.
5. Statistical analysis + all reported figures/numbers R 3.6.1, R_Notebook_Analyses_main.Rmd IN SCOPE base-R analysis on the shipped mined tables; regenerates every headline number

In-scope target results (reproduced)

All headline quantitative claims are produced by stage 5 on the shipped data: 2,015 / 2,656 study counts, 75.9% INSDC deposition, 7.2% not-public, the V3–V4 subset (n=635) fate split (34% reusable / 40.3% not-available / 25.5% partial), and the data-formatting/labeling error rates (11.8% one-fasta, 12% mislabeled amplicon, 16.8% mislabeled paired-end). See original/claims.tsv.

Method

Re-ran the authors' exact data-prep logic from R_Notebook_Analyses_main.Rmd (verbatim transforms, only the ggplot() calls skipped) on the shipped RDS/txt inputs, using R 3.6.3 (paper used 3.6.1) on a «our HPC» compute node (SLURM «job» + 2194300, partition std). Scripts: reproduction/outputs/reproduce*.R.

Not claimed

Stages 1–4 are not reproduced. We do not re-verify that the mined values (mining_output.txt, the TF flags in the RDS files) faithfully reflect the 26,927 PDFs — that would require the copyrighted corpus. We reproduce that the published numbers follow from the shipped data + code, which they do, to the decimal. Whether the shipped flags themselves are correct is an upstream question outside this deposit.

C2
Reported
2,015 16S studies with accessions
Reproduced
2015
exact
C3
Reported
2,656 studies employing 16S amplicon seq
Reproduced
2656
exact
C4
Reported
75.9% deposited to INSDC
Reproduced
75.87%
exact
C5
Reported
7.2% (n=146) not made public
Reproduced
7.25% (n=146)
exact
C6
Reported
V3-V4 subset n=635
Reproduced
635
exact
C7
Reported
34% (n=216) fully reusable
Reproduced
34.0% (216)
exact
C8
Reported
40.3% (n=256) not available/reusable
Reproduced
40.5% (257)
within tolerance
C9
Reported
25.5% (n=162) partially available
Reproduced
25.5% (162)
exact
C11
Reported
11.8% (n=52) single sequence file
Reproduced
11.8% (52)
exact
C12
Reported
12% (n=53) mislabeled (not 'Amplicon')
Reproduced
12% (53)
exact
C13
Reported
16.8% (n=74) mislabeled paired-end
Reproduced
16.8% (74)
exact
C1
Reported
26,927 WoS articles (corpus)
Reproduced
21,572 unique DOIs in deposit
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is a near-perfect 1:1 reproduction of a meta-research/text-mining paper: the authors' analysis code and mined data are openly deposited on Zenodo and MD5-verified, and re-running their R pipeline regenerates all 11 in-scope headline numbers (10 exact, 1 within-tolerance). The only deviation, C8's off-by-one study (n=257 vs 256), is a rounding/classification slack present in the paper itself (its three fate categories sum to 634 against a 635 universe), and C1's mismatch is merely an upstream Web-of-Science corpus count (26,927) that was never deposited — context-only, not a reproduced claim. No discrepancy lies on the authors' computation side and there is no fabrication signal; the deviations are technical/expected (R version, rounding).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

181.9 k
tokens (I/O) · 12.6 M incl. cache
17 min
runtime · 0 CPU-h
0.1 GB
peak RAM
2
HPC jobs
hummel
machine