The archives are half-empty: an assessment of the availability of microbial community sequencing data.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough; 1:1 reproduction. Jurburg et al. is a meta-research/text-mining paper; its code+data are openly deposited on Zenodo (10.5281/zenodo.3953307 = github a-h-b/Data_availability_study @8db57c1, MD5-verified exact). The in-scope part is stage 5: the R statistical analysis on the shipped mined data. Re-running the authors' exact data-prep logic on a «our HPC» compute node (R 3.6.3; paper used 3.6.1) reproduced every headline number to the decimal: 2,015 / 2,656 study counts, 75.9% INSDC deposition, 7.2% not-public (n=146), the V3-V4 fate split 34% reusable (n=216) / 40.3% not-available (n=256) / 25.5% partial (n=162), and the error rates 11.8% one-fasta (n=52), 12% mislabeled amplicon (n=53), 16.8% mislabeled paired-end (n=74). 10/12 claims exact, 1 within-tol (C8 off by 1 study -- and the paper's own fate categories sum to 634 vs the 635 universe, a 1-study slack in the paper itself), 1 partial (C1 = upstream 26,927-PDF Web-of-Science corpus count, which is NOT in the deposit; the deposit ships 21,572 unique mined DOIs). NOT attempted (out of scope): the GROBID 0.5.4 text-mining of the copyrighted PDF corpus (data not shippable -- only its output is) and the read-level FastQC/cutadapt QC of the V3-V4 subset (would require re-downloading fastqs from SRA; its outputs are shipped in raw_data.txt/Metadata.RDS). No fabrication signal: every published number is derivable from the shipped data+code.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 95assessed: 2026-06-18 ⛓ 5e67b4866f5b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper tests whether 16S rRNA gene amplicon sequencing data archived in public genetic repositories (SRA, ENA, DRA) is actually available and reusable for future meta-analyses and synthesis efforts, despite reported deposition.
- ★ More than half of surveyed amplicon sequencing studies were affected by lack of data deposition, improper file formatting, or inconsistent labeling that impede reuse. finding
- ★ Only 34% of the V3-V4 subset studies contained fully reusable datasets, while 40.3% contained data that was not available or not reusable and 25.5% were partially available. finding
- ★ 7.2% of studies listed accession numbers correctly but had not made their sequence data public at the time of publication. finding
- ★ Major contributors to data loss are loss due to data location, errors in data deposition, errors in data formatting, and errors in data labeling. finding
- Deposition to non-INSDC alternative databases (Qiita, MG-RAST, figshare) risks long-term information loss because they are not designed for long-term archiving of amplicon data. mechanism
- ★ A custom pattern-based text extraction algorithm combined with manual curation was used to screen publications and identify 16S rRNA amplicon sequencing studies with INSDC accession numbers. method
- The study provides concrete recommendations (Table 1) for improving data archiving practices. resource
- Legacy non-demultiplexed single-file deposits and missing quality scores render sequence data unusable on INSDC platforms. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Literature text-mining / bibliometric survey | 17 microbiology journals, 26,927 articles (Jan 2015–Mar 2019) | none | Identification of 16S rRNA amplicon sequencing studies and INSDC accession numbers | Custom pattern-based text extraction algorithm |
| Manual curation/inspection of articles | 150 randomly selected articles mentioning 16S rRNA without detected accession numbers | none | Verification of missed accession numbers and data deposition status | — |
| Repository data availability assessment | INSDC databases (SRA, ENA, DRA) for 2015 V3-V4 subset studies | none | Findability/accessibility of accession numbers and public vs private data status | SRA, ENA (ENA), DRA |
| Sequence file format inspection | 441 V3-V4 subset studies with public data (45,440 samples) | none | Demultiplexing status, single-file deposits, quality scores, primer presence | Illumina / 454 pyrosequencing platforms |
| Metadata/labeling inspection | V3-V4 subset studies with accession numbers | none | Correctness of 'Amplicon' labeling and paired-end read file labeling | — |
- – Of 2015 studies with accession numbers, 7.2% (n=146) had not made sequence data public despite correct accession numbers 7.2% (n=146)
- – In the V3-V4 subset, 40.3% (n=256) contained data not available or not reusable 40.3% (n=256)
- – In the V3-V4 subset, only 34% (n=216) contained fully reusable datasets 34% (n=216)
- – 25.5% (n=162) of V3-V4 subset studies contained partially available datasets 25.5% (n=162)
- – Of 2,656 studies employing 16S amplicon sequencing, 75.9% deposited data to an INSDC database 75.9%
- ▲ Proportion of V3-V4 studies claiming INSDC deposition rose from 33/56 (2015) to 172/214 (2018) 33/56 to 172/214
- ▼ 11.8% (n=52) of public V3-V4 studies uploaded a single sequence file despite multiple samples; this decreased from 24.5% to 9.5% over time 24.5% to 9.5%
- – 12% (n=53) of V3-V4 studies incorrectly labeled sequences using terms other than 'Amplicon'; 16.8% (n=74) mislabeled paired-end files 12% (n=53); 16.8% (n=74)
- count 26,927 articles screened; 2015 16S studies identified; 145,203 samples (Initial literature survey of 17 journals)
- count 635 studies in V3-V4 subset (515-806 region) (Focused reusability subset)
- pvalue χ2 = 6.6, p = 0.01 (Increase in INSDC deposition over time (V3-V4))
- pvalue χ2 = 14.04, p < 0.001 (Decrease in deposition to alternative databases over time)
- pvalue χ2 = 16.92, p < 0.001 (Decline in single-sequence-file deposits over time)
- pvalue χ2 = 9.18, p < 0.001 (Increase in incorrect accession numbers from 1.3% (2015) to 5.3% (2019))
- other 2.2% (n=45) listed incorrect accession numbers; 2.5% (n=51) did not make metadata public (Data deposition errors)
- count 441 V3-V4 studies with public data representing 45,440 samples (Format inspection subset)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper is a cross-sectional audit of 26,927 publications across 17 microbiology journals (January 2015–March 2019), combining automated text-parsing with manual curation to identify 2,015 studies reporting INSDC accession numbers for 16S rRNA amplicon sequencing, with a focused sub-analysis of 635 studies targeting the V3–V4 region. Data availability, formatting, and labeling were categorized descriptively as proportions. Temporal trends in those proportions across years were evaluated with chi-squared tests for trend, with results reported as percentages and chi-squared statistics accompanied by p-values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Chi-squared test for trend in proportions | Proportion of studies depositing data to INSDC databases over time (2015–2019); χ²=6.6, p=0.01 | V3–V4 subset, n=635 studies | not stated |
| Chi-squared test for trend in proportions | Proportion of studies depositing to alternative (non-INSDC) databases over time; χ²=14.04, p<0.001 | V3–V4 subset, n=635 studies | not stated |
| Chi-squared test for trend in proportions | Proportion of studies without any publicly available data over time; χ²<0.28, p=0.6 | V3–V4 subset, n=635 studies | not stated |
| Chi-squared test for trend in proportions | Proportion of studies with incorrect/placeholder accession numbers over time; χ²=9.18, p<0.001 | Studies with reported accession numbers, n=2015 | not stated |
| Chi-squared test for trend in proportions | Proportion of studies with data not yet made public; χ²=3.9, p=0.05; and with non-public metadata; χ²=14.83, p<0.001 | Studies with reported accession numbers, n=2015 | not stated |
| Chi-squared test for trend in proportions | Proportion of studies using single-file deposition (χ²=16.92, p<0.001), Illumina platforms (χ²=10.96, p<0.001), 454 pyrosequencing (χ²=10.46, p=0.001), primer presence (χ²=2.33, p=0.13), amplicon mislabeling (χ²=1.41, p=0.24), and paired-end labeling errors (χ²=2.09, p=0.15) over time | V3–V4 studies with public INSDC data, n=441 | not stated |
-
Temporal trends in proportions across years were assessed with chi-squared tests↳ Could also: The Cochran-Armitage test for trend could also be used — The Cochran-Armitage test is designed specifically for testing a monotonic linear trend in binomial proportions across ordered categories (years), yielding a single directional statistic that makes the ordinal nature of time explicit, whereas the general chi-squared test treats year categories as unordered
-
Approximately 12 chi-squared tests were conducted across multiple outcome proportions and overlapping subsets without a multiplicity adjustment↳ Could also: A Benjamini-Hochberg false discovery rate (FDR) correction or a Bonferroni adjustment could also be applied to the family of trend tests — Explicitly controlling the FDR or family-wise error rate when testing many related hypotheses is a common practice that makes the probability of spurious findings transparent; its absence or presence is a choice worth noting in an audit study with many parallel comparisons
-
Key proportions (e.g., 40.3% not available or not reusable; 18% not depositing data at all) were reported as point estimates without uncertainty bounds↳ Could also: Wilson score or Clopper-Pearson confidence intervals for each proportion could also be reported — Confidence intervals convey the precision of proportion estimates, which is especially informative for the inferred 18% figure derived from a 150-article random subsample, where sampling variability is non-negligible
-
The proportion of studies that performed sequencing but provided no data link was estimated by manual review of a convenience sample of 150 articles flagged by the algorithm↳ Could also: A pre-specified stratified random sampling scheme with a formal sample-size calculation (targeting a given margin of error) could also be used — A sample-size-justified design would make the precision target of the extrapolated estimate explicit and reproducible, and would support a formal confidence interval around the 18% figure
-
Each data-quality dimension (deposition, formatting, labeling) was analyzed with separate trend tests↳ Could also: Logistic regression with year as a continuous predictor could also be used for each outcome — Treating year as continuous in a logistic model yields an odds ratio per year as an interpretable effect size, allows adjustment for covariates (e.g., journal, sequencing platform), and is consistent with modeling year as an ordinal rather than categorical variable
-
Study-level data-quality categories (reusable / partially usable / not available) were reported as counts and percentages without uncertainty quantification↳ Could also: A multinomial confidence interval (e.g., Sison-Glaz or bootstrap-based) for the joint distribution across the three categories could also be reported — Simultaneous confidence intervals for a multinomial partition make the joint uncertainty of the three-category breakdown explicit, which is useful when the categories are mutually exclusive and collectively exhaustive and readers wish to compare proportions across them
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32859925
Paper: Jurburg, Konzack, Eisenhauer, Heintz-Buschart (2020). The archives are half-empty: an assessment of the availability of microbial community sequencing data. Commun Biol 3:474. DOI 10.1038/s42003-020-01204-9. PMCID PMC7455719.
Type: meta-research / bibliometric text-mining study. The authors mined the literature (16S rRNA amplicon-sequencing studies, 17 journals, 2015–Mar 2019) for INSDC accession numbers, then checked whether the deposited data was actually present and reusable.
Artifacts
- Code + data: Zenodo
10.5281/zenodo.3953307(v1.0.1) = GitHuba-h-b/Data_availability_study@ commit8db57c1. One 5.6 MB zip, MD5c6518e1eb4f0246cfd9d3c9173ee8869(verified — exact match). - The
code_urlin the registry pointed atkermitt2/grobid; that is only the upstream dependency (GROBID v0.5.4, used to turn PDFs into TEI XML). The authors' own analysis code + the mined data live in the Zenodo deposit. Per brief rule P16, applying/here re-running the authors' shipped pipeline on the shipped data is fully valid.
Pipeline stages and scope decision
| Stage | Tool | In scope? | Why |
|---|---|---|---|
| 1. Literature corpus (26,927 WoS articles → PDFs) | Web of Science | OUT | the raw PDFs are copyrighted and are NOT in the deposit; cannot be re-fetched at scale |
| 2. PDF → TEI XML → accession/primer mining | GROBID 0.5.4 + custom python (pdf2seqs.ipynb) |
OUT | needs the stage-1 PDF corpus; only the output of this stage (mining_output.txt, the RDS tables) is shipped |
| 3. Run-level metadata mining | NCBI Entrez Direct | OUT (output shipped) | network-/time-dependent; outputs captured in shipped RDS |
| 4. Read-level QC of V3–V4 subset (1000 reads/fastq) | FastQC 0.11.3 + cutadapt 1.18 | OUT (not attempted) | would require re-downloading fastqs from SRA; its outputs are in raw_data.txt/Metadata.RDS. A possible future extension. |
| 5. Statistical analysis + all reported figures/numbers | R 3.6.1, R_Notebook_Analyses_main.Rmd |
IN SCOPE | base-R analysis on the shipped mined tables; regenerates every headline number |
In-scope target results (reproduced)
All headline quantitative claims are produced by stage 5 on the shipped data:
2,015 / 2,656 study counts, 75.9% INSDC deposition, 7.2% not-public, the V3–V4
subset (n=635) fate split (34% reusable / 40.3% not-available / 25.5% partial),
and the data-formatting/labeling error rates (11.8% one-fasta, 12% mislabeled
amplicon, 16.8% mislabeled paired-end). See original/claims.tsv.
Method
Re-ran the authors' exact data-prep logic from R_Notebook_Analyses_main.Rmd
(verbatim transforms, only the ggplot() calls skipped) on the shipped RDS/txt
inputs, using R 3.6.3 (paper used 3.6.1) on a «our HPC» compute node (SLURM «job» + 2194300, partition std). Scripts: reproduction/outputs/reproduce*.R.
Not claimed
Stages 1–4 are not reproduced. We do not re-verify that the mined values
(mining_output.txt, the TF flags in the RDS files) faithfully reflect the
26,927 PDFs — that would require the copyrighted corpus. We reproduce that the
published numbers follow from the shipped data + code, which they do, to the
decimal. Whether the shipped flags themselves are correct is an upstream
question outside this deposit.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a near-perfect 1:1 reproduction of a meta-research/text-mining paper: the authors' analysis code and mined data are openly deposited on Zenodo and MD5-verified, and re-running their R pipeline regenerates all 11 in-scope headline numbers (10 exact, 1 within-tolerance). The only deviation, C8's off-by-one study (n=257 vs 256), is a rounding/classification slack present in the paper itself (its three fate categories sum to 634 against a 635 universe), and C1's mismatch is merely an upstream Web-of-Science corpus count (26,927) that was never deposited — context-only, not a reproduced claim. No discrepancy lies on the authors' computation side and there is no fabrication signal; the deviations are technical/expected (R version, rounding).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.