Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Current status of use of high throughput nucleotide sequencing in rheumatology.

RMD Open · 2021
L1 83/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
83/100
Reproducibility score
0.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 61% of all assessed papers rank 430 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduces 1:1 for the core. This is a bibliometric/literature-review paper (PubMed review + SRA-metadata mining), so reproduction = independently recomputing the abstract's headline numbers from the shipped curated tables and rerunning the shipped Python aggregator. ALL headline abstract numbers reproduce EXACTLY: 699 identified articles, RA 182 (26%), SLE 161 (23%), OA 152 (22%), RNA-Seq 457 (65%). The SRA pipeline (parse_SRA_Metadata.py) rerun on the shipped *_all_old_meta.tsv reproduces the Figure-4 sample totals byte-exactly for 13 of 14 diseases across all 8 aggregation tables (assay/instrument/layout/source/tissue/organism/phenotype/studies). Two honest gaps: (1) SLE's SRA total (shipped figure 3785 vs 1789 from the single SLE_all_old_meta.tsv shipped) — the published SLE aggregate used a fuller SLE metadata download than the file committed to the repo; (2) the live PubMed step counts (1097/1162) are date-stamped 2020 searches that drift upward over time (date-pinned recompute 1213/1381, live today 4578/5474) and are not bit-reproducible for any PubMed query. NOT attempted (out of scope, non-pipeline): the manual curation itself (per-paper assay assignment, the ~190-PMID delete list, manual SRA additions) and the SRA Run Selector download step — these are human/manual, not code; we verified the shipped curated outputs are internally consistent rather than re-deriving the curation. No fabrication signal: every checkable number is exactly derivable from the shipped data.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 83
    assessed: 2026-06-18 ⛓ ef989b45d9bc
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper assesses the extent to which high throughput sequencing (HTS) has been adopted in rheumatic research and the availability of public HTS data of rheumatic samples, hypothesizing that rheumatology is becoming a major disease domain for HTS.

Core claims
  • RNA-Seq is the most represented HTS assay in rheumatology research (n=457, 65%), used for biomarker identification in blood or synovial tissue finding
  • Rheumatoid arthritis, systemic lupus erythematosus and osteoarthritis are the rheumatic diseases with the most reported use of HTS assays finding
  • Use of HTS in rheumatology research is growing exponentially, indicating rheumatic diseases are becoming the third disease domain for HTS after cancer and Mendelian diseases finding
  • The quality of clinical characterisation accompanying sequenced patients in public HTS data differs dramatically and is often incomplete finding
  • A semiautomated literature review combining an R-script (easyPubMed) with manual curation plus manual SRA search was used to quantify HTS adoption method
  • The authors propose a minimal set of clinical data necessary to accompany rheumatological-relevant HTS data resource
  • Public raw HTS data are frequently unavailable or poorly documented, with many publications providing no information on raw data availability finding
  • A curated list of all identified publications is available on GitHub resource
Experimental setups
Assay System Perturbation Readout Platform
Semiautomated PubMed literature review PubMed bibliographic database none Number of HTS publications by disease, assay, year, journal R V.3.6.1 with easyPubMed V.2.13
Public sequencing data survey Sequence Read Archive (SRA) none Number of HTS samples/projects, tissue source, platform, phenotype annotation SRA Run Selector, custom python script, pysradb
Disease-name filtering ICD-11 official disease names intersected with PubMed keywords none Identified rheumatic disease names
Clinical metadata quality assessment SLE RNA-Seq SRA projects and associated publications none Completeness/quality of clinical patient characterisation
Key results
  • 699 unique publications (813 total records) identified using HTS in rheumatic diseases 699
  • RNA-Seq is the most represented assay n=457, 65%
  • RA most common disease (n=182, 26%), then SLE (n=161, 23%) and OA (n=152, 22%) 182 (26%)
  • HTS publications increased from 18 in 2014 to 123 in 2018 and 189 in 2019 18 to 189
  • 17 023 HTS samples found in 296 SRA projects across rheumatic diseases 17023 samples
  • RA samples dominate SRA (n=8483, 50%), then SLE (n=3785, 22%) and OA (n=1386, 8%) 8483 (50%)
  • For 6305 (37%) SRA samples no phenotype or disease state was defined in metadata 6305 (37%)
  • Of 70 manually examined SLE RNA-Seq publications, 32 provided no information on raw data availability 32/70
Key statistics
  • count 457 (65%) (RNA-Seq studies among unique publications)
  • count 182 (26%) (RA publications)
  • count 161 (23%) (SLE publications)
  • count 152 (22%) (OA publications)
  • count 17023 (HTS samples in SRA across 296 projects)
  • count 15414 (90.5%) (SRA samples from human biomaterial)
  • count 13063 (77%) (samples sequenced on Illumina HiSeq series)
  • count 40 (9%) (RNA-Seq studies using scRNA-Seq)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a descriptive bibliometric and scoping review quantifying the adoption of high-throughput sequencing (HTS) in rheumatological research. A two-step semi-automated PubMed search using an R script was followed by manual curation; the Sequence Read Archive (SRA) was searched manually with a custom Python script. Results were reported exclusively as counts and percentages, with one median stated for SRA sample counts per study. No inferential statistical tests were performed.

Replicationna Sample size699 unique publications after two-step automated search and manual curation; 17,023 SRA samples across 296 projects; no formal sample-size or power calculation (not applicable for a bibliometric review) GroupsRheumatic diseases and HTS assay types compared by publication count and SRA sample count; journals compared by frequency of HTS publication Pairingna Randomization/blindingnot stated Dispersionnone
Statistical tests used
Test Applied to n Assumptions
Descriptive counts and percentages Publications per disease, assay type, journal, and year; SRA samples per disease and assay type 699 unique publications; 17,023 SRA samples across 296 projects na
Median Number of HTS samples per SRA study within each disease category 296 SRA projects na
Informal exponential growth projection Projected total HTS publications for full-year 2020 extrapolated from partial-year count (supplemental figure S1) 180 publications as of 4 September 2020 not stated
Approaches that could also have been used
  • Year-on-year growth in HTS publications was described narratively as 'exponential' and a single informal end-of-year projection was made for 2020
    Could also: A formal regression model — e.g., Poisson or negative-binomial regression fitted to annual publication counts, or a log-linear least-squares fit — could be used to estimate a growth rate with a confidence interval — A formal model would quantify the growth rate, provide uncertainty bounds on the projection, and make the trend reproducible by other reviewers
  • The median number of samples per SRA project within each disease category was reported without any measure of spread
    Could also: The median could be accompanied by the interquartile range (IQR) or the full range — The text notes that sample counts 'vary dramatically'; an IQR would formally quantify that variability and make the distribution summary more complete, which is standard practice when reporting medians
  • Manual curation classified publications by disease, assay, and other attributes, but no inter-rater reliability statistic was reported
    Could also: A random subset of records could be independently classified by a second reviewer with Cohen's kappa or percentage agreement reported — Reporting inter-rater reliability is a standard element of systematic and scoping review methodology that documents the reproducibility of the classification step
  • Publication counts across diseases and assay types were compared using raw numbers and percentages only
    Could also: Proportions could also be compared with chi-square tests or Fisher's exact tests (for sparse cells), potentially adjusted for overall disease burden or research output — Formal tests would allow inference about whether certain diseases or assays are disproportionately represented beyond what chance variation would predict
  • Clinical-data reporting quality for SLE RNA-Seq studies was characterised using four informal descriptive tiers (none, rudimentary, medium, detailed) applied to 23 projects
    Could also: A standardised completeness score — e.g., percentage of predefined checklist items present per study — could be computed and summarised with median and IQR across all studies — A quantitative completeness metric would be more reproducible, allow comparisons across diseases or time periods, and provide a measurable baseline for future audits
  • The linkage analysis between SRA-deposited projects and PubMed publications was performed in depth only for one disease-assay combination (SLE/RNA-Seq)
    Could also: A structured data-sharing audit applying the same criteria across all major disease-assay combinations could be performed and data-sharing rates reported with proportions and confidence intervals — Extending the audit beyond the single illustrative example would allow a quantitative, generalisable estimate of data-sharing practices across the full field of rheumatological HTS research
Software: R/easyPubMed R V.3.6.1; easyPubMed V.2.13 · R/ggplot2 V.3.2.1 · Python/pysradb · Python (custom SRA metadata extraction scripts)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 33408124

Title: Current status of use of high throughput nucleotide sequencing in rheumatology. Authors: Boegel S, Castle JC, Schwarting A. RMD Open 2021;7:e001324. DOI 10.1136/rmdopen-2020-001324. PMCID PMC7789458. Code: https://github.com/sebboegel/pubmed_rheuma_HTS (commit 7107940b8f7cff11f204f02211eec7e2819838f7, 2020-09-17). Repo type: authors' own code + all input/intermediate/result files + figure PDFs shipped.

What kind of paper this is

This is not a wet-lab sequencing study. It is a bibliometric / systematic literature review + SRA-metadata-mining paper. Two pipelines:

  1. PubMed literature pipeline (pubmed_search.R, R + easyPubMed):
    • Step 1 query (2020-08-15) → keyword/ICD-11 intersection → disease list.
    • Step 2 disease-specific query (2020-09-04) → automatic title/abstract annotation (disease, assay, journal, year) → exclude reviews → heavy manual curation (hundreds of hand-coded result_df[pmid==...]$assay=... overrides + a hand-built deletepmids list of ~190 PMIDs + manual SRA-found additions) → final result table main_results.csv.
  2. SRA-metadata pipeline (parse_SRA_Metadata.py, Python): aggregates the per-disease SRA "old Run Selector" metatables (<Disease>_all_old_meta.tsv) into sra_*.tsv count tables (assay / instrument / layout / source / tissue / organism / phenotype / samples-per-study) that feed Figures 4, 5, S5–S11.

In scope (pipeline-derived, attempted)

# Result Pipeline Reproducibility
C1 699 identified articles R + manual curation → main_results.csv shipped final table → recount unique PMIDs (deterministic)
C2 Per-disease publication counts (RA 182/26%, SLE 161/23%, OA 152/22%) same recount rows per disease (deterministic)
C3 Per-assay unique-pub counts (RNA-Seq 457/65%, etc.) same (R lines 460-465) recompute from shipped table (deterministic)
C4 Figure 4 / S5–S11 SRA sample counts per disease parse_SRA_Metadata.py rerun script on shipped inputs, diff vs shipped outputs (fully deterministic, no network)
C5 Step-1 PubMed count 1097 (2020-08-15) live easyPubMed query live query — time-dependent, not bit-reproducible
C6 Step-2 PubMed count 1162 (2020-09-04) live easyPubMed query live query — time-dependent, not bit-reproducible

Out of scope (not pipeline-derived → not attempted)

  • The manual curation itself (assay assignment by reading each paper, the delete-list, manual SRA additions). This is human judgement, not code; we can only verify the shipped curated table is internally consistent, not re-derive the curation. The headline numbers (699, 182, 161, 152, 457) live in the curated result_df, so we reproduce them from the shipped curated table, not from the raw PubMed hits.
  • The SRA Run Selector download step (manual portal navigation; the metatables are shipped, so we use them as given).
  • Figure 1 funnel narrative (1097→253 keywords→1162→699), the clinical "minimal data set" proposal, and all prose — not computational.

Key reproduction decision

The strongest, fully-deterministic evidence is C4 (rerun the shipped Python on the shipped SRA metatables) and C1–C3 (independently recompute the abstract's headline numbers from the shipped curated main_results.csv). C5/C6 are live PubMed counts that drift upward over time by design (PubMed adds records with back-dated publication dates); we report today's count + a date-pinned recompute to document the drift, not as a 1:1 match.

Figures / tables: Figure 1Figure 4
C1
Reported
699 identified articles
Reproduced
699 unique PMIDs in main_results.csv (813 rows)
exact
C2a
Reported
RA n=182 (26%)
Reproduced
182 (26%)
exact
C2b
Reported
SLE n=161 (23%)
Reproduced
161 (23%)
exact
C2c
Reported
OA n=152 (22%)
Reproduced
152 (22%)
exact
C3
Reported
RNA-Seq n=457 (65%)
Reproduced
457 (65%)
exact
C4
Reported
Figure 4 SRA sample totals (RA 8483, OA 1386, JIA 773, SysSclerosis 753, SpA 641, GPA 331, Sjoegren 326, MyoPolyDerma 245, Vasculitis 121, Uveitis 89, PsA 48, Synovitis 34, AutoSyn 8)
Reproduced
identical for all 13 listed diseases via rerun of parse_SRA_Metadata.py
exact
C4n
Reported
Figure 4 SLE = 3785 samples
Reproduced
1789 (from shipped SLE_all_old_meta.tsv)
partial
C5
Reported
Step-1 PubMed count 1097 (2020-08-15)
Reproduced
1213 date-pinned recompute; 4578 live 2026-06-18
partial
C6
Reported
Step-2 PubMed count 1162 (2020-09-04)
Reproduced
1381 date-pinned recompute; 5474 live 2026-06-18
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 83/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This bibliometric/literature-review paper reproduces 1:1 for its entire core: 699 articles, RA 182 (26%), SLE 161 (23%), OA 152 (22%), RNA-Seq 457 (65%), and 13/14 Fig.4 SRA disease totals all recompute byte-exactly from the shipped curated tables and parse_SRA_Metadata.py. Two peripheral deviations are fully explained and not fabrication: Fig.4 SLE=3785 vs 1789 because the authors did not deposit the fuller SLE metadata download (data-availability gap, authors' side), and the Step-1/Step-2 PubMed counts (1097/1162) drift upward by design and are not bit-reproducible from the live database. Neither touches the central conclusion, so the result is solid with explainable, input-side deviations rather than a critical discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

157 k
tokens (I/O) · 10.2 M incl. cache
15 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.