Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Assessing the impacts of COVID-19 vaccination programme's timing and speed on health benefits, cost-effectiveness, and relative affordability in 27 African coun

BMC Med · 2023
L1 89/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL (all in-scope pipeline-derived headline results reproduced exact/within-tol). Ran the authors' OWN aggregation scripts (6_R2R_R1_key_stat.R, 6_R2R_R1_fig4_update.R) on their shipped intermediate result data_upload/ICER_all.rds (Zenodo 7618749, commit 5400d14) on «our HPC». RESULTS: C5 dataset structure 972=27x36 EXACT; C6 time-horizon sensitivity mRNA 19.65% EXACT, viral-vector 34.77% vs 34.88% (one boundary combo); C1 income-stratified cost-effectiveness all four numbers within ~0.6pp (UMIC -2.20 vs -2.52, 20.21 vs 20; LIC 83.19 vs 82.59, 313.14 vs 313.52). Two documented caveats on C1: (a) the iso3c->income-group map is in an UNSHIPPED private Dropbox xlsx, reconstructed from public World Bank classification (the ~0.2-0.6pp residual); (b) the deposited key_stat.R's active filter date_start<2021-07-01 does NOT reproduce the abstract numbers -- only using all 12 start dates does. No fabrication indicators: every value is data-derivable. NOT attempted (honest, out of reach from deposit): C3 deaths averted (needs impact.rds simulation deaths), C4 affordability (needs private health-expenditure denominator).

💻 Code ↗ 🗄 Data: 10.5281/zenodo.7618749

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Given that COVID-19 vaccine roll-out in African countries was delayed and slow relative to high-income countries, the paper asks whether vaccination remains an impactful and cost-effective strategy, and how programme start timing and roll-out speed affect health benefits, cost-effectiveness, and affordability.

Core claims
  • Vaccination programmes with earlier start dates yield the most health benefits and lowest ICERs compared to late-starting programmes finding
  • Fast vaccine roll-out produces the most health benefits but does not always result in the lowest ICERs finding
  • The highest marginal effectiveness within vaccination programmes is found among older adults (60+) finding
  • High country income group, high proportion of population over 60, or high proportion non-susceptible at programme start are associated with low ICERs relative to GDP per capita finding
  • Most vaccination programmes with small ICERs relative to GDP per capita were also relatively affordable finding
  • Programmes starting late in 2021 may still generate low ICERs and manageable affordability despite significant ICER increases with delay finding
  • An adapted age-specific dynamic transmission model (CovidM) fitted to country-level reported COVID-19 deaths was used to approximate pre-vaccination infection-induced immunity across 27 African countries method
  • Unit cost of delivering mRNA vaccines is substantially higher than viral vector vaccines, and faster roll-out rates are associated with lower vaccine unit costs due to fixed costs spread over more doses finding
Experimental setups
Assay System Perturbation Readout Platform
Dynamic transmission model fitting (maximum likelihood estimation with differential evolutionary algorithm) Population-level model of 27 African Union member states none (fitted to observed pre-vaccination epidemic data) R0, infection introduction dates, COVID-19 death reporting rate, variant-of-concern introduction dates R (4.1.0), adapted CovidM model
Epidemiological simulation/projection of vaccine roll-out scenarios Age-structured population models of 27 African countries 36 vaccine roll-out scenarios (12 start dates Jan-Dec 2021 x 3 roll-out rates: slow/medium/fast) x 2 vaccine types (mRNA, viral vector) Symptomatic infections, severe cases, critical cases, deaths (Jan 2021-Dec 2022)
DALY calculation (health economics modelling) Same 27-country cohort Same vaccine roll-out scenarios Years of life lost (YLL), years lived with disability (YLD), DALYs averted, discounted at 3%/year
Cost-effectiveness analysis (ICER calculation) Same 27-country cohort, healthcare payer perspective Vaccine roll-out scenarios vs. no-vaccination comparator Incremental cost-effectiveness ratios (ICERs) vs. GDP per capita
Ingredient-based (itemised) costing study Ethiopia, Nigeria, South Africa (extrapolated to other countries) Vaccine type (mRNA vs. viral vector), roll-out rate, programme duration Vaccine delivery unit costs (purchasing + planning/coordination + cold chain + transport + waste disposal)
Relative affordability analysis Same 27-country cohort Same vaccine roll-out scenarios Nonmarginal budget impact / relative affordability measure of vaccination programmes
Key results
  • Early programme start dates yielded the most health benefits and lowest ICERs vs. late starts
  • Fast roll-out gave the most health benefits but not always the lowest ICERs
  • Highest marginal effectiveness within programmes found among older adults
  • mRNA vaccine unit cost substantially higher than viral vector at medium roll-out (01 Aug 2021 start) median $16.15 (mRNA) vs $6.92 (viral vector)
  • Faster roll-out rate associated with lower vaccine unit cost e.g. viral vector medium $6.92 vs fast $4.49 median
  • By February 2022, vaccine coverage remained below 10% in over 35% of African Union member states >35% of countries <10% coverage
  • ICERs increased significantly as vaccination programmes were delayed
  • UK had vaccinated 74.6% of population by 31 Aug 2021 vs. ~1% coverage in most African Union states by Aug 2021 74.6% vs ~1%
Key statistics
  • other Roll-out rates: slow 275, medium 826, fast 2066 doses/million population-day (Derived from tertiles of observed African Union roll-out trajectories)
  • count 27 countries included (28 excluded) (Countries with sufficient death-reporting data for model fitting out of 55 African Union member states)
  • mean Median vaccine unit cost: mRNA $16.15, viral vector $6.92 (medium roll-out, 01 Aug 2021 start) (Country-level extrapolated vaccine delivery unit costs)
  • other Vaccine efficacy against infection (2nd dose): mRNA 0.85, viral vector 0.75 (base case) (Vaccine efficacy parameter set used in transmission model)
  • other Maximum population-level vaccine coverage capped at 70% (Consistent with WHO/Africa CDC target)
  • other Discount rate of 3% annually (Applied to DALYs and costs per WHO guidelines)
  • other UK vaccine coverage 74.6% by 31 August 2021 (Comparator to African Union member states' coverage)
  • other Maximum vaccine uptake: 80% among older adults, 60% among other adults (Assumed uptake rate ceilings by age group)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a combined epidemiological and health-economic modelling study, not a classical experimental/observational hypothesis-testing paper. An age-specific dynamic transmission model was fitted to country-level daily reported COVID-19 deaths using maximum likelihood estimation via a differential evolutionary algorithm, to back-estimate infection-induced immunity in 27 African Union member states. Outputs (symptomatic infections, severe/critical cases, deaths, DALYs, ICERs, and a relative affordability measure) were then simulated across 36 vaccine roll-out scenarios (12 start dates × 3 roll-out rates, by 2 vaccine types) and summarized descriptively (e.g., medians and quartile/limit ranges) rather than through inferential significance testing. Sensitivity analyses (alternative vaccine efficacy estimates, extended time horizon) were used to explore result robustness.

Replicationunclear Sample size27 of 55 African Union member states were included in model fitting after excluding 28 for data sparsity (≤10 deaths/day throughout fitting period, n=26) or reporting artefacts (single day >5% of cumulative deaths, n=2); scenarios spanned 12 programme start dates × 3 roll-out rates × 2 vaccine types (36 combinations), with additional country-level cost estimates for 3 countries (Ethiopia, Nigeria, South Africa) GroupsVaccine roll-out scenarios (varying start date, roll-out speed, vaccine type) versus a no-vaccination comparator, and across country income groups Pairingna Randomization/blindingna DispersionIQR
Statistical tests used
Test Applied to n Assumptions
Maximum likelihood estimation with a differential evolutionary algorithm (model fitting) Fitting the dynamic transmission model to country-level daily reported COVID-19 deaths (2020–2022) to estimate R0, infection introduction dates, death reporting rate, and variant introduction dates 27 of 55 African Union member states with sufficient death-reporting data not stated
Univariable linear regression (cumulative doses per million population-day ~ date) Deriving slow/medium/fast vaccine roll-out rate levels from observed uptake trajectories Observed vaccine uptake trajectories among African Union member states not stated
Approaches that could also have been used
  • Model parameters (R0, introduction dates, reporting rate, variant timing) were estimated via maximum likelihood with a differential evolutionary algorithm, yielding point estimates.
    Could also: A Bayesian fitting approach (e.g., MCMC) could also be used — This would generate full posterior distributions for fitted parameters, which can be propagated through to give uncertainty ranges around downstream outcomes like DALYs and ICERs, complementing the point-estimate approach used here.
  • Uncertainty in vaccine efficacy was explored through a single alternative ('lower bound') sensitivity-analysis scenario rather than a distributional approach.
    Could also: Probabilistic sensitivity analysis (PSA) using Monte Carlo simulation over parameter distributions could also be used — PSA is a standard approach in health-economic evaluations for jointly propagating uncertainty across many parameters (efficacy, costs, epidemiological inputs) and typically produces cost-effectiveness acceptability curves and uncertainty intervals around ICERs, which can add further texture to the deterministic/scenario-based sensitivity analyses reported here.
  • Vaccine unit costs and health-service costs are summarized with five-number summaries (lower limit, quartiles, upper limit) across countries.
    Could also: Reporting means with 95% confidence or credible intervals could also be used — This would allow comparison to parametric summary statistics and facilitate propagation of cost uncertainty into formal statistical uncertainty intervals around the ICERs, alongside the quantile-based summaries presented.
  • Roll-out rate levels (slow/medium/fast) were derived from tertiles of a fitted univariable linear model of cumulative doses over time.
    Could also: A mixed-effects or hierarchical regression model (allowing country-specific slopes/intercepts) could also be used — This could account for between-country correlation structure and heterogeneity in roll-out trajectories more explicitly than a pooled univariable linear fit, potentially refining the derived tertile thresholds.
  • The primary outputs (health impacts, ICERs, affordability) are presented as scenario comparisons without inferential hypothesis tests or p-values in the text provided.
    Could also: Formal statistical comparison (e.g., bootstrapped confidence intervals for differences in ICERs between scenarios) could also be used — This would let readers quantify the precision of comparative statements (e.g., 'ICERs increased significantly as programmes were delayed') alongside the descriptive scenario-based presentation used in the study.
Software: R 4.1.0

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

C1a
Reported
UMIC viral-vector ICER = -2.52% of GDP/capita
Reproduced
-2.20%
within tolerance
C1b
Reported
UMIC mRNA ICER = 20% of GDP/capita
Reproduced
20.21%
within tolerance
C1c
Reported
LIC viral-vector ICER = 82.59% of GDP/capita
Reproduced
83.19%
within tolerance
C1d
Reported
LIC mRNA ICER = 313.52% of GDP/capita
Reproduced
313.14%
within tolerance
C5
Reported
972 = 27 countries x 36 scenarios
Reproduced
972 (27 x 36)
exact
C6a
Reported
viral-vector: 34.88% of 972 combos flip ICER category under extended time horizon
Reproduced
34.77% (338/972)
within tolerance
C6b
Reported
mRNA: 19.65% flip
Reproduced
19.65% (191/972)
exact
C3
Reported
deaths averted 11.06/24.19/31.29% (slow/medium/fast)
Reproduced
NOT_ATTEMPTED
m.public.grade.not-attempted
C4
Reported
affordability % of health expenditure (mRNA 3.85/12.42/26.13; VV 1.09/3.24/5.28)
Reproduced
NOT_ATTEMPTED
m.public.grade.not-attempted

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

All in-scope, pipeline-derived headline results reproduce from the authors' shipped intermediate ICER_all.rds by running their own aggregation scripts: C5 972=27×36 exact, C6b 19.65% exact, C6a 34.88%→34.77% (one bin-boundary combo), and C1 within 0.2–0.6pp (the residual attributable to reconstructing the authors' unshipped private income-group map from public World Bank classification). No fabrication indicators — every checked value is data-derivable and the income-gradient/cost-effectiveness conclusion holds. Two fair caveats keep the overall at yellow: the deposited key_stat.R has an active date_start<2021-07-01 filter that does not reproduce the abstract numbers (only all-12-dates does), a non-runnable-verbatim code mismatch on the authors' side; and C3 (deaths averted) / C4 (affordability) are unreachable because their inputs (impact.rds deaths, private Dropbox health-expenditure) were never deposited — a data-availability gap, not a defect in the reproduced values.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

113.2 k
tokens (I/O) · 5.6 M incl. cache
16 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.