Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Projecting contact matrices in 177 geographical regions: An update and comparison with empirical data for the COVID-19 era.

PLoS Comput Biol · 2021
L1 98/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
98/100
Reproducibility score
1.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 93% of all assessed papers rank 65 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduced 1:1. Prem et al.'s synthetic-contact-matrices repo (= Zenodo v2.0) is self-contained for the deterministic results: it ships the demographic inputs, the Bayesian-model intermediate outputs (lambdas/deltas, household age matrices), the final 177-region synthetic matrices, and the empirical comparison surveys. On «infra» «our HPC» (R 4.3.3, 6-second SLURM job) all 8 pinned claims reproduced exactly/within-tol: 177 regions; 97.2% world population; 16x16 (5-year) matrices; the published home and other-location contact matrices regenerate bit-for-bit (max_abs_diff=0 over all 177 countries) from the shipped lambdas/deltas via the published combination code; the long-format synthetic_contacts_2021.csv (675840 rows) reproduces to floating-point precision; the observed-vs-modelled household-age-matrix correlation = 0.9596 (IQR 0.9369-0.9717) vs reported 0.96 (0.94-0.97); and the normalised synthetic-vs-empirical correlation = 0.8399 (IQR 0.7563-0.8767) vs reported 0.84 (0.76-0.88). NOT attempted: (1) the JAGS Bayesian hierarchical fit that produces lambdas/deltas - its script codes/contact_1_world.r is ABSENT from the repo and the Zenodo zip (404), and the MCMC is non-deterministic, so shipped lambdas/deltas were consumed; (2) work + school matrices from scratch - require input/polymod_pworkandsch.rdata, also absent (an output of the missing script), so only home+others were rebuilt from raw and contact_all uses shipped work/school; (3) the downstream COVID-19 transmission-model impact figures (Sections 4-5), out of scope. No fabrication signal: every checked value is fully derivable from the shipped data+code and recomputes to stated precision. C8 uses 10 survey locations (paper text says 11; Peru+Russia NA-masked by the authors' code) yet the median/IQR match exactly.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.8383778

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 98
    assessed: 2026-06-18 ⛓ 87b408dfef8c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can synthetic contact matrices—constructed from setting-specific demographic data combined with European POLYMOD contact data—reproduce empirically-measured contact patterns and yield similar epidemic modelling results, such that they can substitute for empirical contact surveys in settings lacking them?

Core claims
  • Updated synthetic contact matrices were generated for 177 geographical locations covering 97.2% of the world's population (up from 152 locations/95.9% in 2017). resource
  • Synthetic contact matrices show qualitative similarities to empirically-constructed contact matrices. finding
  • Transmission models parameterised with empirical versus synthetic matrices generated similar findings, with few differences except in age groups where empirical matrices had missing or aggregated ages. finding
  • A Bayesian hierarchical modelling framework estimates age- and location-specific contact rates from POLYMOD data, projected onto other countries using household, school and workplace composition. method
  • Contact patterns were stratified by rural and urban settings using stratified population, household, labour force, and school pupil-to-teacher data. method
  • Synthetic contact matrices may be used to model outbreaks in settings where empirical contact studies have not yet been conducted. finding
  • The 'Other' (non-home/work/school) location proportion of contacts was compared between empirical and synthetic matrices as the only setting not informed by local contact data. method
Experimental setups
Assay System Perturbation Readout Platform
Bayesian hierarchical Poisson regression of contact rates 8 POLYMOD European countries (Belgium, Germany, Finland, UK, Italy, Luxembourg, Netherlands, Poland) none age- and location-specific contact rate (lambda) per day R version 3.6.2
Synthetic contact matrix construction/projection 177 geographical regions (countries and subnational regions) none age- and location-specific (home/work/school/other) contact matrices R (socialmixr where applicable)
Household age structure modelling / leave-one-out validation 43 DHS countries (incl. India ~3M individuals, ~600,000 households) none household age matrices reverse-engineered for validation Demographic Household Surveys (DHS)
Rural/urban stratification of contact matrices 177 geographical regions; 36 countries with OECD pupil-to-teacher data; 43 DHS countries rural vs urban setting mean total contacts (children 0-9, older adults 60-69) and basic reproduction number UN Population Division, ILO, OECD data
Empirical contact matrix construction and comparison 11 locations: Shanghai & Hong Kong (China), France, Kenya, Peru, Russia, South Africa, Vietnam, Zambia, Zimbabwe none element-wise comparison of empirical vs synthetic matrices; proportion of 'Other' contacts socialmixr R package; Zenodo social contact database
COVID-19 transmission model with physical distancing age-structured population model physical distancing interventions epidemic dynamics under empirical vs synthetic matrices
Key results
  • Synthetic matrices reproduce the main qualitative features of empirically-measured contact patterns.
  • Models using empirical vs synthetic matrices produced similar findings, differing mainly where empirical matrices had missing/aggregated age groups.
  • Coverage extended to 177 geographical locations representing 97.2% of the world's population. 177 locations / 97.2%
  • Country-specific household data increased from 17 countries (2017) to 51 countries (2020), adding 34 low/lower-middle-income countries. 17 to 51 countries
Key statistics
  • count 177 geographical locations (coverage of 2020 synthetic matrices)
  • other 97.2% of the world's population (population covered by 2020 matrices)
  • count 152 geographical locations covering 95.9% (2017 synthetic matrices coverage)
  • count ~3 million individuals from about 600 000 households (India DHS, largest dataset)
  • count 51 countries (34 additional LMICs) (country-specific household data in 2020)
  • count 4 studies in LMICs vs 54 in high-income countries (availability of empirical contact surveys)
  • count 11 geographical locations (empirical contact matrices available for comparison)
  • count 36 countries (rural/urban pupil-to-teacher ratio differences from OECD)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study applied a Bayesian hierarchical Poisson regression model to estimate age- and location-specific contact rates from the eight-country POLYMOD diary survey, incorporating individual-level random effects for heterogeneity in social activity. These estimated contact rates were projected to 177 geographical regions by combining country-specific demographic, household, school, and labour-force data, with 14 z-scored country-level indicators used to derive similarity-based projection weights for countries lacking household survey data. Internal validation used leave-one-out cross-validation on countries with known household structures; external validation compared synthetic matrix elements descriptively to empirically-constructed matrices from 11 out-of-sample locations and compared modelled basic reproduction numbers across settings.

Replicationunclear Sample size177 geographical regions for projection; 8 POLYMOD countries as source contact data; 43 DHS countries for household data; 11 out-of-sample locations for external validation GroupsSynthetic vs empirical contact matrices; 2017 vs 2020 synthetic matrices; rural vs urban contact patterns Pairingna Randomization/blindingna Dispersionnone Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Bayesian hierarchical Poisson regression with individual random effects Estimation of age- and location-specific daily contact rates from POLYMOD survey data across 8 European countries not stated
Leave-one-out cross-validation Internal validation of projected household age matrices for POLYMOD and DHS countries not stated
Z-score standardisation and weighted mean projection Deriving country-similarity weights for projecting household age structures to regions without empirical household data not stated
Element-wise descriptive matrix comparison Comparison of synthetic vs empirical contact matrices across 11 out-of-sample geographical locations na
Descriptive comparison of mean total contacts by age group Rural vs urban contact patterns in children (0–9 years) and older adults (60–69 years) not stated
Basic reproduction number (R0) calculation and comparison Comparison of transmission model outputs parameterised with empirical vs synthetic matrices and rural vs urban matrices na
Approaches that could also have been used
  • Contact counts were modelled with a Poisson distribution for the Bayesian hierarchical model
    Could also: A negative binomial distribution could also have been used as the likelihood for contact counts — Social contact data are frequently overdispersed relative to the Poisson assumption (variance exceeding the mean); a negative binomial model adds a dispersion parameter that can accommodate this and may yield better-calibrated posterior uncertainty for the estimated contact rates
  • Validation of synthetic matrices against empirical matrices was qualitative and visual, comparing matrix elements across 11 locations
    Could also: Quantitative scalar agreement metrics such as mean absolute error, root mean squared error, or Pearson/Spearman correlation of matrix elements could also have been reported — Numerical summary metrics would allow objective ranking of synthetic-vs-empirical agreement across validation locations and enable direct comparison of the 2017 and 2020 matrices on the same scale
  • Posterior uncertainty from the Bayesian hierarchical model was estimated but not propagated to the final reported synthetic contact matrices
    Could also: Full posterior predictive distributions or element-wise credible intervals for synthetic matrix entries could also have been reported — Propagating Bayesian uncertainty to matrix outputs would allow downstream epidemic models to incorporate contact-rate uncertainty into probabilistic projections and sensitivity analyses
  • Country-level similarity for projection weights used z-scored Euclidean similarity across 14 indicators treated as equally informative
    Could also: Principal components analysis (PCA) or regularised regression (e.g. LASSO) on the 14 indicators could also have been used before deriving similarity weights — PCA removes multicollinearity among correlated development indicators and extracts orthogonal dimensions of variation, which may produce more stable and parsimonious similarity weights; LASSO could select the most predictive subset of indicators
  • Internal validation used leave-one-out cross-validation applied once to countries with known household data
    Could also: Repeated k-fold cross-validation or bootstrap-based cross-validation could also have been applied — Repeated k-fold cross-validation yields a distribution of validation errors with lower variance than a single leave-one-out pass, providing a more stable and generalisable estimate of out-of-sample prediction error for the household projection step
  • Rural versus urban contact differences were summarised for two specific age bands (0–9 and 60–69 years) and as R0 point estimates
    Could also: A structured quantitative summary across all age groups — such as age-group-specific absolute differences or a Frobenius-norm comparison of rural vs urban matrices — could also have been reported — Full-age-range summaries would reveal whether rural-urban heterogeneity is concentrated in specific age groups or is broadly distributed, supporting more targeted age-structured policy analyses
Software: R 3.6.2 · socialmixr (R package)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34310590

Paper: Prem et al. 2021, PLoS Comput Biol 17(7):e1009098. "Projecting contact matrices in 177 geographical regions: an update and comparison with empirical data for the COVID-19 era."

Code: https://github.com/kieshaprem/synthetic-contact-matrices (own code, P16 not needed) Data/Code archive: Zenodo 10.5281/zenodo.8383778 (v2.0) — a single zip that is a snapshot of the GitHub repo (no separate data deposit). The repo ships its own input and output data inside generate_synthetic_matrices/{input,output} and compare_contact_matrices/input.

What kind of paper this is

A computational/statistical modelling paper. It builds synthetic age-structured social contact matrices (16 five-year age bands, 0–4 … 75+) for 177 geographical regions by:

  1. Fitting a Bayesian hierarchical model (JAGS / rjags) to the POLYMOD + DHS contact surveys → per-setting scaling factors lambdas and intercepts deltas.
  2. Combining those with demographic inputs (UN WPP population by age, DHS household age matrices, school enrolment, labour-force participation) to project home/work/school/ other/all contact matrices per country (overall + urban + rural).
  3. Comparing the synthetic matrices to out-of-sample empirical contact surveys from the COVID-19 era, and feeding both into an age-structured COVID-19 transmission model.

In scope (pipeline-derived, deterministic, reproducible)

id result pipeline / script
C1 177 geographical regions contact_2_world.r country list = names(HAM_WORLD) ∩ school$iso
C2 97.2% of world population covered contact_2_world.r (sum poptotal / 7794798729)
C3 16 five-year age groups (16×16 matrices) matrix dims of shipped outputs
C4 home contact matrices (177) re-run contact_2_world.r HOME block from shipped lambdas/deltas/HAMs
C5 other-location contact matrices (177) re-run contact_2_world.r OTHERS block
C6 published synthetic_contacts_2021.csv re-run contact_4_world_storage.r reshape
C7 HAM correlation, median 0.96, IQR 0.94–0.97 home_validation_internal_HAM.r cor(obs HAM, modelled HAM)
C8 synthetic vs empirical correlation, median 0.84, IQR 0.76–0.88 2_normalise_urbanrural.r + 3_plotComparison_urbanrural.r cor loop

Out of scope / not attempted (with reason)

  • JAGS Bayesian fit producing lambdas/deltas — the master sources codes/contact_1_world.r, which is absent from the repo (404 on GitHub master and the Zenodo zip). The MCMC fit is also non-deterministic. We therefore consume the shipped lambdas.rdata/deltas.rdata rather than regenerate them. → NOT attempted.
  • work / school contact matricesgetWorkPop/getSchoolPop require input/polymod_pworkandsch.rdata, which is absent from the repo (produced by the missing contact_1_world.r). The home/work/school sum (contact_all) therefore can only be partially regenerated from scratch. → work/school NOT attempted; home+others ARE.
  • COVID-19 transmission-model impact (Figs on R0, attack rate, intervention effectiveness; Sections 4–5) — downstream epidemic-simulation modelling, not a data pipeline output; large, separate model. → NOT attempted (would be a separate effort).
  • Urban/rural full regeneration — same polymod_pworkandsch.rdata blocker; we use the shipped urban/rural matrices for the comparison (C8) but do not regenerate them.

Compute

All on «infra» «our HPC»-2 via «host» ssh, SLURM job(s). Repo + data cloned to «infra»; analysis = base R 4.3.3 (reused conda prefix env repro-rbase). «host» holds results only.

C1
Reported
177 geographical regions
Reproduced
177
exact
C2
Reported
97.2% of world population
Reproduced
97.2%
exact
C3
Reported
16 five-year age bands (16x16 matrices)
Reproduced
16x16
exact
C4
Reported
published home contact matrices (177 regions)
Reproduced
regenerated from shipped lambdas/deltas: max_abs_diff=0, 177/177 exact
exact
C5
Reported
published other-location contact matrices (177 regions)
Reproduced
regenerated: max_abs_diff=0, 177/177 exact
exact
C6
Reported
published synthetic_contacts_2021.csv
Reproduced
675840 rows match; value max_abs_diff=1.78e-15
within tolerance
C7
Reported
HAM correlation median 0.96 (IQR 0.94-0.97)
Reproduced
median 0.9596 (IQR 0.9369-0.9717), n=51
exact
C8
Reported
synthetic-vs-empirical correlation median 0.84 (IQR 0.76-0.88)
Reproduced
median 0.8399 (IQR 0.7563-0.8767), n=10 locations (2 NA-masked)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 98/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

184.5 k
tokens (I/O) · 15.4 M incl. cache
19 min
runtime · 0 CPU-h
0.3 GB
peak RAM
1
HPC jobs
hummel
machine