Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Widespread nitrous oxide undersaturation in farm waterbodies creates an unexpected greenhouse gas sink.

Proc Natl Acad Sci U S A · 2019
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

P16 reproduction: the GitHub repo simpson-lab/dugout-n2o (Zenodo 10.5281/zenodo.2636099) is a DATA-ONLY deposit (SI_data.csv, 102 obs; no analysis code), so we re-ran the documented mgcv GAM pipeline + deterministic stats on the paper's own data on «our HPC» (R 4.5.3, mgcv 1.9.4). DESCRIBED WELL ENOUGH for the descriptive results, which reproduce 1:1 or within-tol: [N2O] range exact (1.14-109.79 vs 1.14-110), median [N2O] 6.63 vs 6.55 nM, median flux -3.62 vs -4.03, and the sink/source/equilibrium split is EXACT at 69/20/12 (= the paper's 67%/21%/12%; 1 of 102 rows has NA flux -> 101 classified sites = paper n). The headline GAM, however, is only PARTIALLY reproducible: the paper's model has 7 smooth terms but SI_data.csv ships only 4 of the predictors (BF, DIN, surface DO, Chla) -- DeepDO, sediment C:N, surface pH, and N:P are NOT in the deposit. Re-fitting the reducible 4-predictor model with the paper's exact transforms (log Chla/log DIN/sqrt BF), bases (tp k=9; te cubic k=4), Gamma log link, REML, double-penalty select=TRUE gives 58.5% deviance explained vs the reported ~85%; the headline BF x log(DIN) tensor interaction IS recovered as strongly significant (p<1e-3). This is a data-completeness gap (the 85% figure is not independently verifiable from the public deposit), flagged per HARD RULE 5 -- not asserted as fabrication. NOT ATTEMPTED (hard 20%): exact %-saturation 67/21/12 from first principles (needs water temperature + atmospheric N2O for Weiss-Price solubility -- not in CSV; flux-sign split used as the derivable proxy and it matches); the 7.5-33x and 11x overestimates of prior/IPCC models and regional GHG-sink upscaling (need external models + the >110k-dugout inventory, not in shipped data).

💻 Code ↗ 🗄 Data: 10.5281/zenodo.2636099

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 87
    assessed: 2026-06-16 ⛓ 036f2b242d74
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Whether small artificial agricultural reservoirs—assumed to be sources of nitrous oxide (N2O) because they are N-enriched and eutrophic—actually act as N2O sources or sinks to the atmosphere, and what environmental controls govern their N2O concentrations.

Core claims
  • The majority (67%) of small agricultural farm reservoirs are undersaturated and act as atmospheric N2O sinks despite being highly eutrophic and N-rich. finding
  • In situ N2O concentrations are strongly and nonlinearly related to water-column stratification strength (buoyancy frequency) interacting with dissolved inorganic nitrogen, plus surface O2 and Chl-a. finding
  • Elevated DIN does not invariably produce N2O supersaturation; it raises N2O only under weak or no stratification, while strong stratification promotes undersaturation via complete denitrification below the thermocline. mechanism
  • Previously published empirical models (IPCC, DelSontro, Deemer) overestimate measured N2O fluxes from these reservoirs by 7.5- to 33-fold, challenging the view that N-enriched eutrophic waters are strong N2O sources. finding
  • Generalized additive models (GAMs) capable of capturing nonlinear/nonmonotonic relationships were used to predict reservoir N2O concentrations. method
  • This study adds 101 sites (~32%) to the sparse global lake/reservoir N2O dataset, the largest dataset to date on N2O in small waterbodies. resource
  • Inclusion of small reservoirs may provide a means of anthropogenic N retention and reduce net greenhouse gas emissions in agriculture. finding
Experimental setups
Assay System Perturbation Readout Platform
Dissolved N2O concentration measurement (and flux calculation via gas transfer velocity) 101 small constructed agricultural reservoirs, southern Saskatchewan, Northern Great Plains, Canada none (observational spatial survey) N2O concentration (nM), saturation status, and air-water N2O flux (μmol N2O·m−2·d−1)
Water chemistry / nutrient analysis 101 agricultural reservoirs none TDN, DIN, NOx, SRP, chlorophyll-a concentrations, pH
Physical limnology profiling 101 agricultural reservoirs none surface and bottom-water O2 saturation, maximum buoyancy frequency (stratification strength)
Sediment analysis reservoir sediments none sediment C:N ratio
Generalized additive modeling (GAM) statistical analysis reservoir dataset (n=101) none deviance explained in N2O concentrations; partial effects of predictor variables
Key results
  • 67% of reservoirs were undersaturated (N2O sinks), 21% supersaturated, 12% near atmospheric equilibrium 67% undersaturated
  • GAM explained ~85% of deviance in N2O concentrations; N2O predicted by buoyancy frequency × DIN, surface DO, and Chl-a ~85% deviance explained
  • Lowest N2O concentrations occurred under strong stratification and high phytoplankton abundance; DIN boosted N2O only when unstratified/weakly stratified
  • Published models (IPCC, SPW/DelSontro, Deemer) overestimated measured N2O fluxes 7.5- to 33-fold overestimate
  • Median calculated N2O flux was a small sink; sink sites ranged −12 to −2 μmol·m−2·d−1 median −4.03 μmol·m−2·d−1
  • IPCC-predicted Saskatchewan reservoir emissions (10,530 t CO2-eq/y) vastly exceeded actual measured (968 t CO2-eq/y) ~11-fold overestimate
  • Agricultural ponds/wetlands have significantly lower N2O emissions than flowing waters three- to ninefold lower
Key statistics
  • other N2O concentration range 1.14–110 nM, median 6.55 nM (surveyed reservoir N2O concentrations)
  • count 67% undersaturated, 21% supersaturated, 12% at equilibrium (N2O saturation status of sites)
  • other ~85% deviance explained (GAM model fit for N2O concentrations)
  • pvalue P < 0.001 (buoyancy frequency + DIN); P < 0.05 (surface DO; Chl-a) (GAM predictor significance)
  • fold_change 7.5- to 33-fold overestimate (published models vs measured N2O fluxes)
  • other median flux −4.03 μmol·m−2·d−1; sinks −12 to −2; sources 2.21 to 166 (calculated N2O fluxes (69 sinks, 20 sources, 12 equilibrium))
  • mean Chl-a 99 ± 289 µg·L−1; TDN avg 3,082 µg N·L−1 (eutrophic/nutrient status of reservoirs)
  • count 16 million small artificial reservoirs globally; study adds 101 sites (~32% of global dataset) (global context and dataset contribution)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an observational regional survey in which 101 small agricultural reservoirs were sampled once during a 5-week late-summer period, and dissolved N2O concentrations were related to environmental predictors. The primary analysis used generalized additive models (GAMs) to capture nonlinear and nonmonotonic relationships between N2O and predictors such as buoyancy frequency, DIN, surface O2, and Chl-a, with model fit summarized by deviance explained (~85%) and term significance reported as P-value thresholds. Results were also reported descriptively (medians, ranges, means with SD) and compared against fluxes predicted by previously published empirical models.

Replicationunclear Sample sizeStated as a spatial survey of 101 constructed reservoirs selected for relatively even spatial distribution and ease of access; no formal power/sample-size justification described GroupsSites categorized post hoc as N2O sinks/sources/near-equilibrium; primarily continuous predictor–response modeling rather than group comparison Pairingna Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Generalized additive model (GAM) with smooth terms and a buoyancy-frequency × DIN interaction Modeling in situ N2O concentration as a function of environmental predictors (Fig. 2; SI Table S2) 101 surveyed reservoirs not stated
Significance testing of GAM smooth terms (reported as P-value thresholds, e.g., P<0.001, P<0.05) Buoyancy frequency + DIN (P<0.001), surface dissolved O2 (P<0.05), Chl-a (P<0.05) 101 reservoirs not stated
Approaches that could also have been used
  • Nonlinear relationships were modeled with GAMs and term importance was reported using P-value thresholds (P<0.001, P<0.05).
    Could also: Reporting exact P-values alongside effective degrees of freedom and approximate confidence/credible bands for each smooth, or using AIC/cross-validation for term selection. — Exact P-values and information-criterion comparisons convey the strength of evidence on a continuous scale and document model-selection decisions, which can aid reproducibility for readers.
  • Multiple smooth predictors were evaluated within the GAM without a described multiplicity adjustment.
    Could also: Applying a multiple-comparison or shrinkage approach (e.g., double-penalty/select=TRUE shrinkage in GAMs, or FDR control across terms). — Shrinkage or FDR control offers a transparent way to handle the family of predictors simultaneously and can help distinguish well-supported terms from incidental ones.
  • Partial effects were summarized with 95% credible intervals from the fitted GAM.
    Could also: Reporting out-of-sample predictive performance (e.g., k-fold cross-validated R²/RMSE) in addition to in-sample deviance explained. — Cross-validated metrics characterize how the predictive model generalizes to new reservoirs, complementing the in-sample fit for a model intended to predict N2O.
  • Summary statistics were presented as means with SD alongside medians and ranges for highly skewed, multi-order-of-magnitude variables.
    Could also: Emphasizing medians with IQR or geometric means with 95% CIs for the skewed concentration variables. — For variables spanning two to three orders of magnitude, robust summaries can convey central tendency and spread in a way that is less sensitive to extreme values.
  • Sites were sampled once during a single late-summer window across a broad spatial extent.
    Could also: Incorporating spatial structure explicitly (e.g., spatial smooths or spatial autocorrelation terms) and/or repeated seasonal sampling. — Accounting for spatial dependence among nearby sites and adding temporal replication would let the model address spatial autocorrelation and seasonal variability, which the authors themselves note as a direction for future work.

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
106
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31036633

Paper: Webb et al. 2019, PNAS. "Widespread nitrous oxide undersaturation in farm waterbodies creates an unexpected greenhouse gas sink." DOI 10.1073/pnas.1820389116.

Code repo: https://github.com/simpson-lab/dugout-n2o @ commit c126be0dd8c643c3a0bde7416f653fca857c0886data-only deposit: ships SI_data.csv (6823 B, 102 obs) + README + LICENSE. No analysis code. Zenodo 10.5281/zenodo.2636099 is the archived copy of the same repo.

This is therefore a P16 reproduction: we re-run the documented pipeline (mgcv GAM, per the paper's Methods) on the paper's own shipped data. Per the BRIEF, applying the described tool to the paper's data is fully valid.

Data columns (SI_data.csv): Site_ID, latitude, longitude, N2O.nM (concentration), N2O.umol.m2.d (flux), DO.mg.L, Chla.ug.L, DIN.ug.N.L, b.f.max (max buoyancy frequency). The 4 GAM predictors named as significant in the paper (max buoyancy frequency, DIN, surface DO, chl-a) + the response all ship in this file.

IN SCOPE (pipeline-derived, attempted)

id result reported where how to reproduce
C1 GAM deviance explained ~85% Results / Methods mgcv::gam, Gamma family, te(b.f.max, DIN) tensor + s(DO) + s(Chla) thin-plate, select=TRUE (double penalty), REML, on SI_data.csv
C2 median [N2O] 6.55 nM Results median(N2O.nM) — deterministic
C3 [N2O] range 1.14–110 nM Results range(N2O.nM) — deterministic
C4 median N2O flux −4.03 µmol m⁻² d⁻¹ Results median(N2O.umol.m2.d) — deterministic
C5 site flux split 69 sink / 20 source / 12 equilibrium Results sign of flux column (sink<0, source>0, ~equilibrium) — deterministic-ish

OUT OF SCOPE / not attempted (the hard last 20%)

  • Saturation status 67% under / 21% super / 12% equilibrium — requires %N2O saturation = measured vs atmospheric-equilibrium concentration, which needs water temperature + atmospheric N2O mole fraction (Weiss-Price solubility). Water temperature is not in the shipped CSV → exact % not derivable. The flux-sign split (C5) is the closest derivable proxy. Noted, not graded as exact.
  • "7.5- to 33-fold" / "11-fold" overestimates of prior empirical & IPCC models — require external emission models + areal upscaling; the comparison inputs are not in the shipped data. Out of scope.
  • Areal/regional GHG-sink upscaling (per-region flux totals) — needs the

    110,000-dugout inventory + areas; not in shipped data. Out of scope.

Compute

All compute on «our HPC» («infra» SLURM) per HARD RULE 1, even though the dataset is tiny — single R/mgcv job. Repo cloned + data kept on «infra».

C1
Reported
~85% deviance explained (GAM)
Reproduced
58.53% (reduced 4-predictor model, paper's exact transforms/bases/link)
partial
C2
Reported
median [N2O] 6.55 nM
Reproduced
6.63 nM
within tolerance
C3
Reported
[N2O] range 1.14-110 nM
Reproduced
1.14-109.79 nM
exact
C4
Reported
median flux -4.03 umol m-2 d-1
Reproduced
-3.62 umol m-2 d-1
within tolerance
C5
Reported
69 sink / 20 source / 12 equilibrium sites
Reproduced
69 / 20 / 12
exact
C6
Reported
BF x log(DIN) tensor interaction significant
Reproduced
p<1e-3
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

The descriptive results and the central ecological claim reproduce essentially 1:1: the 69/20/12 sink/source/equilibrium split is exact, the concentration range matches, medians are within-tol, and the BF×log(DIN) driver interaction is recovered (p<1e-3). The single real deviation is the headline ~85% deviance explained, which drops to 58.5% — but only because the public deposit ships 4 of 7 model predictors, so the figure is not independently verifiable from shared data (authors'-side data-completeness gap), not evidence of fabrication. Severity is moderate and explainable; the paper's core conclusion holds.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

90.4 k
tokens (I/O) · 5.4 M incl. cache
11 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
2
HPC jobs
hummel
machine