Temporal trends in methane emissions from a small eutrophic reservoir: the key role of a spring burst.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (partial, strong). USEPA/actonEC @ ceeab99 (R 3.6.1) ships ALL raw inputs + Level-2 products incl. the ANN-gap-filled EC FCH4 series (dataL2/gapFilledEC_results.csv, 32805 half-hours, with L95/U95 bands) and per-figure CSVs. I reproduced the headline pipeline-derived numbers in Python on «our HPC» compute node n137, directly from the shipped Level-2 data using the authors' OWN documented formulas (cumulative factor 603016/1e6 verified to equal the shipped ch4_cumulative; ecoQ10=10^(10b) of lm(log10(flux*24)~sedT) per fluxTmprPlotsFig4.R), avoiding the multi-hour stochastic ANN AND the fragile archived-CRAN R env. Result: EC cumulative annual emissions 2017/2018 = 40.69/71.47 g m-2 reproduce Table 1 (40.7/71.4) EXACTLY incl. ~bands (5.18/3.73 vs 5.9/4.2); EC warm-season mean total flux 9.60/17.49 vs 9.73/17.5 (within-tol/exact); the title 'spring burst' peak 61.96 mg m-2 h-1 reproduces 62.0 exactly (peak day 30 vs 29 May, 1-day timestamp-boundary). The 12-d burst cumulative 12.55 vs 10.8 g m-2 is ~16% high (partial; window-edge definition). ecoQ10 reproduces in magnitude/order (AFT 30-35 >> EC 5-9) with shallow-2017 exact (35.49 vs 35.1) and EC-2018 close, but 4/6 Table-3 values diverge 14-87% because the shipped Fig4data snapshot does not exactly reproduce the runtime-rebuilt regression inputs (zero-flux/Inf filtering). 12 claims graded: 4 exact, 3 within-tol, 4 partial, 1 mismatch. NOT attempted: AFT/GRTS/hybrid cumulative (Table 1) and 2DKS threshold (Table 3) -> need the trap/chamber/spsurvey R pipeline; documented as stretch, not fabricated. Dataset is open CC-BY, self-contained, grade B (deducted for missing ANN bootstrap RData + Table-3 ecoQ10 snapshot inconsistency). All grades PROVISIONAL pending human audit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study investigates the temporal patterns and biophysical drivers of CH4 emissions from a small eutrophic reservoir (Acton Lake), organized around two questions: how important interannual and intra-lake variability in CH4 emissions can be and what causes it, and how limited monitoring resources can best be used to constrain reservoir methane emissions.
- ★ Acton Lake had cumulative areal CH4 emission rates of 45.6 ± 8.3 g CH4 m-2 in 2017 and 51.4 ± 4.3 g CH4 m-2 in 2018 (109 ± 14 and 123 ± 10 Mg CH4 whole-reservoir, respectively) finding
- ★ The main difference between years was a spring burst of elevated emissions lasting less than 2 weeks in 2018, contributing 17% of annual emissions in the shallow region finding
- ★ The spring burst coincided with a phytoplankton bloom likely driven by favorable precipitation and temperature conditions in 2018 compared to 2017 mechanism
- ★ The relationship between CH4 emissions and sediment temperature depended on location within the reservoir finding
- ★ There was a clear spatiotemporal offset in maximum CH4 emissions as a function of reservoir depth finding
- ★ This study is only the second to report pseudo-continuous, multi-year eddy covariance CH4 flux results over open water, and the first for a temperate region, eutrophic system, and reservoir resource
- ★ An artificial neural network was used to gap-fill the eddy covariance time series and explore the relative importance of biophysical drivers at the interannual timescale method
- Spatially integrated GHG emission methods report higher F_CH4 than survey methods (per Deemer et al., 2016) finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Eddy covariance flux monitoring | whole-reservoir open water, Acton Lake | none | CH4, CO2, and water vapor fluxes | LI-7700 and LI-7500A IRGAs, R.M. Young Model 81000 sonic anemometer, LI-7550 datalogger, EddyPro v6.2 |
| Active funnel trap ebullition monitoring | sediment-water interface, shallow (U-14) and deep (U-12) sites, Acton Lake | none | ebullitive CH4 flux and bubble gas composition | differential pressure sensor funnel traps; Bruker 450 GC with flame ionization detector |
| Floating chamber diffusion measurements | water surface, two biweekly sites, Acton Lake | none | diffusive CH4 flux from change in headspace CH4 mixing ratio | CSIRO-design floating chamber with ABB Los Gatos UGGA (PN 915-0011) |
| Spatially extensive emission surveys | whole reservoir, Acton Lake | none | spatial distribution of CH4 emissions | — |
| Water temperature depth profiling | shallow (U-14) and deep (U-12) water columns, Acton Lake | none | water temperature profile, buoyancy frequency/water column stability | RBRsoloT thermistor string; YSI ProODO sondes |
| Multiparameter water quality sampling | shallow and deep sites, Acton Lake | none | temperature, specific conductivity, dissolved oxygen, pH, chlorophyll a | YSI multiparameter sonde |
| Chlorophyll a analysis | water samples from reservoir inlet, Acton Lake | none | chlorophyll a concentration | TD-700 fluorometer (Turner Designs) |
- ▲ Cumulative areal CH4 emissions were higher in 2018 (51.4 ± 4.3 g CH4 m-2) than 2017 (45.6 ± 8.3 g CH4 m-2) ~13% increase
- ▲ Whole-reservoir CH4 emissions were 109 ± 14 Mg (2017) and 123 ± 10 Mg (2018) across the 2.4 km2 lake
- ▲ A spring burst lasting <2 weeks in 2018 accounted for 17% of annual emissions in the shallow region 17%
- – Overall EC CH4 data acceptance rate over the 2-year monitoring period was 31.3%, lower in 2017 (23.4%) than 2018 (39.8%) 31.3%
- – EC S-2 (mid-reservoir, May-Nov 2018) data coverage was 52.8% 52.8%
- – At EC S-2, a u_star threshold of 0.07 m/s was used for low-turbulence filtering based on site-specific CH4/CO2 flux relationships 0.07 m/s
- mean 45.6 ± 8.3 g CH4 m-2 (2017 cumulative areal CH4 emission)
- mean 51.4 ± 4.3 g CH4 m-2 (2018 cumulative areal CH4 emission)
- mean 109 ± 14 Mg CH4 (2017 whole-reservoir CH4 emission)
- mean 123 ± 10 Mg CH4 (2018 whole-reservoir CH4 emission)
- other 17% (spring burst share of annual shallow-region emissions)
- other 31.3% (overall EC data acceptance rate, 2017-2018)
- other 23.4% (2017) vs 39.8% (2018) (annual EC data acceptance rates)
- other r2 > 0.9 (threshold for retaining chamber flux models)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper is a methods-focused environmental monitoring study combining eddy covariance (EC), active funnel trap ebullition monitoring, floating-chamber diffusion measurements, and spatial surveys to characterize methane emissions from a reservoir over two years. Data reduction relied on established micrometeorological QA/QC procedures (stationarity/turbulence tests, u*-threshold filtering), an artificial neural network for gap-filling and driver-importance exploration, and Akaike information criterion (AIC) for choosing between linear and nonlinear models of chamber gas accumulation. Results were reported as point estimates with a ± dispersion value (e.g., cumulative emissions of 45.6 ± 8.3 g CH4 m-2), though the provided text (which appears to be truncated before the Results/Discussion sections) does not describe formal hypothesis tests, p-values, or multiplicity corrections.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Akaike information criterion (AIC) model selection between linear and nonlinear models | fitting rate of change in chamber headspace CH4 mixing ratio (dχ/dt) for floating chamber diffusive flux calculation | — | not stated |
| Model goodness-of-fit threshold (r² > 0.9) for retaining chamber flux models | floating chamber diffusive flux measurements | — | not stated |
| Artificial neural network (ANN) | gap-filling the eddy covariance CH4 flux time series and exploring relative importance of biophysical drivers on interannual timescales | — | not stated |
| Site-specific u* (friction velocity) threshold filtering | quality control of EC S-2 flux data, based on relationship between u* and CH4/CO2 fluxes | — | not stated |
| Two-dimensional flux footprint prediction model | characterizing source area contributing to EC flux measurements at both tower sites | — | not stated |
-
An artificial neural network was used to gap-fill the eddy covariance CH4 flux time series and assess driver importance.↳ Could also: Marginal distribution sampling (MDS) or look-up-table gap-filling, as implemented in tools like REddyProc, could also be used — These methods are widely used in the flux-tower community, provide established uncertainty quantification for gap-filled values, and allow direct comparison with other long-term EC datasets that use the same standard gap-filling approach.
-
Linear vs. nonlinear models for chamber headspace CH4 accumulation were selected using AIC together with a fixed r² > 0.9 retention threshold.↳ Could also: Cross-validation-based model selection, or reporting AIC weights/AICc alongside the r² criterion, could also be used — Cross-validation can guard against overfitting in small per-deployment datasets, and reporting AIC weights conveys the relative support for competing models rather than relying on a single hard-cutoff criterion.
-
Annual and cumulative emission estimates are reported as a point estimate with a ± value without specifying whether it reflects SD, SE, or propagated uncertainty.↳ Could also: Explicitly stating the dispersion measure, and/or providing a bootstrapped or analytically propagated 95% confidence interval for the upscaled annual emission estimate, could also be used — Because these are upscaled values combining multiple measurement streams (EC, chambers, traps, surveys), a clearly labeled and appropriately propagated uncertainty interval would make the combined uncertainty easier to interpret and compare across studies.
-
A single site-specific u* threshold was used to filter EC S-2 data for low-turbulence periods.↳ Could also: A bootstrapped range of candidate u* thresholds (e.g., following Papale et al., 2006) could also be used — Testing multiple candidate thresholds and propagating the resulting variability into the flux uncertainty estimate is a common way to characterize how sensitive the filtered flux record is to the threshold choice.
-
Differences between 2017 and 2018 annual emissions (including the 2018 spring burst) are described narratively based on the magnitude of point estimates.↳ Could also: A formal statistical comparison, such as a t-test, permutation test, or comparison of overlapping confidence/bootstrap intervals, could also be used — A formal test or interval-overlap comparison would provide an explicit statistical statement about whether the interannual difference exceeds what would be expected from measurement and upscaling uncertainty alone, complementing the descriptive comparison.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35126532
Paper: Waldo, Beaulieu, Barnett, Balz, Vanni, Williamson, Walker (2021). "Temporal trends in methane emissions from a small eutrophic reservoir: the key role of a spring burst." Biogeosciences 18, 5291–5311. DOI 10.5194/bg-18-5291-2021. PMID 35126532 · PMCID PMC8815417.
Code: https://github.com/USEPA/actonEC — public, not archived. Pinned commit
ceeab99bae03f9cdd5dcbd27865b43bfee9ecf07 (master HEAD, committed 2020-11-29).
R pipeline (R 3.6.1). Entry point scripts/masterScript.R; env in
scripts/masterLibrary.R.
Data: Zenodo 10.5281/zenodo.4540271 = actonEC-1.0.zip (50 MB) — a tagged
snapshot of the SAME repo (title: "R Code for: …"). The repo itself ships ALL raw
inputs under data/ (~150 MB of CSV/xlsx: EddyPro output epOutOrder.csv 43 MB,
LGR gga.csv 24 MB, HOBO trap hobo.csv 24 MB, met vanni30min.csv, thermistor
strings, GC master files, ebullition xlsx, GRTS survey shapefiles) AND precomputed
Level-2 products under dataL2/ (per-figure CSVs + gapFilledEC_results.csv, the
ANN-gapfilled EC FCH4 time series). The only thing NOT in the repo is the ANN
bootstrap resamples BestANNsResampleNN.RData (NN=01..20), said to be on Zenodo.
What this study reports (candidate claims)
See original/claims.tsv (28 rows). Headline pipeline-derived results:
- Cumulative annual CH4 emissions 2017 & 2018 by method (EC, shallow/deep AFT+chamber, GRTS survey, hybrid upscaled) — Table 1.
- Spring-burst magnitude (peak FCH4, 12-day cumulative, % of annual) — Sec 3.1.
- ecoQ10 temperature sensitivities (EC, AFT shallow/deep, 2017/2018) — Table 3.
- GRTS whole-lake warm-season mean diffusive/ebullitive/total flux ± 95% CI — Table 1.
- Ebullition vs diffusion partitioning (% of total FCH4) — Sec 3.3.
IN SCOPE (pipeline-derived, attempted)
| Result group | Pipeline / script | Inputs | Notes |
|---|---|---|---|
| EC cumulative emissions, spring burst, time series (Fig 2, 9) | qcECfluxes.R → ANN gapfill → plotCumulativeTSFig9.R / plotTimeSeriesFig2.R | data/epOutOrder.csv, dataL2/gapFilledEC_results.csv |
Downstream figs read the SHIPPED gapfilled CSV → reproducible WITHOUT rerunning the multi-hour ANN. Re-running the ANN itself is a stretch goal (stochastic; bootstrap RData not in repo). |
| GRTS whole-lake means + 95% CI (Table 1) | grtsReadSiteData.R → grtsCalcEmissions.R → grtsLakeCalcs.R | data/survey/, data/grtsEqArea/ |
spsurvey-style GRTS variance estimator coded in masterLibrary.R. |
| Chamber emissions | plotCleanLgr.R → calculateChamberEmissions.R | data/gga.csv, data/chamberBiweekly.xlsx |
|
| Ebullition time series + uncertainty | calculateEbullition.R | data/hobo.csv, ebullition xlsx |
|
| Dissolved/saturated gas | dissolvedGasCalc.R (def.calc.sdg.R, GLEON) | data/ GC master files |
|
| ecoQ10 temperature sensitivity (Table 3) | fluxTmprPlotsFig4.R | dataL2/Fig4data.csv |
OUT OF SCOPE (not attempted)
- Wet-lab / field measurement steps (GC analysis, trap deployment, sonde calibration).
- The ANN bootstrap uncertainty resamples (
BestANNsResampleNN.RData) — not shipped in the repo; on Zenodo per README but the Zenodo deposit is only the 50 MB code zip, so their availability is uncertain (will verify on «our HPC»). The point estimate gap-filled series IS shipped (dataL2/gapFilledEC_results.csv), so EC cumulative totals are reproducible; the ± uncertainty half-widths may not be. - External reference datasets (NADP/NTN deposition
NTN-OH09-d.csv, Miami Univ. hydrology/chl) are inputs, not reproduced.
Strategy
Primary target (fast 80%): regenerate the reported numeric values by running the
data-processing + visualization scripts against the shipped data/ + dataL2/
inputs and reading off Table 1 / Table 3 / Sec 3.1 quantities. This avoids the
multi-hour stochastic ANN while still reproducing the headline cumulative-emission,
Q10, GRTS, and partitioning numbers from the authors' own data + code.
Stretch: re-run the ANN gap-fill
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.