Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Temporal trends in methane emissions from a small eutrophic reservoir: the key role of a spring burst.

Biogeosciences · 2021
72/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
72/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 688 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (partial, strong). USEPA/actonEC @ ceeab99 (R 3.6.1) ships ALL raw inputs + Level-2 products incl. the ANN-gap-filled EC FCH4 series (dataL2/gapFilledEC_results.csv, 32805 half-hours, with L95/U95 bands) and per-figure CSVs. I reproduced the headline pipeline-derived numbers in Python on «our HPC» compute node n137, directly from the shipped Level-2 data using the authors' OWN documented formulas (cumulative factor 603016/1e6 verified to equal the shipped ch4_cumulative; ecoQ10=10^(10b) of lm(log10(flux*24)~sedT) per fluxTmprPlotsFig4.R), avoiding the multi-hour stochastic ANN AND the fragile archived-CRAN R env. Result: EC cumulative annual emissions 2017/2018 = 40.69/71.47 g m-2 reproduce Table 1 (40.7/71.4) EXACTLY incl. ~bands (5.18/3.73 vs 5.9/4.2); EC warm-season mean total flux 9.60/17.49 vs 9.73/17.5 (within-tol/exact); the title 'spring burst' peak 61.96 mg m-2 h-1 reproduces 62.0 exactly (peak day 30 vs 29 May, 1-day timestamp-boundary). The 12-d burst cumulative 12.55 vs 10.8 g m-2 is ~16% high (partial; window-edge definition). ecoQ10 reproduces in magnitude/order (AFT 30-35 >> EC 5-9) with shallow-2017 exact (35.49 vs 35.1) and EC-2018 close, but 4/6 Table-3 values diverge 14-87% because the shipped Fig4data snapshot does not exactly reproduce the runtime-rebuilt regression inputs (zero-flux/Inf filtering). 12 claims graded: 4 exact, 3 within-tol, 4 partial, 1 mismatch. NOT attempted: AFT/GRTS/hybrid cumulative (Table 1) and 2DKS threshold (Table 3) -> need the trap/chamber/spsurvey R pipeline; documented as stretch, not fabricated. Dataset is open CC-BY, self-contained, grade B (deducted for missing ANN bootstrap RData + Table-3 ecoQ10 snapshot inconsistency). All grades PROVISIONAL pending human audit.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.4540271

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study investigates the temporal patterns and biophysical drivers of CH4 emissions from a small eutrophic reservoir (Acton Lake), organized around two questions: how important interannual and intra-lake variability in CH4 emissions can be and what causes it, and how limited monitoring resources can best be used to constrain reservoir methane emissions.

Core claims
  • Acton Lake had cumulative areal CH4 emission rates of 45.6 ± 8.3 g CH4 m-2 in 2017 and 51.4 ± 4.3 g CH4 m-2 in 2018 (109 ± 14 and 123 ± 10 Mg CH4 whole-reservoir, respectively) finding
  • The main difference between years was a spring burst of elevated emissions lasting less than 2 weeks in 2018, contributing 17% of annual emissions in the shallow region finding
  • The spring burst coincided with a phytoplankton bloom likely driven by favorable precipitation and temperature conditions in 2018 compared to 2017 mechanism
  • The relationship between CH4 emissions and sediment temperature depended on location within the reservoir finding
  • There was a clear spatiotemporal offset in maximum CH4 emissions as a function of reservoir depth finding
  • This study is only the second to report pseudo-continuous, multi-year eddy covariance CH4 flux results over open water, and the first for a temperate region, eutrophic system, and reservoir resource
  • An artificial neural network was used to gap-fill the eddy covariance time series and explore the relative importance of biophysical drivers at the interannual timescale method
  • Spatially integrated GHG emission methods report higher F_CH4 than survey methods (per Deemer et al., 2016) finding
Experimental setups
Assay System Perturbation Readout Platform
Eddy covariance flux monitoring whole-reservoir open water, Acton Lake none CH4, CO2, and water vapor fluxes LI-7700 and LI-7500A IRGAs, R.M. Young Model 81000 sonic anemometer, LI-7550 datalogger, EddyPro v6.2
Active funnel trap ebullition monitoring sediment-water interface, shallow (U-14) and deep (U-12) sites, Acton Lake none ebullitive CH4 flux and bubble gas composition differential pressure sensor funnel traps; Bruker 450 GC with flame ionization detector
Floating chamber diffusion measurements water surface, two biweekly sites, Acton Lake none diffusive CH4 flux from change in headspace CH4 mixing ratio CSIRO-design floating chamber with ABB Los Gatos UGGA (PN 915-0011)
Spatially extensive emission surveys whole reservoir, Acton Lake none spatial distribution of CH4 emissions
Water temperature depth profiling shallow (U-14) and deep (U-12) water columns, Acton Lake none water temperature profile, buoyancy frequency/water column stability RBRsoloT thermistor string; YSI ProODO sondes
Multiparameter water quality sampling shallow and deep sites, Acton Lake none temperature, specific conductivity, dissolved oxygen, pH, chlorophyll a YSI multiparameter sonde
Chlorophyll a analysis water samples from reservoir inlet, Acton Lake none chlorophyll a concentration TD-700 fluorometer (Turner Designs)
Key results
  • Cumulative areal CH4 emissions were higher in 2018 (51.4 ± 4.3 g CH4 m-2) than 2017 (45.6 ± 8.3 g CH4 m-2) ~13% increase
  • Whole-reservoir CH4 emissions were 109 ± 14 Mg (2017) and 123 ± 10 Mg (2018) across the 2.4 km2 lake
  • A spring burst lasting <2 weeks in 2018 accounted for 17% of annual emissions in the shallow region 17%
  • Overall EC CH4 data acceptance rate over the 2-year monitoring period was 31.3%, lower in 2017 (23.4%) than 2018 (39.8%) 31.3%
  • EC S-2 (mid-reservoir, May-Nov 2018) data coverage was 52.8% 52.8%
  • At EC S-2, a u_star threshold of 0.07 m/s was used for low-turbulence filtering based on site-specific CH4/CO2 flux relationships 0.07 m/s
Key statistics
  • mean 45.6 ± 8.3 g CH4 m-2 (2017 cumulative areal CH4 emission)
  • mean 51.4 ± 4.3 g CH4 m-2 (2018 cumulative areal CH4 emission)
  • mean 109 ± 14 Mg CH4 (2017 whole-reservoir CH4 emission)
  • mean 123 ± 10 Mg CH4 (2018 whole-reservoir CH4 emission)
  • other 17% (spring burst share of annual shallow-region emissions)
  • other 31.3% (overall EC data acceptance rate, 2017-2018)
  • other 23.4% (2017) vs 39.8% (2018) (annual EC data acceptance rates)
  • other r2 > 0.9 (threshold for retaining chamber flux models)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper is a methods-focused environmental monitoring study combining eddy covariance (EC), active funnel trap ebullition monitoring, floating-chamber diffusion measurements, and spatial surveys to characterize methane emissions from a reservoir over two years. Data reduction relied on established micrometeorological QA/QC procedures (stationarity/turbulence tests, u*-threshold filtering), an artificial neural network for gap-filling and driver-importance exploration, and Akaike information criterion (AIC) for choosing between linear and nonlinear models of chamber gas accumulation. Results were reported as point estimates with a ± dispersion value (e.g., cumulative emissions of 45.6 ± 8.3 g CH4 m-2), though the provided text (which appears to be truncated before the Results/Discussion sections) does not describe formal hypothesis tests, p-values, or multiplicity corrections.

Replicationunclear Sample sizeDescribed qualitatively in terms of monitoring duration and data acceptance/coverage rates (e.g., 31.3% overall EC data acceptance rate over the 2-year period, 23.4% in 2017, 39.8% in 2018), rather than as a sample-size or power calculation for hypothesis testing. Groups2017 vs. 2018 annual emissions; shallow vs. deep reservoir sites Pairingunclear Randomization/blindingnot stated Dispersionunclear
Statistical tests used
Test Applied to n Assumptions
Akaike information criterion (AIC) model selection between linear and nonlinear models fitting rate of change in chamber headspace CH4 mixing ratio (dχ/dt) for floating chamber diffusive flux calculation not stated
Model goodness-of-fit threshold (r² > 0.9) for retaining chamber flux models floating chamber diffusive flux measurements not stated
Artificial neural network (ANN) gap-filling the eddy covariance CH4 flux time series and exploring relative importance of biophysical drivers on interannual timescales not stated
Site-specific u* (friction velocity) threshold filtering quality control of EC S-2 flux data, based on relationship between u* and CH4/CO2 fluxes not stated
Two-dimensional flux footprint prediction model characterizing source area contributing to EC flux measurements at both tower sites not stated
Approaches that could also have been used
  • An artificial neural network was used to gap-fill the eddy covariance CH4 flux time series and assess driver importance.
    Could also: Marginal distribution sampling (MDS) or look-up-table gap-filling, as implemented in tools like REddyProc, could also be used — These methods are widely used in the flux-tower community, provide established uncertainty quantification for gap-filled values, and allow direct comparison with other long-term EC datasets that use the same standard gap-filling approach.
  • Linear vs. nonlinear models for chamber headspace CH4 accumulation were selected using AIC together with a fixed r² > 0.9 retention threshold.
    Could also: Cross-validation-based model selection, or reporting AIC weights/AICc alongside the r² criterion, could also be used — Cross-validation can guard against overfitting in small per-deployment datasets, and reporting AIC weights conveys the relative support for competing models rather than relying on a single hard-cutoff criterion.
  • Annual and cumulative emission estimates are reported as a point estimate with a ± value without specifying whether it reflects SD, SE, or propagated uncertainty.
    Could also: Explicitly stating the dispersion measure, and/or providing a bootstrapped or analytically propagated 95% confidence interval for the upscaled annual emission estimate, could also be used — Because these are upscaled values combining multiple measurement streams (EC, chambers, traps, surveys), a clearly labeled and appropriately propagated uncertainty interval would make the combined uncertainty easier to interpret and compare across studies.
  • A single site-specific u* threshold was used to filter EC S-2 data for low-turbulence periods.
    Could also: A bootstrapped range of candidate u* thresholds (e.g., following Papale et al., 2006) could also be used — Testing multiple candidate thresholds and propagating the resulting variability into the flux uncertainty estimate is a common way to characterize how sensitive the filtered flux record is to the threshold choice.
  • Differences between 2017 and 2018 annual emissions (including the 2018 spring burst) are described narratively based on the magnitude of point estimates.
    Could also: A formal statistical comparison, such as a t-test, permutation test, or comparison of overlapping confidence/bootstrap intervals, could also be used — A formal test or interval-overlap comparison would provide an explicit statistical statement about whether the interannual difference exceeds what would be expected from measurement and upscaling uncertainty alone, complementing the descriptive comparison.
Software: EddyPro 6.2 · R (including rLakeAnalyzer package and custom scripts) · Online two-dimensional flux-footprint prediction tool (Kljun et al., 2015)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35126532

Paper: Waldo, Beaulieu, Barnett, Balz, Vanni, Williamson, Walker (2021). "Temporal trends in methane emissions from a small eutrophic reservoir: the key role of a spring burst." Biogeosciences 18, 5291–5311. DOI 10.5194/bg-18-5291-2021. PMID 35126532 · PMCID PMC8815417.

Code: https://github.com/USEPA/actonEC — public, not archived. Pinned commit ceeab99bae03f9cdd5dcbd27865b43bfee9ecf07 (master HEAD, committed 2020-11-29). R pipeline (R 3.6.1). Entry point scripts/masterScript.R; env in scripts/masterLibrary.R.

Data: Zenodo 10.5281/zenodo.4540271 = actonEC-1.0.zip (50 MB) — a tagged snapshot of the SAME repo (title: "R Code for: …"). The repo itself ships ALL raw inputs under data/ (~150 MB of CSV/xlsx: EddyPro output epOutOrder.csv 43 MB, LGR gga.csv 24 MB, HOBO trap hobo.csv 24 MB, met vanni30min.csv, thermistor strings, GC master files, ebullition xlsx, GRTS survey shapefiles) AND precomputed Level-2 products under dataL2/ (per-figure CSVs + gapFilledEC_results.csv, the ANN-gapfilled EC FCH4 time series). The only thing NOT in the repo is the ANN bootstrap resamples BestANNsResampleNN.RData (NN=01..20), said to be on Zenodo.

What this study reports (candidate claims)

See original/claims.tsv (28 rows). Headline pipeline-derived results:

  • Cumulative annual CH4 emissions 2017 & 2018 by method (EC, shallow/deep AFT+chamber, GRTS survey, hybrid upscaled) — Table 1.
  • Spring-burst magnitude (peak FCH4, 12-day cumulative, % of annual) — Sec 3.1.
  • ecoQ10 temperature sensitivities (EC, AFT shallow/deep, 2017/2018) — Table 3.
  • GRTS whole-lake warm-season mean diffusive/ebullitive/total flux ± 95% CI — Table 1.
  • Ebullition vs diffusion partitioning (% of total FCH4) — Sec 3.3.

IN SCOPE (pipeline-derived, attempted)

Result group Pipeline / script Inputs Notes
EC cumulative emissions, spring burst, time series (Fig 2, 9) qcECfluxes.R → ANN gapfill → plotCumulativeTSFig9.R / plotTimeSeriesFig2.R data/epOutOrder.csv, dataL2/gapFilledEC_results.csv Downstream figs read the SHIPPED gapfilled CSV → reproducible WITHOUT rerunning the multi-hour ANN. Re-running the ANN itself is a stretch goal (stochastic; bootstrap RData not in repo).
GRTS whole-lake means + 95% CI (Table 1) grtsReadSiteData.R → grtsCalcEmissions.R → grtsLakeCalcs.R data/survey/, data/grtsEqArea/ spsurvey-style GRTS variance estimator coded in masterLibrary.R.
Chamber emissions plotCleanLgr.R → calculateChamberEmissions.R data/gga.csv, data/chamberBiweekly.xlsx
Ebullition time series + uncertainty calculateEbullition.R data/hobo.csv, ebullition xlsx
Dissolved/saturated gas dissolvedGasCalc.R (def.calc.sdg.R, GLEON) data/ GC master files
ecoQ10 temperature sensitivity (Table 3) fluxTmprPlotsFig4.R dataL2/Fig4data.csv

OUT OF SCOPE (not attempted)

  • Wet-lab / field measurement steps (GC analysis, trap deployment, sonde calibration).
  • The ANN bootstrap uncertainty resamples (BestANNsResampleNN.RData) — not shipped in the repo; on Zenodo per README but the Zenodo deposit is only the 50 MB code zip, so their availability is uncertain (will verify on «our HPC»). The point estimate gap-filled series IS shipped (dataL2/gapFilledEC_results.csv), so EC cumulative totals are reproducible; the ± uncertainty half-widths may not be.
  • External reference datasets (NADP/NTN deposition NTN-OH09-d.csv, Miami Univ. hydrology/chl) are inputs, not reproduced.

Strategy

Primary target (fast 80%): regenerate the reported numeric values by running the data-processing + visualization scripts against the shipped data/ + dataL2/ inputs and reading off Table 1 / Table 3 / Sec 3.1 quantities. This avoids the multi-hour stochastic ANN while still reproducing the headline cumulative-emission, Q10, GRTS, and partitioning numbers from the authors' own data + code. Stretch: re-run the ANN gap-fill

Figures / tables: Table
EC_cum_2017
Reported
40.7 +/- 5.9 g CH4 m-2
Reproduced
40.69 (band +/-5.18)
exact
EC_cum_2018
Reported
71.4 +/- 4.2 g CH4 m-2
Reproduced
71.47 (band +/-3.73)
exact
EC_warm_2017
Reported
9.73 +/- 0.67 mg CH4 m-2 h-1
Reproduced
9.60
within tolerance
EC_warm_2018
Reported
17.5 +/- 0.38 mg CH4 m-2 h-1
Reproduced
17.49
exact
burst_peak
Reported
62.0 mg CH4 m-2 h-1 on 29 May 2018
Reproduced
61.96 on 30 May 2018
within tolerance
burst_cum12d
Reported
10.8 g CH4 m-2 (15% of 2018)
Reproduced
12.55 (17.6%)
partial
q10_ec_2017
Reported
6.96 (R2 0.85)
Reproduced
9.50 (R2 0.70)
partial
q10_ec_2018
Reported
5.64 (R2 0.83)
Reproduced
5.43 (R2 0.80)
within tolerance
q10_shal_2017
Reported
35.1 (R2 0.48)
Reproduced
35.49 (R2 0.49)
exact
q10_shal_2018
Reported
35.8 (R2 0.85)
Reproduced
42.36 (R2 0.89)
partial
q10_deep_2017
Reported
30.4 (R2 0.60)
Reproduced
26.17 (R2 0.66)
partial
q10_deep_2018
Reported
30.7 (R2 0.38)
Reproduced
57.47 (R2 0.53)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.