Sediment Resuspension as a System-Wide Driver of Legacy and Bioavailable Phosphorus Release in Lake Erie.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (1:1). The paper ships a single self-contained Jupyter notebook (authors' own code) + a complete data deposit (Zenodo v1.0.0 / Git-LFS). Re-running it end-to-end on «our HPC» reproduced every headline numeric claim: total mobilized TP 40,314 kg and bioavailable SRP 12,665 kg reproduced TO THE KILOGRAM; SPM-algorithm NRMSE 13.93% vs reported 13.9%; R2(TP~SPM) 0.734/0.650 vs 0.73/0.66; the ANOVA structure (P fractions non-significant by depth, highly significant by station; metals largely stable 2016-2023) matches SI Tables 4 and 5; Db mixing rates match the shipped values exactly. 9/14 claims exact, 3 within-tol, 2 partial (western-basin sediment share 91.4% vs 89.2%; a few metals significant by year), 0 mismatch. NOT attempted: the wet-lab/field measurements that are the shipped inputs (7Be counting, SEDEX fractionation, ICP metals, in-situ sampling) - not recomputable. 8 figure-only cells initially failed purely on library-version drift (pandas delim_whitespace, matplotlib cm.get_cmap, pandas-2.x str dtype); after three documented one-line compat shims the whole notebook runs. Only data friction: the Zenodo GitHub-source zip ships Git-LFS pointer stubs, so a separate 'git lfs pull' is required to obtain the satellite/coastline binaries.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 90assessed: 2026-06-21 ⛓ 37e74d7f748e
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether episodic sediment resuspension is an overlooked but significant pathway of internal legacy phosphorus loading in Lake Erie, capable of releasing bioavailable phosphorus at magnitudes that rival or exceed steady-state diffusive fluxes and known tributary loading.
- ★ Sediment resuspension is a major episodic internal phosphorus source releasing bioavailable P far exceeding previously reported aerobic diffusive fluxes. finding
- ★ Beryllium-7 depth profiles show repeated sediment mixing reworks multiyear deposits and remobilizes legacy phosphorus during resuspension. finding
- ★ A single-wavelength semiempirical SPM algorithm applied to Sentinel-3 OLCI reflectance enables a mechanistic, spatially resolved, basin-wide framework linking benthic sediment traits to satellite-derived SPM for P release estimation. method
- Bioavailable phosphorus (BioP) is defined as the labile TP fraction comprising loosely bound, redox-sensitive, and Al/Fe oxide-bound P that can be released during resuspension and contribute to dissolved SRP. method
- ★ The single observed May 2023 resuspension event contributed ~4.8% and ~7.0% of Maumee River spring total P and soluble reactive P target loads, respectively. finding
- Nearshore sites (e.g., WLE13) show frequent resuspension and preferential loss of fine 7Be-bearing particles, while offshore sites (WLE14, WLE15) act as depositional/focusing zones with deeper 7Be penetration. finding
- ★ Episodic internal P loading from sediment resuspension has not previously been discretely characterized at the whole-system scale in large lakes like Lake Erie. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| sediment core collection and depth sectioning | western Lake Erie benthic sediment (stations WLE1, WLE13, WLE14, WLE15) | none (seasonal in situ sampling 2023-2024) | sediment depth profiles (0-5 cm, 1 cm intervals) | GMX-25 Gomex box corer |
| sequential phosphorus fractionation (5-step) | sediment cores | none | P fractions: loosely bound, redox-sensitive, Al/Fe oxide-bound, Ca-bound, residual organic P | Seal AQ2+ Automated Discrete Analyzer (EPA Methods 118A/134A) |
| Beryllium-7 gamma spectrometry | sediment cores and suspended particulate matter | none | 7Be activity/inventory, mixing rate, mixing depth, particle residence time, settling and accumulation rates | EG&G Ortec lithium-drifted germanium detector / intrinsic germanium detector with multichannel analyzer |
| water column TP and SRP measurement | Lake Erie surface water (1 m depth), stations WLE13/14/16 | natural wind-driven resuspension event (May 2023) | TP and SRP concentrations before/after resuspension | Seal AQ2+ Analyzer (EPA Method 365.1) |
| SPM gravimetric filtration | lake water samples | none | suspended particulate matter concentration | precombusted glass fiber filters (0.7 µm), 105°C drying |
| metal extraction and ICP-MS analysis | sediment | none | metal concentrations (compared to 2016 data) | ICP-MS (USEPA Method 3050B) |
| satellite remote sensing SPM retrieval | western Lake Erie basin-wide | natural wind-driven resuspension event (May 18 vs May 26, 2023) | ΔSPM and spatially resolved internal P loading estimates | Sentinel-3A/B OLCI, POLYMER atmospheric correction, Nechad et al. algorithm |
| in situ hyperspectral radiometry validation | Great Lakes (117 sites, 78 in Lake Erie, 2023-2024) | none | validation of satellite SPM algorithm accuracy | SpectraVista Corporation HR-512i spectroradiometer |
- ▲ Sediments released bioavailable P during the observed resuspension event 2.3-11 x 10^-2 g m-2
- ▲ Bioavailable P release during resuspension greatly exceeded previously reported aerobic diffusive fluxes 22-256x
- ▲ Single resuspension event's contribution to Maumee River spring P target loads ~4.8% (TP), ~7.0% (SRP)
- – 7Be detected below the uppermost sediment layer, indicating vertical redistribution by mixing
- ▼ WLE13 (nearshore) showed lower but more variable 7Be inventories, including periods of minimal activity in May 2023
- ▲ Offshore stations WLE14 and WLE15 exhibited higher 7Be activities reaching 4-5 cm depth, indicating focusing/depositional zones
- – Mass-specific SPM 7Be activity was low at WLE13, consistent with limited new material deposition 6.62 dpm g-1
- fold_change 22-256x (bioavailable P release from resuspension vs. previously reported aerobic diffusive fluxes)
- other 2.3-11 x 10^-2 g m-2 (bioavailable P released during the observed resuspension event)
- other ~4.8% (event contribution to Maumee River spring total P target load)
- other ~7.0% (event contribution to Maumee River spring soluble reactive P target load)
- mean 6.62 dpm g-1 (mass-specific SPM 7Be activity at station WLE13)
- count n = 12 (surface sediment samples collected via Ponar grab, May and August 2022)
- other T1/2 ~ 53 days (half-life of 7Be radionuclide tracer used for sediment mixing analysis)
- other 1.1-2.0 x 10^6 kg TP; 2.2-4.0 x 10^5 kg SRP (Maumee River spring P inputs during 2017-2021)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combines field-based sediment and water-column sampling with satellite remote sensing to quantify phosphorus release during a sediment resuspension event in Lake Erie. Statistical treatment centers on one-way ANOVA with Tukey HSD post-hoc comparisons to assess temporal (across months) and spatial (across stations) variation in sediment phosphorus fractions, alongside descriptive metrics (e.g., coefficient of variation) and an algorithm-validation error metric (NRMSE) for the satellite SPM retrieval. Most other quantitative results (fluxes, mixing rates, accumulation rates, mass balance estimates) are derived from mechanistic/empirical equations (e.g., 7Be decay modeling, SPM-based erosion depth) rather than inferential hypothesis tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| One-way ANOVA | Temporal variation in sediment P fractions across months (April–June) within each station | n = 10 to 20 | not stated |
| One-way ANOVA | Spatial variation in P concentrations among core stations within each month | n = 20 | not stated |
| Tukey HSD post-hoc pairwise comparisons | Specific month and station pair comparisons following the ANOVAs | — | not stated |
-
Temporal and spatial variation in sediment P fractions were assessed with separate one-way ANOVAs (one for month-within-station, one for station-within-month)↳ Could also: A two-way (or mixed-effects) ANOVA with station and month as crossed factors, potentially with station treated as a random effect — This would allow direct testing of a station-by-month interaction (i.e., whether temporal patterns differ by location) in a single model, and a mixed-effects framework could account for repeated sampling at the same stations over time
-
Tukey HSD was used for post-hoc pairwise comparisons after the ANOVAs↳ Could also: Alternative multiple-comparison procedures such as Bonferroni, Šidák, or a false-discovery-rate method (e.g., Benjamini-Hochberg) — These offer different balances of power versus family-wise/false-discovery control and are commonly chosen depending on how many pairwise comparisons are made across the full set of stations, months, and P fractions
-
Spatial variability in P fractions was summarized using the coefficient of variation (SD divided by mean) across stations↳ Could also: Reporting a 95% confidence interval or standard error alongside the mean — CIs directly convey the precision of the estimated mean and are often preferred for communicating uncertainty in small-n environmental sampling designs
-
The SPM satellite algorithm's performance was validated using a single normalized RMSE (NRMSE) threshold to classify resuspension pixels↳ Could also: Reporting complementary validation metrics such as bias, mean absolute error, or R² alongside RMSE — Multiple accuracy metrics together give a fuller picture of algorithm performance (e.g., distinguishing systematic bias from random scatter) beyond a single normalized error threshold
-
Sample sizes for the ANOVA comparisons (n = 10–20) were reported without an accompanying power analysis↳ Could also: An a priori or post-hoc power analysis, or reporting effect sizes (e.g., eta-squared) alongside the ANOVA results — This would help contextualize the practical magnitude of detected differences and the sensitivity of the design to detect smaller effects, complementing significance testing
-
Quality assurance relied on duplicate/triplicate analyses and recovery checks against certified reference materials↳ Could also: Explicitly reporting relative standard deviation (%RSD) or CV across replicate measurements in the main text — This would make measurement precision directly visible alongside the reported concentrations, complementing the recovery-based QA already described
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41960750
Paper: Dhiman et al. (2026) "Sediment Resuspension as a System-Wide Driver of Legacy and Bioavailable Phosphorus Release in Lake Erie." Environ Sci Technol. DOI 10.1021/acs.est.5c17601 · PMCID PMC13130958.
Code: https://github.com/Carbon-and-Optics/wle_p_analysis (commit 57e4dfc;
Zenodo v1.0.0 snapshot commit d1db5a8, DOI 10.5281/zenodo.19489452).
Data: shipped inside the repo under data/ (Git-LFS for the Sentinel-3
.nc scenes + GSHHS coastlines; the GitHub source zip on Zenodo ships only LFS
pointers, so LFS objects must be pulled separately — done).
Nature of the work
A single self-contained Jupyter notebook
analysis/manuscript_and_SI_figures_and_analysis.ipynb (112 cells) regenerates
all manuscript + SI figures and the statistical tables from the shipped inputs.
No external downloads at run time. This is a third-party-tool-free, authors'-own
pipeline reproduction: run the notebook end-to-end and compare its emitted
numbers/figures to the paper.
IN SCOPE (pipeline-derived, reproduced by running the notebook)
| Result | Notebook location | Output kind |
|---|---|---|
| Fig 1a/b SPM maps (Nechad SPM from Sentinel-3, GSHHS land mask) | cells 1–13 | map PNG + derived .nc |
| Fig 1c ΔSPM (May26−May18) + 13.9% NRMSE threshold mask | cells 14–24 | map PNG, masked area |
| Fig 5 P estimates from satellite ΔSPM (TP / bioP release kg) | cells 25–42 | numbers + map |
| Fig 2 Be-7 distribution 2023/2024 | cells 43–49 | figure |
| Fig 3 P fractionation depth profiles + boxplots | cells 50–56 | figure |
| SI Fig 1 Db mixing rates from 7Be | cells 61–62 | figure + Db values |
| SI Fig 2 measured/modeled 7Be vertical profiles | cells 63–68 | figure |
| SI Fig 3 7Be inventories over time | cells 69–70 | figure |
| SI Fig 4 SPM vs TP / SPM vs Chl linear fits (R²) | cells 71–79 | R² values |
| SI Fig 5 SPM algorithm validation | cells 80–82 | figure + error |
| SI Table 4 one-way ANOVA, P fractions across stations & months | cells 83–104 | F/p tables |
| SI Table 5 one-way ANOVA, metals across years (2016–2023) | cells 105–110 | F/p table |
Primary numeric claims to grade (highest auditability): NRMSE 13.9%; R²(TP~SPM) 0.73 (2023) / 0.66 (2024); total P 40,314 kg & bioP 12,665 kg per May-2023 event; erosion depths (e.g. WLE13 1.39 cm); Db ranges; SI Table 4/5 ANOVA significance.
OUT OF SCOPE (not pipeline-derived; not attempted)
- Wet-lab/field measurements themselves (7Be gamma counting, SEDEX P fractionation chemistry, ICP metals, in-situ SPM/TP/SRP sampling) — these are the inputs shipped as CSVs, not recomputable.
- Sediment accumulation/residence-time interpretive numbers stated in text that derive from external models not in the notebook (flagged if encountered).
Compute plan
Run the whole notebook via jupyter nbconvert --execute as a SLURM job on «our HPC»
(reads inputs from «infra»). Capture the executed notebook (all stdout/tables) +
regenerated figures, then grade printed numbers vs the paper.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Clean 1:1 reproduction. The paper ships the authors' own self-contained Jupyter notebook plus a complete Zenodo/Git-LFS data deposit; re-running end-to-end on «our HPC» reproduced both headline phosphorus masses to the kilogram (40,314 kg TP; 12,665 kg SRP), the SPM NRMSE to 0.03 pp, and the full ANOVA significance structure. The only deviations are on our/technical side: C13 western-basin share 91.39% vs 89.2% (~2 pp, likely a minor lat-cutoff/preprocessing difference) and C8 three metals (Mo/Mn/Co) significant by year against a 'stable' claim — both minor, same-direction, and not touching the central conclusion. No fabrication signal: every reported value is derivable from the shared data, and figure-cell failures were pure library-version drift fixed with no logic change.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.