Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Electron Energy Partition across Interplanetary Shocks. I. Methodology and Data Product.

Astrophys J Suppl Ser · 2019
L1 92/100 PQI 93
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
How its reproducibility compares
92/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 83% of all assessed papers rank 179 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1. The paper IS a methodology + data-product paper; its headline Table 2 (electron VDF fit exponents) and Sec 3.4 (densities, reduced chi-squared, percent deviation) summary statistics were recomputed DIRECTLY from the authors' shipped Zenodo data product (doi:10.5281/zenodo.2875806, file Wind_ip_shock_3dp_fit_results_electrons.txt, md5 727d923f0e5b86f32b929bae41c3b464). 17 claims graded: 10 exact, 5 within-tol, 2 partial. The exact reproduction of the reported population counts (stable halo fits 13871, beam fits 9567) independently confirms the column mapping. Only real discrepancy: deltaR mean (14.6 vs 12.7) -- its median/Q1/Q3 reproduce, so the paper's mean almost certainly excludes the FFlag=0 fits whose %R is capped at 100; not a fabrication signal. Core-exponent quartiles are within-tol because our component sub-counts run ~2% low. NOT attempted (80/20 tail): (1) re-running the authors' IDL fitter wind_3dp_pros -- needs an IDL license + the raw Wind/3DP telemetry, and the data product already is its output, so recomputing from it is the faithful auditable route; (2) the upstream/downstream/Mach/theta_Bn sub-column splits of Table 2 -- require per-VDF shock-region classification from shock arrival times. No fabrication concern: every headline number regenerates from the public product to printed precision.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.2875806

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 92
    assessed: 2026-06-14 ⛓ 258a02e15d13
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How is electron kinetic energy partitioned across interplanetary shocks, and what model velocity distribution functions best describe the cold core, hot halo, and field-aligned beam/strahl electron components observed near 1 au?

Core claims
  • Solar wind electron VDFs below ~1.2 keV are best modeled as the sum of three components: a cold dense core, a hot tenuous halo, and a field-aligned beam/strahl. method
  • This is the first statistical study to show the core electron distribution is better fit by a self-similar VDF than a bi-Maxwellian under all conditions. finding
  • The core is best modeled as a bi-kappa or symmetric/asymmetric bi-self-similar VDF, while halo and beam/strahl are best fit by a bi-kappa VDF. method
  • The self-similar distribution's deviation from a Maxwellian is a measure of inelasticity in particle scattering from waves and/or turbulence. mechanism
  • A data product of fit parameters for 15,314 electron VDFs around 52 IP shocks observed by Wind near 1 au is produced and documented. resource
  • The asymmetric bi-self-similar function produces the best core fit in the downstream of strong (⟨M_f⟩up ≳ 2.5) IP shocks. finding
  • An improved semianalytic relationship between spacecraft potential and ion number density was derived. method
Experimental setups
Assay System Perturbation Readout Platform
Electron velocity distribution function measurement and model fitting (nonlinear least-squares, Levenberg-Marquardt/MPFIT) Solar wind plasma near 1 au around 52 interplanetary shocks none (observational, IP shock crossings) Electron VDF fit parameters (kappa/self-similar exponents, density, thermal speeds, drift velocities) for core, halo, beam/strahl Wind/3DP EESA Low electron electrostatic analyzer
Quasi-static magnetic field measurement Solar wind near 1 au none Magnetic field vectors (B_o) defining parallel/perpendicular directions Wind/MFI dual triaxial fluxgate magnetometers (3 s cadence)
Proton and alpha-particle velocity moment determination (nonlinear least-squares fitting) Solar wind ions near 1 au none Ion velocity moments / number density Wind/SWE Faraday cups
Key results
  • Stable core fits obtained for 14,847 VDFs (~98%), halo fits for 13,871 (~91%), beam/strahl fits for 9567 (~63%) of VDFs analyzed ~98%/~91%/~63%
  • ~80.5% of core VDFs fit to a symmetric bi-self-similar function satisfied 2.0 ⩽ s_ec ⩽ 2.05, nearly indistinguishable from bi-Maxwellian ~80.5%
  • Core kappa exponent interquartile range κ_ec ~ 5.40–10.2
  • Halo kappa exponent interquartile range κ_eh ~ 3.58–5.34
  • Beam/strahl kappa exponent interquartile range κ_eb ~ 3.40–5.16
  • Symmetric bi-self-similar core exponent interquartile range s_ec ~ 2.00–2.04
  • Asymmetric bi-self-similar core parallel exponent interquartile range p_ec ~ 2.20–4.00
  • Asymmetric bi-self-similar core perpendicular exponent interquartile range q_ec ~ 2.00–2.46
Key statistics
  • count 15,314 electron VDFs analyzed (VDFs within ±2 hr of 52 IP shocks)
  • count 52 interplanetary shocks (Wind shock database events examined)
  • count 16 quasi-parallel (θ_Bn ⩽ 45°) (shock geometry classification)
  • count 36 quasi-perpendicular (θ_Bn > 45°) (shock geometry classification)
  • count 45 low Mach number (⟨M_f⟩up < 3) (Mach number classification)
  • count 7 high Mach number (⟨M_f⟩up ⩾ 3) (Mach number classification)
  • count 15,210 VDFs progressed to fit analysis (of 15,314 observed VDFs)
  • other ΔE/E ~ 20%, Δϕ ~ 5°–22°.5 (EESA Low energy/angular resolution)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methodology paper that fits 15,314 electron velocity distribution functions (VDFs) observed within ±2 hr of 52 interplanetary shocks to a sum of three model functions (bi-Maxwellian, bi-kappa, and symmetric/asymmetric bi-self-similar) using a nonlinear least-squares Levenberg–Marquardt algorithm (MPFIT). Model selection among candidate functions was guided by comparing reduced chi-squared values, and the resulting fit-parameter distributions are summarized descriptively using lower and upper quartile ranges. The work reports the fraction of VDFs yielding stable fits per component and defers formal statistical comparison to companion Papers II and III.

Replicationunclear Sample sizeSample sizes reported as raw counts of VDFs and shocks (15,314 VDFs within ±2 hr of 52 IP shocks); events selected by burst-mode 3DP availability; no power analysis described GroupsElectron VDF components (core, halo, beam/strahl); shocks categorized by quasi-parallel/perpendicular and Mach number Pairingna Randomization/blindingna DispersionIQR Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Nonlinear least-squares fitting (Levenberg–Marquardt algorithm, MPFIT) Fitting each electron VDF core/halo/beam component to bi-Maxwellian, bi-kappa, or bi-self-similar model functions 15,314 VDFs observed; 15,210 progressed to fit; stable fits for 14,847 core (~98%), 13,871 halo (~91%), 9567 beam/strahl (~63%) not stated
Reduced chi-squared comparison for model selection Choosing symmetric vs asymmetric bi-self-similar vs bi-Maxwellian core model (self-similar yielded lower reduced chi-squared) not stated
Approaches that could also have been used
  • Parameter spreads were summarized with lower-to-upper quartile (interquartile) ranges.
    Could also: Full distributions could also be summarized with medians plus bootstrap or analytic confidence intervals, or with kernel-density/histogram displays. — Reporting a central estimate with an interval would additionally convey the precision of the summary statistic, complementing the quartile range that conveys spread.
  • Competing core model functions were selected by comparing reduced chi-squared values.
    Could also: Information criteria such as AIC or BIC, or formal F-tests / likelihood-ratio tests for nested models, could also guide selection among candidate functions. — These criteria explicitly penalize added parameters (e.g., the extra self-similar exponent), which can help balance goodness of fit against model complexity when comparing nested forms.
  • Fit parameters were obtained via nonlinear least-squares (Levenberg–Marquardt) point estimates.
    Could also: Bayesian or Monte Carlo (e.g., MCMC) parameter estimation could also be used to characterize each fit. — A sampling-based approach would yield full posterior distributions and correlated parameter uncertainties, which can be informative when fit parameters are degenerate or constrained.
  • Per-component fit success was reported as percentages of VDFs yielding stable parameters.
    Could also: These proportions could also be accompanied by confidence intervals on the success fractions (e.g., Wilson or Clopper–Pearson intervals). — Interval estimates on a proportion would communicate the sampling uncertainty around the reported success rates, especially where rates differ across components.
  • Goodness of fit was characterized through reduced chi-squared values.
    Could also: Residual diagnostics (e.g., residual distributions, Q–Q plots, or runs tests) could also be examined alongside reduced chi-squared. — Residual-based diagnostics can reveal structured deviations between model and data that a single summary statistic may not capture.
Software: MPFIT (Levenberg–Marquardt least-squares package)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Authors · 13
1Lynn B. Wilson III 2Li-Jen Chen 3Shan Wang 4Steven J. Schwartz 5Drew L. Turner 6Michael L. Stevens 7Justin C. Kasper 8Adnane Osmane 9Damiano Caprioli 10Stuart D. Bale 11Marc P. Pulupa 12Chadi S. Salem 13Katherine A. Goodrich
Citations
75
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

10.5281/zenodo.2875806 DOI in References (http://purl.org/orb/References)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31806920

Paper: Wilson, L. B. III, et al. 2019, "Electron Energy Partition across Interplanetary Shocks. I. Methodology and Data Product." ApJS 245, 24. DOI 10.3847/1538-4365/ab22bd.

Type: Methodology + data product paper. The pipeline takes Wind/3DP electron velocity distribution functions (VDFs) near 52 interplanetary shocks and fits each VDF to a 3-component model (core: bi-kappa / bi-self-similar; halo: bi-kappa; beam/strahl: bi-kappa) using MPFIT (Levenberg–Marquardt). The shipped Zenodo data product (doi:10.5281/zenodo.2875806) is the per-VDF table of fit parameters. The paper's headline numbers (Table 2, Sec 3.4) are summary statistics of that table.

In scope (pipeline-derived, attempted)

The summary statistics are deterministically derivable from the shipped data product — no re-fitting needed. We recompute them directly:

  • Table 2 — exponent parameters over all VDFs: κ_ec, s_ec, p_ec, q_ec (core), κ_eh (halo), κ_eb (beam): min/max/mean/median/Q1/Q3.
  • Sec 3.4 density summary — n_ec, n_eh, n_eb: mean/median/Q1/Q3.
  • Sec 3.4 fit-quality — δR (%R), reduced χ² per component + total.
  • Population counts — # stable halo/beam fits.

This is the valid third-party-on-authors'-data route (BRIEF P16): we treat the fit-results file as ground truth and reproduce the reported statistics from it.

Out of scope / not attempted (the hard ~20%)

  • Re-running the IDL fitter (wind_3dp_pros): the analysis code is a large general-purpose IDL library requiring an IDL license + the raw Wind/3DP L0/L1 telemetry (not the small derived product). Re-fitting 15k VDFs is the heavy, poorly-pinned tail; the data product already is the fit output, so recomputing statistics from it is the faithful, auditable reproduction.
  • Upstream/downstream/Mach/θ_Bn sub-splits (Table 2 sub-columns): require classifying each VDF's region from per-shock arrival times. Doable but the 80/20 tail; we reproduce the headline "ALL VDFs" column only and say so.
  • Wet-lab / external / manual values: none (space-physics paper).
Figures / tables: Table
C05
Reported
kappa_eh min/max/mean/median/Q1/Q3 = 1.51/19.7/4.62/4.38/3.58/5.34
Reproduced
1.51/19.7/4.62/4.38/3.58/5.34 (N=13871)
exact
C06
Reported
kappa_eb = 1.52/20.0/4.57/4.17/3.40/5.16
Reproduced
1.52/20.0/4.57/4.17/3.40/5.16 (N=9567)
exact
C07
Reported
n_ec mean/median/Q1/Q3 = 13.7/11.3/6.44/19.5 cm^-3
Reproduced
13.72/11.27/6.44/19.52
exact
C08
Reported
n_eh = 0.52/0.36/0.21/0.63 cm^-3
Reproduced
0.52/0.36/0.20/0.63
exact
C09
Reported
n_eb = 0.21/0.16/0.09/0.27 cm^-3
Reproduced
0.21/0.16/0.09/0.27
exact
C11
Reported
chi2_core = 6.47/1.94/0.90/4.28
Reproduced
6.47/1.94/0.90/4.28
exact
C12
Reported
chi2_halo = 2.11/0.72/0.41/1.59
Reproduced
2.11/0.72/0.41/1.59
exact
C13
Reported
chi2_beam = 1.50/0.66/0.36/1.28
Reproduced
1.50/0.66/0.36/1.28
exact
C15
Reported
# stable halo fits = 13871
Reproduced
13871
exact
C16
Reported
# stable beam/strahl fits = 9567
Reproduced
9567
exact
C01
Reported
kappa_ec = 2.14/100/9.15/7.92/5.40/10.2
Reproduced
2.14/100/9.15/7.91/5.36/10.18
within tolerance
C02
Reported
s_ec = 2.00/3.00/2.03/2.00/2.00/2.04
Reproduced
2.00/3.00/2.03/2.00/2.00/2.03
within tolerance
C03
Reported
p_ec = 2.00/5.43/3.09/3.00/2.20/4.00
Reproduced
2.00/5.43/3.09/3.00/2.17/4.00
within tolerance
C04
Reported
q_ec = 2.00/3.29/2.24/2.00/2.00/2.46
Reproduced
2.00/3.29/2.24/2.00/2.00/2.43
within tolerance
C14
Reported
chi2_total = 1459/4.92/2.85/9.39
Reproduced
1447/4.88/2.82/9.36
within tolerance
C10
Reported
deltaR(%) mean/median/Q1/Q3 = 12.7/10.7/6.8/16.3
Reproduced
14.60/10.91/6.84/16.77 (median+quartiles match; mean differs)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 92/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4

Reproduction recomputed all headline statistics directly from the authors' public Zenodo data product (identical file, matching md5), yielding 10 exact + 5 within-tol matches plus exact population-count confirmation (13871/9567) that independently validates the column mapping. The only real discrepancy is δR mean (14.6 vs 12.7); its median/Q1/Q3 reproduce, so the paper's mean almost certainly excludes the FFlag=0 fits capped at 100% — an underspecified inclusion criterion on our methodology side, not an authors' or data defect. Core-exponent quartiles drift in the last digit because our component sub-counts run ~2% low. No fabrication concern; severity negligible and the central methodology/data-product claim is fully confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

123.7 k
tokens (I/O) · 6.6 M incl. cache
11 min
runtime · 0 CPU-h
0.1 GB
peak RAM
2
HPC jobs
hummel
machine