Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Turbulent dynamo in the terrestrial magnetosheath.

Nat Commun · 2026
L1 70/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
70/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 37% of all assessed papers rank 732 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> PARTIAL 1:1 reproduction. Zenodo CodesData.zip ships, per figure, the authors' MATLAB plotting scripts + .mat files holding the ALREADY-DERIVED quantities as irfu-matlab TSeries/EpochTT objects (the heavy MMS->derived step is done by the authors, not shipped as a runnable pipeline). Reproduced the headline scalar results WITHOUT MATLAB by writing a custom MATLAB-v7 MCOS decoder (analog of the v7.3 decoder from pmid-35105885) to read the TSeries data + TT2000 epochs, on «our HPC». Epoch->UTC lands exactly on the paper's 2015-11-30 00:21:44-00:26:43 UT interval. C1 Eq.1 balance reproduces EXACTLY (r=1.0) but is an algebraic identity encoded in the data, not an independent empirical agreement (flagged for auditor). C2 length-scale anisotropy 37-44 vs reported 40+/-20 (within tol). C5 Case I dynamo excursion -6.13 vs -6+/-0.7 (within tol). C6 Case II fall -8.33 vs -9+/-0.9 (within tol) but rise +4.79 vs +6+/-0.5 below tol (boundary-sensitive). NOT attempted: raw MMS CDF->TSeries re-derivation (c_4_grad etc.), gyroradii (need raw FPI moments; flagged 95 km paper vs 99 km script discrepancy), Pm scaling estimate, pitch-angle spectrograms (MATLAB-only) -- the hard ~20%. No fabrication detected; one internal 95-vs-99 km annotation inconsistency and the Eq.1-as-identity caveat noted.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.17780770

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 70
    assessed: 2026-06-14 ⛓ 989e4a66ec9c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The authors test whether a turbulent small-scale dynamo (SSD) operates in the terrestrial magnetosheath, asking if high-resolution multi-point MMS observations capture the predicted dynamo signatures—stretched-folded magnetic field topology and pressure-anisotropy instabilities that amplify magnetic fields in collisionless plasma turbulence.

Core claims
  • Evidence for a turbulent small-scale dynamo is present in the terrestrial magnetosheath, captured by high-resolution MMS observations. finding
  • The predicted stretched-and-folded magnetic field topology arises naturally in magnetosheath turbulence, shown by anticorrelation between magnetic field magnitude and field line curvature. finding
  • Pressure-anisotropy instabilities (mirror and firehose) that break adiabatic invariance and enable field amplification are abundant in the magnetosheath. mechanism
  • Magnetic field changes (d lnB/dt) can be evaluated in-situ from the magnetic induction equation using only the stretching (bb:∇Vi) and compression (−∇·Vi) terms computable from multi-point data. method
  • A large magnetic Prandtl number regime (Pm ≫ 1) indicates the plasma behaves as a viscous-scale 'stretch-and-fold' dynamo even without frequent collisions. mechanism
  • Earth's magnetosheath can serve as a natural testbed/laboratory for validating dynamo theories and simulations. resource
Experimental setups
Assay System Perturbation Readout Platform
Multi-point in-situ magnetic field measurements (tetrahedron gradient estimation) terrestrial magnetosheath plasma (Earth) none (natural observation) magnetic field magnitude/components and spatial gradients (bb:∇Vi, −∇·Vi, d lnB/dt) Magnetospheric Multiscale (MMS) mission, 4 spacecraft (MMS1-4)
In-situ ion velocity / plasma moment measurements terrestrial magnetosheath plasma (Earth) none ion velocity (150 ms cadence), plasma beta, anisotropic ion temperatures MMS plasma instruments (GSE coordinates)
Magnetic field curvature / characteristic wavenumber analysis terrestrial magnetosheath (tetrahedron barycenter) none field line curvature K, parallel/perpendicular length scales l∣∣ and l⊥ MMS tetrahedron (separations <20 km)
Key results
  • Anticorrelation between normalized squared magnetic field magnitude and normalized field line curvature, consistent with stretched-folded geometry.
  • Ratio of parallel to perpendicular magnetic length scales l∣∣/l⊥ measured for the interval. 40 ± 20
  • Inferred magnetic Prandtl number range from length-scale ratio, indicating Pm ≫ 1 stretch-and-fold dynamo regime. Pm ∈ (400, 3600)
  • Pressure-anisotropy points exceed mirror (Ti⊥/Ti∣∣ > 1) and firehose (Ti⊥/Ti∣∣ < 1) instability thresholds, with subintervals I1 (firehose) and I2 (mirror) identified.
  • Positive stretching term and negative compression (converging flows) lead to magnetic field growth; opposite signs to field decrease.
  • Spacecraft sampled convected plasma volume over ~40,000 km along-trajectory extent during the interval. ~40,000 km
Key statistics
  • other l∣∣/l⊥ = 40 ± 20 (ratio of parallel to perpendicular magnetic length scales)
  • other Pm ∈ (400, 3600) (inferred magnetic Prandtl number range)
  • mean 95 ± 84 km (average ± standard deviation of ion gyroradius)
  • mean 0.9 ± 0.63 km (average ± standard deviation of electron gyroradius)
  • other <20 km (MMS inter-spacecraft separation (sub-ion, near-electron scales))
  • other 150 ms (ion velocity measurement cadence)
  • other β ≫ 1 (typically β ≳ 1) (plasma beta in magnetosheath promoting instabilities)
  • other 200 km s−1 (assumed convection speed; interval >4 min on 2015-11-30 00:21:45–00:26:43 UT)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper presents observational evidence for turbulent small-scale dynamo action in the terrestrial magnetosheath using ~4 minutes of high-resolution four-spacecraft MMS data from 2015-11-30. The statistical approach is primarily descriptive: physical quantities derived from multi-point spacecraft gradient estimates (dynamo terms, magnetic curvature, pressure anisotropy, characteristic length scales) are visualized as 2D scatter/histogram plots and compared against deterministic theoretical instability thresholds. Results are reported for the full interval alongside two representative case-study subintervals; dispersion is conveyed as mean ± standard deviation and rms normalization, with no formal inferential hypothesis tests.

Replicationunclear Sample sizeSingle opportunistically selected ~4-minute magnetosheath interval (2015-11-30 00:21:45–00:26:43 UT); four-spacecraft MMS constellation used for multi-point gradient estimation; no formal power calculation or sample-size justification stated GroupsFull interval statistical analysis vs. two representative case-study subintervals (I1: firehose-unstable; I2: mirror-unstable) Pairingna Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
2D histogram / scatter anticorrelation analysis (no formal test statistic reported) Normalized squared magnetic field magnitude vs. normalized magnetic field curvature (Fig. 2a); color-coded by number of points per bin All data points from the single ~4-minute interval across 4 spacecraft; exact n not stated not stated
Comparison of measured β_i∥ vs T_i⊥/T_i∥ against theoretical instability growth-rate threshold curves Pressure-anisotropy instability diagram (Fig. 2b); mirror-mode and firehose thresholds from refs 33–35 MMS1–4 time series from the ~4-minute interval; exact n not stated not stated
Scatter plot distribution with mean ± SD summary Perpendicular vs. parallel magnetic length scales l⊥ and l∥ (Fig. 2c); average ion and electron gyroradii reported as mean ± SD not stated
Signed scatter plot comparison of dynamo terms Stretching (bb:∇V_i) and compression (−∇·V_i) terms vs. d ln B/dt (Fig. 2d) not stated
Approaches that could also have been used
  • The anticorrelation between magnetic field magnitude and curvature (Fig. 2a) is shown as a 2D scatter/histogram without a quantitative correlation coefficient
    Could also: A Spearman rank correlation coefficient (ρ) with a bootstrap 95% confidence interval could also be reported alongside the scatter plot — A numerical correlation coefficient enables direct quantitative comparison with dynamo and turbulence simulations that report the same anticorrelation, and a CI conveys the precision of the estimate from the available data
  • Physical quantities (gyroradii, l∥/l⊥) are summarized with mean ± SD; standard deviations are large relative to the means (e.g., ion gyroradius SD ≈ 88% of mean), suggesting skewed distributions
    Could also: Median with interquartile range (IQR) or a log-normal characterization could also be reported for heavily skewed turbulence data — For the heavy-tailed or skewed distributions typical of turbulent plasma parameters, the median and IQR are more robust location and spread estimates than the mean and SD
  • The study is based on a single, selected ~4-minute event interval
    Could also: A statistical survey across multiple magnetosheath intervals meeting pre-defined selection criteria could also be conducted — A multi-event survey would allow estimation of how frequently the reported dynamo signatures occur and whether the selected interval is representative, extending the generalizability of the findings
  • The proportion of data points falling in pressure-anisotropy unstable regions (Fig. 2b) is not quantified
    Could also: The fraction of points beyond each instability threshold, with a bootstrap confidence interval, could also be reported — Quantifying the prevalence of instability-unstable states as a reproducible percentage provides a metric directly comparable across different intervals, missions, or simulations
  • Case-study subintervals I1 and I2 are highlighted as representative examples without a stated objective selection criterion
    Could also: Pre-specified quantitative selection criteria (e.g., intervals where |d ln B/dt| exceeds a defined threshold) could also be used to identify representative subintervals — Objective, pre-specified criteria make case-study selection reproducible and reduce the potential for confirmation bias in illustrative example choice
  • Correlation between measured magnetic field/velocity fluctuations and the computed dynamo terms is described qualitatively ('strongly correlated') for the highlighted subintervals
    Could also: A cross-correlation function or Pearson/Spearman coefficient with lag analysis could also be computed to quantify the degree of correspondence between the measured and modeled fluctuations — A quantitative correlation measure with uncertainty bounds would allow assessment of how well the kinematic dynamo model captures the observed fluctuation structure across different subintervals
Software: Not stated in provided text

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41708617

Paper: Vörös Z. et al., "Turbulent dynamo in the terrestrial magnetosheath", Nat Commun 2026. PMID 41708617 · PMCID PMC13032041 · DOI 10.1038/s41467-026-69469-y. Manuscript ref NCOMMS-25-30764A.

Code: irfu-matlab (https://github.com/irfu/irfu-matlab, also zenodo:11550090) — a third-party MATLAB toolbox for MMS space-physics analysis. The authors' own figure-generating MATLAB scripts ship in the Zenodo data package (see below). Per BRIEF rule P16, applying / reading the shipped scripts on the shipped data is a fully valid reproduction.

Data: zenodo:10.5281/zenodo.17780770 → single archive CodesData.zip (13.8 MB, sha256 ba3f588ad7d1081b48f323d1861df3e3a1a7e0ad7a6a6902128fe55582f3431f). Contents:

codes/Figure{1,2,3,4}_script.m      MATLAB plotting scripts (irf_plot etc.)
data/Figure{1,2,3,4}_data.mat       MATLAB v7 .mat, pre-computed derived quantities
readme.txt

The .mat files contain the already-derived physical quantities as irfu TSeries/EpochTT objects plus plain numeric arrays. The scripts only plot them — there is NO heavy pipeline that runs from raw MMS CDFs in this package; the heavy MMS→derived-quantity step (c_4_grad gradients, rotate_tensor, resampling) was done by the authors and its outputs shipped. Raw MMS brst CDFs (fgm, fpi dis-moms/dis-dist, 2015-11-30) live at the LASP SDC.

In scope (pipeline-derived, reproducible from shipped data, NO MATLAB needed)

Decode the v7 .mat with scipy.io.loadmat (these are NOT v7.3/HDF5, so the HDF5-MCOS decoder from pmid-35105885 is not needed; scipy reads v7 directly). Then recompute / verify the reported scalar results:

  • C1 — Equation 1 internal consistency. Eq.1: d lnB/dt = bb:∇Vi − ∇·Vi. Data ships dBts (measured d lnB/dt), bbgradVts (bb:∇Vi), divVts (∇·Vi). Verify dBts ≈ bbgradVts − divVts pointwise (this is the paper's central balance, Fig.1d / Fig.2d).
  • C2 — anisotropy of magnetic length scales l∥/l⊥ = 40 ± 20 (Results / Fig.2c). From lpar, lperp TSeries arrays in Figure2_data.mat.
  • C3 — mean ion gyroradius ⟨ρi⟩ ≈ 95 km (paper) — NB the shipped Fig.2c script annotates "~99 km". Record both; flag the 95-vs-99 mismatch.
  • C4 — mean electron gyroradius ⟨ρe⟩ ≈ 0.9 km (paper & script agree).
  • C5 — Case I (folded field, I1 = 2015-11-30T00:24:06.8–00:24:12.8): cs(d lnB/dt) decreases by −6 ± 0.7 s⁻¹ (Fig.3). From btimes_sum over I1.
  • C6 — Case II (mirror mode, I2 = 2015-11-30T00:25:22.9–00:25:27.2): cs(d lnB/dt) increases ~+6 ± 0.5 then decreases −9 ± 0.9 s⁻¹ (Fig.4). From btimes_sum over I2.

Out of scope (not attempted, and why)

  • Raw MMS CDF → derived gradient quantities (c_4_grad on the tetrahedron, rotate_tensor, 150 ms barycentre resampling). The authors' derived outputs are shipped; re-deriving from raw L2 CDFs would require pulling the full brst MMS dataset + a MATLAB irfu-matlab run — the hard ~20%, explicitly skipped.
  • Prandtl-number estimate Pm ∈ (400, 3600) — a scaling argument from l∥/l⊥ and assumed diffusivities, not a direct pipeline output; checked only for plausibility.
  • Pitch-angle spectrograms (Fig.3/4 panels) — qualitative figure panels, no scalar to grade; need iPDistpitch PDist objects + irf_spectrogram (MATLAB).
  • Instability threshold curves (Fig.2b) are analytic formulae (Hellinger/Gary), not data — out of scope as a "reproduction".

Method

scipy.io.loadmat decode of the 4 shipped v7 .mat files, on «our HPC» (compute node, pip venv with numpy+scipy). Data staged on «infra»; only small derived scalars + this scope come back to «host». Compute is trivially light (<1 s, <15 MB).

Figures / tables: Fig.1dFig.2dFig.2cFig.3Fig.4
C1
Reported
Eq.1 d lnB/dt = bb:gradVi - divVi (Fig.1d/2d)
Reproduced
dBts == bbgradVts-divVts to machine precision (r=1.0, RMS resid=0, n=1999)
exact
C2
Reported
l_par/l_perp = 40 +/- 20 (Fig.2c)
Reproduced
ratio-of-means 37.25, median 44.4 (n=1976)
within tolerance
C5
Reported
Case I cs(d lnB/dt) decreases -6 +/- 0.7 s^-1 (Fig.3)
Reproduced
net -6.13, fall -6.47 over I1 (n=40)
within tolerance
C6
Reported
Case II cs(d lnB/dt) +6 +/- 0.5 then -9 +/- 0.9 s^-1 (Fig.4)
Reproduced
rise +4.79 (below tol), fall -8.33 (within tol) over I2 (n=29)
partial
C3
Reported
<rho_i> = 95 +/- 84 km
Reproduced
not recomputed (needs raw FPI moments); shipped script annotates ~99 km
partial
C4
Reported
<rho_e> = 0.9 +/- 0.63 km
Reproduced
not recomputed; shipped script annotation 0.9 km agrees
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 70/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

Reproduction used the authors' already-derived quantities deposited 1:1 on Zenodo, decoded without MATLAB; 4 of 6 gradeable headline scalars reproduce exactly or within tolerance (C2 37–44 vs 40±20, C5 -6.13 vs -6±0.7, C6 fall -8.33 vs -9±0.9). The single out-of-tol deviation (C6 rise +4.79 vs +6±0.5) is on our side — boundary-sensitive to the self-chosen I2 window — and direction/magnitude still hold. The core turbulent-dynamo conclusion is confirmed; remaining caveats are minor and explainable: Eq.1's balance is a definitional identity in the shipped data (not independent empirical proof), a 95-vs-99 km gyroradius annotation inconsistency, and gyroradii/Pm left un-recomputed (raw moments/scaling, out of scope). No fabrication detected.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

258 k
tokens (I/O) · 15.8 M incl. cache
27 min
runtime · 0 CPU-h
0.1 GB
peak RAM
6 (2 failed)
HPC jobs
hummel
machine