Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Geometric Reliability of Super-Resolution Reconstructed Images from Clinical Fetal MRI in the Second Trimester.

Neuroinformatics · 2023
L1 No data access 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • No relevant deviation in data/preprocessing
  • Any deviation was negligible
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
No data access Data access not granted

This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.

Reproduction agent’s raw note

PARTIAL. Paper (Ciceri et al. 2023) evaluates geometric reliability of super-resolution (SR) reconstruction of clinical 2nd-trimester fetal-brain MRI by three public toolkits (NiftyMIC v0.8, MIALSRTK v2.0, SVRTK v0.2). (A) The paper's specific reported numbers (C10-C15: 40 SR volumes/algorithm, discard rates, Gwet AC1=0.83, ANOVA p=0.027, cLLD p-values, Passing-Bablok slopes) are NOT independently reproducible -- the 17 clinical exams are ethics-restricted/on-request, every value is downstream of MANUAL 3D-Slicer biometry + MANUAL expert quality scoring, and no analysis code is shipped (R, no scripts). The registry data accession zenodo:10.5281/zenodo.4290209 is a FALSE POSITIVE (MIALSRTK SOFTWARE, not a dataset). (B) We nonetheless EXECUTED the paper's one in-scope computational pipeline -- NiftyMIC SR reconstruction -- end-to-end on «our HPC» (SLURM «job») on FULLY PUBLIC data: degraded the public STA30 Gholipour fetal-brain atlas template into 3 orthogonal motion-corrupted 3mm LR stacks, reconstructed a 0.8mm isotropic SR volume, and quantified recovery vs ground truth: brain-mask Dice 0.984, NCC 0.772, PSNR 17.8 dB. Positive control: the NiftyMIC SR tool is functional and geometrically reliable (consistent with the paper's thesis), but it reproduces no specific paper value. (C) As auditability/fabrication evidence we recomputed the paper's reported aggregate statistics from its OWN detail Tables 2/4/5: all 9 aggregates match to within rounding (ICC averages <=0.003; overall %error mean/SD; win-counts 11/15 and 9/15 EXACT) -- reported summaries are internally self-consistent, no fabrication signal. DID NOT attempt (genuinely blocked): SR on the paper's clinical subjects, manual biometry, manual quality scoring, ANOVA/Gwet/Passing-Bablok/Bland-Altman statistics from raw data. Provisional verdict; a human signs off.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.4290209

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-16 ⛓ 8661bded74bc
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether Super-Resolution (SR) reconstruction toolkits (NiftyMIC, MIALSRTK, SVRTK) applied to clinical fetal MRI in the narrow 20-21 week second-trimester window produce geometrically reliable brain volumes for biometric assessment, and which acquisition sequence (TSE vs b-FFE) yields more robust reconstructions.

Core claims
  • NiftyMIC and MIALSRTK provide reliable SR reconstructed volumes suitable for biometric assessments finding
  • NiftyMIC improves operator intraclass correlation coefficient (ICC) on quantitative biometric measures compared to the acquired 2D images finding
  • TSE sequences lead to more robust fetal brain reconstructions against intensity artifacts compared to b-FFE sequences finding
  • b-FFE sequences exhibit more defined anatomical details than TSE sequences finding
  • SVRTK reconstructions had substantially more 'bad' quality outputs (44% discarded) than NiftyMIC and MIALSRTK (15% each) finding
  • Automatic SR toolkits should be adopted for fetal brain reconstruction to enable biometry evaluation on common clinical MR at an early pregnancy stage method
  • Inter-rater agreement on reconstruction quality rating between two blinded experts was Good (Gwet's AC1=0.83) finding
Experimental setups
Assay System Perturbation Readout Platform
T2w TSE MRI acquisition + biometric measurement fetal brain, second-trimester singleton pregnancies (GA 20.24±0.44 weeks, n=17 exams) none 15 biometric measures (mm/degrees) on 2D images Achieva d-Stream 3T Philips scanner, phased-array abdominal coil
b-FFE MRI acquisition + biometric measurement same fetal cohort as above none 15 biometric measures on 2D images Achieva d-Stream 3T Philips scanner
Super-Resolution volume reconstruction fetal brain MR image subsets (orthogonal TSE/b-FFE series) SR algorithm (toolkit) as independent variable reconstructed isotropic 3D brain volume; biometric measures; quality rating NiftyMIC v0.8
Super-Resolution volume reconstruction fetal brain MR image subsets SR algorithm (toolkit) reconstructed isotropic 3D brain volume; biometric measures; quality rating MIALSRTK v2.03
Super-Resolution volume reconstruction fetal brain MR image subsets SR algorithm (toolkit) reconstructed isotropic 3D brain volume; quality rating SVRTK v0.2
Blinded Likert-scale qualitative image quality rating (1-4) 40 reconstructed SR fetal brain volumes per toolkit none quality category (bad/poor/acceptable/excellent) by 2 raters
ICC operator reliability analysis 9 fetal brain reconstructions (NiftyMIC, MIALSRTK) and corresponding 2D images, 3 operators operator (rater) as variable Intraclass Correlation Coefficient of biometric measures 3D Slicer
Passing-Bablok regression / Bland-Altman agreement analysis 34 fetal brain SR reconstructions (NiftyMIC + MIALSRTK) vs corresponding 2D images none agreement/correlation between 2D and SR biometric measures R software v4.0.5
Key results
  • NiftyMIC and MIALSRTK produced SR reconstructions suitable for biometric assessment, while SVRTK was excluded from biometric analysis due to high proportion of unusable ('bad') volumes
  • Operator ICC on biometric measures was higher for NiftyMIC reconstructions (0.93, 95% CI [0.81-0.98]) than for 2D images (0.90, 95% CI [0.85-0.94]); MIALSRTK was 0.88 [0.70-0.97] 0.93 vs 0.90
  • TSE sequences showed greater robustness to intensity artifacts than b-FFE during reconstruction
  • b-FFE sequences showed more defined anatomical details than TSE sequences
  • Percentage of reconstructions rated 'bad' (unusable) and discarded: 15% NiftyMIC, 15% MIALSRTK, 44% SVRTK 15% vs 15% vs 44%
  • Inter-rater agreement (Gwet's AC1) between the two quality raters was 0.83, classified as 'Good' with 98.8% probability AC1=0.83
Key statistics
  • other Gwet's AC1 = 0.83 (Good, probability 98.8%) (inter-rater reliability of blinded SR quality ratings)
  • other ICC = 0.90, 95% CI [0.85-0.94] (operator ICC of biometric measures on 2D images)
  • other ICC = 0.93, 95% CI [0.81-0.98] (operator ICC of biometric measures on NiftyMIC SR reconstructions)
  • other ICC = 0.88, 95% CI [0.70-0.97] (operator ICC of biometric measures on MIALSRTK SR reconstructions)
  • count 6 (15%), 6 (15%), 17.5 (44%) volumes discarded as 'bad' out of 40 (discarded SR reconstructions per toolkit (NiftyMIC, MIALSRTK, SVRTK))
  • mean 20.24 ± 0.44 weeks (gestational age of study cohort)
  • mean 3.35 sequences per subject (average number of series used for SR reconstruction after visual inspection)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper evaluated the geometric reliability of super-resolution (SR) reconstructed fetal brain MRI volumes from three toolkits (NiftyMIC, MIALSRTK, SVRTK) by comparing 15 biometric measurements derived from 2D acquired images with those from SR reconstructions in 17 second-trimester fetal exams (GA 20–21 weeks). Agreement between 2D and SR measurements was assessed via Passing-Bablok regression, Bland-Altman analysis, and intraclass correlation coefficients (ICC); inter-rater reliability of blinded qualitative quality scoring was quantified with Gwet's AC1. Normality of measure distributions was evaluated with Shapiro-Wilk tests prior to parametric comparisons, and all analyses were conducted in R v4.0.5.

Replicationbiological Sample size17 fetal MR exams stated; no formal power calculation or sample size justification reported Groups2D acquired images vs SR reconstructions (NiftyMIC, MIALSRTK, SVRTK); TSE vs b-FFE acquisition sequences; also reduced vs wide FOV variants within each sequence type Pairingpaired Randomization/blindingstated DispersionSD Effect sizesno Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Passing-Bablok regression with Pearson's correlation coefficient Agreement between biometric measures on 2D acquired images (reference) and SR reconstructed volumes (NiftyMIC, MIALSRTK); also used to compare slope coefficients and intercepts across toolkits and acquisition sequences Up to 34 SR reconstructions per toolkit (NiftyMIC or MIALSRTK) after exclusion of bad-quality volumes, from 17 subjects not stated
Bland-Altman plot analysis Agreement between biometric measures on 2D images and SR reconstructed volumes for NiftyMIC and MIALSRTK Up to 34 SR reconstructions per toolkit from 17 subjects not stated
Intraclass Correlation Coefficient (ICC) Inter-operator reliability of biometric measures on 2D images and on NiftyMIC and MIALSRTK SR reconstructions; also used as a reliability index between 2D and SR measures 3 operators measured 9 fetal brain SR reconstructions and their corresponding 2D images not stated
One-way ANOVA Exploring significant differences in ICC values according to image type (2D images vs SR reconstructions) null not stated
Shapiro-Wilk normality test Testing normality of biometric measure distributions prior to parametric comparisons null na
Paired two-tailed t-test Comparing mean biometric values between 2D images and SR reconstructions; comparing Passing-Bablok slope coefficients and intercepts across toolkits and sequences null not stated
F-test Comparing standard deviations of biometric measures between 2D images and SR reconstructions null not stated
Gwet's AC1 agreement coefficient with Altman's benchmarking scale Inter-rater reliability of blinded qualitative Likert-scale (1–4) rating of SR reconstruction quality by two expert raters 40 volumes per SR algorithm rated by 2 raters not stated
Approaches that could also have been used
  • Multiple paired t-tests and F-tests were applied across 15 biometric measures and multiple toolkit/sequence comparisons without a stated correction for multiple comparisons
    Could also: Apply a false discovery rate correction (e.g., Benjamini-Hochberg) or a family-wise correction (e.g., Bonferroni) across the simultaneous tests — When many tests are conducted in parallel, a multiplicity correction controls the expected proportion of false discoveries; FDR methods are often preferred over Bonferroni in exploratory imaging studies because they are less conservative while still limiting erroneous findings across the full test family
  • The ICC model type (one-way random, two-way mixed, two-way random) used for the inter-operator analysis is not explicitly stated, even though Koo and Li (2016) — which the paper cites for ICC interpretation — recommends specifying the model
    Could also: Explicitly report the ICC model form (e.g., ICC(2,1) if raters are treated as a random sample; ICC(3,1) if raters are fixed) and justify the choice — Different ICC formulations partition variance differently among subjects, raters, and their interaction; stating the model makes results reproducible and interpretable in light of whether the three operators are considered a random sample of all possible raters or a fixed set
  • Qualitative inter-rater reliability was assessed with Gwet's AC1 on a 4-point ordinal Likert scale
    Could also: Use weighted Cohen's kappa (with linear or quadratic weights) or Krippendorff's alpha, which are also widely used for ordinal categorical agreement between two raters — Weighted kappa explicitly penalizes larger disagreements proportionally to their magnitude on the ordinal scale, and is more commonly reported in imaging reliability literature, facilitating direct comparison with published studies using those metrics
  • A one-way ANOVA was used to compare ICC values across image types (2D images vs SR reconstructions)
    Could also: Use a linear mixed-effects model with subject as a random effect, since the same subjects contributed ICC-based measurements under each image type condition — When the same individuals contribute data to multiple conditions, a mixed-effects model accounts for within-subject correlation; in a small sample of 17 subjects this may yield more precise estimates of the image-type effect than a one-way ANOVA that treats observations as independent
  • Pearson's correlation coefficient was reported alongside the non-parametric Passing-Bablok regression
    Could also: Report Spearman's rank correlation as a complement, particularly for measures where normality was not confirmed by the Shapiro-Wilk test — Passing-Bablok regression is itself a non-parametric method chosen in part for robustness to distributional assumptions; pairing it with Spearman's ρ rather than Pearson's r would be internally consistent and robust to the same departures from normality
  • Variability of biometric measurements was summarized with standard deviation (SD)
    Could also: Also report 95% confidence intervals around mean differences in the agreement analyses — In a sample of 17 subjects, SD describes the distribution of individual measurements, but 95% CIs around the mean differences directly communicate estimation uncertainty and precision of the group-level comparison, which are distinct pieces of information relevant to clinical interpretation of agreement
Software: R 4.0.5 · 3D Slicer (used for biometric measurement extraction) null

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
11
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 37284977

Paper: Ciceri et al. (2023) Geometric Reliability of Super-Resolution Reconstructed Images from Clinical Fetal MRI in the Second Trimester. Neuroinformatics. DOI 10.1007/s12021-023-09635-5. PMCID PMC10406722.

Study design (what produces the numbers)

17 clinical fetal-brain MR exams (singleton, GA ~20 wk, IRCCS Ca' Granda, Milan). Per subject, orthogonal T2w-TSE and b-FFE 2D stacks were reconstructed into isotropic Super-Resolution (SR) volumes by three public toolkits: NiftyMIC v0.8, MIALSRTK v2.0, SVRTK v0.2. Then 15 biometric distances/angles were measured manually in 3D Slicer by expert operators on both the 2D images and the SR volumes, plus a manual visual quality score by 2 experts. All reported results are statistics computed on those manual measurements.

Pipeline vs non-pipeline classification

Step Type In scope?
SR reconstruction (NiftyMIC / MIALSRTK / SVRTK) computational pipeline YES in principle — but needs the restricted input data
Biometric measurements (3D Slicer, by experts) manual NO (out of scope: manual)
Visual quality scoring (2 experts, bad/poor/acc/excellent) manual NO (out of scope: manual)
Statistical analysis (Passing-Bablok, Bland-Altman, ICC, ANOVA, Gwet's AC1, %error) in R v4.0.5 computational YES in principle — but no code shipped and inputs are the manual measures

Blockers (why nothing pipeline-derived is reproducible)

  1. Data restricted. Data Availability statement: "Owing to ethics and privacy limitations, the data will be made available by request which includes a formal project outline and an agreement of data sharing." The 17 clinical fetal MR exams are NOT public (on-request, ethics agreement). → data_restricted.
  2. Recorded "data accession" is a false positive. The registry's zenodo:10.5281/zenodo.4290209 resolves to MIALSRTK — MIAL Super-Resolution Toolkit software (latest version record 7612119), i.e. one of the SR tools, NOT the paper's data. No paper dataset is deposited anywhere.
  3. Every reported number is downstream of manual measurement. Even with the raw images, reproducing any value requires re-doing manual 3D-Slicer biometry and manual expert quality scoring — out of scope and not shippable.
  4. No analysis code. The GitHub link is the NiftyMIC SR tool, not the paper's statistical pipeline; no analysis code/scripts are provided.
  5. Supplement ships only summary tables (Passing-Bablok S1/S3/S4, Bland-Altman S2) — no raw per-subject measurements to recompute statistics from.

What we DID do (auditable, honest, not a pipeline reproduction)

Internal-consistency recomputation: the paper reports aggregate statistics (average ICCs, overall % error mean/SD, win-counts) that should equal simple aggregates of its own per-measure detail tables (Tables 2, 4, 5). We recomputed those aggregates and checked them. See AUDIT.md / agreement.json. This is a fabrication-detection check on the reported summaries, NOT an independent reproduction of the pipeline.

Verdict

DROP — data_restricted. The one genuine computational pipeline (SR reconstruction) cannot be run on the paper's data because that data is on-request/ethics-restricted, and all reported quantities depend on manual measurement with no shipped analysis code. Honest partial value delivered: the reported aggregates are internally consistent with the detail tables.

Figures / tables: Table
CX-SR
Reported
SR reconstruction is geometrically reliable (paper thesis; NiftyMIC SR pipeline)
Reproduced
Ran NiftyMIC SR pipeline on «our HPC» on PUBLIC data (degrade STA30 fetal template -> 3 orthogonal LR stacks -> 0.8mm isotropic SR recon); brain recovered with Dice 0.984, NCC 0.772, PSNR 17.8 dB
partial
C1
Reported
avg 2D-SRR ICC NiftyMIC 0.82
Reproduced
0.817 (mean of paper Table 4)
within tolerance
C2
Reported
avg 2D-SRR ICC MIALSRTK 0.79
Reproduced
0.787 (mean of paper Table 4)
within tolerance
C3
Reported
operator ICC 2D 0.90
Reproduced
0.897 (mean of paper Table 2)
within tolerance
C4
Reported
operator ICC NiftyMIC 0.93
Reproduced
0.932 (mean of paper Table 2)
within tolerance
C5
Reported
operator ICC MIALSRTK 0.88
Reproduced
0.877 (mean of paper Table 2)
within tolerance
C6
Reported
overall %error NiftyMIC -0.1%+/-4.9%
Reproduced
-0.120%+/-4.91% (Table 5)
within tolerance
C7
Reported
overall %error MIALSRTK -0.7%+/-5.1%
Reproduced
-0.691%+/-5.12% (Table 5)
within tolerance
C8
Reported
NiftyMIC smaller |mean error| 11/15
Reproduced
11/15
exact
C9
Reported
NiftyMIC smaller SD 9/15
Reproduced
9/15
exact
C10-C15
Reported
40 SR volumes/algo; discard rates; Gwet AC1 0.83; ANOVA p=0.027; cLLD p-values; Passing-Bablok S1
Reproduced
NOT independently reproducible (restricted 17-subject clinical MRI + manual 3D-Slicer biometry + manual expert scoring + no shipped analysis code)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 44/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟢3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

465.5 k
tokens (I/O) · 50.6 M incl. cache
174 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.