Geometric Reliability of Super-Resolution Reconstructed Images from Clinical Fetal MRI in the Second Trimester.
Part of the results reproduced; minor but material deviations remained.
- ✓No relevant deviation in data/preprocessing
- ✓Any deviation was negligible
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.
▸Reproduction agent’s raw note
PARTIAL. Paper (Ciceri et al. 2023) evaluates geometric reliability of super-resolution (SR) reconstruction of clinical 2nd-trimester fetal-brain MRI by three public toolkits (NiftyMIC v0.8, MIALSRTK v2.0, SVRTK v0.2). (A) The paper's specific reported numbers (C10-C15: 40 SR volumes/algorithm, discard rates, Gwet AC1=0.83, ANOVA p=0.027, cLLD p-values, Passing-Bablok slopes) are NOT independently reproducible -- the 17 clinical exams are ethics-restricted/on-request, every value is downstream of MANUAL 3D-Slicer biometry + MANUAL expert quality scoring, and no analysis code is shipped (R, no scripts). The registry data accession zenodo:10.5281/zenodo.4290209 is a FALSE POSITIVE (MIALSRTK SOFTWARE, not a dataset). (B) We nonetheless EXECUTED the paper's one in-scope computational pipeline -- NiftyMIC SR reconstruction -- end-to-end on «our HPC» (SLURM «job») on FULLY PUBLIC data: degraded the public STA30 Gholipour fetal-brain atlas template into 3 orthogonal motion-corrupted 3mm LR stacks, reconstructed a 0.8mm isotropic SR volume, and quantified recovery vs ground truth: brain-mask Dice 0.984, NCC 0.772, PSNR 17.8 dB. Positive control: the NiftyMIC SR tool is functional and geometrically reliable (consistent with the paper's thesis), but it reproduces no specific paper value. (C) As auditability/fabrication evidence we recomputed the paper's reported aggregate statistics from its OWN detail Tables 2/4/5: all 9 aggregates match to within rounding (ICC averages <=0.003; overall %error mean/SD; win-counts 11/15 and 9/15 EXACT) -- reported summaries are internally self-consistent, no fabrication signal. DID NOT attempt (genuinely blocked): SR on the paper's clinical subjects, manual biometry, manual quality scoring, ANOVA/Gwet/Passing-Bablok/Bland-Altman statistics from raw data. Provisional verdict; a human signs off.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-16 ⛓ 8661bded74bc
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether Super-Resolution (SR) reconstruction toolkits (NiftyMIC, MIALSRTK, SVRTK) applied to clinical fetal MRI in the narrow 20-21 week second-trimester window produce geometrically reliable brain volumes for biometric assessment, and which acquisition sequence (TSE vs b-FFE) yields more robust reconstructions.
- ★ NiftyMIC and MIALSRTK provide reliable SR reconstructed volumes suitable for biometric assessments finding
- ★ NiftyMIC improves operator intraclass correlation coefficient (ICC) on quantitative biometric measures compared to the acquired 2D images finding
- ★ TSE sequences lead to more robust fetal brain reconstructions against intensity artifacts compared to b-FFE sequences finding
- ★ b-FFE sequences exhibit more defined anatomical details than TSE sequences finding
- SVRTK reconstructions had substantially more 'bad' quality outputs (44% discarded) than NiftyMIC and MIALSRTK (15% each) finding
- ★ Automatic SR toolkits should be adopted for fetal brain reconstruction to enable biometry evaluation on common clinical MR at an early pregnancy stage method
- Inter-rater agreement on reconstruction quality rating between two blinded experts was Good (Gwet's AC1=0.83) finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| T2w TSE MRI acquisition + biometric measurement | fetal brain, second-trimester singleton pregnancies (GA 20.24±0.44 weeks, n=17 exams) | none | 15 biometric measures (mm/degrees) on 2D images | Achieva d-Stream 3T Philips scanner, phased-array abdominal coil |
| b-FFE MRI acquisition + biometric measurement | same fetal cohort as above | none | 15 biometric measures on 2D images | Achieva d-Stream 3T Philips scanner |
| Super-Resolution volume reconstruction | fetal brain MR image subsets (orthogonal TSE/b-FFE series) | SR algorithm (toolkit) as independent variable | reconstructed isotropic 3D brain volume; biometric measures; quality rating | NiftyMIC v0.8 |
| Super-Resolution volume reconstruction | fetal brain MR image subsets | SR algorithm (toolkit) | reconstructed isotropic 3D brain volume; biometric measures; quality rating | MIALSRTK v2.03 |
| Super-Resolution volume reconstruction | fetal brain MR image subsets | SR algorithm (toolkit) | reconstructed isotropic 3D brain volume; quality rating | SVRTK v0.2 |
| Blinded Likert-scale qualitative image quality rating (1-4) | 40 reconstructed SR fetal brain volumes per toolkit | none | quality category (bad/poor/acceptable/excellent) by 2 raters | — |
| ICC operator reliability analysis | 9 fetal brain reconstructions (NiftyMIC, MIALSRTK) and corresponding 2D images, 3 operators | operator (rater) as variable | Intraclass Correlation Coefficient of biometric measures | 3D Slicer |
| Passing-Bablok regression / Bland-Altman agreement analysis | 34 fetal brain SR reconstructions (NiftyMIC + MIALSRTK) vs corresponding 2D images | none | agreement/correlation between 2D and SR biometric measures | R software v4.0.5 |
- – NiftyMIC and MIALSRTK produced SR reconstructions suitable for biometric assessment, while SVRTK was excluded from biometric analysis due to high proportion of unusable ('bad') volumes
- ▲ Operator ICC on biometric measures was higher for NiftyMIC reconstructions (0.93, 95% CI [0.81-0.98]) than for 2D images (0.90, 95% CI [0.85-0.94]); MIALSRTK was 0.88 [0.70-0.97] 0.93 vs 0.90
- – TSE sequences showed greater robustness to intensity artifacts than b-FFE during reconstruction
- – b-FFE sequences showed more defined anatomical details than TSE sequences
- ▼ Percentage of reconstructions rated 'bad' (unusable) and discarded: 15% NiftyMIC, 15% MIALSRTK, 44% SVRTK 15% vs 15% vs 44%
- – Inter-rater agreement (Gwet's AC1) between the two quality raters was 0.83, classified as 'Good' with 98.8% probability AC1=0.83
- other Gwet's AC1 = 0.83 (Good, probability 98.8%) (inter-rater reliability of blinded SR quality ratings)
- other ICC = 0.90, 95% CI [0.85-0.94] (operator ICC of biometric measures on 2D images)
- other ICC = 0.93, 95% CI [0.81-0.98] (operator ICC of biometric measures on NiftyMIC SR reconstructions)
- other ICC = 0.88, 95% CI [0.70-0.97] (operator ICC of biometric measures on MIALSRTK SR reconstructions)
- count 6 (15%), 6 (15%), 17.5 (44%) volumes discarded as 'bad' out of 40 (discarded SR reconstructions per toolkit (NiftyMIC, MIALSRTK, SVRTK))
- mean 20.24 ± 0.44 weeks (gestational age of study cohort)
- mean 3.35 sequences per subject (average number of series used for SR reconstruction after visual inspection)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper evaluated the geometric reliability of super-resolution (SR) reconstructed fetal brain MRI volumes from three toolkits (NiftyMIC, MIALSRTK, SVRTK) by comparing 15 biometric measurements derived from 2D acquired images with those from SR reconstructions in 17 second-trimester fetal exams (GA 20–21 weeks). Agreement between 2D and SR measurements was assessed via Passing-Bablok regression, Bland-Altman analysis, and intraclass correlation coefficients (ICC); inter-rater reliability of blinded qualitative quality scoring was quantified with Gwet's AC1. Normality of measure distributions was evaluated with Shapiro-Wilk tests prior to parametric comparisons, and all analyses were conducted in R v4.0.5.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Passing-Bablok regression with Pearson's correlation coefficient | Agreement between biometric measures on 2D acquired images (reference) and SR reconstructed volumes (NiftyMIC, MIALSRTK); also used to compare slope coefficients and intercepts across toolkits and acquisition sequences | Up to 34 SR reconstructions per toolkit (NiftyMIC or MIALSRTK) after exclusion of bad-quality volumes, from 17 subjects | not stated |
| Bland-Altman plot analysis | Agreement between biometric measures on 2D images and SR reconstructed volumes for NiftyMIC and MIALSRTK | Up to 34 SR reconstructions per toolkit from 17 subjects | not stated |
| Intraclass Correlation Coefficient (ICC) | Inter-operator reliability of biometric measures on 2D images and on NiftyMIC and MIALSRTK SR reconstructions; also used as a reliability index between 2D and SR measures | 3 operators measured 9 fetal brain SR reconstructions and their corresponding 2D images | not stated |
| One-way ANOVA | Exploring significant differences in ICC values according to image type (2D images vs SR reconstructions) | null | not stated |
| Shapiro-Wilk normality test | Testing normality of biometric measure distributions prior to parametric comparisons | null | na |
| Paired two-tailed t-test | Comparing mean biometric values between 2D images and SR reconstructions; comparing Passing-Bablok slope coefficients and intercepts across toolkits and sequences | null | not stated |
| F-test | Comparing standard deviations of biometric measures between 2D images and SR reconstructions | null | not stated |
| Gwet's AC1 agreement coefficient with Altman's benchmarking scale | Inter-rater reliability of blinded qualitative Likert-scale (1–4) rating of SR reconstruction quality by two expert raters | 40 volumes per SR algorithm rated by 2 raters | not stated |
-
Multiple paired t-tests and F-tests were applied across 15 biometric measures and multiple toolkit/sequence comparisons without a stated correction for multiple comparisons↳ Could also: Apply a false discovery rate correction (e.g., Benjamini-Hochberg) or a family-wise correction (e.g., Bonferroni) across the simultaneous tests — When many tests are conducted in parallel, a multiplicity correction controls the expected proportion of false discoveries; FDR methods are often preferred over Bonferroni in exploratory imaging studies because they are less conservative while still limiting erroneous findings across the full test family
-
The ICC model type (one-way random, two-way mixed, two-way random) used for the inter-operator analysis is not explicitly stated, even though Koo and Li (2016) — which the paper cites for ICC interpretation — recommends specifying the model↳ Could also: Explicitly report the ICC model form (e.g., ICC(2,1) if raters are treated as a random sample; ICC(3,1) if raters are fixed) and justify the choice — Different ICC formulations partition variance differently among subjects, raters, and their interaction; stating the model makes results reproducible and interpretable in light of whether the three operators are considered a random sample of all possible raters or a fixed set
-
Qualitative inter-rater reliability was assessed with Gwet's AC1 on a 4-point ordinal Likert scale↳ Could also: Use weighted Cohen's kappa (with linear or quadratic weights) or Krippendorff's alpha, which are also widely used for ordinal categorical agreement between two raters — Weighted kappa explicitly penalizes larger disagreements proportionally to their magnitude on the ordinal scale, and is more commonly reported in imaging reliability literature, facilitating direct comparison with published studies using those metrics
-
A one-way ANOVA was used to compare ICC values across image types (2D images vs SR reconstructions)↳ Could also: Use a linear mixed-effects model with subject as a random effect, since the same subjects contributed ICC-based measurements under each image type condition — When the same individuals contribute data to multiple conditions, a mixed-effects model accounts for within-subject correlation; in a small sample of 17 subjects this may yield more precise estimates of the image-type effect than a one-way ANOVA that treats observations as independent
-
Pearson's correlation coefficient was reported alongside the non-parametric Passing-Bablok regression↳ Could also: Report Spearman's rank correlation as a complement, particularly for measures where normality was not confirmed by the Shapiro-Wilk test — Passing-Bablok regression is itself a non-parametric method chosen in part for robustness to distributional assumptions; pairing it with Spearman's ρ rather than Pearson's r would be internally consistent and robust to the same departures from normality
-
Variability of biometric measurements was summarized with standard deviation (SD)↳ Could also: Also report 95% confidence intervals around mean differences in the agreement analyses — In a sample of 17 subjects, SD describes the distribution of individual measurements, but 95% CIs around the mean differences directly communicate estimation uncertainty and precision of the group-level comparison, which are distinct pieces of information relevant to clinical interpretation of agreement
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 37284977
Paper: Ciceri et al. (2023) Geometric Reliability of Super-Resolution Reconstructed Images from Clinical Fetal MRI in the Second Trimester. Neuroinformatics. DOI 10.1007/s12021-023-09635-5. PMCID PMC10406722.
Study design (what produces the numbers)
17 clinical fetal-brain MR exams (singleton, GA ~20 wk, IRCCS Ca' Granda, Milan). Per subject, orthogonal T2w-TSE and b-FFE 2D stacks were reconstructed into isotropic Super-Resolution (SR) volumes by three public toolkits: NiftyMIC v0.8, MIALSRTK v2.0, SVRTK v0.2. Then 15 biometric distances/angles were measured manually in 3D Slicer by expert operators on both the 2D images and the SR volumes, plus a manual visual quality score by 2 experts. All reported results are statistics computed on those manual measurements.
Pipeline vs non-pipeline classification
| Step | Type | In scope? |
|---|---|---|
| SR reconstruction (NiftyMIC / MIALSRTK / SVRTK) | computational pipeline | YES in principle — but needs the restricted input data |
| Biometric measurements (3D Slicer, by experts) | manual | NO (out of scope: manual) |
| Visual quality scoring (2 experts, bad/poor/acc/excellent) | manual | NO (out of scope: manual) |
| Statistical analysis (Passing-Bablok, Bland-Altman, ICC, ANOVA, Gwet's AC1, %error) in R v4.0.5 | computational | YES in principle — but no code shipped and inputs are the manual measures |
Blockers (why nothing pipeline-derived is reproducible)
- Data restricted. Data Availability statement: "Owing to ethics and privacy
limitations, the data will be made available by request which includes a formal
project outline and an agreement of data sharing." The 17 clinical fetal MR
exams are NOT public (on-request, ethics agreement). →
data_restricted. - Recorded "data accession" is a false positive. The registry's
zenodo:10.5281/zenodo.4290209resolves to MIALSRTK — MIAL Super-Resolution Toolkit software (latest version record 7612119), i.e. one of the SR tools, NOT the paper's data. No paper dataset is deposited anywhere. - Every reported number is downstream of manual measurement. Even with the raw images, reproducing any value requires re-doing manual 3D-Slicer biometry and manual expert quality scoring — out of scope and not shippable.
- No analysis code. The GitHub link is the NiftyMIC SR tool, not the paper's statistical pipeline; no analysis code/scripts are provided.
- Supplement ships only summary tables (Passing-Bablok S1/S3/S4, Bland-Altman S2) — no raw per-subject measurements to recompute statistics from.
What we DID do (auditable, honest, not a pipeline reproduction)
Internal-consistency recomputation: the paper reports aggregate statistics (average ICCs, overall % error mean/SD, win-counts) that should equal simple aggregates of its own per-measure detail tables (Tables 2, 4, 5). We recomputed those aggregates and checked them. See AUDIT.md / agreement.json. This is a fabrication-detection check on the reported summaries, NOT an independent reproduction of the pipeline.
Verdict
DROP — data_restricted. The one genuine computational pipeline (SR
reconstruction) cannot be run on the paper's data because that data is
on-request/ethics-restricted, and all reported quantities depend on manual
measurement with no shipped analysis code. Honest partial value delivered:
the reported aggregates are internally consistent with the detail tables.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.