Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Expanding the clinical spectrum of COL2A1 related disorders by a mass like phenotype.

· 2022
PubMed 35296718 ↗ pmid-35296718
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Total score +2
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Clinical case report (Sci Rep 12:4544, 2022) - 4 patients/3 families with a MASS-like phenotype and 3 COL2A1 missense variants. Described well enough to reproduce its computational sliver but NOT a 1:1 pipeline paper: there is no analysis-code repo and no deposited patient sequencing data, so the primary variant-detection pipeline and all clinical phenotyping are out of scope. The three in-silico supporting analyses reproduce cleanly from public data: (C1) SWISS-Model C-propeptide identity 68.72% reproduced as ~69% (within <1 pp; explained by alignment-method/template-span differences), run on «our HPC»; (C2) gnomAD v2.1.1 rarity - K1312N AF 3.98E-06, D65N & S1338N absent, all below the reported MPF 5.15E-06 - exact; (C3) ClinVar deposition SUB9514030 confirmed via consecutive SCV001571667/668/669 with matching per-variant classifications (LP, VUS; D65N aggregate drifted to Conflicting post-publication) - exact. Did NOT attempt: variant calling (no data), clinical/imaging findings (wet-lab), Chimera clash analysis (qualitative), and the MPF threshold derivation (inputs unstated). Overall: partial - honest reproduction of the publicly-supported in-silico results of a fundamentally clinical paper.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-19 ⛓ e86d354b7d75
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates whether patients presenting with a MASS (mitral valve, aorta, skeletal, skin)-like connective tissue phenotype who test negative for FBN1 variants carry pathogenic variants in COL2A1, thereby expanding the clinical spectrum of COL2A1-related type II collagenopathies.

Core claims
  • Four FBN1-negative patients from three families with a MASS-like phenotype carry likely pathogenic or uncertain-significance missense variants in the propeptide-coding regions of COL2A1 finding
  • The identified COL2A1 variants localize to propeptide domains: the N-propeptide VWFC repeat (Asp65Asn) and the C-propeptide fibrillar collagen NC1 domain (Lys1312Asn, Ser1338Asn) finding
  • A MASS-like phenotype has not previously been associated with COL2A1 variants, expanding the phenotypic spectrum of type II collagenopathies finding
  • Molecular modeling predicts structural disruption from Lys1312Asn and Ser1338Asn in the NC1 domain and altered surface hydrophobicity from Asp65Asn in the VWFC domain mechanism
  • COL2A1 sequencing should be recommended in FBN1-negative patients with a MASS/Marfan-like phenotype lacking aortopathy method
  • A 62-gene targeted next-generation sequencing panel covering aortopathy and connective tissue disease genes was used to screen 200 probands method
  • Variants were classified using ACMG/AMP guidelines aided by automated VarSome in-silico analysis method
  • The Lys1312Asn variant segregates with the MASS-like phenotype within family 1 (present in affected mother and son, absent in unaffected relatives) finding
Experimental setups
Assay System Perturbation Readout Platform
targeted next-generation sequencing (gene panel) blood-derived DNA from 200 probands with suspected heritable connective tissue disease none identification of variants across 62 aortopathy/connective tissue genes tNGS
clinical phenotyping / physical examination four patients (subjects 1A, 1B, 2, 3) none MASS-like clinical features (height, arachnodactyly, skeletal/spinal deformities) and revised Ghent nosology systemic score
transthoracic echocardiography patients none aortic root diameter and Z-score
spinal MRI patients (3 of 4 individuals) none presence of dural ectasia
molecular/structural modeling COL2A1 protein domains (VWFC repeat, fibrillar collagen NC1 domain), in silico amino acid substitution (Asp65Asn, Lys1312Asn, Ser1338Asn) surface hydrophobicity and van der Waals contacts/clashes UCSF Chimera v1.2
variant pathogenicity classification identified COL2A1 missense variants none ACMG/AMP classification (LPV/VUS) and damaging/uncertain/tolerated prediction counts VarSome v10.2
segregation genotyping family members (blood/buccal swab DNA) none presence or absence of variant in relatives
population allele frequency lookup gnomAD database none allele frequency and allele count for each variant gnomAD v2.1.1
Key results
  • Subjects 1A and 1B (mother/son) carry heterozygous COL2A1 c.3936G>T p.(Lys1312Asn), classified as likely pathogenic
  • Subject 2 carries heterozygous COL2A1 c.193G>A p.(Asp65Asn), classified as likely pathogenic
  • Subject 3 carries heterozygous COL2A1 c.4013G>A p.(Ser1338Asn), classified as variant of uncertain significance
  • All four patients met revised Ghent nosology criteria for a MASS-like phenotype, with systemic scores ranging from 5 to 9
  • None of the four patients showed significant aortic root dilatation despite meeting other MASS-like criteria
  • Structural modeling of Lys1312Asn showed van der Waals clashes (≥0.6 Å) replacing normal contacts seen with wild-type Lys1312, indicating structural disruption of the NC1 domain
  • The Lys1312Asn variant was absent in the unaffected sister and father of subject 1A, consistent with segregation with the phenotype
  • The c.4013G>A p.(Ser1338Asn) variant was also identified in the father of subject 3 via buccal swab; parents were not clinically examined
Key statistics
  • other 0.000003976 (allele count 1/251,478, 0 homozygotes) (gnomAD population allele frequency for Lys1312Asn)
  • other 0 (allele count 0/280,928, 0 homozygotes) (gnomAD population allele frequency for Asp65Asn)
  • other 0 (allele count 0/251,496, 0 homozygotes) (gnomAD population allele frequency for Ser1338Asn)
  • other 5.15E-06 (calculated maximum population frequency (MPF) threshold for COL2A1 variants)
  • count 200 (probands undergoing gene panel sequencing in the study)
  • count 62 genes (genes analyzed on the targeted NGS panel)
  • other systemic scores 7, 5, 7, 9 (revised Ghent nosology systemic scores for subjects 1A, 1B, 2, and 3 respectively)
  • other VarSome damaging/uncertain/tolerated: 13/1/2, 13/1/2, 11/1/4, 3/2/11 (in-silico pathogenicity predictions per variant for subjects 1A, 1B, 2, and 3)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a descriptive case series/report of four patients with a MASS-like connective tissue phenotype, drawn from a larger clinical cohort of 200 individuals screened by targeted gene panel sequencing. Statistical treatment is limited to population-genetics filtering (comparison of variant allele frequencies in gnomAD against a calculated maximum population frequency threshold) and standardized clinical scoring (Ghent nosology systemic score, height expressed as SD from reference growth data). No formal inferential hypothesis testing (e.g., group comparisons, p-values) is reported; variant pathogenicity was assessed via structured qualitative classification (ACMG/AMP criteria) supported by in-silico prediction tools and molecular modeling.

Replicationunclear Sample sizeA cohort of 200 probands was screened by gene panel sequencing; 4 patients from 3 families with COL2A1 variants are the focus of this report. No power calculation or formal sample-size justification is described, consistent with a descriptive case series. GroupsNo comparator group; individual patients described relative to reference population/database values (gnomAD, growth charts) Pairingna Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Allele frequency threshold comparison (observed gnomAD AF vs. calculated maximum population frequency, MPF) Filtering of COL2A1 (and panel gene) variants prior to classification gnomAD v2.1.1 population database (up to 280,928 alleles depending on gene/variant); patient cohort of 200 probands screened not stated
ACMG/AMP qualitative variant classification (rule-based criteria combination, e.g., PM1, PM2, PP2, PP3, BP4) via automated in-silico tool (VarSome v10.2) Classification of the four reported COL2A1 variants as LPV or VUS 4 index patients (plus segregation data in available relatives) not stated
Standard deviation (SD)/percentile scoring against reference height distributions Reporting patient height (e.g., +1.8 SD, >95th percentile) Individual patient values relative to an external reference growth dataset (cited reference 19) not stated
Revised Ghent nosology systemic scoring (clinical criteria point system) Diagnosis of MASS-like phenotype in each of the 4 subjects 4 index patients na
Approaches that could also have been used
  • Variant frequency evaluation was performed by comparing gnomAD allele frequency against a single calculated maximum population frequency (MPF) threshold per variant/gene.
    Could also: A formal statistical test of frequency difference (e.g., Fisher's exact test comparing observed vs. expected allele counts) or a Bayesian likelihood-ratio framework for pathogenicity — Such approaches can provide a quantitative probability or confidence interval around the frequency comparison in addition to the threshold-based filter already used.
  • Height was expressed as SD units and percentiles relative to reference growth data for each patient individually.
    Could also: Reporting a 95% confidence interval or z-score with an accompanying reference range table — This would let readers directly gauge the precision of each height measurement relative to the reference distribution, complementing the SD/percentile framing already provided.
  • Variant pathogenicity was determined using a rule-based ACMG/AMP criteria combination via an automated in-silico tool (VarSome).
    Could also: Multiple independent in-silico predictors (e.g., REVEL, CADD, AlphaMissense) reported individually alongside a meta-predictor consensus score — Presenting several predictor outputs side-by-side can illustrate the degree of concordance among tools, which is complementary to the single automated classification already used.
  • The report describes four patients without a formal comparison group or statistical hypothesis test.
    Could also: A case-control or cohort comparison (e.g., comparing phenotype scores or variant burden between COL2A1-positive and COL2A1-negative individuals within the 200-patient screened cohort) analyzed with an appropriate test such as Fisher's exact test or logistic regression — This design, common in case-series-to-cohort extensions, could quantify how distinctive the observed phenotype/variant association is relative to the broader screened population, complementing the descriptive case presentation.
  • Segregation analysis was performed qualitatively (presence/absence of variant in relatives) to support variant classification.
    Could also: A quantitative LOD-score-based linkage/segregation analysis where family size permits — This could offer a formal statistical measure of co-segregation strength, complementing the qualitative presence/absence assessment already reported.
Software: VarSome (in-silico ACMG/AMP variant classification) 10.2 · UCSF Chimera (molecular modeling/visualization) 1.2 · gnomAD database v2.1.1

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35296718

Title: Expanding the clinical spectrum of COL2A1 related disorders by a mass like phenotype. Authors: Demal TJ, Scholz T, Schüler H, Olfe J, Fröhlich A, Speth F, von Kodolitsch Y, Mir TS, Reichenspurner H, Kubisch C, Hempel M, Rosenberger G. Journal: Scientific Reports 12, 4544 (2022). DOI: 10.1038/s41598-022-08476-7. PMID 35296718 · PMCID PMC8927422. Study type: Clinical case report — four patients from three families with a MASS-like connective-tissue phenotype.

Data / Code availability (verbatim from the paper)

"All data generated or analysed during this study are included in this published article [and its supplementary information files]. All variants and associated phenotypes have been submitted to the ClinVar database (submission name SUB9514030)."

No code-availability statement. No repository. No raw/processed sequencing data deposited (clinical patient panel data).

In scope (pipeline-derived / in-silico results — reproduced here)

# Result Method named in paper Reproducible?
C1 COL2A1 vs COL1A1 C-propeptide 68.72 % sequence identity (homology model on PDB 5K31 template) SWISS-Model homology modeling Yes — re-aligned the two C-propeptide sequences
C2 The three COL2A1 variants are rare/absent in gnomAD v2.1.1 (filtering below the computed MPF threshold 5.15E-06) gnomAD v2.1.1 allele-frequency filter Yes — gnomAD API lookup
C3 Variants deposited in ClinVar (SUB9514030) with ACMG/AMP classifications (2× likely pathogenic, 1× VUS) VarSome v10.2 + ACMG/AMP, ClinVar submission Yes — ClinVar API lookup
C4 MPF (max credible population frequency) for COL2A1 = 5.15E-06 Disease-specific maxAF calculation Partial — value documented; not independently re-derivable without the authors' prevalence/penetrance/heterogeneity inputs (not stated)

Out of scope (not pipeline-derived OR data not available)

  • Variant detection from the 62-gene targeted NGS panel in the 4 patients — raw/processed patient sequencing data is not deposited (clinical, restricted). Cannot re-run variant calling. (data_restricted for that step.)
  • Clinical phenotyping (tall stature, arachnodactyly, dural ectasia, echocardiography, radiology, family pedigrees) — wet-lab/clinical/manual.
  • UCSF Chimera VDW clash/contact analysis on the homology models — feasible in principle (build SWISS-Model model + run findclash), but the reported output is qualitative (mutations introduce steric clashes), not a pinnable number; not run. Noted as a further step.

Heavy compute

None required — the paper has no heavy bioinformatic pipeline. The reproducible results are light in-silico steps (pairwise alignment, public-DB lookups). The pairwise alignment was nonetheless executed on «our HPC» («host») for provenance.

Figures / tables: Table
C1
Reported
68.72% COL2A1-COL1A1 C-propeptide identity (SWISS-Model, PDB 5K31)
Reproduced
69.1-69.8% (vs 5K31); 70.0-70.6% (vs UniProt COL1A1)
within tolerance
C2
Reported
3 COL2A1 variants rare/absent in gnomAD v2.1.1, below MPF 5.15E-06
Reproduced
K1312N exome AF=3.98E-06; D65N absent; S1338N absent (all < 5.15E-06)
exact
C3
Reported
variants deposited ClinVar SUB9514030; 2x likely pathogenic, 1x VUS
Reproduced
SCV001571667/668/669; K1312N Likely pathogenic, S1338N VUS, D65N authors' submission present (aggregate now Conflicting)
exact
C4
Reported
MPF threshold 5.15E-06
Reproduced
not re-derived (authors' prevalence/penetrance inputs unstated; downstream effect reproduced via C2)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Total score +2

This is a clinical case report with no analysis-code repo and no deposited patient sequencing data, so it is fundamentally not a 1:1 pipeline paper. The three publicly-supported in-silico claims reproduce cleanly: C1 homology-model identity 68.72% → ~69% (<1pp, alignment-method noise), C2 gnomAD v2.1.1 rarity exact (all variants <5.15E-06), and C3 ClinVar deposition SUB9514030 with matching classifications. Limitations are on the data-availability side (restricted patient NGS) and a methodology gap (MPF threshold not re-derivable from the paper's stated inputs), not on the authors' computational integrity — no fabrication indicators. Overall a solid partial reproduction with only explainable, negligible deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

141.5 k
tokens (I/O) · 7.1 M incl. cache
15 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.