Exome sequencing in 38 patients with intracranial aneurysms and subarachnoid hemorrhage.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the ANNOTATION LAYER, but the headline result is irreproducible by anyone. The core WES variant-calling pipeline (the result that PRODUCED Table 2) CANNOT be reproduced: no patient exome data was deposited (no EGA/dbGaP/SRA/ENA accession, no on-request offer) - a genuine data-unavailable drop for that result. What IS reproducible is the downstream annotation/stat layer on the 20 reported risk-gene variants, and it reproduces strongly: gnomAD r2.1.1 MAF 20/20 (9 exact, 11 within 1% rel); VEP GRCh37 protein/transcript 20/20 exact; CADD v1.6 14/20 within +/-1 PHRED (rest = predictor version drift, same direction); GTEx v8 confirms EDIL3 highest in arteries+brain (Fig S3); and the EDIL3 Fisher's exact p reproduces to 4 sig figs (0.006237 vs 0.00624). No fabrication signal: every reported annotation value is independently re-derivable from the variant IDs + public DBs. Verdict: PARTIAL - annotation/stat layer reproduced 1:1, primary WES pipeline a data-unavailable drop.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 78assessed: 2026-06-22 ⛓ 98b264fd66bb
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study aimed to verify recently reported genetic risk genes (ADAMTS15, ANGPTL6, ARHGEF17, LOXL2, PCNT, RNF213, THSD1, TMEM132B) and to identify novel sequence variants involved in the etiology of unruptured intracranial aneurysms (UIA) and aneurysmal subarachnoid hemorrhage (aSAH) using exome sequencing.
- ★ Sequence variants in PCNT, RNF213 and THSD1 support a role as susceptibility factors for cerebrovascular disease (UIA/aSAH) finding
- ★ EDIL3 is proposed as a novel valid candidate disease gene for UIA/aSAH based on pathogenicity, population genetics and vascular biology relevance finding
- ★ Exome sequencing was performed in 35 unrelated individuals and 3 affected family members to screen known risk genes and discover novel ones method
- No MAF≤5% variants were detected in ARHGEF17 or LOXL2 among the 38 exome-sequenced patients finding
- ★ Prioritization of variants shared among three affected relatives yielded five novel putative risk genes finding
- ★ Screening of an additional 37 individuals identified a further very rare EDIL3 variant in two unrelated sporadic patients finding
- Study cohort was enriched for young patients with an elevated number of aneurysms and above-average family history relative to literature method
- Patients with known vascular/connective tissue disorder genes (e.g. polycystic kidney disease, Ehlers-Danlos, Loeys-Dietz, Marfan) were clinically excluded prior to genetic analysis method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Exome sequencing (ES) | 35 unrelated patients with UIA and/or aSAH | none | rare/low-frequency sequence variants (MAF ≤5%) in 8 reported risk genes | — |
| Exome sequencing (ES) | 3 affected members of one family | none | shared unknown (MAF=0) variants for novel candidate gene discovery | — |
| Sanger sequencing | 37 additional patients with UIA and/or aSAH | none | targeted screening of EDIL3 candidate gene variants | — |
| Gene/variant association test | 75-patient clinical cohort | none | statistical association of EDIL3 variants with UIA/aSAH | — |
| Molecular modeling | EDIL3 protein structure | none | predicted structural/localization impact of identified variant | — |
| In silico pathogenicity and splice-site prediction | identified variants in candidate genes (PCNT, RNF213, THSD1, ANGPTL6, ADAMTS15, TMEM132B) | none | predicted pathogenicity (CADD, REVEL, M-CAP, ClinPred) and splicing effects (HSF, NetGene2, MaxEntScan, BDGP) | CADD/REVEL/M-CAP/ClinPred; HSF/NetGene2/MaxEntScan/BDGP |
- – 20 rare (MAF≤5%) heterozygous missense variants identified in 18 of 38 exome-sequenced patients across 6 of 8 candidate genes 20 variants/18 patients
- – PCNT variants identified in 9 patients 9 variants/9 patients
- – RNF213 variants identified in 3 patients 4 variants/3 patients
- – THSD1 variants identified in 6 patients 3 variants/6 patients
- – No variants detected in ARHGEF17 or LOXL2
- – Family-based variant prioritization yielded five novel putative risk genes; EDIL3 selected as top candidate 5 genes
- – Additional very rare EDIL3 variant found in 2 unrelated sporadic patients among 37 further screened individuals 2/37
- – ANGPTL6 variants found in 3 patients; ADAMTS15 and TMEM132B each carried 1 variant in 1 patient 2 variants/3 patients (ANGPTL6); 1/1 each (ADAMTS15, TMEM132B)
- count 75 subjects (53 female, 70.7%) (overall study cohort composition)
- count 48 aSAH (64.0%), 27 UIA (36.0%) (disease subgroup breakdown)
- mean PHASES score 4.8 (range 0-13) (aneurysm rupture risk score across cohort)
- count 30 individuals (40.0%) with positive family history (family history prevalence)
- count MAF ≤0.05 (5%) cutoff yielded 20 variants across 6 of 8 screened genes (candidate risk gene variant screen)
- other gnomAD MAF 0.006092% for ADAMTS15 p.Gly420Ser (population frequency of an example novel variant)
- mean Age at SAH 46.9 ± 12.1 years (range 9-71) (age distribution of aSAH patients)
- count 9 PCNT variants in 9 patients; 4 RNF213 variants in 3 patients; 3 THSD1 variants in 6 patients (gene-level variant and patient counts among exome-sequenced individuals)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is an observational exome-sequencing case-series study of 75 patients with unruptured intracranial aneurysms (UIA) and/or aneurysmal subarachnoid hemorrhage (aSAH), with clinical characteristics summarized descriptively (means ± SD with range, and counts/percentages) in Table 1. Genetic analysis involved exome sequencing in 38 individuals, filtering variants in eight candidate genes by minor allele frequency (MAF) thresholds against gnomAD/dbSNP reference population data, classifying variants per ACMG/AMP criteria, and using in silico pathogenicity predictors (CADD, REVEL, M-CAP, ClinPred) plus splice-site predictors. The text references 'gene/variant association tests' supporting the candidate gene EDIL3, but the specific statistical test(s) are detailed in Supplementary Methods, which are not included in the provided text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Descriptive summary statistics (mean ± SD, range; counts/percentages) | Table 1, clinical characteristics of the cohort and subgroups (aSAH, UIA, ES IND, FAM IND, SPO IND) | n = 75 total cohort; subgroup n as listed (aSAH n=48, UIA n=27, ES IND n=35, FAM IND n=30, SPO IND n=45) | not stated |
| Gene/variant association test (specific test unspecified in provided text) | supports EDIL3 as candidate risk gene (per Conclusions); detailed in Supplementary Methods, not included here | — | not stated |
| In silico pathogenicity prediction scoring (CADD, REVEL, M-CAP, ClinPred) with stated thresholds | Table 2, classification of individual sequence variants in reported risk genes | per variant (20 variants across 18 individuals) | stated (thresholds: CADD ≥20, REVEL ≥0.5, M-CAP ≥0.025, ClinPred ≥0.5) |
-
Clinical characteristics of subgroups (e.g., aSAH vs UIA, familial vs sporadic) are summarized descriptively with mean ± SD (range) and percentages, without an accompanying inferential comparison in the provided text.↳ Could also: Formal between-group hypothesis tests could also be used, such as Student's t-test or the Mann-Whitney U test for continuous variables (e.g., age, aneurysm diameter, PHASES score) and chi-square or Fisher's exact test for categorical variables (e.g., sex, family history). — These tests would quantify whether the observed descriptive differences between subgroups are statistically distinguishable, complementing the summary statistics already presented.
-
Candidate gene variants were evaluated by comparing observed MAF in patients against gnomAD/dbSNP population frequencies, using frequency thresholds to flag rare or unknown variants.↳ Could also: A formal case-control burden or association test (e.g., Fisher's exact test on carrier counts, or a gene-based burden test such as SKAT/SKAT-O) could also be applied to compare variant carrier frequency between the patient cohort and a reference population. — This would yield an explicit p-value and effect size (e.g., odds ratio) for enrichment of variants in candidate genes, in addition to the frequency-threshold-based filtering approach.
-
Eight candidate genes (and multiple variants within them) were screened for association with UIA/aSAH.↳ Could also: A multiple-testing correction such as Bonferroni adjustment or Benjamini-Hochberg false discovery rate control could also be applied across the panel of genes/variants tested. — This would help control the family-wise error rate or false discovery rate when many genes are evaluated in parallel, which can be a useful complement when reporting significance across a multi-gene screen.
-
Variant pathogenicity was assessed by combining several in silico predictors (CADD, REVEL, M-CAP, ClinPred) against individual thresholds, integrated qualitatively per ACMG/AMP criteria.↳ Could also: A combined ensemble or meta-predictor score with an associated quantitative probability and confidence interval could also be used. — This could provide a single quantified estimate of pathogenicity likelihood rather than a qualitative combination of independent threshold-based calls.
-
The candidate gene EDIL3 was prioritized from shared rare variants in one family and then found in 2 additional sporadic patients out of 37 screened, described in count terms.↳ Could also: A formal case-control statistical comparison (e.g., Fisher's exact test comparing EDIL3 variant carrier rate in the patient cohort versus its rate in gnomAD) with a reported p-value or odds ratio could also be presented. — This would give a quantified statistical measure of association strength for the candidate gene, complementing the descriptive carrier counts.
-
The exome-sequenced cohort size (38 patients, later extended to 75) was determined by clinical recruitment criteria rather than a stated power calculation.↳ Could also: A post-hoc power or sample-size sensitivity analysis could also be reported. — This would clarify the probability of detecting true gene-disease associations of a given effect size in a cohort of this size, providing additional context for interpreting negative or borderline findings.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32367296
Paper: Sauvigny T, …, Rosenberger G. "Exome sequencing in 38 patients with intracranial aneurysms and subarachnoid hemorrhage." Journal of Neurology 267(9):2533–2545, 2020. DOI 10.1007/s00415-020-09865-6 · PMC7419486.
What kind of paper
Clinical germline whole-exome sequencing (WES) study of an intracranial aneurysm (IA) / aneurysmal subarachnoid hemorrhage (aSAH) cohort. 38 individuals exome-sequenced (35 unrelated + 1 three-member family). Candidate-gene screen of known IA risk genes + family-based discovery → EDIL3 proposed as a novel candidate.
Data availability (critical)
- No data-availability statement and NO deposit of patient sequence data. No EGA / dbGaP / SRA / ENA accession anywhere in the article or PMC record.
- The only "data" that are public are the summary tables (variant lists) in the paper + Word supplement. The PMC supplement is behind a proof-of-work download wall; Europe PMC mirror has no supplement file. Table 2 (the 20 risk-gene variants) was captured from the article HTML.
- GTEx (phs000424) is cited as external reference data, not a deposit of this study's data.
In scope (pipeline-derived, reproducible WITHOUT the raw exomes)
These are deterministic annotation/stat outputs derivable from the reported variants + public reference databases — a faithful re-run of the paper's annotation/filtering layer:
- R1 — gnomAD MAF of the 20 reported risk-gene variants (paper: gnomAD v2.1/v2.1.1).
- R2 — CADD score of the 20 variants (in-silico pathogenicity predictor used in Methods).
- R3 — variant→protein mapping (HGVS p.) of the 20 variants (confirms variant identity/transcript).
- R4 — EDIL3 tissue expression (Suppl. Fig S3, GTEx) — qualitative reproduction.
- R5 — EDIL3 case/control association p-value (reported p=0.0152) — attempted.
OUT of scope (cannot be reproduced)
- The core WES pipeline (FASTQ → BAM → variant calls): NOT reproducible. The raw/processed patient exome data were never deposited and are not available on request in the paper. Alignment, variant calling, and the genome-wide filtering that produced the variant list cannot be re-run. This is the primary pipeline-derived result and it is a DROP for lack of input data.
- Wet-lab work: Sanger validation, cosegregation genotyping, molecular modelling of EDIL3 (structural), patient phenotyping — manual/experimental, out of scope.
- SKAT / gene-burden test internals (Table S5) — depend on per-sample genotypes not available; only the headline p-value could be checked (R5).
Pipeline tools named (Methods/Suppl.)
gnomAD v2.1/v2.1.1, dbSNP, ClinVar, in-house IKMB/PopGen controls; predictors CADD (≥20), REVEL (≥0.5), M-CAP (≥0.025), ClinPred (≥0.5); splice tools HSF, NetGene2, MaxEntScan, BDGP; ACMG/AMP classification. Aligner/variant-caller names are relegated to the Supplement (not in main text).
Reproduction strategy
Re-derive R1–R3 by annotating the 20 variants against GRCh37 via Ensembl VEP REST (coordinates + protein change), gnomAD r2.1.1 GraphQL API (MAF), and CADD v1.6 (remote-tabix), all on «our HPC». R4 via GTEx v8 API. R5 via exact Fisher test on the reported counts. Compare to the paper's Table 2 / Suppl values.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.