Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Accelerating rare disease diagnostics by linking DNA and RNA through an explainable and interactive RNA-guided workflow.

NAR Genom Bioinform · 2026
L1 No data access 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • No authors-side cause for any deviation
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
No data access Data access not granted

This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.

Reproduction agent’s raw note

DROP / data_restricted. The paper (MOLGENIS VIP v7.9.1 DNA pipeline + FRASER v1.99.4 / OUTRIDER v1.20.1 RNA outlier detection on the UDN cohort) is described well enough to reproduce in principle, and the analysis CODE is fully public and version-pinnable (github.com/molgenis/vip, tag v7.9.1 = commit 28a7bade1a, LGPL-3.0, not archived). The blocker is purely the DATA: the only accession given for this RU, SRR10516820 (run-pair SRR10520582; exp SRX7202307; BioProject PRJNA350185; dbGaP phs001232/GRU = Undiagnosed Diseases Network), is dbGaP controlled-access WGS DNA. NCBI SDL locator returns HTTP 403 'Access denied - please request permission to access phs001232 / GRU in dbGaP'; ENA marks fastq as 'Protected file(s). Go to dbGap'; SRA download_path is the @dbgap@ protected namespace. No public FASTQ is obtainable on «our HPC» without an approved institutional Data Access Request. The cohort-level RNA claims (OUTRIDER/FRASER auPR, diagnostic yield) additionally require the full UDN cohort, not a single sample, and that cohort is likewise controlled-access. NOT ATTEMPTED: I deliberately did NOT run VIP's bundled test data and report it as a reproduction, because that would demonstrate only that the tool executes, not that any paper claim reproduces (no fabricated result). No «our HPC» compute was spent. To reproduce 1:1 one must obtain dbGaP authorization for phs001232, then run VIP v7.9.1@28a7bade1a + FRASER v1.99.4 / OUTRIDER v1.20.1 on the released reads against GRCh38_no_alt_analysis_set (Gencode v45 / v29).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-16 ⛓ 21664f79e1e3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a complete, explainable RNA-guided workflow that corrects biological and technical variation in RNA-seq data—integrating outlier expression and splicing analysis with DNA variant interpretation—accelerate the prioritization and reclassification of coding and non-coding variants in rare disease diagnostics, even when RNA is from accessible tissues sequenced across different centres, protocols, and timepoints?

Core claims
  • An integrated RNA-guided variant interpretation workflow combining OUTRIDER, FRASER, MOLGENIS VIP, and Borzoi enhances clinical variant interpretation and reclassification of VUS in rare disease cases. resource
  • The workflow handles biological and technical variation from RNA-seq generated by different centres, protocols, and timepoints, enabling outlier detection in realistic multi-source cohorts. method
  • RNA outlier (aberrant expression and splicing) analysis adds evidence that aids prioritization of coding and non-coding SNVs and reclassification of clinically relevant VUS. finding
  • Borzoi, a sequence-based model, links DNA SNVs near the TSS to predicted gene expression changes and can prioritize regulatory non-coding variants. method
  • The workflow produces self-contained interactive reports visualizing outlier genes and prioritized patient-level variants for immediate clinical interpretation. resource
  • A customized VIP decision tree classifies coding and non-coding SNVs as B, LB, VUS, or LP/P following ACMG/VKGL guidelines. method
  • A Borzoi threshold for distinguishing potentially pathogenic from benign expression-affecting variants was derived using ClinVar and gnomAD benchmark variants. method
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome / whole-exome sequencing (DNA variant calling & annotation via MOLGENIS VIP) UDN rare disease patients (144 patients / 165 samples) none SNV classification (B/LB/VUS/LP/P) and annotation MOLGENIS VIP v7.9.1; DeepVariant v1.6.1; GRCh38 reference
Bulk RNA-seq aberrant expression analysis (OUTRIDER) Whole blood (PAXgene) from UDN patients none Outlier gene expression z-scores (P-adjusted < 0.05) TruSeq stranded mRNA/total; STAR v2.7.9a; OUTRIDER v1.20.1; Nextflow v24.04.2
Bulk RNA-seq aberrant expression analysis (OUTRIDER) Skin fibroblasts from UDN patients none Outlier gene expression z-scores (P-adjusted < 0.05) TruSeq stranded total RNA; STAR v2.7.9a; OUTRIDER v1.20.1
Bulk RNA-seq aberrant splicing analysis (FRASER) Whole blood and skin fibroblasts from UDN patients none Outlier splicing via intron Jaccard index, z-scores (P-adjusted < 0.05) FRASER v1.99.4; Nextflow v24.04.2
Splice-altering variant prediction (SpliceAI) SNVs in patient WGS/WES data none Predicted splicing effect (cut-off 0.42) SpliceAI within MOLGENIS VIP
Sequence-based gene expression effect prediction (Borzoi) SNVs within 5 kb of outlier-gene transcripts; ClinVar/gnomAD benchmark variants near TSS none Predicted expression change score / Boolean threshold pass Borzoi model
Phenotype and segregation analysis UDN patients with HPO terms none Gene-phenotype match and segregation/inheritance classification of candidate SNVs HPO database (release 13 Aug 2024); OMIM
Key results
  • Cohort of 165 samples from 144 UDN patients assembled across three centres (BCM, UCLA, SU) with WGS/WES, RNA-seq, and HPO data
  • For 21 patients, both whole blood and skin fibroblast RNA-seq data were available
  • 227 additional whole-blood and 93 additional skin fibroblast RNA-seq samples included as background to correct biological/technical variation
  • 10 RNA-first-diagnosed validation samples contained 17 known variants (9 SNVs) in genes with splicing/expression outliers and phenotype match
  • Borzoi threshold chosen by how well it identified LP promoter SNVs near the TSS, distinguishing potentially pathogenic from benign expression-affecting variants
  • Prior RNA-seq studies reported additional molecular diagnoses (Frésard 7.5%, Murdock 12%), motivating the RNA-guided approach 7.5%–12%
Key statistics
  • count 144 patients, 165 samples (Total UDN cohort analysed)
  • count 77 patients with WB RNA-seq; 88 with skin FB RNA-seq (Affected patients by tissue type)
  • count 21 (Patients with both WB and skin FB RNA-seq)
  • count 227 WB + 93 FB additional RNA-seq samples (Background samples for variation correction)
  • other P-adjusted < .05 (Benjamini–Yekutieli FDR) (Significance cut-off for outlier expression/splicing)
  • other 0.42 (SpliceAI cut-off used in VIP)
  • other pHaplo score > 0.4; MAF > 0.1%; 600 bp upstream window (Variant selection criteria for Borzoi threshold benchmark)
  • other 7.5% (Frésard) and 12% (Murdock) additional molecular diagnoses (Prior UDN RNA-seq diagnostic yields cited)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper presents a clinical RNA-guided variant interpretation workflow integrating OUTRIDER (negative binomial–based outlier expression detection) and FRASER (beta-binomial–based splicing outlier detection) applied to a multi-centre cohort of 144 rare disease patients with whole-blood or fibroblast RNA-seq data. Outlier significance was determined using the Benjamini-Yekutieli FDR method at an adjusted P < 0.05, with results expressed as z-scores. Autoencoder performance for variation correction was optimised using the area under the precision-recall curve (auPR), and a Borzoi score threshold for predicting expression-altering non-coding SNVs was derived from a ClinVar/gnomAD benchmark. Results were delivered as patient-level interactive HTML reports for clinical interpretation.

Replicationmixed Sample size144 patients (165 RNA-seq samples: 77 WB, 88 skin FB, 21 with both); additional 227 WB and 93 FB background samples included for variation correction; 10 previously diagnosed samples used as positive controls; no formal power calculation described GroupsAffected patients (WGS/WES + RNA-seq + HPO) vs. background cohort (RNA-seq only); WB vs. skin FB as parallel analyses; three sequencing centres (BCM, UCLA, Stanford) Pairingmixed Randomization/blindingnot stated Dispersionnone Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Yekutieli (BY) false discovery rate
Statistical tests used
Test Applied to n Assumptions
OUTRIDER negative binomial Wald-type test; z-scores reported for outlier expression Outlier gene expression detection for each patient vs. background cohort (whole blood and fibroblast separately) 304 BAM files stated for full pipeline run; composed of 77 affected WB patients + 227 background WB samples, and 88 affected FB patients + 93 background FB samples not stated
FRASER beta-binomial test on intron Jaccard index; z-scores reported for outlier splicing Outlier splicing detection for each patient vs. background cohort (whole blood and fibroblast separately) 304 BAM files (as stated for full pipeline run) not stated
Area under the precision-recall curve (auPR) on automatically imputed outlier values Optimization and evaluation of OUTRIDER and FRASER built-in autoencoders for biological/technical variation correction na
Borzoi score threshold determination by benchmark against ClinVar LP/P vs LB/B variants and gnomAD variants (MAF > 0.1%); optimization metric stated as identifying LP promoter SNVs, specific criterion not stated Threshold for predicted expression-change annotation of TSS-proximal and 5' UTR SNVs in outlier genes not stated
Approaches that could also have been used
  • Multiple testing correction used the Benjamini-Yekutieli (BY) FDR method, which controls FDR under arbitrary dependency structures between tests.
    Could also: The Benjamini-Hochberg (BH) FDR method could also have been used; it is the standard in genomics outlier pipelines (including OUTRIDER's defaults in published literature) and assumes positive regression dependency (PRDS), a condition often considered reasonable for gene expression data. — BH yields higher power (more discoveries) than BY in the same dataset; explicitly comparing both thresholds, or reporting which is used by default in the tool, helps readers calibrate sensitivity relative to published benchmarks.
  • The Borzoi score threshold was derived by benchmarking against ClinVar variants, but the specific criterion used to select the threshold (e.g., maximising F1, achieving sensitivity ≥ 0.8 at fixed specificity, or Youden index) is not stated in the available text.
    Could also: A full precision-recall or ROC curve with reported AUC, alongside the explicit optimisation criterion and its resulting sensitivity/specificity on the benchmark set, could also characterise threshold performance. — Reporting AUC and the selection criterion makes the threshold reproducible and allows users to adapt the cut-off to their own false-positive tolerance, which may differ across clinical settings.
  • Technical and biological variation across three centres, two tissue types, and multiple protocols was addressed using the OUTRIDER and FRASER built-in autoencoders.
    Could also: Explicit upstream batch-correction methods — such as ComBat, limma's removeBatchEffect, or RUVSeq — applied before outlier detection could also model centre and protocol effects. — Making the batch-effect model explicit and separate from the outlier-detection step can improve interpretability and allow independent assessment of how much variance each centre or protocol contributes, which is especially relevant in multi-centre diagnostic settings.
  • Workflow validation relied on 10 samples with previously published molecular diagnoses as positive controls, representing a small benchmark for estimating sensitivity.
    Could also: Leave-one-out cross-validation on those known-positive cases, or a larger simulated spike-in benchmark (e.g., randomly down-regulating genes in controls to mimic expression outliers), could also be used to estimate sensitivity and specificity more stably. — Performance estimates based on n = 10 known cases carry wide uncertainty; a cross-validated or simulated benchmark provides more stable sensitivity/specificity metrics to guide prospective clinical adoption decisions.
  • Outlier magnitude is reported as z-scores, without accompanying fold-change or absolute expression values.
    Could also: Log2 fold-change (observed vs. predicted expression from OUTRIDER's fitted model) could also be reported alongside z-scores and adjusted P-values. — Fold-change conveys biological magnitude independently of sample size and is more directly interpretable clinically; pairing it with z-scores and adjusted P-values is common practice in differential expression reporting and aids in gauging whether an outlier's expression change is large enough to be functionally meaningful.
  • Autoencoder optimisation used auPR computed on automatically imputed outlier values within the training data.
    Could also: Held-out sample cross-validation (e.g., k-fold or leave-one-out) or evaluation on an independent cohort could also assess generalisation of the corrected model to unseen samples. — Training-set auPR may overestimate performance on prospective samples; cross-validated or independent-cohort auPR/AUC-ROC provides a less optimistic estimate of how well variation correction generalises across new patients or centres.
Software: OUTRIDER (R/Bioconductor) 1.20.1 · FRASER (R/Bioconductor) 1.99.4 · MOLGENIS VIP 7.9.1 · Nextflow 24.04.2 · STAR 2.7.9a · DeepVariant 1.6.1 · TrimGalore 0.6.7 · Borzoi · SpliceAI (integrated in VIP) · SRA toolkit 3.1.1

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GCA_000001405.15 GCA in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
phs001232 dbGaP in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRR10762563 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR12358183 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR13431041 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR13431517 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR13431547 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR13431587 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR13431959 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR13433024 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR22091353 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR22091410 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR22092000 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR6132800 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR6132839 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR6133613 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR8061687 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR8062654 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR8068094 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR8068341 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR8086648 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR8099971 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR8112700 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41685349

Title: Accelerating rare disease diagnostics by linking DNA and RNA through an explainable and interactive RNA-guided workflow. Journal: NAR Genom Bioinform 2026 · DOI 10.1093/nargab/lqag016 · PMCID PMC12891912 Code: https://github.com/molgenis/vip (MOLGENIS VIP) — public, LGPL-3.0, tag v7.9.1 = commit 28a7bade1a (the exact version used in the paper). Data (as given in RU): sra:SRR10516820.

Pipelines named in the paper (Methods)

  • MOLGENIS VIP v7.9.1 — DNA variant interpretation (WGS/WES). DeepVariant v1.6.1 (GQ ≥ 20), SpliceAI cutoff 0.42, Nextflow v24.04.2 orchestration. Reference GRCh38_no_alt_analysis_set_GCA_000001405.15, Gencode v45 for variant annotation.
  • OUTRIDER v1.20.1 and FRASER v1.99.4 — RNA expression / splicing outlier detection (whole blood & skin fibroblast cohorts). Gencode v29 for RNA alignment. Borzoi (14.85-fold) and P-adjusted < .05 (Benjamini–Yekutieli FDR).

In-scope (pipeline-derived) results

Result Pipeline Reproducible in principle
Per-patient avg 505.96 VUS SNVs prioritised [95% CI 485.67–526.25] VIP v7.9.1 needs cohort WGS/WES
MRE11 splicing-outlier case (validation): compound het NM_005591.4 c.1726C>T + c.1500+1153_1563+1027del; AR; splicing outlier in whole blood FRASER + VIP this is the sample SRR10516820's case
OUTRIDER auPR 0.55 (WB) / 0.50 (skin FB) OUTRIDER needs full RNA cohort
FRASER auPR 0.74 (WB) / 0.80 (skin FB) FRASER needs full RNA cohort
Diagnostic yield 4% (undiagnosed); 5/10 previously diagnosed reproduced VIP+RNA needs full cohort

Out-of-scope / not attempted

  • All cohort-level outlier metrics (OUTRIDER/FRASER auPR, diagnostic yield): FRASER/OUTRIDER are cohort methods — a single sample cannot yield an outlier call. The cohort is the UDN dataset, not shipped as a single accession.
  • Wet-lab validation, manual curation, interactive VIP-report inspection.

DROP — data_restricted (the binding blocker)

The only data accession provided for this RU, SRR10516820 (and its run-pair SRR10520582, same experiment SRX7202307 / biosample SAMN12913659), is dbGaP controlled-access, not public:

  • NCBI SRA download_path = @dbgap@:reads/... (dbGaP-protected namespace).
  • ENA filereport submitted_ftp / fastq_ftp = empty; note: "Protected file(s). Go to dbGap."
  • NCBI SDL data locator (locate.ncbi.nlm.nih.gov/sdl/2/retrieve?acc=SRR10516820): HTTP 403 — "Access denied - please request permission to access phs001232 / GRU in dbGaP."
  • dbGaP study phs001232 = Undiagnosed Diseases Network (UDN), consent group GRU (General Research Use) → requires an approved institutional Data Access Request; not downloadable on «our HPC»/anywhere without authorization.
  • Entire BioProject PRJNA350185 (UDN) is GENOMIC and shows no public FASTQ.

The sample is WGS DNA ("DNA sample from Blood of a human male"), not the RNA data; the RNA cohort needed for the FRASER/OUTRIDER claims is likewise UDN controlled-access.

Conclusion: the analysis code is fully public and even version-pinned (VIP v7.9.1 @ 28a7bade1a), but no public input data exists to run it against any reported paper claim. Per the controlled drop vocabulary this is data_restricted (controlled-access, distinct from data_unavailable). No compute spent on «our HPC»; running VIP's own bundled test data would demonstrate the tool executes but would not reproduce any paper claim, so it was not done (no fabricated result).

C1_avg_VUS_SNVs
Reported
505.96 SNVs/patient [95% CI 485.67, 526.25]
Reproduced
not attempted
m.public.grade.error
C2_MRE11_case
Reported
MRE11 compound het NM_005591.4:c.1726C>T + c.1500+1153_1563+1027del; splicing outlier whole blood (sample SRR10516820)
Reproduced
not attempted
m.public.grade.error
C3_OUTRIDER_auPR
Reported
0.55 whole blood / 0.50 skin fibroblast
Reproduced
not attempted
m.public.grade.error
C4_FRASER_auPR
Reported
0.74 whole blood / 0.80 skin fibroblast
Reproduced
not attempted
m.public.grade.error
C5_diagnostic_yield
Reported
4% diagnostic yield; 5/10 previously diagnosed reproduced
Reproduced
not attempted
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 44/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

56.8 k
tokens (I/O) · 2.7 M incl. cache
7 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.