Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genetic Variants in ARHGEF6 Cause Congenital Anomalies of the Kidneys and Urinary Tract in Humans, Mice, and Frogs.

· 2023
PubMed 36414417 ↗ pmid-36414417
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Total score +9
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL — described well enough; in-silico annotation reproduces, cohort discovery is restricted-data (different/not-attempted). The paper's Table-1 in-silico annotation of the 6 ARHGEF6 variants reproduces strongly on the paper's own variant list (transcript NM_004840.3, GRCh37): gnomAD v2.1.1 (2 exact, 1 within-tol on AN denominator only), CADD v1.6 (20.3 & 23.1 EXACT, version pinned), REVEL & MutationTaster exact. SIFT/PolyPhen-2 were INDEPENDENTLY RE-COMPUTED on a «our HPC» compute node (SLURM 2218562, Ensembl VEP 110.1 offline, SIFT 5.2.2 + PolyPhen 2.2.2): A5124 SIFT deleterious=Del (exact) and PolyPhen-2 0.938~=0.94 (within-tol, best match of any method); B3089 PolyPhen-2 benign (partial); B3089 SIFT 'tolerated' — a genuine mismatch vs paper 'Del' now confirmed by TWO independent SIFT engines (flag for human, not a standalone fabrication signal). VEP also confirms all variant consequences + HGVS on NM_004840.3. NOT attempted: upstream exome cohort discovery (1265 families) — raw WES restricted by patient consent (on-request, no accession); and wet-lab functional work (cells/mouse/frog) — out of scope. All grades provisional; a human signs off.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-19 ⛓ 84a9243814d2
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Because known CAKUT genes explain only ~20% of cases, the authors test whether deleterious variants in ARHGEF6 (an X-linked guanine nucleotide exchange factor acting downstream of integrin-linked kinase and parvin proteins) are a novel monogenic cause of congenital anomalies of the kidneys and urinary tract.

Core claims
  • Hemizygous variants in the X-linked gene ARHGEF6 cause X-linked CAKUT in humans finding
  • Deleterious ARHGEF6 variants may cause disease via dysregulation of integrin-parvin-RAC1/CDC42 signaling mechanism
  • Wild-type ARHGEF6, but not proband-derived mutant ARHGEF6, increases active CDC42/RAC1, induces lamellipodia, and stimulates PARVA-dependent cell spreading finding
  • ARHGEF6-mutant proteins show loss of interaction with PARVA mechanism
  • Arhgef6 deficiency in mouse and frog models recapitulates features of human CAKUT finding
  • 3D MDCK cultures expressing ARHGEF6-mutant proteins show reduced lumen formation and polarity defects finding
  • Exome sequencing in an international CAKUT cohort can identify novel monogenic disease genes method
Experimental setups
Assay System Perturbation Readout Platform
Exome sequencing International cohort of 1265 families with CAKUT none Pathogenic/deleterious genetic variants
CDC42/RAC1 activation assay Kidney cells Overexpression of wild-type vs mutant ARHGEF6 Active levels of CDC42/RAC1
Lamellipodia formation / cell spreading assay Kidney cells Overexpression of wild-type vs mutant ARHGEF6 Lamellipodia formation and PARVA-dependent cell spreading
Protein-protein interaction assay Cells expressing ARHGEF6 Wild-type vs mutant ARHGEF6 ARHGEF6-PARVA interaction
3D cyst/lumen culture Madin-Darby canine kidney (MDCK) cells Expression of ARHGEF6-mutant proteins Lumen formation and cell polarity
In vivo developmental phenotyping Mouse (Arhgef6 deficiency) Arhgef6 deficiency/KO CAKUT-like renal/urinary tract phenotypes
In vivo developmental phenotyping Frog (Xenopus, Arhgef6 deficiency) Arhgef6 deficiency/knockdown CAKUT-like renal phenotypes
Key results
  • Six different hemizygous ARHGEF6 variants detected in eight individuals from six families with CAKUT
  • Wild-type ARHGEF6 overexpression increased active CDC42/RAC1 levels, whereas mutant did not
  • Wild-type ARHGEF6 induced lamellipodia formation and stimulated PARVA-dependent cell spreading; mutant did not
  • ARHGEF6-mutant proteins lost interaction with PARVA
  • 3D MDCK cultures expressing ARHGEF6-mutant proteins showed reduced lumen formation and polarity defects
  • Arhgef6 deficiency in mouse and frog recapitulated features of human CAKUT
Key statistics
  • count 1265 families (International CAKUT cohort exome-sequenced)
  • count six different hemizygous variants (ARHGEF6 variants detected)
  • count eight individuals from six families (Individuals with CAKUT carrying ARHGEF6 variants)
  • other about 40 disease genes (Known isolated CAKUT disease genes to date)
  • other 20% (Proportion of CAKUT cases explained by known genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper is a genetics/functional-genomics study combining exome sequencing of an international CAKUT cohort with cellular (kidney cell overexpression, 3D MDCK culture) and animal (mouse, frog) models. The provided text (title, author/affiliation block, and structured abstract with Background/Methods/Results/Conclusions) does not include a dedicated statistical methods section, so specific tests, sample-size justifications, or reporting details are not stated in the excerpt available.

Replicationunclear Sample sizeA cohort of 1265 families with CAKUT was exome sequenced, yielding six variants in ARHGEF6 in eight individuals from six families; no statement of biological/technical replicate numbers for the cellular or animal experiments is present in this excerpt GroupsWild-type vs. proband-derived mutant ARHGEF6 in cellular assays; Arhgef6-deficient vs. control mouse and frog models Pairingunclear Randomization/blindingnot stated Dispersionunclear
Approaches that could also have been used
  • The study identifies candidate disease variants via exome sequencing across a large multi-family cohort (1265 families).
    Could also: A formal statistical burden or case-control enrichment test (e.g., gene-based burden testing against a population reference such as gnomAD) could also be reported — This would give a quantitative estimate (e.g., odds ratio or p-value) of how unlikely the observed variant clustering in ARHGEF6 is under a null model, complementing the qualitative segregation/functional evidence described.
  • Functional differences between wild-type and mutant ARHGEF6 were described qualitatively (e.g., 'increased active levels of CDC42/RAC1,' 'reduced lumen formation') based on cellular and 3D culture assays.
    Could also: Quantifying these readouts (e.g., percentage of spheroids with a single lumen, active GTPase levels by densitometry) and comparing wild-type vs. mutant groups with a standard test such as an unpaired t-test, Mann-Whitney U, or a proportion test (e.g., Fisher's exact or chi-square) could also be used — Reporting a summary statistic with a measure of dispersion and a p-value or confidence interval would let readers gauge the magnitude and precision of the wild-type vs. mutant difference, alongside the descriptive phenotype.
  • Multiple independent models (human cohort, kidney cell overexpression, 3D MDCK culture, mouse, and frog) are used to support the same conclusion about ARHGEF6 function.
    Could also: A pre-specified analysis plan noting the number of biological/technical replicates per model and, where multiple phenotype endpoints are tested within a model, a multiple-comparison correction (e.g., Benjamini-Hochberg FDR or Bonferroni) could also be applied and reported — This would make explicit how many independent comparisons were run per model and control the chance of a false-positive finding when several endpoints (e.g., lamellipodia formation, cell spreading, lumen formation, polarity markers) are assessed together.
  • Sample sizes/replicate counts for the in vitro and animal experiments are not stated in this excerpt.
    Could also: Reporting the exact n for each experiment (e.g., number of independent transfections, embryos, or litters) and, where feasible, a power calculation or effect-size estimate could also be included — Explicit n and effect sizes help readers assess whether the study was adequately powered to detect the reported phenotypic differences between wild-type and mutant conditions.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Table
R1_gnomAD_GM2
Reported
0/0/4/205,321
Reproduced
0/0/4/205,321
exact
R1_gnomAD_A5124
Reported
0/4/10/183,437
Reproduced
0/4/10/183,437
exact
R1_gnomAD_B3089
Reported
0/7/19/204,748
Reproduced
0/7/19/204,031
within tolerance
R2_CADD_B3089
Reported
20.3
Reproduced
20.3 (CADD v1.6)
exact
R2_CADD_A5124
Reported
23.1
Reproduced
23.1 (CADD v1.6)
exact
R3_REVEL_B3089
Reported
0.092
Reproduced
0.092
exact
R3_REVEL_A5124
Reported
0.44
Reproduced
0.44
exact
R3_MutTaster_B3089
Reported
Dis
Reproduced
D (disease_causing)
exact
R3_MutTaster_A5124
Reported
Dis
Reproduced
D (disease_causing)
exact
R3_SIFT_A5124
Reported
Del
Reproduced
deleterious(0) [VEP-«our HPC» SIFT 5.2.2; dbNSFP D(0.001) concordant]
exact
R3_PolyPhen2_A5124
Reported
0.94
Reproduced
0.938 [VEP-«our HPC» PolyPhen 2.2.2; ~exact]
within tolerance
R3_SIFT_B3089
Reported
Del
Reproduced
tolerated(0.09) [VEP-«our HPC» + dbNSFP 0.098 both 'tolerated']
did not match
R3_PolyPhen2_B3089
Reported
0.36
Reproduced
0.072 (benign) [VEP-«our HPC»]
partial
VEP_consequence_HGVS
Reported
GM1 nonsense c.571C>T; GM2 splice c.2135+4A>G; B3089/A5124 missense (NM_004840.3)
Reproduced
GM1 stop_gained c.571C>T; GM2 splice_donor_region c.2135+4A>G; missense confirmed
exact
OUT_cohort_discovery
Reported
6 variants/8 individuals/6 families from 1265 exomes
Reproduced
not attempted (restricted data)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Total score +9

The in-silico annotation of Table 1 reproduces strongly on the paper's own 6-variant list (10 exact, 2 within-tol via gnomAD v2.1.1, CADD v1.6, REVEL, MutationTaster). Real deviations are confined to annotation-tool version/transcript drift — an honest SIFT flip for B3089 (Del→tolerated 0.098) and PolyPhen-2 0.36→0.092 (both still benign-range) — compounded by the paper not stating tool versions. The central discovery (1265 exomes→6 variants) and wet-lab work are not reproducible, but this is a data-availability/consent limit on the authors-vs-data axis, not a fabrication or computation defect. Overall a solid, explainable partial reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

348.9 k
tokens (I/O) · 28.1 M incl. cache
121 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.