Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification of common genetic characteristics of rheumatoid arthritis and major depressive disorder by bioinformatics analysis and machine learning.

Front Immunol · 2023
L1 58/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
58/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 18% of all assessed papers rank 950 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the two headline pipeline outputs 1:1 even though the paper ships NO own code (registry 'code' link slowkow/ggrepel is a generic plotting library; reproduced per P16 with the named third-party tools DESeq2/limma/pROC on the paper's own GEO data). STRONG near-exact hits: (1) RA DEG counts on GSE169082 via DESeq2 = 2166/1444/722 vs reported 2177/1458/719 (~1% off); (2) MDD hub-gene ROC AUCs on GSE38206 = EAF1 0.9105, SDCBP 0.8951, RNF19B 0.9074 vs reported 0.911/0.895/0.907 (match to 3 decimals) - confirms the 3 hub genes EAF1/SDCBP/RNF19B and the MDD diagnostic analysis are genuine. DIVERGENCES, all explainable: MDD DEG counts (1600/1364 or 1040/626 vs 592/441) - GSE38206 has 36 samples (9+9 patients/controls x 0w/8w) and the paper's design (timepoints/paired/probe-collapse) is unspecified, so the exact count is not pinnable (same order of magnitude). RA per-gene AUCs (0.978/0.6122/0.942) are NOT derivable from GSE169082 (n=7 gives degenerate AUC=1.0; 0.6122+CI impossible at that n) - they require the large RA cohort GSE97476 (246+30), so the RA-AUC dataset attribution is ambiguous in the text (flagged, not fabrication). NOT attempted (80/20): WGCNA modules, the stochastic LASSO+RandomForest gene-selection lists (seeds unstated; but its endpoint - the 3-gene AUC - IS verified), PPI/MCODE counts, GO/KEGG terms, nomogram AUC, and GSE97476 RA ROC. No fabrication detected; the two near-exact matches argue the core pipeline is real.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 58
    assessed: 2026-06-14 ⛓ 7475b262ec8a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Rheumatoid arthritis (RA) and major depressive disorder (MDD) share overlapping symptoms and comorbidity, so the study tests whether RA and MDD have common genetic characteristics/pathogenesis that could serve as diagnostic biomarkers to distinguish the two conditions.

Core claims
  • EAF1, SDCBP and RNF19B are common genetic characteristics (hub genes) shared by RA and MDD finding
  • Monocyte infiltration is a shared immune mechanism connecting RA and MDD finding
  • A nomogram built from EAF1, SDCBP and RNF19B has high diagnostic value for both RA and MDD finding
  • WGCNA combined with DEG analysis identifies disease-associated gene modules that intersect between RA and MDD datasets method
  • LASSO regression and random forest machine learning were combined to filter candidate hub genes for RA and MDD method
  • The shared gene sets between RA and MDD are enriched in immune response and inflammation pathways mechanism
  • TIMER 2.0 database and Cibersort algorithm were used to correlate hub gene expression with immune cell infiltration method
Experimental setups
Assay System Perturbation Readout Platform
WGCNA (weighted gene co-expression network analysis) PBMC, RA dataset GSE97476 RA disease vs control gene co-expression modules correlated with RA trait GPL10904; R WGCNA package
WGCNA PBMC, MDD dataset GSE38206 MDD disease vs control gene co-expression modules correlated with MDD trait GPL13607; R WGCNA package
Differential gene expression (RNA-seq, DESeq2) PBMC, RA dataset GSE169082 RA disease vs control differentially expressed genes (|log2FC|>1, adj.p<0.05) GPL20795; DESeq2
Differential gene expression (microarray, limma) PBMC, MDD dataset GSE38206 MDD disease vs control differentially expressed genes (|log2FC|>0.4, adj.p<0.05) GPL13607; limma
GO/KEGG functional enrichment analysis intersected DEG/module gene sets (194, 135, 42 genes) none enriched biological pathways/ontologies ClusterProfiler R package
Protein-protein interaction network construction with MCODE clustering candidate genes from RA/MDD intersections none interacting node proteins/subnetworks STRING database v11.5; Cytoscape
Machine learning feature selection (LASSO regression, random forest) candidate RA and MDD genes (PBMC datasets) none ranked candidate hub genes R packages glmnet and randomForest
ROC curve analysis and nomogram construction PBMC, RA and MDD datasets disease vs control diagnostic AUC/95% CI for EAF1, SDCBP, RNF19B R packages pROC and rms
Key results
  • EAF1, SDCBP and RNF19B identified as the 3-gene intersection of RA (9 candidate genes) and MDD (7 candidate genes) from LASSO/RF machine learning
  • Combined 3-gene nomogram diagnostic AUC for RA and MDD RA AUC 0.994 (95%CI 0.986-1.000); MDD AUC 0.969 (95%CI 0.925-1.000)
  • Individual gene diagnostic AUC in RA EAF1 0.978 (0.961-0.995); SDCBP 0.612 (0.523-0.702); RNF19B 0.942 (0.909-0.975)
  • Individual gene diagnostic AUC in MDD EAF1 0.911 (0.819-1.000); SDCBP 0.895 (0.789-1.000); RNF19B 0.907 (0.804-1.000)
  • EAF1, SDCBP and RNF19B were significantly upregulated in both RA and MDD patients versus controls (Wilcoxon test)
  • Monocyte levels were higher in RA and MDD patients than in controls
  • DEGs identified in RA (GSE169082) and MDD (GSE38206) datasets RA: 2177 DEGs (1458 up, 719 down); MDD: 592 up, 441 down
  • Gene sets from WGCNA/DEG intersections were enriched for immune response and inflammation pathways
Key statistics
  • fold_change RA nomogram AUC 0.994 (95%CI 0.986-1.000) (combined 3-gene diagnostic model for RA)
  • fold_change MDD nomogram AUC 0.969 (95%CI 0.925-1.000) (combined 3-gene diagnostic model for MDD)
  • correlation lightcyan module r=0.60, p=2e-26 (WGCNA module-trait correlation with RA)
  • correlation blue module r=-0.69, p=8e-37 (WGCNA module-trait correlation with RA)
  • correlation turquoise module r=0.61, p=1e-04 (WGCNA module-trait correlation with MDD)
  • correlation pink module r=-0.84, p=8e-10 (WGCNA module-trait correlation with MDD)
  • count 2177 DEGs (1458 up, 719 down) (GSE169082 RA DEGs via DESeq2)
  • other RA patients have 47% higher risk of depression than controls (meta-analysis of 11 cohort studies cited in introduction)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study reanalyzed three public GEO PBMC gene-expression datasets (two RA, one MDD) using a multi-step pipeline: differential expression analysis (DESeq2 on GSE169082; limma on GSE38206), weighted gene co-expression network analysis (WGCNA on GSE97476 and GSE38206), PPI network construction with MCODE filtering, and two machine-learning algorithms (LASSO and Random Forest) whose intersecting outputs nominated three hub genes (EAF1, SDCBP, RNF19B). These genes were assessed by Wilcoxon group comparisons and ROC/AUC analysis with 95% CIs, and a nomogram was constructed. Immune cell infiltration was quantified by CIBERSORT and compared between groups via Wilcoxon tests, with linear-fit scatterplots relating hub gene expression to immune cell fractions.

Replicationbiological Sample sizeSample sizes stated per GEO dataset in Table 1 (GSE97476: 30 controls, 246 RA; GSE169082: 3 controls, 4 RA; GSE38206: 18 controls, 18 MDD); no formal power calculation described GroupsRA patients vs. healthy controls; MDD patients vs. healthy controls; shared hub gene identification across both disease comparisons Pairingunpaired Randomization/blindingnot stated Dispersionnone Effect sizesyes Confidence intervalsyes Multiplicity correctionFDR-adjusted p < 0.05 as implemented by DESeq2 and limma for DEGs and by ClusterProfiler for GO/KEGG enrichment; no adjustment stated for the family of Wilcoxon tests across 22 immune cell types
Statistical tests used
Test Applied to n Assumptions
DESeq2 (negative binomial GLM / Wald test) Differentially expressed genes in RA dataset GSE169082 7 (3 controls, 4 RA) not stated
limma linear model (moderated t-statistic) Differentially expressed genes in MDD dataset GSE38206 36 (18 controls, 18 MDD) not stated
WGCNA Pearson module-trait correlation Module-trait associations in GSE97476 (RA) and GSE38206 (MDD) GSE97476: 276 (30 controls, 246 RA); GSE38206: 36 not stated
GO and KEGG enrichment analysis (ClusterProfiler, adjusted p < 0.05) Functional annotation of gene sets at four filtering stages (194, 135, 42, and 17 genes) null na
LASSO regression (glmnet) Candidate hub gene selection for RA and MDD null not stated
Random Forest importance ranking (randomForest) Candidate hub gene selection and importance ranking for RA and MDD null not stated
Wilcoxon rank-sum test Comparison of 22 CIBERSORT immune cell proportions between patient and control groups in RA and MDD; expression comparison of three hub genes between groups RA (GSE97476): 276; MDD (GSE38206): 36 not stated
ROC / AUC analysis with 95% CI (pROC) Diagnostic value of EAF1, SDCBP, RNF19B individually and as a nomogram-combined set for RA and MDD RA: 276; MDD: 36 na
Approaches that could also have been used
  • Wilcoxon tests were applied to compare proportions of 22 immune cell types between patient and control groups in both RA and MDD without a stated adjustment for the resulting family of approximately 44 simultaneous comparisons
    Could also: Apply Benjamini-Hochberg FDR adjustment across all Wilcoxon tests within each disease comparison — BH-FDR adjustment would quantify the expected proportion of false discoveries among the significant immune cell findings, which is informative when many simultaneous tests are run on correlated outcomes such as CIBERSORT fractions that sum to one
  • Hub gene selection combined LASSO and Random Forest by taking the intersection of each method's top output with an ad hoc threshold (top-15/top-25 for RA; top-10/top-20 for MDD)
    Could also: Use stability selection (bootstrap-based selection probabilities) or elastic net regularization (alpha between 0 and 1 in glmnet) to produce formal, threshold-independent selection probabilities — Stability selection provides calibrated false-discovery control for variable selection and reduces sensitivity to the specific top-N cutoff, which is otherwise an arbitrary choice
  • DEG fold-change thresholds differed between datasets (|log2FC| > 1 for GSE169082; |log2FC| > 0.4 for GSE38206 and GSE97476)
    Could also: Apply a uniform log2FC threshold across all datasets, or combine the two RA datasets (GSE97476 and GSE169082) via a fixed-effects or random-effects meta-analysis of effect sizes before intersection with MDD DEGs — A consistent threshold or meta-analytic combination reduces the risk that the intersected gene list reflects dataset-specific cutoff choices, and pooling the RA datasets would leverage the larger sample size of GSE97476 (n=276) alongside GSE169082
  • The entire pipeline — from DEG and WGCNA discovery through machine-learning selection to ROC evaluation — was performed within the same datasets with no held-out partition or independent validation cohort
    Could also: Apply k-fold cross-validation within the larger RA dataset (GSE97476, n=276) or validate the three hub genes in an independent publicly available GEO dataset — Internal cross-validation or external replication would provide an estimate of generalization performance less susceptible to optimistic bias, which is a recognized concern when the same data inform both feature selection and performance evaluation
  • Diagnostic performance of the hub genes and nomogram was summarized exclusively by AUC
    Could also: Supplement AUC with calibration curves and decision curve analysis — AUC measures discrimination but not calibration; calibration curves show whether predicted probabilities correspond to observed event rates, and decision curve analysis evaluates net clinical benefit across a range of decision thresholds, together providing a more complete picture of diagnostic utility
  • Correlations between hub gene expression and immune cell fractions were visualized with linear-fit scatterplots via ggstatsplot without reporting quantified correlation coefficients or their uncertainty in the main text
    Could also: Report Spearman (or Pearson) correlation coefficients with 95% CIs and FDR-adjusted p-values alongside the scatterplots — Quantified correlation statistics and their uncertainty allow readers to assess the strength and reliability of the gene–immune cell associations independently of the visual scale of the scatterplot
Software: R 4.2.1 · DESeq2 · limma · WGCNA · ClusterProfiler · ggplot2 · glmnet · randomForest · pROC · rms · ggstatsplot · String database 11.5 · Cytoscape / MCODE · CIBERSORT · TIMER 2.0

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
3
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GPL13607 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GPL10904 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GPL20795 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE169082 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE38206 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE97476 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37415981

Paper: Jiang W, Wang X, Tao D, Zhao X. Identification of common genetic characteristics of rheumatoid arthritis and major depressive disorder by bioinformatics analysis and machine learning. Front Immunol 2023;14:1183115. PMID 37415981 · PMCID PMC10320004 · DOI 10.3389/fimmu.2023.1183115.

Code-availability reality (P16 note)

The registry "Code" link is github.com/slowkow/ggrepel — a generic R ggplot2 text-label library, NOT the authors' analysis code. The paper ships no own analysis repository. Per brief rule 2 / P16, this does not down-rank the paper: we reproduce by applying the standard third-party tools the paper names (DESeq2/limma, glmnet/randomForest, pROC) to the paper's own GEO data, following the parameters stated in Methods.

Datasets (all public GEO)

GEO role platform design in scope?
GSE169082 RA DEGs (THIS RU's assigned accession) GPL20795 HiSeq X Ten, RNA-seq raw counts 3 RA vs 4 control PBMC YES — primary
GSE38206 MDD DEGs GPL13607 Agilent, microarray 18 MDD vs 18 control PBMC YES
GSE97476 RA WGCNA GPL10904 246 RA vs 30 control partial (WGCNA)

Pipeline-derived results (candidate claims)

id result pipeline grade-ability
C1 RA DEGs from GSE169082: 2177 total (1458 up, 719 down), |log2FC|>1 & adjP<0.05 DESeq2 on raw counts deterministic — primary target
C2 MDD DEGs from GSE38206: 592 up, 441 down, |log2FC|>0.4 & adjP<0.05 limma on microarray deterministic
C3 Common up DEGs = 131, common down DEGs = 4 (intersection RA∩MDD) set intersection on C1∩C2 derivable (depends on C1,C2 + symbol mapping)
C4 Hub-gene ROC AUC (EAF1/SDCBP/RNF19B) in RA: 0.978 / 0.612 / 0.942 pROC on GSE169082 expr deterministic
C5 Hub-gene ROC AUC in MDD: 0.911 / 0.895 / 0.907 pROC on GSE38206 expr deterministic

IN SCOPE (attempted)

  • C1 RA DEG counts (GSE169082, DESeq2) — primary, cleanest 1:1.
  • C2 MDD DEG counts (GSE38206, limma).
  • C4/C5 ROC AUC of the 3 named hub genes EAF1, SDCBP, RNF19B in both datasets.
  • C3 up/down DEG intersection (if symbol harmonization is clean).

OUT OF SCOPE (80/20 — not attempted, why)

  • WGCNA module assignment (GSE97476, GSE38206): module colors/sizes (19 modules; lightcyan/salmon/purple; β=7/6) are sensitive to soft-threshold, cut height, and merge params not fully pinned → not a clean 1:1; skip per rule 3.
  • LASSO + RandomForest feature selection (9 RA / 7 MDD candidates → 3 common): RF is stochastic (seed unstated) and LASSO λ-selection (lambda.min vs 1se) unstated → the specific 9/7/3 gene lists are not deterministically reproducible. We DO test the endpoint (the 3 hub genes' AUC, C4/C5) which is the headline ML claim. The intermediate gene-count path is the optional last 20%.
  • PPI/MCODE counts (30 / 17 genes), GO/KEGG term lists, nomogram AUC (0.994/0.969): downstream of the above, depend on STRING version + multivariate model details → not attempted.
  • Wet-lab / manual steps: none (fully computational paper); n/a.

Compute plan

All heavy compute on «our HPC» (SLURM, conda-in-job). R env: r-base + bioconductor-deseq2/limma/geoquery + r-proc + r-edger. Data fetched inside the compute node (has internet) onto «infra»; only small result numbers returned.

Figures / tables: Fig4CFig4D
C1a
Reported
2177 RA DEGs (GSE169082)
Reproduced
2166
within tolerance
C1b
Reported
1458 RA up
Reproduced
1444
within tolerance
C1c
Reported
719 RA down
Reproduced
722
within tolerance
C5a
Reported
MDD AUC EAF1 0.911
Reproduced
0.9105
within tolerance
C5b
Reported
MDD AUC SDCBP 0.895
Reproduced
0.8951
within tolerance
C5c
Reported
MDD AUC RNF19B 0.907
Reproduced
0.9074
within tolerance
C2a
Reported
592 MDD up (GSE38206)
Reproduced
1600 (18v18) / 1040 (9v9 0w)
did not match
C2b
Reported
441 MDD down
Reproduced
1364 / 626
did not match
C3a
Reported
131 common up
Reproduced
166
partial
C3b
Reported
4 common down
Reproduced
8
partial
C4b
Reported
RA AUC SDCBP 0.6122 (95%CI 0.523-0.702)
Reproduced
1.0 on GSE169082 (n=7) - not comparable; value requires GSE97476
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 58/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Solid partial reproduction with explainable deviations. Two headline outputs reproduced near-exactly — RA DEG counts (2166/1444/722 vs 2177/1458/719, ~1%) and the MDD hub-gene ROC AUCs (EAF1/SDCBP/RNF19B matching 0.911/0.895/0.907 to three decimals) — which strongly argue the core EAF1/SDCBP/RNF19B pipeline is genuine and not fabricated. The divergences sit on the input side and on dataset attribution, not in the core statistics: MDD DEG counts (1600/1364 vs 592/441) reflect an unspecified GSE38206 design (timepoints/probe-collapse) that is partly our methodology and partly paper underspecification, and the RA per-gene AUCs (e.g. SDCBP 0.6122) are not derivable from the assigned GSE169082 (n=7 → degenerate AUC=1.0) and require the unstated GSE97476. Severity is moderate and the central diagnostic-gene conclusion holds in limited form (confirmed on the MDD side, unverified on the RA side), warranting an overall yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

135.5 k
tokens (I/O) · 8.2 M incl. cache
14 min
runtime · 0.02 CPU-h
1.6 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine