Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.

Front Immunol · 2022
L1 90/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1, exemplary). Authors' own repo (commit 1f8ebb6) on fully-public GEO data. B-HOT+ random-forest classifier reproduced end-to-end on «our HPC» («job»): GSE98320 nested 10x3 CV and independent GSE129166 validation. CV metrics match the paper essentially exactly (accuracy 0.921->0.919; AUC NR/ABMR/TCMR 0.980/0.976/0.994 -> 0.979/0.977/0.994, all |diff|<=0.001; F1 NR/ABMR/TCMR 0.949/0.874/0.828 -> 0.944/0.872/0.846). Independent validation also matches (accuracy 0.883->0.870; AUC ABMR 0.982->0.984; AUC NR 0.965->0.959; F1 0.938->0.929, 0.742->0.722) - all within ~0.02. Preprocessing label splits reproduce the paper EXACTLY: GSE98320 774 NR/81 TCMR/326 ABMR, GSE129166 60 NR/15 ABMR/2 TCMR. Two documented faithful deviations (neither changes the conclusion): (1) probe->Entrez annotation sourced from PUBLIC GEO platform tables GPL15207/GPL570 (ENTREZ_GENE_ID, ' /// ' multimap) instead of the login-gated Affymetrix NetAffx na36 CSVs the repo assumes - explains the sub-0.02 metric shifts; (2) recent R/Bioconductor (tidyverse+sva+GEOquery) instead of pinned R 4.1.2 since the 2022 Bioconductor build no longer solves on bioconda. Python env matched README pins exactly (numpy1.20.3/pandas1.3.4/sklearn0.24.2/matplotlib3.4.3/joblib1.1.0). One real-world snag handled: GEO RE-DEPOSITED GSE98320, renaming the archetype characteristic to 'archetype cluster: N' - identical clustering (reproduces the exact 774/81/326 split). NOT attempted: C14 recursive feature-selection model (optional hard ~20%) and MMDx/histological labels (wet-lab/proprietary, out of scope). No fabrication signal: reported metrics are the deterministic output of the shipped pipeline on public data, which is what we ran.

💻 Code ↗ 🗄 Data: GSE98320

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-23
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

A classifier for kidney transplant rejection built using only the genes of the standardized, decentralized Banff-Human Organ Transplant (B-HOT) panel can perform as well as one using genome-wide feature selection, enabling a multi-platform-compatible molecular diagnostic tool for rejection classification.

Core claims
  • A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR. finding
  • Adding 6 highly predictive non-B-HOT genes (CST7, KLRC4-KLRK1, TRBC1, TRBV6-5, TRBV19, ZFX) to the B-HOT panel (B-HOT+ Model) outperforms both the pure B-HOT Model and the genome-wide Feature Selection Model. finding
  • The B-HOT+ Model generalizes to an independent validation dataset from a different microarray platform after batch correction. finding
  • ComBat batch-effect correction successfully removes platform-driven variation between two microarray datasets, leaving biologically meaningful (diagnosis-driven) variation. method
  • The authors propose adding the 6 identified genes to the B-HOT panel to optimize the commercially available panel. resource
  • A nested cross-validation (10 outer, 3 inner folds) scheme with grid search and sequential forward feature selection was used to develop and tune random forest classifiers. method
Experimental setups
Assay System Perturbation Readout Platform
microarray gene expression profiling (Affymetrix hgu219 PrimeView) kidney transplant biopsies, human (GSE98320, 1,181 samples post-filtering) none (diagnostic classification by Banff/MMDx category) gene expression levels used to train random forest rejection classifiers (NR/ABMR/TCMR) Affymetrix hgu219 PrimeView microarray
microarray gene expression profiling (Affymetrix GeneChip Human Genome U133 Plus 2.0) kidney transplant biopsies, human (GSE129166, 77 samples post-filtering, blood samples excluded) none (independent external validation) predicted rejection class probabilities (NR/ABMR/TCMR) validated against histological diagnosis Affymetrix GeneChip Human Genome U133 Plus 2.0
machine learning classification (random forest, nested cross-validation) gene expression matrix filtered to B-HOT panel genes (762 genes) none accuracy, precision, recall, AUC, F1-score per class scikit-learn, Python 3.9
sequential forward feature selection (k-nearest neighbors wrapper) gene expression matrix, all 18,945 overlapping genes none 100 most predictive genes selected; downstream random forest performance metrics scikit-learn, Python 3.9
principal component analysis combined GSE98320/GSE129166 gene expression data none (pre/post ComBat batch correction comparison) variance explained by dataset origin vs. diagnosis label along PC1/PC2 PCAtools (R)
batch effect correction combined microarray gene expression datasets none removal of platform-driven variance ComBat package (R)
Key results
  • B-HOT+ Model achieved the highest mean cross-validation accuracy among the three models 92.1%
  • B-HOT+ Model AUC on external validation set for NR and ABMR 0.965 (NR), 0.982 (ABMR)
  • B-HOT Model cross-validation average accuracy 91.3%
  • Feature Selection Model cross-validation average accuracy (lower than B-HOT-based models) 90.9%
  • Six genes (CST7, KLRC4-KLRK1, TRBC1, TRBV6-5, TRBV19, ZFX) added to B-HOT panel improved model performance
  • PCA showed variation dominated by dataset origin before ComBat, and by class label after ComBat
  • 18,945 genes overlapped between the two microarray platforms after probe matching/aggregation 18,945 genes
  • B-HOT Model AUCs across classes during cross-validation 0.980 (NR), 0.976 (ABMR), 0.995 (TCMR)
Key statistics
  • mean 92.1% average accuracy (B-HOT+ Model nested cross-validation accuracy)
  • other AUC 0.965 (B-HOT+ Model external validation AUC for NR class)
  • other AUC 0.982 (B-HOT+ Model external validation AUC for ABMR class)
  • mean 91.3% average accuracy (B-HOT Model cross-validation accuracy)
  • mean 90.9% average accuracy (Feature Selection Model cross-validation accuracy)
  • count 762 genes for 1,181 samples (final B-HOT training feature set/sample size)
  • count 18,945 genes (genes overlapping between GSE98320 and GSE129166 after preprocessing)
  • other Cohen's kappa 0.2-0.4 (interobserver disagreement in histological Banff diagnosis (background/motivation, not a study result))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study describes machine-learning classifier development rather than a classical hypothesis-testing analysis: gene expression data from two public microarray datasets were preprocessed (batch correction with ComBat, robust scaling, PCA for visualization), and three random forest classifiers (B-HOT, Feature Selection, B-HOT+) were trained and compared using nested cross-validation (10 outer / 3 inner folds) with hyperparameters tuned by grid search. Model performance was reported as accuracy, precision, recall, F1-score, and AUC per class during cross-validation, and the best-performing model was then evaluated on an independent external validation dataset using the same metrics.

Replicationbiological Sample sizeSample sizes given as biopsy/patient counts per dataset and per diagnostic class (Table 1, Table 2); no formal power calculation described GroupsNon-rejection (NR) vs antibody-mediated rejection (ABMR) vs T-cell-mediated rejection (TCMR) biopsy classes, across two independent microarray datasets Pairingna Randomization/blindingnot stated Dispersionrange Exact p-valuesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Nested cross-validation (10 outer folds, 3 inner folds) with grid-search hyperparameter tuning Evaluation and comparison of the B-HOT, Feature Selection, and B-HOT+ random forest models 1,181 training biopsies (GSE98320) not stated
Sequential forward feature selection using a k-nearest-neighbors (k=3) classifier Selection of top 100 predictive genes for the Feature Selection Model 1,181 training biopsies not stated
External validation on an independent dataset (accuracy, precision, recall, AUC, F1) Best-performing model (B-HOT+) tested on GSE129166 77 validation biopsies not stated
Principal component analysis Visualizing batch effects/class separation before and after ComBat correction 18,945 overlapping genes across all samples not stated
Approaches that could also have been used
  • Model performance (accuracy, AUC, F1, etc.) is reported as single point estimates from nested cross-validation for each model.
    Could also: Bootstrap resampling or repeated nested cross-validation could also be used to generate confidence intervals around these performance metrics. — This would convey the precision/variability of the performance estimates in addition to their central values.
  • The three candidate models (B-HOT, Feature Selection, B-HOT+) are compared by their average cross-validation accuracy and other point-estimate metrics to select the best-performing one.
    Could also: A paired statistical comparison across cross-validation folds (e.g., a paired t-test, Wilcoxon signed-rank test, or DeLong's test for AUC) could also be applied. — This would provide a formal statistical basis for concluding whether observed performance differences between models exceed what might be expected from fold-to-fold variability.
  • TCMR is a notably smaller class than NR and ABMR in both datasets (e.g., 81/1,181 in training; 2/77 in validation).
    Could also: Metrics such as balanced accuracy or Matthews correlation coefficient (MCC) could also be reported alongside AUC/F1. — These metrics are often used as complementary summaries for imbalanced multiclass classification tasks.
  • Feature selection for the second model used a wrapper approach (sequential forward selection with a KNN classifier).
    Could also: Embedded regularization-based feature selection methods, such as LASSO- or elastic-net-penalized logistic regression, could also be used. — These methods provide a regularization path and can offer an alternative, computationally efficient way to rank or select predictive genes.
  • Batch effect removal (ComBat) was assessed qualitatively via PCA biplots before and after correction.
    Could also: Quantitative batch-mixing metrics (e.g., kBET or LISI) could also be used alongside the PCA visualization. — Quantitative metrics can complement visual assessment by providing a numerical measure of how well batches are mixed after correction.
  • Demographic variables such as age are summarized with the mean and range.
    Could also: The median and interquartile range (IQR) could also be reported for these variables. — Medians and IQRs are often preferred for variables with skewed distributions or wide ranges, as they are less influenced by extreme values.
Software: R 4.1 · ComBat (batch effect correction) · quantable (robustscale function) · PCAtools · Python 3.9 · scikit-learn

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35619722

Paper: van Baardwijk et al. 2022, Front Immunol — "A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel." DOI 10.3389/fimmu.2022.841519.

Code: https://github.com/ErasmusMC-Bioinformatics/KidneyRejectionClassifier (authors' own repo). preprocessing.R + banff_randomforest.py. Ships Data/BHOT_entrez_mapping.csv, Data/BHOT_plus_entrez_mapping.csv, and two pre-trained joblib models.

Data (public, GEO):

  • GSE98320 — training. 1,208 biopsies / 1,045 patients, Affymetrix PrimeView (hgu219). After filtering → 1,181 samples (774 NR, 326 ABMR, 81 TCMR). Labels from MMDx archetypes.
  • GSE129166 — independent validation. 117 blood + 95 biopsy on HG-U133 Plus 2. After filtering to biopsies w/o mixed/borderline → 77 (60 NR, 15 ABMR, 2 TCMR). Labels histological.

In scope (pipeline-derived computational results)

The whole paper IS a bioinformatic pipeline: GEO download → probe→Entrez mapping (Affymetrix na36 annotation) → median aggregation → ComBat batch correction + robust scaling → random-forest classifier with nested CV (10 outer / 3 inner, stratified, random_state=1) → 3-class prediction (NR / ABMR / TCMR).

Primary reproduction target = the B-HOT+ model (best performer, 768 genes), which is exactly what the README run-command produces (--features bhot+):

  1. Cross-validation performance on GSE98320 (Table — B-HOT+): overall accuracy, per-class AUC / precision / recall / F1.
  2. Independent validation performance on GSE129166 (Table — B-HOT+): accuracy, per-class AUC / precision / recall / F1.

Secondary (if cheap): B-HOT (762 genes) and feature-selection model CV numbers — same script, different --features flag. Will report if the bhot+ run succeeds.

Out of scope / not attempted

  • Wet-lab / histological Banff scoring (input labels, not computed here).
  • MMDx archetype assignment for GSE98320 labels (proprietary MMDx system; we use the labels as shipped in the GEO metadata / repo phenotype step).
  • Any figure that is purely descriptive (cohort tables, etc.).

Determinism note

All seeds are fixed (random_state=1 for outer CV, inner CV, and the RandomForestClassifier). With identical preprocessed inputs the script is deterministic up to scikit-learn/numpy version differences → near-exact reproduction expected if preprocessing matches.

Key risk

preprocessing.R requires Affymetrix PrimeView.na36.annot.csv and HG-U133_Plus_2.na36.annot.csv (NetAffx) which are NOT in the repo and may need a Thermo Fisher login. Mitigations: (a) download from a public mirror inside the compute job; (b) fall back to Bioconductor primeview.db (GPL15207) / hgu133plus2.db (GPL570) to build an equivalent probe→Entrez CSV. If neither works → partial reproduction or docs_insufficient/env_unresolvable note.

C1
Reported
0.921
Reproduced
0.9187
exact
C2
Reported
0.980
Reproduced
0.9793
exact
C3
Reported
0.976
Reproduced
0.9767
exact
C4
Reported
0.994
Reproduced
0.9945
exact
C5
Reported
0.949
Reproduced
0.9442
exact
C6
Reported
0.874
Reproduced
0.8717
exact
C7
Reported
0.828
Reproduced
0.846
within tolerance
C8
Reported
0.883
Reproduced
0.8701
within tolerance
C9
Reported
0.965
Reproduced
0.9591
within tolerance
C10
Reported
0.982
Reproduced
0.9843
exact
C11
Reported
0.742
Reproduced
0.7222
within tolerance
C12
Reported
0.938
Reproduced
0.9286
within tolerance
C13
Reported
0.913
Reproduced
running (bhot model)
m.public.grade.pending
C15
Reported
768 (762 BHOT + 6)
Reproduced
shipped BHOT_plus_entrez_mapping = 772 distinct Entrez; model features = intersection with combat-scaled genes
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4

Exemplary 1:1 reproduction. The authors' own repo (commit 1f8ebb6) was run on fully-public GEO data (GSE98320, GSE129166), reproducing the label splits exactly (774/81/326; 60/15/2) and all 12 graded B-HOT+ metrics within ~0.02 — 6 of 7 CV metrics exact. The only deviations (max ~0.0198, e.g. F1 TCMR 0.828->0.846) sit on the input/preprocessing side and are attributable to our methodology choice of substituting the public probe->Entrez annotation for the login-gated NetAffx na36 files, not to any authors' defect. The central classifier claim holds fully and there is no fabrication signal; values are the deterministic output of the shipped pipeline.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

394.3 k
tokens (I/O) · 26.4 M incl. cache
165 min
runtime · 1.69 CPU-h
9.2 GB
peak RAM
1
HPC jobs
hummel
machine