Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A temporal classifier predicts histopathology state and parses acute-chronic phasing in inflammatory bowel disease patients.

Commun Biol · 2023
L1 83/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
83/100
Reproducibility score
0.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 61% of all assessed papers rank 430 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to verify 1:1 against the authors' published source data. We did NOT rerun the from-FASTQ pipeline (STAR/DESeq2/limma/Mfuzz/random-forest over 340 mouse samples) - the hard 20%, heavily parameterized, and the authors' code ships as a supplementary ZIP (the cited GitHub repo LosicLab/losiclab.github.io is just the lab's 2019 Jekyll website, NOT analysis code). Instead we recomputed the reported headline numbers from the authors' figshare numerical source data (doi:10.6084/m9.figshare.21706202.v2) on «our HPC». Result: 1:1 EXACT reproduction of the DESeq2 DE gene counts (2040/5740/2771), the 725-gene dynamic temporal signature, the Mfuzz cluster-size set, and the cluster sum; WITHIN-ROUNDING reproduction of the 87% DE-overlap (87.4%) and all V(D)J statistics (DSS 99.4%/median 497.5; AT 44.2%/median ~9). No fabrication detected in any checked value. PARTIAL on two items: (C2) differential-splicing counts are not derivable from the shipped gene-level DE table; (C7) the random-forest histology classifier - the shipped fit object contains only whole-colon Janssen models (OOB Spearman rho 0.63-0.84 across signatures, consistent with the paper's 'rho ~0.8' colon range), so the paper's exact region-specific values (proximal colon ~0.8, blood ~0.5) cannot be confirmed 1:1. All 3 GEO accessions (GSE214600/GSE186507/GSE193677) are now public despite the paper's stale 'data private until Oct 2024/2025' statement. NOT attempted: full pipeline rerun, human MSCCR Nancy/GHAS histology prediction, Bayesian causal-network overlap, wet-lab/qPCR validations.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 83
    assessed: 2026-06-15 ⛓ 4b9a875a97f8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can disease-specific temporal (acute vs chronic) gene expression and splicing signatures, rather than ordinary time-specific differential expression, be derived from murine colitis models and used via machine learning to predict histopathological state across tissues, models, and human IBD patients?

Core claims
  • The DSS phenotype-by-time interaction defines parsimonious temporal (dynamic) expression and splicing signatures of acute and chronic colitis distinct from time-specific differential expression. finding
  • Disease-specific expression and splicing signatures are largely orthogonal, affecting different genes. finding
  • A random forest classifier trained on temporal signatures predicts histopathology scores across lab sites, tissues (blood and colon), colitis models (DSS and AT), and an independent IBD patient cohort. method
  • Repeated DSS exposure cycles interspersed with recovery induce chronic colitis pathology. finding
  • Projecting predictive signatures onto human IBD causal networks identifies acute and chronic subnetworks marking acute-to-chronic transition points. method
  • DSS colitis induces overexpression of short, soluble TIR-lacking Il1rl1 isoform and a short LN-domain-lacking Lama3 isoform. mechanism
  • LAMA3 is differentially spliced in human IBD biopsies (CD and UC) whereas IL1RL1 is not. finding
  • Immunoglobulin-gene cluster (cluster B) progressively increases from day 5 to day 36, indicating growing B-cell involvement in chronic disease. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (gene expression, differential exon-usage/splicing, V(D)J/CDR3 assembly) mouse intestine/whole colon and whole blood, DSS-induced colitis model DSS (dextran sodium sulfate) cycling exposure vs control differential gene expression, differential splicing, adaptive immune repertoire magnitude/clonality
bulk RNA-seq mouse colon and blood, adoptive T-cell transfer (AT) colitis models (TC1, TC2) adoptive T-cell transfer gene expression / histology-predictive features
phenotypic endpoint measurement DSS colitis mice (6–24 per group) DSS cycling vs control body weight, colon length, histopathology scores
H&E histology mouse distal colon slices DSS vs control edema, inflammation, gland loss (×100 magnification)
qPCR isoform-specific validation mouse colon (same DSS and control samples) DSS vs control relative expression of TIR-containing vs TIR-lacking Il1rl1 and long vs short Lama3 transcripts (RQ, 36B4 endogenous control)
bulk RNA-seq differential splicing analysis intestinal biopsies (ileum, cecum, right colon, rectum) of IBD (CD/UC) patients none (IBD vs reference) differential splicing of IL1RL1 and LAMA3
random forest machine learning prediction DSS whole-colon training set; validation in DSS Janssen blood, MSSM blood, MSSM proximal/distal colon, AT colon/blood, IBD patient cohort none (computational) predicted histological scoring (Spearman correlation)
Key results
  • Mice in DSS disease group lost weight after cycles one and three; no weight loss after cycle two; weight gain in all DSS-free phases 20% loss after cycle 1; 10% loss after cycle 3
  • Differential gene expression rose from day 5 to peak at day 17, then declined at day 36 2040 (day5), 5740 (day17), 2771 (day36) DE genes
  • Differential splicing signal much weaker than expression, peaking at day 12 125 DS genes at day 12 (FDR<0.05)
  • Temporal expression and splicing trajectories are orthogonal, with only 2 genes showing both 2 genes overlap; 141 dynamic DS genes; 725 dynamic DE genes
  • Il1rl1 short TIR-lacking isoform overexpressed in DSS, increasing TIR-lacking/TIR-containing ratio ratio from 1:1 to 3.5:1
  • Lama3 short LN-domain-lacking isoform overexpressed with reduced long isoform in colitis short:long ratio from 1:1 to 3:1
  • Pathological lymphocyte aggregate counts correlate with VDJ measurements in first DSS phase spearman rho ~0.86, p~0.05
  • VDJ clones detected in >99% of DSS samples vs sparse detection in AT models median 500 clones/sample (DSS) vs ~8 clones in 44% of AT samples
Key statistics
  • count 2040, 5740, 2771 DE genes (days 5, 17, 36) (DSS vs control DE genes, logFC>0 FDR<0.05)
  • count 125 DS genes at day 12; 246 DS genes total (differential splicing genes FDR<0.05)
  • pvalue FDR 2.39e-18 (early DSS-specific Lama3 isoform differential splicing)
  • pvalue FDR 1.4e-3 (early DSS-specific Il1rl1 isoform differential splicing)
  • correlation spearman rho ~0.86, p~0.05 (lymphocyte aggregate counts vs VDJ measurements, first DSS phase)
  • fold_change 3.5:1 (Il1rl1 TIR-lacking:TIR-containing); 3:1 (Lama3 short:long) (qPCR isoform ratios in colitis vs 1:1 baseline)
  • count 725 dynamic DE genes; 141 dynamic DS genes; clusters of 113–194 genes (disease:time interaction signatures, FDR<0.05)
  • count median 500 clones/sample (DSS) vs ~8 clones (AT, 44% of samples) (VDJ clonal detection across models)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study characterizes the transcriptional landscape of acute and chronic murine colitis (DSS and adoptive transfer models) using RNAseq of blood and colon tissue at four timepoints (days 5, 12, 17, 36). A linear model framework identifies both fixed (timepoint-specific) and dynamic (disease × time interaction) differential expression and splicing signatures, with FDR-based multiple-testing correction applied throughout. Random forest classifiers trained with 10-fold cross-validation predict continuous histopathology scores from molecular features, with predictive performance evaluated by Spearman correlation. Resulting signatures are projected onto human IBD causal networks to delineate acute and chronic disease subnetworks.

Replicationbiological Sample size6–24 mice per group stated in figure legend; no formal power calculation or sample-size justification mentioned in the available text GroupsDSS disease vs. control at days 5, 12, 17, 36; adoptive transfer (AT) colitis vs. control (TC1 and TC2); cross-site (Janssen vs. MSSM); blood vs. colon tissue; independent IBD patient cohort Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFDR (specific algorithm, e.g., Benjamini-Hochberg, not explicitly named in the available text)
Statistical tests used
Test Applied to n Assumptions
Linear model with disease × time interaction term Dynamic differential gene expression signatures across four DSS timepoints; 725 genes identified at FDR < 0.05 6–24 mice per group (stated in Fig. 1 legend) not stated
Linear model (timepoint-specific disease vs. control) Fixed differential expression at each sacrifice day (days 5, 12, 17, 36); threshold logFC > 0 with FDR < 0.05 6–24 mice per group (stated in Fig. 1 legend) not stated
Wilcoxon rank-sum test Histopathology score comparisons across DSS cycles (Fig. 1c) 6–24 mice per group (stated in Fig. 1 legend) not stated
Spearman correlation VDJ measurements vs. pathological lymphocyte aggregate counts (rho ~0.86, p ~0.05); random forest model performance evaluation across validation sets null not stated
Pearson correlation (R > 0.8 and R > 0.7 thresholds) Isoform-specific co-expression analysis for Il1rl1 and Lama3 to construct gene lists for pathway enrichment null not stated
Random forest with 10-fold cross-validation Prediction of continuous histopathology scores from expression, splicing, and VDJ features; evaluated across DSS Janssen, DSS MSSM, and AT model validation sets null na
Approaches that could also have been used
  • Differential expression was identified with a linear model framework; the specific software tool is not named in the available text
    Could also: DESeq2 or edgeR, which use negative binomial models explicitly parameterized for the overdispersion of RNAseq read counts — Count-specific models account for the discrete and overdispersed nature of RNA-seq data and represent the current community standard; they also offer shrinkage estimators for log fold-change that stabilize estimates from small samples
  • Temporal expression clusters were derived using the Mfuzz soft-clustering algorithm applied to median disease trajectories
    Could also: Trajectory-aware methods such as ImpulseDE2 or splineTimeR, or dynamic time warping-based hierarchical clustering — Methods designed for ordered time-series data can capture non-monotone dynamics and provide model-based uncertainty estimates per cluster; comparing multiple clustering approaches can also assess robustness of the reported cluster assignments
  • Multiple separate random forest models were trained for different signatures and evaluated on multiple validation sets
    Could also: Elastic-net regularized regression (e.g., glmnet) or gradient-boosted trees (e.g., XGBoost), also with cross-validation — Penalized linear models provide explicit coefficient estimates that aid biological interpretation and allow inference on feature importance; comparing multiple learner types also permits assessment of whether findings are robust to model choice
  • Predictive model performance was summarized exclusively with Spearman correlation between predicted and observed histopathology scores
    Could also: Root mean squared error (RMSE) and mean absolute error (MAE) alongside rank correlation — Spearman rho captures rank agreement but not the absolute magnitude of prediction error; RMSE/MAE provide complementary information about practical predictive accuracy on the original histopathology score scale
  • Phenotype measurement dispersion (body weight, colon length, histopathology) was shown as ±1 SD with group sizes that varied from 6 to 24 mice
    Could also: 95% confidence intervals or SEM, particularly given the variable and sometimes small group sizes — CI or SEM convey uncertainty about the group mean rather than spread of individual observations, which can be informative when comparing timepoints with substantially different numbers of animals
  • Isoform co-expression gene lists for pathway enrichment were constructed using fixed Pearson R thresholds (> 0.8 and > 0.7)
    Could also: Weighted gene co-expression network analysis (WGCNA) or soft-thresholding approaches that use the full correlation distribution — Continuous weighting reduces sensitivity to the arbitrary threshold value and can improve stability of enrichment results, particularly when sample sizes are modest; it also allows integration of weaker co-expression signals that a hard cutoff would exclude
Software: Mfuzz (soft-clustering) · Random forest (specific package or implementation not stated)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
9
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

NM_010743.3 RefSeq in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
XM_006525689.3 RefSeq in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
XM_006525690.2 RefSeq in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36694043

Paper: Peters LA et al. A temporal classifier predicts histopathology state and parses acute-chronic phasing in inflammatory bowel disease patients. Commun Biol 2023. PMID 36694043 · PMCID PMC9873918 · DOI 10.1038/s42003-023-04469-y

Code / data availability (verbatim from the paper)

  • Code availability: "Code is provided in the supplementary data (Supplementary code 1). Open-source R code from publicly available packages that was exclusively used in this study is available here https://github.com/LosicLab/losiclab.github.io"
    • ⚠️ The GitHub URL the RU was seeded with (LosicLab/losiclab.github.io) is the lab's Jekyll website, last pushed 2019-01-15 (before the 2023 paper). It contains no analysis code for this paper (only generic open-source R packages, per the statement). The actual analysis code is Supplementary code 1 attached to the article. This is a no_own_repo situation — per brief rule P16 this does not down-rank the paper.
  • Data availability:
    • Mouse DSS/AT models (this paper's primary data): GEO GSE214600Public on Dec 31 2022, Series_pubmed_id = 36694043. ✅ available.
    • Human MSCCR blood: GSE186507Public on Sep 16 2022. ✅
    • Human MSCCR biopsy: GSE193677Public on Sep 16 2022. ✅
    • The paper text says "Data are private and will be released in October 2024 and 2025" — that embargo statement is out of date; all three series are public now (verified via GEO E-utilities 2026-06-15).
    • Numerical source data for graphs/charts: figshare DOI 10.6084/m9.figshare.21706202.v2 — public, 36 files (CSV / RData / txt). This is the per-figure source data.

Reproduction strategy (80/20)

The full pipeline (STAR 2.4.0g1 alignment → featureCounts → DESeq2/limma-voom DE+DS → Mfuzz temporal clustering → random-forest classifier) over 340 mouse + ~1000 human RNA-seq samples is the hard ~20%: heavy compute, many unspecified degrees of freedom (exact covariate set, CPM/FDR thresholds, batch correction), and the authors' own analysis code is a supplementary ZIP rather than a runnable repo. We do not attempt a full from-FASTQ rerun.

Instead — and equally valid per the brief — we perform an honest 1:1 verification of the paper's reported pipeline-derived numbers against the authors' own published numerical source data (figshare). This directly tests whether the headline numbers are actually derivable from the shipped data (a fabrication check), which is the project's goal. All download + computation runs on «our HPC»/«infra».

In scope (pipeline-derived, attempted)

# Reported result Pipeline Source-data file (figshare)
C1 DE gene counts per DSS timepoint: 2,040 (d5), 5,740 (d17), 2,771 (d36) DESeq2 (disease-time interaction) fig2/fig3 master DE+DS table
C2 Differentially spliced genes up to 125 (day 12) limma-voom master DE+DS table
C3 725 dynamic temporal-signature genes (87% also fixed DE) LRT disease×time cluster master table / DE table
C4 Mfuzz soft clusters A–E: 113–194 genes each; cluster B (late) enriched immunoglobulin Mfuzz fig2_cycling..._cluster_master_table
C5 DE/cycling signature list sizes (per-day & cluster human-symbol lists) DESeq2 + ortholog map fig3_*_human_symbols.txt
C6 V(D)J: detection >99% of DSS samples; median ~500 clones (AT: 44%, median 8) RNA-seq VDJ assembly s4_vdj_and_umi_summary_stats...csv
C7 Classifier predictive correlation rho ~0.8 (DSS proximal colon), ~0.5 (blood) random forest, 10-fold CV fig4_master_model_training_and_validation_list.RData

Out of scope (not pipeline / not attempted, with reason)

  • Wet-lab: body-weight/colon-length phenotyping, H&E histopathology scoring, qPCR isoform validation (IL1RL1 3.5:1, LAMA3 3:1 ratios) — bench measurements, not pipeline outputs.
  • Full from-FASTQ RNA-seq alignment & quantification rerun — hard 20%, heavy
Figures / tables: Fig 2aFig 2bFig 2cFig 4a
C1a
Reported
2040 DE genes DSS day5 (FDR<0.05)
Reproduced
2040
exact
C1b
Reported
5740 DE genes DSS day17
Reproduced
5740
exact
C1c
Reported
2771 DE genes DSS day36
Reproduced
2771
exact
C2
Reported
125 DS genes peak day12
Reproduced
not derivable from shipped gene-level table
partial
C3a
Reported
725 dynamic temporal-signature genes
Reproduced
725
exact
C3b
Reported
87% of 725 also DE in >=1 timepoint
Reproduced
87.4% (634/725)
within tolerance
C4set
Reported
Mfuzz cluster sizes {113,125,136,157,194}
Reproduced
{113,125,136,157,194} (A/B/D/E letters permuted)
within tolerance
C4sum
Reported
725 (sum of clusters)
Reproduced
725
exact
C6a
Reported
>99% DSS samples V(D)J detected
Reproduced
99.4% (178/179)
within tolerance
C6b
Reported
median 500 clones DSS
Reproduced
497.5
within tolerance
C6c
Reported
44% AT samples V(D)J detected
Reproduced
44.2% (99/224)
within tolerance
C6d
Reported
median 8 clones AT
Reproduced
9 (T1=5,T2=11)
within tolerance
C7a
Reported
RF sig D, DSS proximal colon rho~0.8
Reproduced
0.732 OOB (whole-colon Janssen; ols R2=0.779); proximal-colon model not in shipped object
partial
C7b
Reported
RF sig B, DSS blood rho~0.5
Reproduced
0.76 OOB (whole-colon Janssen); blood model not in shipped object
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 83/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

Strong, largely 1:1 reproduction: the headline DESeq2 DE counts (2040/5740/2771), the 725-gene dynamic temporal signature, the Mfuzz cluster-size set {113,125,136,157,194} and cluster sum all reproduce exactly from the authors' own figshare source data, and the V(D)J and 87%-overlap statistics match within rounding — no fabrication detected. The deviations that remain are on the authors'/data-availability side: the DS=125 splicing count and the region-specific random-forest rho (proximal colon ~0.8, blood ~0.5) are not derivable from the shipped data (which carries only gene-level DE and whole-colon Janssen models), and the cited code repo is the lab's website rather than analysis code. The central classifier conclusion is therefore qualitatively but not 1:1 confirmed (whole-colon rho 0.63–0.84 ≈ paper's ~0.8), making this a solid reproduction with explainable, deposit-driven gaps rather than a critical discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

186 k
tokens (I/O) · 15.7 M incl. cache
20 min
runtime · 0 CPU-h
0.6 GB
peak RAM
3
HPC jobs
hummel
machine