Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

A temporal classifier predicts histopathology state and parses acute-chronic phasing in inflammatory bowel disease patients.

Commun Biol · 2023
L1 83/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
83/100
Reproducibility score
0.5 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 61% of all assessed papers rank 431 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to verify 1:1 against the authors' published source data. We did NOT rerun the from-FASTQ pipeline (STAR/DESeq2/limma/Mfuzz/random-forest over 340 mouse samples) - the hard 20%, heavily parameterized, and the authors' code ships as a supplementary ZIP (the cited GitHub repo LosicLab/losiclab.github.io is just the lab's 2019 Jekyll website, NOT analysis code). Instead we recomputed the reported headline numbers from the authors' figshare numerical source data (doi:10.6084/m9.figshare.21706202.v2) on «our HPC». Result: 1:1 EXACT reproduction of the DESeq2 DE gene counts (2040/5740/2771), the 725-gene dynamic temporal signature, the Mfuzz cluster-size set, and the cluster sum; WITHIN-ROUNDING reproduction of the 87% DE-overlap (87.4%) and all V(D)J statistics (DSS 99.4%/median 497.5; AT 44.2%/median ~9). No fabrication detected in any checked value. PARTIAL on two items: (C2) differential-splicing counts are not derivable from the shipped gene-level DE table; (C7) the random-forest histology classifier - the shipped fit object contains only whole-colon Janssen models (OOB Spearman rho 0.63-0.84 across signatures, consistent with the paper's 'rho ~0.8' colon range), so the paper's exact region-specific values (proximal colon ~0.8, blood ~0.5) cannot be confirmed 1:1. All 3 GEO accessions (GSE214600/GSE186507/GSE193677) are now public despite the paper's stale 'data private until Oct 2024/2025' statement. NOT attempted: full pipeline rerun, human MSCCR Nancy/GHAS histology prediction, Bayesian causal-network overlap, wet-lab/qPCR validations.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 83
    assessed: 2026-06-15 ⛓ 4b9a875a97f8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether temporal (dynamic) gene expression and splicing signatures, distinct from ordinary fixed-timepoint differential expression, can parse acute versus chronic phases of colitis in murine models and predict histopathological outcomes in both mice and IBD patients.

Core claims
  • Disease-specific temporal (dynamic) gene expression and splicing signatures, distinct from fixed timepoint differential expression, can be derived from DSS and adoptive transfer colitis models to capture acute-chronic disease dynamics finding
  • Disease-specific differential expression and differential splicing signatures are largely orthogonal, affecting different genetic bodies finding
  • Machine learning models built from these temporal signatures predict histopathological measures across blood and intestinal tissue in murine colitis models and in an independent IBD patient cohort finding
  • DSS colitis induces overexpression of minor splice isoforms of Il1rl1 (TIR-domain-lacking) and Lama3 (LN-domain-lacking) finding
  • Sub-networks of patient-derived causal networks enriched in temporal signatures can distinguish acute and chronic disease components within the broader IBD molecular landscape finding
  • Repeated cycles of DSS exposure interspersed with recovery periods induce chronic disease pathology with incomplete colonic healing finding
  • RNAseq-derived VDJ/CDR3 read quantification can serve as a proxy for adaptive immune clonality and magnitude over time method
  • Random forest models combining expression, splicing, and VDJ features predict histological scoring across labs, tissues, and colitis model types method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (gene expression and differential exon usage) mouse colon (intestine) DSS differential gene expression and exon usage across timepoints (days 5,12,17,36)
bulk RNA-seq mouse whole blood DSS temporal gene expression and splicing signatures
qPCR isoform-specific validation mouse colon tissue DSS vs control relative expression (RQ) of TIR-containing vs TIR-lacking Il1rl1 transcripts
qPCR isoform-specific validation mouse colon tissue DSS vs control relative expression of long vs short Lama3 transcripts
histopathology scoring and H&E staining mouse distal colon DSS histopathology score, gland loss, edema, inflammation
VDJ alignment / CDR3 de novo assembly from RNAseq mouse blood and colon (DSS and adoptive transfer models) DSS or adoptive T-cell transfer immune clonotype counts and clonality
bulk RNA-seq differential splicing analysis IBD patient intestinal biopsies (ileum, cecum, right colon, rectum) CD/UC vs control differential splicing of IL1RL1 and LAMA3
random forest machine learning modeling mouse blood/colon (DSS Janssen, DSS MSSM, adoptive transfer TC1/TC2) and IBD patient cohort DSS/AT/none predicted histological score from expression, splicing, VDJ features
Key results
  • DSS mice lost 20% and 10% body weight after DSS cycles one and three respectively 20%; 10%
  • Differentially expressed genes (FDR<0.05) increased from day 5 to day 17 then decreased by day 36 2040 (day5), 5740 (day17), 2771 (day36) genes
  • Differential splicing (DS) signal peaked at day 12 125 genes at FDR<0.05
  • 725 genes showed disease-specific dynamic (temporal) expression signature 725 genes, FDR<0.05
  • 141 genes showed disease-specific dynamic splicing signature, largely non-overlapping with dynamic expression genes 141 genes, FDR<0.05; only 2 genes overlap with expression
  • Il1rl1 TIR-lacking/TIR-containing transcript ratio increased in DSS colitis, confirmed by qPCR ratio 1:1 to 3.5:1
  • Lama3 short (LN-domain-lacking)/long isoform ratio increased in DSS colitis, confirmed by qPCR ratio 1:1 to 3:1
  • VDJ read-based measurements correlated with pathological lymphocyte aggregate counts in first DSS phase spearman rho ~0.86, p~0.05
Key statistics
  • count 2040 DE genes (day5), 5740 DE genes (day17), 2771 DE genes (day36) (timepoint-specific DSS vs control differential expression, FDR<0.05)
  • count 125 DS genes at day 12 (highest number of significant differential splicing genes, FDR<0.05)
  • count 725 genes (disease-specific temporal (dynamic) expression signature, FDR<0.05)
  • count 141 genes (disease-specific temporal (dynamic) splicing signature, FDR<0.05)
  • pvalue FDR=2.39e-18 (Lama3); FDR=1.4e-3 (Il1rl1) (early DSS-specific differential splicing at day 5)
  • fold_change Il1rl1 TIR-lacking:TIR-containing ratio 1:1 to 3.5:1 (qPCR validation, colon tissue)
  • fold_change Lama3 short:long isoform ratio 1:1 to 3:1 (qPCR validation, colon tissue)
  • correlation spearman rho ~0.86, p val ~0.05 (VDJ measurements vs pathological lymphocyte aggregate counts, first DSS phase)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study characterizes the transcriptional landscape of acute and chronic murine colitis (DSS and adoptive transfer models) using RNAseq of blood and colon tissue at four timepoints (days 5, 12, 17, 36). A linear model framework identifies both fixed (timepoint-specific) and dynamic (disease × time interaction) differential expression and splicing signatures, with FDR-based multiple-testing correction applied throughout. Random forest classifiers trained with 10-fold cross-validation predict continuous histopathology scores from molecular features, with predictive performance evaluated by Spearman correlation. Resulting signatures are projected onto human IBD causal networks to delineate acute and chronic disease subnetworks.

Replicationbiological Sample size6–24 mice per group stated in figure legend; no formal power calculation or sample-size justification mentioned in the available text GroupsDSS disease vs. control at days 5, 12, 17, 36; adoptive transfer (AT) colitis vs. control (TC1 and TC2); cross-site (Janssen vs. MSSM); blood vs. colon tissue; independent IBD patient cohort Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFDR (specific algorithm, e.g., Benjamini-Hochberg, not explicitly named in the available text)
Statistical tests used
Test Applied to n Assumptions
Linear model with disease × time interaction term Dynamic differential gene expression signatures across four DSS timepoints; 725 genes identified at FDR < 0.05 6–24 mice per group (stated in Fig. 1 legend) not stated
Linear model (timepoint-specific disease vs. control) Fixed differential expression at each sacrifice day (days 5, 12, 17, 36); threshold logFC > 0 with FDR < 0.05 6–24 mice per group (stated in Fig. 1 legend) not stated
Wilcoxon rank-sum test Histopathology score comparisons across DSS cycles (Fig. 1c) 6–24 mice per group (stated in Fig. 1 legend) not stated
Spearman correlation VDJ measurements vs. pathological lymphocyte aggregate counts (rho ~0.86, p ~0.05); random forest model performance evaluation across validation sets null not stated
Pearson correlation (R > 0.8 and R > 0.7 thresholds) Isoform-specific co-expression analysis for Il1rl1 and Lama3 to construct gene lists for pathway enrichment null not stated
Random forest with 10-fold cross-validation Prediction of continuous histopathology scores from expression, splicing, and VDJ features; evaluated across DSS Janssen, DSS MSSM, and AT model validation sets null na
Approaches that could also have been used
  • Differential expression was identified with a linear model framework; the specific software tool is not named in the available text
    Could also: DESeq2 or edgeR, which use negative binomial models explicitly parameterized for the overdispersion of RNAseq read counts — Count-specific models account for the discrete and overdispersed nature of RNA-seq data and represent the current community standard; they also offer shrinkage estimators for log fold-change that stabilize estimates from small samples
  • Temporal expression clusters were derived using the Mfuzz soft-clustering algorithm applied to median disease trajectories
    Could also: Trajectory-aware methods such as ImpulseDE2 or splineTimeR, or dynamic time warping-based hierarchical clustering — Methods designed for ordered time-series data can capture non-monotone dynamics and provide model-based uncertainty estimates per cluster; comparing multiple clustering approaches can also assess robustness of the reported cluster assignments
  • Multiple separate random forest models were trained for different signatures and evaluated on multiple validation sets
    Could also: Elastic-net regularized regression (e.g., glmnet) or gradient-boosted trees (e.g., XGBoost), also with cross-validation — Penalized linear models provide explicit coefficient estimates that aid biological interpretation and allow inference on feature importance; comparing multiple learner types also permits assessment of whether findings are robust to model choice
  • Predictive model performance was summarized exclusively with Spearman correlation between predicted and observed histopathology scores
    Could also: Root mean squared error (RMSE) and mean absolute error (MAE) alongside rank correlation — Spearman rho captures rank agreement but not the absolute magnitude of prediction error; RMSE/MAE provide complementary information about practical predictive accuracy on the original histopathology score scale
  • Phenotype measurement dispersion (body weight, colon length, histopathology) was shown as ±1 SD with group sizes that varied from 6 to 24 mice
    Could also: 95% confidence intervals or SEM, particularly given the variable and sometimes small group sizes — CI or SEM convey uncertainty about the group mean rather than spread of individual observations, which can be informative when comparing timepoints with substantially different numbers of animals
  • Isoform co-expression gene lists for pathway enrichment were constructed using fixed Pearson R thresholds (> 0.8 and > 0.7)
    Could also: Weighted gene co-expression network analysis (WGCNA) or soft-thresholding approaches that use the full correlation distribution — Continuous weighting reduces sensitivity to the arbitrary threshold value and can improve stability of enrichment results, particularly when sample sizes are modest; it also allows integration of weaker co-expression signals that a hard cutoff would exclude
Software: Mfuzz (soft-clustering) · Random forest (specific package or implementation not stated)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
9
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

NM_010743.3 RefSeq in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
XM_006525689.3 RefSeq in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
XM_006525690.2 RefSeq in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36694043

Paper: Peters LA et al. A temporal classifier predicts histopathology state and parses acute-chronic phasing in inflammatory bowel disease patients. Commun Biol 2023. PMID 36694043 · PMCID PMC9873918 · DOI 10.1038/s42003-023-04469-y

Code / data availability (verbatim from the paper)

  • Code availability: "Code is provided in the supplementary data (Supplementary code 1). Open-source R code from publicly available packages that was exclusively used in this study is available here https://github.com/LosicLab/losiclab.github.io"
    • ⚠️ The GitHub URL the RU was seeded with (LosicLab/losiclab.github.io) is the lab's Jekyll website, last pushed 2019-01-15 (before the 2023 paper). It contains no analysis code for this paper (only generic open-source R packages, per the statement). The actual analysis code is Supplementary code 1 attached to the article. This is a no_own_repo situation — per brief rule P16 this does not down-rank the paper.
  • Data availability:
    • Mouse DSS/AT models (this paper's primary data): GEO GSE214600Public on Dec 31 2022, Series_pubmed_id = 36694043. ✅ available.
    • Human MSCCR blood: GSE186507Public on Sep 16 2022. ✅
    • Human MSCCR biopsy: GSE193677Public on Sep 16 2022. ✅
    • The paper text says "Data are private and will be released in October 2024 and 2025" — that embargo statement is out of date; all three series are public now (verified via GEO E-utilities 2026-06-15).
    • Numerical source data for graphs/charts: figshare DOI 10.6084/m9.figshare.21706202.v2 — public, 36 files (CSV / RData / txt). This is the per-figure source data.

Reproduction strategy (80/20)

The full pipeline (STAR 2.4.0g1 alignment → featureCounts → DESeq2/limma-voom DE+DS → Mfuzz temporal clustering → random-forest classifier) over 340 mouse + ~1000 human RNA-seq samples is the hard ~20%: heavy compute, many unspecified degrees of freedom (exact covariate set, CPM/FDR thresholds, batch correction), and the authors' own analysis code is a supplementary ZIP rather than a runnable repo. We do not attempt a full from-FASTQ rerun.

Instead — and equally valid per the brief — we perform an honest 1:1 verification of the paper's reported pipeline-derived numbers against the authors' own published numerical source data (figshare). This directly tests whether the headline numbers are actually derivable from the shipped data (a fabrication check), which is the project's goal. All download + computation runs on «our HPC»/«infra».

In scope (pipeline-derived, attempted)

# Reported result Pipeline Source-data file (figshare)
C1 DE gene counts per DSS timepoint: 2,040 (d5), 5,740 (d17), 2,771 (d36) DESeq2 (disease-time interaction) fig2/fig3 master DE+DS table
C2 Differentially spliced genes up to 125 (day 12) limma-voom master DE+DS table
C3 725 dynamic temporal-signature genes (87% also fixed DE) LRT disease×time cluster master table / DE table
C4 Mfuzz soft clusters A–E: 113–194 genes each; cluster B (late) enriched immunoglobulin Mfuzz fig2_cycling..._cluster_master_table
C5 DE/cycling signature list sizes (per-day & cluster human-symbol lists) DESeq2 + ortholog map fig3_*_human_symbols.txt
C6 V(D)J: detection >99% of DSS samples; median ~500 clones (AT: 44%, median 8) RNA-seq VDJ assembly s4_vdj_and_umi_summary_stats...csv
C7 Classifier predictive correlation rho ~0.8 (DSS proximal colon), ~0.5 (blood) random forest, 10-fold CV fig4_master_model_training_and_validation_list.RData

Out of scope (not pipeline / not attempted, with reason)

  • Wet-lab: body-weight/colon-length phenotyping, H&E histopathology scoring, qPCR isoform validation (IL1RL1 3.5:1, LAMA3 3:1 ratios) — bench measurements, not pipeline outputs.
  • Full from-FASTQ RNA-seq alignment & quantification rerun — hard 20%, heavy
Figures / tables: Fig 2aFig 2bFig 2cFig 4a
C1a
Reported
2040 DE genes DSS day5 (FDR<0.05)
Reproduced
2040
exact
C1b
Reported
5740 DE genes DSS day17
Reproduced
5740
exact
C1c
Reported
2771 DE genes DSS day36
Reproduced
2771
exact
C2
Reported
125 DS genes peak day12
Reproduced
not derivable from shipped gene-level table
partial
C3a
Reported
725 dynamic temporal-signature genes
Reproduced
725
exact
C3b
Reported
87% of 725 also DE in >=1 timepoint
Reproduced
87.4% (634/725)
within tolerance
C4set
Reported
Mfuzz cluster sizes {113,125,136,157,194}
Reproduced
{113,125,136,157,194} (A/B/D/E letters permuted)
within tolerance
C4sum
Reported
725 (sum of clusters)
Reproduced
725
exact
C6a
Reported
>99% DSS samples V(D)J detected
Reproduced
99.4% (178/179)
within tolerance
C6b
Reported
median 500 clones DSS
Reproduced
497.5
within tolerance
C6c
Reported
44% AT samples V(D)J detected
Reproduced
44.2% (99/224)
within tolerance
C6d
Reported
median 8 clones AT
Reproduced
9 (T1=5,T2=11)
within tolerance
C7a
Reported
RF sig D, DSS proximal colon rho~0.8
Reproduced
0.732 OOB (whole-colon Janssen; ols R2=0.779); proximal-colon model not in shipped object
partial
C7b
Reported
RF sig B, DSS blood rho~0.5
Reproduced
0.76 OOB (whole-colon Janssen); blood model not in shipped object
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 83/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

Strong, largely 1:1 reproduction: the headline DESeq2 DE counts (2040/5740/2771), the 725-gene dynamic temporal signature, the Mfuzz cluster-size set {113,125,136,157,194} and cluster sum all reproduce exactly from the authors' own figshare source data, and the V(D)J and 87%-overlap statistics match within rounding — no fabrication detected. The deviations that remain are on the authors'/data-availability side: the DS=125 splicing count and the region-specific random-forest rho (proximal colon ~0.8, blood ~0.5) are not derivable from the shipped data (which carries only gene-level DE and whole-colon Janssen models), and the cited code repo is the lab's website rather than analysis code. The central classifier conclusion is therefore qualitatively but not 1:1 confirmed (whole-colon rho 0.63–0.84 ≈ paper's ~0.8), making this a solid reproduction with explainable, deposit-driven gaps rather than a critical discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

186 k
tokens (I/O) · 15.7 M incl. cache
20 min
runtime · 0 CPU-h
0.6 GB
peak RAM
3
HPC jobs
hummel
machine