A temporal classifier predicts histopathology state and parses acute-chronic phasing in inflammatory bowel disease patients.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to verify 1:1 against the authors' published source data. We did NOT rerun the from-FASTQ pipeline (STAR/DESeq2/limma/Mfuzz/random-forest over 340 mouse samples) - the hard 20%, heavily parameterized, and the authors' code ships as a supplementary ZIP (the cited GitHub repo LosicLab/losiclab.github.io is just the lab's 2019 Jekyll website, NOT analysis code). Instead we recomputed the reported headline numbers from the authors' figshare numerical source data (doi:10.6084/m9.figshare.21706202.v2) on «our HPC». Result: 1:1 EXACT reproduction of the DESeq2 DE gene counts (2040/5740/2771), the 725-gene dynamic temporal signature, the Mfuzz cluster-size set, and the cluster sum; WITHIN-ROUNDING reproduction of the 87% DE-overlap (87.4%) and all V(D)J statistics (DSS 99.4%/median 497.5; AT 44.2%/median ~9). No fabrication detected in any checked value. PARTIAL on two items: (C2) differential-splicing counts are not derivable from the shipped gene-level DE table; (C7) the random-forest histology classifier - the shipped fit object contains only whole-colon Janssen models (OOB Spearman rho 0.63-0.84 across signatures, consistent with the paper's 'rho ~0.8' colon range), so the paper's exact region-specific values (proximal colon ~0.8, blood ~0.5) cannot be confirmed 1:1. All 3 GEO accessions (GSE214600/GSE186507/GSE193677) are now public despite the paper's stale 'data private until Oct 2024/2025' statement. NOT attempted: full pipeline rerun, human MSCCR Nancy/GHAS histology prediction, Bayesian causal-network overlap, wet-lab/qPCR validations.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 83assessed: 2026-06-15 ⛓ 4b9a875a97f8
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether temporal (dynamic) gene expression and splicing signatures, distinct from ordinary fixed-timepoint differential expression, can parse acute versus chronic phases of colitis in murine models and predict histopathological outcomes in both mice and IBD patients.
- ★ Disease-specific temporal (dynamic) gene expression and splicing signatures, distinct from fixed timepoint differential expression, can be derived from DSS and adoptive transfer colitis models to capture acute-chronic disease dynamics finding
- ★ Disease-specific differential expression and differential splicing signatures are largely orthogonal, affecting different genetic bodies finding
- ★ Machine learning models built from these temporal signatures predict histopathological measures across blood and intestinal tissue in murine colitis models and in an independent IBD patient cohort finding
- ★ DSS colitis induces overexpression of minor splice isoforms of Il1rl1 (TIR-domain-lacking) and Lama3 (LN-domain-lacking) finding
- ★ Sub-networks of patient-derived causal networks enriched in temporal signatures can distinguish acute and chronic disease components within the broader IBD molecular landscape finding
- Repeated cycles of DSS exposure interspersed with recovery periods induce chronic disease pathology with incomplete colonic healing finding
- RNAseq-derived VDJ/CDR3 read quantification can serve as a proxy for adaptive immune clonality and magnitude over time method
- ★ Random forest models combining expression, splicing, and VDJ features predict histological scoring across labs, tissues, and colitis model types method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (gene expression and differential exon usage) | mouse colon (intestine) | DSS | differential gene expression and exon usage across timepoints (days 5,12,17,36) | — |
| bulk RNA-seq | mouse whole blood | DSS | temporal gene expression and splicing signatures | — |
| qPCR isoform-specific validation | mouse colon tissue | DSS vs control | relative expression (RQ) of TIR-containing vs TIR-lacking Il1rl1 transcripts | — |
| qPCR isoform-specific validation | mouse colon tissue | DSS vs control | relative expression of long vs short Lama3 transcripts | — |
| histopathology scoring and H&E staining | mouse distal colon | DSS | histopathology score, gland loss, edema, inflammation | — |
| VDJ alignment / CDR3 de novo assembly from RNAseq | mouse blood and colon (DSS and adoptive transfer models) | DSS or adoptive T-cell transfer | immune clonotype counts and clonality | — |
| bulk RNA-seq differential splicing analysis | IBD patient intestinal biopsies (ileum, cecum, right colon, rectum) | CD/UC vs control | differential splicing of IL1RL1 and LAMA3 | — |
| random forest machine learning modeling | mouse blood/colon (DSS Janssen, DSS MSSM, adoptive transfer TC1/TC2) and IBD patient cohort | DSS/AT/none | predicted histological score from expression, splicing, VDJ features | — |
- ▼ DSS mice lost 20% and 10% body weight after DSS cycles one and three respectively 20%; 10%
- – Differentially expressed genes (FDR<0.05) increased from day 5 to day 17 then decreased by day 36 2040 (day5), 5740 (day17), 2771 (day36) genes
- ▲ Differential splicing (DS) signal peaked at day 12 125 genes at FDR<0.05
- – 725 genes showed disease-specific dynamic (temporal) expression signature 725 genes, FDR<0.05
- – 141 genes showed disease-specific dynamic splicing signature, largely non-overlapping with dynamic expression genes 141 genes, FDR<0.05; only 2 genes overlap with expression
- ▲ Il1rl1 TIR-lacking/TIR-containing transcript ratio increased in DSS colitis, confirmed by qPCR ratio 1:1 to 3.5:1
- ▲ Lama3 short (LN-domain-lacking)/long isoform ratio increased in DSS colitis, confirmed by qPCR ratio 1:1 to 3:1
- ▲ VDJ read-based measurements correlated with pathological lymphocyte aggregate counts in first DSS phase spearman rho ~0.86, p~0.05
- count 2040 DE genes (day5), 5740 DE genes (day17), 2771 DE genes (day36) (timepoint-specific DSS vs control differential expression, FDR<0.05)
- count 125 DS genes at day 12 (highest number of significant differential splicing genes, FDR<0.05)
- count 725 genes (disease-specific temporal (dynamic) expression signature, FDR<0.05)
- count 141 genes (disease-specific temporal (dynamic) splicing signature, FDR<0.05)
- pvalue FDR=2.39e-18 (Lama3); FDR=1.4e-3 (Il1rl1) (early DSS-specific differential splicing at day 5)
- fold_change Il1rl1 TIR-lacking:TIR-containing ratio 1:1 to 3.5:1 (qPCR validation, colon tissue)
- fold_change Lama3 short:long isoform ratio 1:1 to 3:1 (qPCR validation, colon tissue)
- correlation spearman rho ~0.86, p val ~0.05 (VDJ measurements vs pathological lymphocyte aggregate counts, first DSS phase)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study characterizes the transcriptional landscape of acute and chronic murine colitis (DSS and adoptive transfer models) using RNAseq of blood and colon tissue at four timepoints (days 5, 12, 17, 36). A linear model framework identifies both fixed (timepoint-specific) and dynamic (disease × time interaction) differential expression and splicing signatures, with FDR-based multiple-testing correction applied throughout. Random forest classifiers trained with 10-fold cross-validation predict continuous histopathology scores from molecular features, with predictive performance evaluated by Spearman correlation. Resulting signatures are projected onto human IBD causal networks to delineate acute and chronic disease subnetworks.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Linear model with disease × time interaction term | Dynamic differential gene expression signatures across four DSS timepoints; 725 genes identified at FDR < 0.05 | 6–24 mice per group (stated in Fig. 1 legend) | not stated |
| Linear model (timepoint-specific disease vs. control) | Fixed differential expression at each sacrifice day (days 5, 12, 17, 36); threshold logFC > 0 with FDR < 0.05 | 6–24 mice per group (stated in Fig. 1 legend) | not stated |
| Wilcoxon rank-sum test | Histopathology score comparisons across DSS cycles (Fig. 1c) | 6–24 mice per group (stated in Fig. 1 legend) | not stated |
| Spearman correlation | VDJ measurements vs. pathological lymphocyte aggregate counts (rho ~0.86, p ~0.05); random forest model performance evaluation across validation sets | null | not stated |
| Pearson correlation (R > 0.8 and R > 0.7 thresholds) | Isoform-specific co-expression analysis for Il1rl1 and Lama3 to construct gene lists for pathway enrichment | null | not stated |
| Random forest with 10-fold cross-validation | Prediction of continuous histopathology scores from expression, splicing, and VDJ features; evaluated across DSS Janssen, DSS MSSM, and AT model validation sets | null | na |
-
Differential expression was identified with a linear model framework; the specific software tool is not named in the available text↳ Could also: DESeq2 or edgeR, which use negative binomial models explicitly parameterized for the overdispersion of RNAseq read counts — Count-specific models account for the discrete and overdispersed nature of RNA-seq data and represent the current community standard; they also offer shrinkage estimators for log fold-change that stabilize estimates from small samples
-
Temporal expression clusters were derived using the Mfuzz soft-clustering algorithm applied to median disease trajectories↳ Could also: Trajectory-aware methods such as ImpulseDE2 or splineTimeR, or dynamic time warping-based hierarchical clustering — Methods designed for ordered time-series data can capture non-monotone dynamics and provide model-based uncertainty estimates per cluster; comparing multiple clustering approaches can also assess robustness of the reported cluster assignments
-
Multiple separate random forest models were trained for different signatures and evaluated on multiple validation sets↳ Could also: Elastic-net regularized regression (e.g., glmnet) or gradient-boosted trees (e.g., XGBoost), also with cross-validation — Penalized linear models provide explicit coefficient estimates that aid biological interpretation and allow inference on feature importance; comparing multiple learner types also permits assessment of whether findings are robust to model choice
-
Predictive model performance was summarized exclusively with Spearman correlation between predicted and observed histopathology scores↳ Could also: Root mean squared error (RMSE) and mean absolute error (MAE) alongside rank correlation — Spearman rho captures rank agreement but not the absolute magnitude of prediction error; RMSE/MAE provide complementary information about practical predictive accuracy on the original histopathology score scale
-
Phenotype measurement dispersion (body weight, colon length, histopathology) was shown as ±1 SD with group sizes that varied from 6 to 24 mice↳ Could also: 95% confidence intervals or SEM, particularly given the variable and sometimes small group sizes — CI or SEM convey uncertainty about the group mean rather than spread of individual observations, which can be informative when comparing timepoints with substantially different numbers of animals
-
Isoform co-expression gene lists for pathway enrichment were constructed using fixed Pearson R thresholds (> 0.8 and > 0.7)↳ Could also: Weighted gene co-expression network analysis (WGCNA) or soft-thresholding approaches that use the full correlation distribution — Continuous weighting reduces sensitivity to the arbitrary threshold value and can improve stability of enrichment results, particularly when sample sizes are modest; it also allows integration of weaker co-expression signals that a hard cutoff would exclude
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36694043
Paper: Peters LA et al. A temporal classifier predicts histopathology state and parses acute-chronic phasing in inflammatory bowel disease patients. Commun Biol 2023. PMID 36694043 · PMCID PMC9873918 · DOI 10.1038/s42003-023-04469-y
Code / data availability (verbatim from the paper)
- Code availability: "Code is provided in the supplementary data (Supplementary code 1).
Open-source R code from publicly available packages that was exclusively used in this study
is available here https://github.com/LosicLab/losiclab.github.io"
- ⚠️ The GitHub URL the RU was seeded with (
LosicLab/losiclab.github.io) is the lab's Jekyll website, last pushed 2019-01-15 (before the 2023 paper). It contains no analysis code for this paper (only generic open-source R packages, per the statement). The actual analysis code is Supplementary code 1 attached to the article. This is ano_own_reposituation — per brief rule P16 this does not down-rank the paper.
- ⚠️ The GitHub URL the RU was seeded with (
- Data availability:
- Mouse DSS/AT models (this paper's primary data): GEO GSE214600 — Public on Dec 31 2022,
Series_pubmed_id = 36694043. ✅ available. - Human MSCCR blood: GSE186507 — Public on Sep 16 2022. ✅
- Human MSCCR biopsy: GSE193677 — Public on Sep 16 2022. ✅
- The paper text says "Data are private and will be released in October 2024 and 2025" — that embargo statement is out of date; all three series are public now (verified via GEO E-utilities 2026-06-15).
- Numerical source data for graphs/charts: figshare DOI 10.6084/m9.figshare.21706202.v2 — public, 36 files (CSV / RData / txt). This is the per-figure source data.
- Mouse DSS/AT models (this paper's primary data): GEO GSE214600 — Public on Dec 31 2022,
Reproduction strategy (80/20)
The full pipeline (STAR 2.4.0g1 alignment → featureCounts → DESeq2/limma-voom DE+DS → Mfuzz temporal clustering → random-forest classifier) over 340 mouse + ~1000 human RNA-seq samples is the hard ~20%: heavy compute, many unspecified degrees of freedom (exact covariate set, CPM/FDR thresholds, batch correction), and the authors' own analysis code is a supplementary ZIP rather than a runnable repo. We do not attempt a full from-FASTQ rerun.
Instead — and equally valid per the brief — we perform an honest 1:1 verification of the paper's reported pipeline-derived numbers against the authors' own published numerical source data (figshare). This directly tests whether the headline numbers are actually derivable from the shipped data (a fabrication check), which is the project's goal. All download + computation runs on «our HPC»/«infra».
In scope (pipeline-derived, attempted)
| # | Reported result | Pipeline | Source-data file (figshare) |
|---|---|---|---|
| C1 | DE gene counts per DSS timepoint: 2,040 (d5), 5,740 (d17), 2,771 (d36) | DESeq2 (disease-time interaction) | fig2/fig3 master DE+DS table |
| C2 | Differentially spliced genes up to 125 (day 12) | limma-voom | master DE+DS table |
| C3 | 725 dynamic temporal-signature genes (87% also fixed DE) | LRT disease×time | cluster master table / DE table |
| C4 | Mfuzz soft clusters A–E: 113–194 genes each; cluster B (late) enriched immunoglobulin | Mfuzz | fig2_cycling..._cluster_master_table |
| C5 | DE/cycling signature list sizes (per-day & cluster human-symbol lists) | DESeq2 + ortholog map | fig3_*_human_symbols.txt |
| C6 | V(D)J: detection >99% of DSS samples; median ~500 clones (AT: 44%, median 8) | RNA-seq VDJ assembly | s4_vdj_and_umi_summary_stats...csv |
| C7 | Classifier predictive correlation rho ~0.8 (DSS proximal colon), ~0.5 (blood) | random forest, 10-fold CV | fig4_master_model_training_and_validation_list.RData |
Out of scope (not pipeline / not attempted, with reason)
- Wet-lab: body-weight/colon-length phenotyping, H&E histopathology scoring, qPCR isoform validation (IL1RL1 3.5:1, LAMA3 3:1 ratios) — bench measurements, not pipeline outputs.
- Full from-FASTQ RNA-seq alignment & quantification rerun — hard 20%, heavy
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Strong, largely 1:1 reproduction: the headline DESeq2 DE counts (2040/5740/2771), the 725-gene dynamic temporal signature, the Mfuzz cluster-size set {113,125,136,157,194} and cluster sum all reproduce exactly from the authors' own figshare source data, and the V(D)J and 87%-overlap statistics match within rounding — no fabrication detected. The deviations that remain are on the authors'/data-availability side: the DS=125 splicing count and the region-specific random-forest rho (proximal colon ~0.8, blood ~0.5) are not derivable from the shipped data (which carries only gene-level DE and whole-colon Janssen models), and the cited code repo is the lab's website rather than analysis code. The central classifier conclusion is therefore qualitatively but not 1:1 confirmed (whole-colon rho 0.63–0.84 ≈ paper's ~0.8), making this a solid reproduction with explainable, deposit-driven gaps rather than a critical discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.