Predicting position along a looping immune response trajectory.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> reproduced 1:1. The authors' GitLab repo (gitlab.com/prath/resilience2018 @ a46cea3) ships the Looper pipeline AND the pre-processed GSE47122 human-monocyte matrix; re-running it on «our HPC» (Python 2.7 + pandas 0.24.2, SLURM 2176461) reproduced every primary headline number exactly: 18859 genes, top-0.5% = 95 genes, 102 phase-shifted pairs of 4465 possible, the deterministic 26/34 train-test split, and the IL1A-TNIP3 94% (=32/34) time-prediction accuracy; the angle-vs-time Pearson rho=0.988 matches the reported 0.98 (within-tol). The split is index-based (no random seed), so these are exactly reproducible. NOTE: the BRIEF's listed code repo (github.com/bytorres/PlosBio2015) is the predecessor MATLAB methods repo, NOT this paper's code; the real code is the cited GitLab repo, which we used. NOT attempted (optional ~20%): YF17D cohorts GSE13699/GSE13485 (Fig 5, 83/73/65% - shipped but not run), Ayasdi TDA (Fig 4, proprietary software - out of scope), Tableau rendering (visualization only), and the separate predicted-vs-actual-time R2=0.99 regression. No possible-fabrication flags: all reproduced values derive from shipped data+code.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 98assessed: 2026-06-14 ⛓ 0f65dbbf7eed
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether an automated computational method (Looper) can identify gene pairs whose expression traces looping trajectories in phase space, and whether these loops can be used to predict an individual's position (time/stage of perturbation) along a self-resolving immune response trajectory.
- ★ Looper is a computational method that automatically identifies phase-shifted gene pairs (using SAX symbolic representation) that trace loops in longitudinal expression data. method
- ★ IL1A and TNIP3 expression are phase-shifted and trace a loop in human monocyte data, consistent with a feedback mechanism where IL1A induction precedes and is later suppressed by TNIP3-mediated inhibition of NF-kB. finding
- ★ The IL1A-TNIP3 loop predicts time of perturbation in withheld human monocyte test samples with 94% accuracy. finding
- ★ The IL1A-TNIP3 phase-shift and loop was experimentally validated by qRT-PCR in monocytes from independent human donors. finding
- ★ CDC20 and IFI44L expression are phase-shifted and trace a loop in YF17D-vaccinated individuals (Montreal cohort), with IFI44L peaking before CDC20. finding
- ★ The CDC20-IFI44L loop derived from the Montreal cohort predicts perturbation stage with 83% accuracy in withheld Montreal samples, 73% in the independent Lausanne cohort, and 65% in the independent Emory cohort. finding
- Randomly sampled gene pairs yield poor prediction accuracy compared to Looper-identified looping gene pairs. finding
- Looper's predictions remain robust (65% to 63% accuracy) when Gaussian noise is added to out-of-sample data. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| microarray gene expression profiling | human monocytes from 12 donors (in vitro) | sequential treatment with CCL2, LPS, TNFα, IFNγ (inflammatory) then IL10, TGFβ (deactivating) | gene expression across 9 time points (0–48h) | — |
| qRT-PCR | purified human monocytes from 3 independent donors (in vitro) | same inflammatory/resolution stimuli, including LPS dose variants (5, 50, 500 ng/ml) | IL1A and TNIP3 gene expression over time | — |
| microarray gene expression profiling | whole blood, YF17D-vaccinated individuals, Montreal cohort (n=15) and Lausanne cohort (n=13) | Yellow Fever Vaccine 17D vaccination | gene expression at days 0, 3, 7, 10, 14, 28, 60 post-vaccination | — |
| microarray gene expression profiling | PBMCs, YF17D-vaccinated individuals, Emory cohort (n=25, two trials) | Yellow Fever Vaccine 17D vaccination | gene expression at days 0, 3, 7 post-vaccination | — |
| topological data analysis (TDA) | Montreal YF17D cohort gene expression subset (91 genes) | Yellow Fever Vaccine 17D vaccination | network structure/looping return to baseline expression | — |
| in silico sensitivity analysis (Gaussian noise simulation) | Emory cohort CDC20-IFI44L expression data | computationally added noise | change in prediction accuracy | — |
- ▲ IL1A-TNIP3 loop angle (polar transform) correlates linearly with time in human monocyte training data Pearson's ρ = 0.98
- – Perturbation time predicted in 34 withheld human monocyte test samples using IL1A-TNIP3 loop and KNN (K=3) 94% accuracy, R² = 0.99
- – Phase-shift and loop of IL1A-TNIP3 confirmed experimentally across all 3 independent donors by qRT-PCR
- ▲ CDC20-IFI44L loop angle correlates linearly with time in Montreal cohort training data Pearson's ρ = 0.91
- – Perturbation stage predicted in withheld Montreal cohort test samples (4 individuals, 24 data points) using CDC20-IFI44L loop 83% accuracy
- – Perturbation stage predicted in independent Lausanne cohort using same CDC20-IFI44L loop 73% accuracy
- – Perturbation stage predicted in independent Emory cohort using same CDC20-IFI44L loop 65% accuracy
- ▼ Adding Gaussian noise to Emory cohort data reduced but did not eliminate prediction accuracy 65% to 63%
- correlation Pearson's ρ = 0.98 (IL1A-TNIP3 loop angle vs time, human monocyte training data)
- other 94% prediction accuracy, R² = 0.99 (withheld human monocyte test samples, IL1A-TNIP3 loop, KNN prediction)
- correlation Pearson's ρ = 0.91 (CDC20-IFI44L loop angle vs time, Montreal cohort training data)
- other 83% prediction accuracy (withheld Montreal cohort test samples (4 individuals), CDC20-IFI44L loop)
- other 73% prediction accuracy (Lausanne cohort validation using Montreal-derived CDC20-IFI44L loop)
- other 65% prediction accuracy (Emory cohort (25 individuals) validation using Montreal-derived CDC20-IFI44L loop)
- other shift from 65% to 63% accuracy (in silico Gaussian noise sensitivity analysis on Emory cohort)
- count top 0.5% (genes selected by expression range/standard deviation for loop identification and TDA)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper introduces Looper, a computational pipeline applied to publicly available longitudinal microarray datasets (GSE47122, GSE13699, GSE13485). Candidate phase-shifted gene pairs were identified by filtering for top-0.5% expression-range genes and applying Symbolic Aggregate Approximation (SAX) pattern matching; a K-nearest neighbor (K=3) classifier then predicted perturbation stage in withheld samples. Results were reported primarily as classification accuracy (%) and Pearson correlation coefficients (ρ) between polar-coordinate loop angles and time, with R² for linear fits of predicted vs. actual stage.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation (ρ) | Angle derived from IL1A-TNIP3 polar transformation vs. ordinal time (ρ=0.98, training data); CDC20-IFI44L angle vs. time in Montreal YF17D cohort (ρ=0.91) | 9 time-point means (monocyte training data); 6 time-point means across n=11 Montreal training individuals | not stated |
| K-nearest neighbor classification (K=3) | Predicting perturbation stage from IL1A-TNIP3 loop (94% accuracy, 34 monocyte test data points); CDC20-IFI44L loop: Montreal withheld (83%, 24 data points), Lausanne cohort (73%, 33 data points), Emory cohort (65%, 75 data points) | 34 (monocyte test); 24 (Montreal withheld); 33 (Lausanne); 75 (Emory) | not stated |
| Symbolic Aggregate Approximation (SAX) pattern matching with T/4 phase shift | Identifying phase-shifted gene pair candidates among the top 0.5% expression-range genes in training data | top 0.5% of all assayed genes by expression range | not stated |
| Topological Data Analysis (TDA) | Assessing overall temporal network structure of Montreal YF17D cohort (GSE13699) on 91 genes (top 0.5% by SD from day-0 baseline) | 15 individuals, 91 genes | not stated |
| Leave-one-individual-out cross-validation (LOOCV) | Evaluating robustness of CDC20-IFI44L perturbation-stage predictions across all 15 Montreal cohort individuals | 15 individuals | na |
| Gaussian noise sensitivity analysis (in silico) | Testing robustness of CDC20-IFI44L loop predictions on Emory cohort under noise sampled from gene-pair mean/SD distribution | 75 data points (Emory cohort) | not stated |
-
Classification accuracy (%) was the sole reported performance metric for the KNN perturbation-stage predictor across all cohorts↳ Could also: Per-class sensitivity/recall, macro-averaged F1 score, Matthews correlation coefficient, or bootstrap confidence intervals around accuracy could also be reported — Accuracy alone can be misleading when stage class sizes differ; F1/MCC capture per-class error balance, and a bootstrap CI would convey statistical uncertainty in the accuracy estimate given the modest sample sizes
-
Multiple candidate gene pairs were evaluated and the highest-accuracy pair (IL1A-TNIP3; CDC20-IFI44L) was selected and reported, without adjustment for the number of pairs examined↳ Could also: A permutation-based approach — comparing observed best-pair accuracy to the null distribution of best-pair accuracy from randomly shuffled labels or random gene pairs — could also contextualize the reported accuracy — Selecting the top performer from many candidates inflates apparent performance; a permutation test would estimate how much of the accuracy gain is attributable to the selection process itself
-
A K-nearest neighbor classifier (K=3) in the 2-D loop space was used to assign perturbation stage↳ Could also: Ordinal logistic regression, a support vector machine with RBF kernel, or a random forest could also map 2-D loop coordinates to ordered stages — Model-based classifiers make decision boundaries explicit and, in the case of ordinal logistic regression, respect the ordering of perturbation stages; they also allow formal inference on classifier performance and calibrated probability outputs
-
Pearson's ρ was used to summarize the linear relationship between polar-coordinate angle and time using time-point means↳ Could also: Spearman's rank correlation or a simple linear regression with a 95% confidence interval for the slope could also be reported — Pearson's ρ assumes linearity and normality; Spearman's ρ is robust to nonlinearity and outliers, and a regression CI would convey uncertainty in the angle–time relationship given the small number of time-point summary values used
-
Genes were pre-filtered by retaining those in the top 0.5% of expression range, using a fixed arbitrary threshold stated as such in the paper↳ Could also: A data-driven threshold such as a permutation-derived cutoff, an elbow criterion on the ranked range distribution, or variance-stabilizing normalization followed by FDR-controlled selection could also be applied — A data-driven cutoff would generalize more transparently across datasets with different dynamic ranges and reduce sensitivity to the arbitrary 0.5% choice
-
Phase shifts between gene pairs were detected using SAX discretization with a fixed T/4 offset as the target pattern↳ Could also: Pairwise cross-correlation functions or dynamic time warping distances could also quantify the lag between gene expression time series — Cross-correlation provides a continuous lag estimate at all offsets and a natural test statistic for significance; dynamic time warping handles irregular or unequal time intervals; both complement the binary SAX match with graded measures of phase similarity
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-30296270
Paper: Rath P, Allen JA, Schneider DS. Predicting position along a looping immune response trajectory. PLoS One 2018;13(10):e0200147. DOI 10.1371/journal.pone.0200147.
Code artifact (corrected)
The BRIEF listed github.com/bytorres/PlosBio2015 — that is the predecessor
methods repo (Torres BY et al., "Tracking resilience to infections by mapping
disease space", PLOS Biology 2016; MATLAB+R polar transform). It is NOT the code
for this paper.
The actual code for this paper is the GitLab repo cited in the article's
"Code availability": https://gitlab.com/prath/resilience2018
(Poonam Rath, created 2018-01-10, commit a46cea3664e4284cec6833ee2b6e34dd6c947318).
It ships the Python Looper library (looper.py), the SAX implementation
(saxpy.py, N. Hoffman, MIT), the metageo GEO-parsing module, the analysis
Jupyter notebooks, AND the pre-processed input CSVs. Per BRIEF rule 2 (P16),
applying a third-party/own tool to the paper's own data is equally valid — here
we run the authors' own shipped pipeline on their shipped data.
Pipeline (in scope)
Method = SAX (Symbolic Aggregate approXimation) phase-shift loop discovery + a K=3 nearest-neighbour time predictor, in Python 2.7 (pandas/numpy/scipy).
Primary dataset: GSE47122 (human monocytes, 12 donors, 9 time points 0–48 h,
sequential immune elicitors). The processed matrix is shipped as
code/human_mono_gse47122.csv (18859 genes × 60 samples). Reproduction notebook:
code/script_for_human_monocyte_FINAL.ipynb.
In-scope reproducible results (deterministic — split is index-based, not random)
| id | result | paper location |
|---|---|---|
| C1 | 18859 input genes; train 26 / test 34 samples | Methods; Fig 3 |
| C2 | top 0.5% by range → 95 genes (of 18859) | Results / Methods |
| C3 | 102 phase-shifted gene pairs of 4465 possible (=C(95,2)) | Results |
| C4 | IL1A–TNIP3 loop predicts perturbation time at 94% over 34 test samples | Fig 3F; abstract |
| C5 | IL1A–TNIP3 angle vs time Pearson ρ≈0.98, R²≈0.99 | Fig 3E / S2B |
We reproduce C1–C5 by running the shipped looper.py pipeline on the shipped
human_mono_gse47122.csv (and shipped FigS2B polar CSV for C5), in a rebuilt
Python-2.7 conda environment on «our HPC».
Out of scope (the optional hard ~20%)
- YF17D vaccination cohorts (GSE13699 Montreal/Lausanne, GSE13485 Emory):
83% / 73% / 65% accuracies, ρ=0.91 (Fig 5). Reproducible in principle from the
shipped
*.csv+script_for_yellow_fever_FINAL.ipynb, but secondary; attempt only if primary lands cleanly with budget left. - Ayasdi 3.0 Topological Data Analysis (Fig 4): out of scope — proprietary commercial software (Ayasdi), not reproducible.
- Tableau v9.0 figure rendering: out of scope (visualization only, no new number).
- Wet-lab / GEO raw normalization upstream of the shipped matrix: not attempted (we start from the authors' shipped processed matrix, as the notebook does).
Honesty notes
- The train/test split uses a deterministic index-order 50% cut per timepoint
(
index[:split_pt]/index[split_pt:]), so results are reproducible without a random seed — good for auditing. create_composite_profilekeeps theTimecolumn among "genes" due to aset('Time')bug in the original; we replicate the original behaviour verbatim rather than fix it, so counts match the paper's own code path.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
A clean 1:1 reproduction: ran the authors' own shipped Looper pipeline (Python 2.7) on their own shipped GSE47122 human-monocyte matrix. Every headline number is exact — 18859 input genes, 95 top-0.5%-range genes, 102 phase-shifted pairs of 4465 (=C(95,2)), the deterministic 26-train/34-test split, and the IL1A-TNIP3 94% (=32/34) time-prediction accuracy — with the angle-vs-time Pearson rho=0.9884 vs reported 0.98 (rounding). The split is index-based with no random seed, so results are exactly reproducible. The registry code link is a false-positive (predecessor MATLAB repo); the agent correctly used the cited GitLab repo. No fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.