Predicting position along a looping immune response trajectory.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> reproduced 1:1. The authors' GitLab repo (gitlab.com/prath/resilience2018 @ a46cea3) ships the Looper pipeline AND the pre-processed GSE47122 human-monocyte matrix; re-running it on «our HPC» (Python 2.7 + pandas 0.24.2, SLURM 2176461) reproduced every primary headline number exactly: 18859 genes, top-0.5% = 95 genes, 102 phase-shifted pairs of 4465 possible, the deterministic 26/34 train-test split, and the IL1A-TNIP3 94% (=32/34) time-prediction accuracy; the angle-vs-time Pearson rho=0.988 matches the reported 0.98 (within-tol). The split is index-based (no random seed), so these are exactly reproducible. NOTE: the BRIEF's listed code repo (github.com/bytorres/PlosBio2015) is the predecessor MATLAB methods repo, NOT this paper's code; the real code is the cited GitLab repo, which we used. NOT attempted (optional ~20%): YF17D cohorts GSE13699/GSE13485 (Fig 5, 83/73/65% - shipped but not run), Ayasdi TDA (Fig 4, proprietary software - out of scope), Tableau rendering (visualization only), and the separate predicted-vs-actual-time R2=0.99 regression. No possible-fabrication flags: all reproduced values derive from shipped data+code.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 98assessed: 2026-06-14 ⛓ 0f65dbbf7eed
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan an automated computational method identify pairs of genes that are phase-shifted and trace loops in longitudinal immune data, and can these loops be used to predict a host's position (time of perturbation/stage) along an inflammation-resolution trajectory back to health?
- ★ Looper is an automated computational method that identifies gene pairs that trace looping trajectories when plotted against each other in longitudinally sampled data method
- ★ Loops derived from training data can predict the time/stage of perturbation in withheld test samples finding
- ★ In human monocytes, the IL1A-TNIP3 loop predicts time of perturbation in withheld samples with 94% accuracy finding
- ★ In YF17D-vaccinated individuals, the CDC20-IFI44L loop predicts perturbation stage with 65-83% accuracy within and across independent cohorts finding
- ★ The angle derived from polar transformation of a gene-pair loop correlates linearly with time, allowing loops to recapitulate disease trajectory mechanism
- ★ The IL1A-TNIP3 loop was experimentally validated in monocytes from independent human donors by qRT-PCR finding
- Symbolic aggregate approximation (SAX) with a T/4 (90-degree) phase-shift search pattern identifies candidate phase-shifted gene pairs that trace loops method
- Looper-identified looping gene pairs yield far better prediction accuracy than randomly sampled gene pairs finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Microarray gene expression (re-analysis of public data) | Human monocytes from 12 volunteers, in vitro | Sequential stimulation with CCL2, LPS, TNFα, IFNγ (inflammation) then IL10, TGFβ (resolution) | Log2-scaled gene expression across 9 time points (0, 2, 2.5, 3, 3.5, 4, 14, 24, 48 h) | GSE47122 (downloaded via Metageo Python module) |
| Microarray gene expression (re-analysis of public data) | Whole blood, humans (Montreal cohort, 15 individuals; Lausanne cohort, 13 individuals) | YF17D yellow fever vaccination | Gene expression across time points (Montreal: 7 time points days 0-60; Lausanne: days 0,3,7) | GSE13699 |
| Microarray gene expression (re-analysis of public data) | PBMCs, humans (Emory cohort, 25 individuals) | YF17D yellow fever vaccination | Gene expression at days 0, 3, 7 post-vaccination | GSE13485 |
| qRT-PCR (experimental validation) | Monocytes purified from whole blood of 3 independent human donors, in vitro | Inflammatory stimulation including LPS at 5, 50, 500 ng/ml | IL1A and TNIP3 gene expression over time (loop tracing) | — |
| Topological data analysis (TDA) | Montreal cohort whole blood (YF17D) | YF17D vaccination | High-dimensional network structure of top 0.5% (91 genes) by SD from baseline | — |
- – IL1A-TNIP3 loop predicted time of perturbation in 34 withheld human monocyte test samples with 94% accuracy 94% accuracy; R2=0.99
- ▲ Angle from polar transformation of IL1A-TNIP3 loop is positively linearly correlated with time Pearson's ρ=0.98
- – CDC20-IFI44L loop predicted perturbation stage in withheld Montreal cohort test samples 83% accuracy
- – CDC20-IFI44L loop predicted perturbation stage in independent Lausanne cohort 73% accuracy
- – CDC20-IFI44L loop predicted perturbation stage in independent Emory cohort 65% accuracy
- ▲ Angle from polar transformation of CDC20-IFI44L loop is positively linearly correlated with time Pearson's ρ=0.91
- – IFI44L peaks before CDC20 (CDC20 peaks day 3, IFI44L peaks day 7), demonstrating phase-shift
- ▼ Adding Gaussian noise to Emory cohort shifted prediction accuracy only slightly, indicating robustness to noise 65% to 63%
- correlation Pearson's ρ = 0.98 (IL1A-TNIP3 loop angle vs time, monocytes)
- correlation R2 = 0.99 (Predicted vs actual time, monocyte test samples)
- other 94% accuracy (Perturbation time prediction, human monocyte test samples (KNN, K=3))
- correlation Pearson's ρ = 0.91 (CDC20-IFI44L loop angle vs time, Montreal cohort)
- other 83% accuracy (Perturbation stage prediction, withheld Montreal cohort samples)
- other 73% accuracy (Perturbation stage prediction, Lausanne cohort (33 data points))
- other 65% accuracy (Perturbation stage prediction, Emory cohort (75 data points))
- count top 0.5% (91 genes) (Genes by SD from baseline used for TDA, Montreal cohort)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper introduces Looper, a computational pipeline applied to publicly available longitudinal microarray datasets (GSE47122, GSE13699, GSE13485). Candidate phase-shifted gene pairs were identified by filtering for top-0.5% expression-range genes and applying Symbolic Aggregate Approximation (SAX) pattern matching; a K-nearest neighbor (K=3) classifier then predicted perturbation stage in withheld samples. Results were reported primarily as classification accuracy (%) and Pearson correlation coefficients (ρ) between polar-coordinate loop angles and time, with R² for linear fits of predicted vs. actual stage.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation (ρ) | Angle derived from IL1A-TNIP3 polar transformation vs. ordinal time (ρ=0.98, training data); CDC20-IFI44L angle vs. time in Montreal YF17D cohort (ρ=0.91) | 9 time-point means (monocyte training data); 6 time-point means across n=11 Montreal training individuals | not stated |
| K-nearest neighbor classification (K=3) | Predicting perturbation stage from IL1A-TNIP3 loop (94% accuracy, 34 monocyte test data points); CDC20-IFI44L loop: Montreal withheld (83%, 24 data points), Lausanne cohort (73%, 33 data points), Emory cohort (65%, 75 data points) | 34 (monocyte test); 24 (Montreal withheld); 33 (Lausanne); 75 (Emory) | not stated |
| Symbolic Aggregate Approximation (SAX) pattern matching with T/4 phase shift | Identifying phase-shifted gene pair candidates among the top 0.5% expression-range genes in training data | top 0.5% of all assayed genes by expression range | not stated |
| Topological Data Analysis (TDA) | Assessing overall temporal network structure of Montreal YF17D cohort (GSE13699) on 91 genes (top 0.5% by SD from day-0 baseline) | 15 individuals, 91 genes | not stated |
| Leave-one-individual-out cross-validation (LOOCV) | Evaluating robustness of CDC20-IFI44L perturbation-stage predictions across all 15 Montreal cohort individuals | 15 individuals | na |
| Gaussian noise sensitivity analysis (in silico) | Testing robustness of CDC20-IFI44L loop predictions on Emory cohort under noise sampled from gene-pair mean/SD distribution | 75 data points (Emory cohort) | not stated |
-
Classification accuracy (%) was the sole reported performance metric for the KNN perturbation-stage predictor across all cohorts↳ Could also: Per-class sensitivity/recall, macro-averaged F1 score, Matthews correlation coefficient, or bootstrap confidence intervals around accuracy could also be reported — Accuracy alone can be misleading when stage class sizes differ; F1/MCC capture per-class error balance, and a bootstrap CI would convey statistical uncertainty in the accuracy estimate given the modest sample sizes
-
Multiple candidate gene pairs were evaluated and the highest-accuracy pair (IL1A-TNIP3; CDC20-IFI44L) was selected and reported, without adjustment for the number of pairs examined↳ Could also: A permutation-based approach — comparing observed best-pair accuracy to the null distribution of best-pair accuracy from randomly shuffled labels or random gene pairs — could also contextualize the reported accuracy — Selecting the top performer from many candidates inflates apparent performance; a permutation test would estimate how much of the accuracy gain is attributable to the selection process itself
-
A K-nearest neighbor classifier (K=3) in the 2-D loop space was used to assign perturbation stage↳ Could also: Ordinal logistic regression, a support vector machine with RBF kernel, or a random forest could also map 2-D loop coordinates to ordered stages — Model-based classifiers make decision boundaries explicit and, in the case of ordinal logistic regression, respect the ordering of perturbation stages; they also allow formal inference on classifier performance and calibrated probability outputs
-
Pearson's ρ was used to summarize the linear relationship between polar-coordinate angle and time using time-point means↳ Could also: Spearman's rank correlation or a simple linear regression with a 95% confidence interval for the slope could also be reported — Pearson's ρ assumes linearity and normality; Spearman's ρ is robust to nonlinearity and outliers, and a regression CI would convey uncertainty in the angle–time relationship given the small number of time-point summary values used
-
Genes were pre-filtered by retaining those in the top 0.5% of expression range, using a fixed arbitrary threshold stated as such in the paper↳ Could also: A data-driven threshold such as a permutation-derived cutoff, an elbow criterion on the ranked range distribution, or variance-stabilizing normalization followed by FDR-controlled selection could also be applied — A data-driven cutoff would generalize more transparently across datasets with different dynamic ranges and reduce sensitivity to the arbitrary 0.5% choice
-
Phase shifts between gene pairs were detected using SAX discretization with a fixed T/4 offset as the target pattern↳ Could also: Pairwise cross-correlation functions or dynamic time warping distances could also quantify the lag between gene expression time series — Cross-correlation provides a continuous lag estimate at all offsets and a natural test statistic for significance; dynamic time warping handles irregular or unequal time intervals; both complement the binary SAX match with graded measures of phase similarity
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-30296270
Paper: Rath P, Allen JA, Schneider DS. Predicting position along a looping immune response trajectory. PLoS One 2018;13(10):e0200147. DOI 10.1371/journal.pone.0200147.
Code artifact (corrected)
The BRIEF listed github.com/bytorres/PlosBio2015 — that is the predecessor
methods repo (Torres BY et al., "Tracking resilience to infections by mapping
disease space", PLOS Biology 2016; MATLAB+R polar transform). It is NOT the code
for this paper.
The actual code for this paper is the GitLab repo cited in the article's
"Code availability": https://gitlab.com/prath/resilience2018
(Poonam Rath, created 2018-01-10, commit a46cea3664e4284cec6833ee2b6e34dd6c947318).
It ships the Python Looper library (looper.py), the SAX implementation
(saxpy.py, N. Hoffman, MIT), the metageo GEO-parsing module, the analysis
Jupyter notebooks, AND the pre-processed input CSVs. Per BRIEF rule 2 (P16),
applying a third-party/own tool to the paper's own data is equally valid — here
we run the authors' own shipped pipeline on their shipped data.
Pipeline (in scope)
Method = SAX (Symbolic Aggregate approXimation) phase-shift loop discovery + a K=3 nearest-neighbour time predictor, in Python 2.7 (pandas/numpy/scipy).
Primary dataset: GSE47122 (human monocytes, 12 donors, 9 time points 0–48 h,
sequential immune elicitors). The processed matrix is shipped as
code/human_mono_gse47122.csv (18859 genes × 60 samples). Reproduction notebook:
code/script_for_human_monocyte_FINAL.ipynb.
In-scope reproducible results (deterministic — split is index-based, not random)
| id | result | paper location |
|---|---|---|
| C1 | 18859 input genes; train 26 / test 34 samples | Methods; Fig 3 |
| C2 | top 0.5% by range → 95 genes (of 18859) | Results / Methods |
| C3 | 102 phase-shifted gene pairs of 4465 possible (=C(95,2)) | Results |
| C4 | IL1A–TNIP3 loop predicts perturbation time at 94% over 34 test samples | Fig 3F; abstract |
| C5 | IL1A–TNIP3 angle vs time Pearson ρ≈0.98, R²≈0.99 | Fig 3E / S2B |
We reproduce C1–C5 by running the shipped looper.py pipeline on the shipped
human_mono_gse47122.csv (and shipped FigS2B polar CSV for C5), in a rebuilt
Python-2.7 conda environment on «our HPC».
Out of scope (the optional hard ~20%)
- YF17D vaccination cohorts (GSE13699 Montreal/Lausanne, GSE13485 Emory):
83% / 73% / 65% accuracies, ρ=0.91 (Fig 5). Reproducible in principle from the
shipped
*.csv+script_for_yellow_fever_FINAL.ipynb, but secondary; attempt only if primary lands cleanly with budget left. - Ayasdi 3.0 Topological Data Analysis (Fig 4): out of scope — proprietary commercial software (Ayasdi), not reproducible.
- Tableau v9.0 figure rendering: out of scope (visualization only, no new number).
- Wet-lab / GEO raw normalization upstream of the shipped matrix: not attempted (we start from the authors' shipped processed matrix, as the notebook does).
Honesty notes
- The train/test split uses a deterministic index-order 50% cut per timepoint
(
index[:split_pt]/index[split_pt:]), so results are reproducible without a random seed — good for auditing. create_composite_profilekeeps theTimecolumn among "genes" due to aset('Time')bug in the original; we replicate the original behaviour verbatim rather than fix it, so counts match the paper's own code path.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
A clean 1:1 reproduction: ran the authors' own shipped Looper pipeline (Python 2.7) on their own shipped GSE47122 human-monocyte matrix. Every headline number is exact — 18859 input genes, 95 top-0.5%-range genes, 102 phase-shifted pairs of 4465 (=C(95,2)), the deterministic 26-train/34-test split, and the IL1A-TNIP3 94% (=32/34) time-prediction accuracy — with the angle-vs-time Pearson rho=0.9884 vs reported 0.98 (rounding). The split is index-based with no random seed, so results are exactly reproducible. The registry code link is a false-positive (predecessor MATLAB repo); the agent correctly used the cited GitLab repo. No fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.