Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Predicting position along a looping immune response trajectory.

PLoS One · 2018
L1 98/100 PQI 99
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
98/100
Reproducibility score
1.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 93% of all assessed papers rank 65 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> reproduced 1:1. The authors' GitLab repo (gitlab.com/prath/resilience2018 @ a46cea3) ships the Looper pipeline AND the pre-processed GSE47122 human-monocyte matrix; re-running it on «our HPC» (Python 2.7 + pandas 0.24.2, SLURM 2176461) reproduced every primary headline number exactly: 18859 genes, top-0.5% = 95 genes, 102 phase-shifted pairs of 4465 possible, the deterministic 26/34 train-test split, and the IL1A-TNIP3 94% (=32/34) time-prediction accuracy; the angle-vs-time Pearson rho=0.988 matches the reported 0.98 (within-tol). The split is index-based (no random seed), so these are exactly reproducible. NOTE: the BRIEF's listed code repo (github.com/bytorres/PlosBio2015) is the predecessor MATLAB methods repo, NOT this paper's code; the real code is the cited GitLab repo, which we used. NOT attempted (optional ~20%): YF17D cohorts GSE13699/GSE13485 (Fig 5, 83/73/65% - shipped but not run), Ayasdi TDA (Fig 4, proprietary software - out of scope), Tableau rendering (visualization only), and the separate predicted-vs-actual-time R2=0.99 regression. No possible-fabrication flags: all reproduced values derive from shipped data+code.

💻 Code ↗ 🗄 Data: GSE47122

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 98
    assessed: 2026-06-14 ⛓ 0f65dbbf7eed
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can an automated computational method identify pairs of genes that are phase-shifted and trace loops in longitudinal immune data, and can these loops be used to predict a host's position (time of perturbation/stage) along an inflammation-resolution trajectory back to health?

Core claims
  • Looper is an automated computational method that identifies gene pairs that trace looping trajectories when plotted against each other in longitudinally sampled data method
  • Loops derived from training data can predict the time/stage of perturbation in withheld test samples finding
  • In human monocytes, the IL1A-TNIP3 loop predicts time of perturbation in withheld samples with 94% accuracy finding
  • In YF17D-vaccinated individuals, the CDC20-IFI44L loop predicts perturbation stage with 65-83% accuracy within and across independent cohorts finding
  • The angle derived from polar transformation of a gene-pair loop correlates linearly with time, allowing loops to recapitulate disease trajectory mechanism
  • The IL1A-TNIP3 loop was experimentally validated in monocytes from independent human donors by qRT-PCR finding
  • Symbolic aggregate approximation (SAX) with a T/4 (90-degree) phase-shift search pattern identifies candidate phase-shifted gene pairs that trace loops method
  • Looper-identified looping gene pairs yield far better prediction accuracy than randomly sampled gene pairs finding
Experimental setups
Assay System Perturbation Readout Platform
Microarray gene expression (re-analysis of public data) Human monocytes from 12 volunteers, in vitro Sequential stimulation with CCL2, LPS, TNFα, IFNγ (inflammation) then IL10, TGFβ (resolution) Log2-scaled gene expression across 9 time points (0, 2, 2.5, 3, 3.5, 4, 14, 24, 48 h) GSE47122 (downloaded via Metageo Python module)
Microarray gene expression (re-analysis of public data) Whole blood, humans (Montreal cohort, 15 individuals; Lausanne cohort, 13 individuals) YF17D yellow fever vaccination Gene expression across time points (Montreal: 7 time points days 0-60; Lausanne: days 0,3,7) GSE13699
Microarray gene expression (re-analysis of public data) PBMCs, humans (Emory cohort, 25 individuals) YF17D yellow fever vaccination Gene expression at days 0, 3, 7 post-vaccination GSE13485
qRT-PCR (experimental validation) Monocytes purified from whole blood of 3 independent human donors, in vitro Inflammatory stimulation including LPS at 5, 50, 500 ng/ml IL1A and TNIP3 gene expression over time (loop tracing)
Topological data analysis (TDA) Montreal cohort whole blood (YF17D) YF17D vaccination High-dimensional network structure of top 0.5% (91 genes) by SD from baseline
Key results
  • IL1A-TNIP3 loop predicted time of perturbation in 34 withheld human monocyte test samples with 94% accuracy 94% accuracy; R2=0.99
  • Angle from polar transformation of IL1A-TNIP3 loop is positively linearly correlated with time Pearson's ρ=0.98
  • CDC20-IFI44L loop predicted perturbation stage in withheld Montreal cohort test samples 83% accuracy
  • CDC20-IFI44L loop predicted perturbation stage in independent Lausanne cohort 73% accuracy
  • CDC20-IFI44L loop predicted perturbation stage in independent Emory cohort 65% accuracy
  • Angle from polar transformation of CDC20-IFI44L loop is positively linearly correlated with time Pearson's ρ=0.91
  • IFI44L peaks before CDC20 (CDC20 peaks day 3, IFI44L peaks day 7), demonstrating phase-shift
  • Adding Gaussian noise to Emory cohort shifted prediction accuracy only slightly, indicating robustness to noise 65% to 63%
Key statistics
  • correlation Pearson's ρ = 0.98 (IL1A-TNIP3 loop angle vs time, monocytes)
  • correlation R2 = 0.99 (Predicted vs actual time, monocyte test samples)
  • other 94% accuracy (Perturbation time prediction, human monocyte test samples (KNN, K=3))
  • correlation Pearson's ρ = 0.91 (CDC20-IFI44L loop angle vs time, Montreal cohort)
  • other 83% accuracy (Perturbation stage prediction, withheld Montreal cohort samples)
  • other 73% accuracy (Perturbation stage prediction, Lausanne cohort (33 data points))
  • other 65% accuracy (Perturbation stage prediction, Emory cohort (75 data points))
  • count top 0.5% (91 genes) (Genes by SD from baseline used for TDA, Montreal cohort)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper introduces Looper, a computational pipeline applied to publicly available longitudinal microarray datasets (GSE47122, GSE13699, GSE13485). Candidate phase-shifted gene pairs were identified by filtering for top-0.5% expression-range genes and applying Symbolic Aggregate Approximation (SAX) pattern matching; a K-nearest neighbor (K=3) classifier then predicted perturbation stage in withheld samples. Results were reported primarily as classification accuracy (%) and Pearson correlation coefficients (ρ) between polar-coordinate loop angles and time, with R² for linear fits of predicted vs. actual stage.

Replicationmixed Sample sizePer-cohort sample sizes stated: monocyte data 12 donors (9 time points, ~50% random train/test split); Montreal YF17D 15 individuals (11 train / 4 withheld test); Lausanne 13 individuals; Emory 25 individuals across 2 trials. Experimental qRT-PCR validation used 3 independent donors. GroupsPerturbation stages (early/middle/late inflammation-resolution) vs. baseline (pre- and post-perturbation time points), evaluated within and across independent cohorts Pairingmixed Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Pearson correlation (ρ) Angle derived from IL1A-TNIP3 polar transformation vs. ordinal time (ρ=0.98, training data); CDC20-IFI44L angle vs. time in Montreal YF17D cohort (ρ=0.91) 9 time-point means (monocyte training data); 6 time-point means across n=11 Montreal training individuals not stated
K-nearest neighbor classification (K=3) Predicting perturbation stage from IL1A-TNIP3 loop (94% accuracy, 34 monocyte test data points); CDC20-IFI44L loop: Montreal withheld (83%, 24 data points), Lausanne cohort (73%, 33 data points), Emory cohort (65%, 75 data points) 34 (monocyte test); 24 (Montreal withheld); 33 (Lausanne); 75 (Emory) not stated
Symbolic Aggregate Approximation (SAX) pattern matching with T/4 phase shift Identifying phase-shifted gene pair candidates among the top 0.5% expression-range genes in training data top 0.5% of all assayed genes by expression range not stated
Topological Data Analysis (TDA) Assessing overall temporal network structure of Montreal YF17D cohort (GSE13699) on 91 genes (top 0.5% by SD from day-0 baseline) 15 individuals, 91 genes not stated
Leave-one-individual-out cross-validation (LOOCV) Evaluating robustness of CDC20-IFI44L perturbation-stage predictions across all 15 Montreal cohort individuals 15 individuals na
Gaussian noise sensitivity analysis (in silico) Testing robustness of CDC20-IFI44L loop predictions on Emory cohort under noise sampled from gene-pair mean/SD distribution 75 data points (Emory cohort) not stated
Approaches that could also have been used
  • Classification accuracy (%) was the sole reported performance metric for the KNN perturbation-stage predictor across all cohorts
    Could also: Per-class sensitivity/recall, macro-averaged F1 score, Matthews correlation coefficient, or bootstrap confidence intervals around accuracy could also be reported — Accuracy alone can be misleading when stage class sizes differ; F1/MCC capture per-class error balance, and a bootstrap CI would convey statistical uncertainty in the accuracy estimate given the modest sample sizes
  • Multiple candidate gene pairs were evaluated and the highest-accuracy pair (IL1A-TNIP3; CDC20-IFI44L) was selected and reported, without adjustment for the number of pairs examined
    Could also: A permutation-based approach — comparing observed best-pair accuracy to the null distribution of best-pair accuracy from randomly shuffled labels or random gene pairs — could also contextualize the reported accuracy — Selecting the top performer from many candidates inflates apparent performance; a permutation test would estimate how much of the accuracy gain is attributable to the selection process itself
  • A K-nearest neighbor classifier (K=3) in the 2-D loop space was used to assign perturbation stage
    Could also: Ordinal logistic regression, a support vector machine with RBF kernel, or a random forest could also map 2-D loop coordinates to ordered stages — Model-based classifiers make decision boundaries explicit and, in the case of ordinal logistic regression, respect the ordering of perturbation stages; they also allow formal inference on classifier performance and calibrated probability outputs
  • Pearson's ρ was used to summarize the linear relationship between polar-coordinate angle and time using time-point means
    Could also: Spearman's rank correlation or a simple linear regression with a 95% confidence interval for the slope could also be reported — Pearson's ρ assumes linearity and normality; Spearman's ρ is robust to nonlinearity and outliers, and a regression CI would convey uncertainty in the angle–time relationship given the small number of time-point summary values used
  • Genes were pre-filtered by retaining those in the top 0.5% of expression range, using a fixed arbitrary threshold stated as such in the paper
    Could also: A data-driven threshold such as a permutation-derived cutoff, an elbow criterion on the ranked range distribution, or variance-stabilizing normalization followed by FDR-controlled selection could also be applied — A data-driven cutoff would generalize more transparently across datasets with different dynamic ranges and reduce sensitivity to the arbitrary 0.5% choice
  • Phase shifts between gene pairs were detected using SAX discretization with a fixed T/4 offset as the target pattern
    Could also: Pairwise cross-correlation functions or dynamic time warping distances could also quantify the lag between gene expression time series — Cross-correlation provides a continuous lag estimate at all offsets and a natural test statistic for significance; dynamic time warping handles irregular or unequal time intervals; both complement the binary SAX match with graded measures of phase similarity
Software: Metageo (Python module, used for GEO dataset download) · Looper (custom Python computational pipeline, described in paper)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE13485 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE13699 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE13845 GEO in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet
GSE47122 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-30296270

Paper: Rath P, Allen JA, Schneider DS. Predicting position along a looping immune response trajectory. PLoS One 2018;13(10):e0200147. DOI 10.1371/journal.pone.0200147.

Code artifact (corrected)

The BRIEF listed github.com/bytorres/PlosBio2015 — that is the predecessor methods repo (Torres BY et al., "Tracking resilience to infections by mapping disease space", PLOS Biology 2016; MATLAB+R polar transform). It is NOT the code for this paper.

The actual code for this paper is the GitLab repo cited in the article's "Code availability": https://gitlab.com/prath/resilience2018 (Poonam Rath, created 2018-01-10, commit a46cea3664e4284cec6833ee2b6e34dd6c947318). It ships the Python Looper library (looper.py), the SAX implementation (saxpy.py, N. Hoffman, MIT), the metageo GEO-parsing module, the analysis Jupyter notebooks, AND the pre-processed input CSVs. Per BRIEF rule 2 (P16), applying a third-party/own tool to the paper's own data is equally valid — here we run the authors' own shipped pipeline on their shipped data.

Pipeline (in scope)

Method = SAX (Symbolic Aggregate approXimation) phase-shift loop discovery + a K=3 nearest-neighbour time predictor, in Python 2.7 (pandas/numpy/scipy).

Primary dataset: GSE47122 (human monocytes, 12 donors, 9 time points 0–48 h, sequential immune elicitors). The processed matrix is shipped as code/human_mono_gse47122.csv (18859 genes × 60 samples). Reproduction notebook: code/script_for_human_monocyte_FINAL.ipynb.

In-scope reproducible results (deterministic — split is index-based, not random)

id result paper location
C1 18859 input genes; train 26 / test 34 samples Methods; Fig 3
C2 top 0.5% by range → 95 genes (of 18859) Results / Methods
C3 102 phase-shifted gene pairs of 4465 possible (=C(95,2)) Results
C4 IL1A–TNIP3 loop predicts perturbation time at 94% over 34 test samples Fig 3F; abstract
C5 IL1A–TNIP3 angle vs time Pearson ρ≈0.98, R²≈0.99 Fig 3E / S2B

We reproduce C1–C5 by running the shipped looper.py pipeline on the shipped human_mono_gse47122.csv (and shipped FigS2B polar CSV for C5), in a rebuilt Python-2.7 conda environment on «our HPC».

Out of scope (the optional hard ~20%)

  • YF17D vaccination cohorts (GSE13699 Montreal/Lausanne, GSE13485 Emory): 83% / 73% / 65% accuracies, ρ=0.91 (Fig 5). Reproducible in principle from the shipped *.csv + script_for_yellow_fever_FINAL.ipynb, but secondary; attempt only if primary lands cleanly with budget left.
  • Ayasdi 3.0 Topological Data Analysis (Fig 4): out of scope — proprietary commercial software (Ayasdi), not reproducible.
  • Tableau v9.0 figure rendering: out of scope (visualization only, no new number).
  • Wet-lab / GEO raw normalization upstream of the shipped matrix: not attempted (we start from the authors' shipped processed matrix, as the notebook does).

Honesty notes

  • The train/test split uses a deterministic index-order 50% cut per timepoint (index[:split_pt] / index[split_pt:]), so results are reproducible without a random seed — good for auditing.
  • create_composite_profile keeps the Time column among "genes" due to a set('Time') bug in the original; we replicate the original behaviour verbatim rather than fix it, so counts match the paper's own code path.
Figures / tables: Fig 3FFig 3E
C1
Reported
18859 input genes
Reproduced
18859
exact
C1b
Reported
34 test samples
Reproduced
26 train / 34 test
exact
C2
Reported
95 genes (top 0.5% by range)
Reproduced
95
exact
C3
Reported
102 phase-shifted gene pairs of 4465
Reproduced
102 of 4465
exact
C4
Reported
IL1A-TNIP3 predicts time at 94% (34 test)
Reproduced
0.9412 (32/34)
exact
C5
Reported
IL1A-TNIP3 angle-vs-time Pearson rho=0.98
Reproduced
rho=0.9884
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 98/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

A clean 1:1 reproduction: ran the authors' own shipped Looper pipeline (Python 2.7) on their own shipped GSE47122 human-monocyte matrix. Every headline number is exact — 18859 input genes, 95 top-0.5%-range genes, 102 phase-shifted pairs of 4465 (=C(95,2)), the deterministic 26-train/34-test split, and the IL1A-TNIP3 94% (=32/34) time-prediction accuracy — with the angle-vs-time Pearson rho=0.9884 vs reported 0.98 (rounding). The split is index-based with no random seed, so results are exactly reproducible. The registry code link is a false-positive (predecessor MATLAB repo); the agent correctly used the cited GitLab repo. No fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

141.9 k
tokens (I/O) · 8.8 M incl. cache
15 min
runtime · 0.03 CPU-h
1.3 GB
peak RAM
1
HPC jobs
hummel
machine