Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Predicting position along a looping immune response trajectory.

PLoS One · 2018
L1 98/100 PQI 99
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
98/100
Reproducibility score
1.4 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 93% of all assessed papers rank 65 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> reproduced 1:1. The authors' GitLab repo (gitlab.com/prath/resilience2018 @ a46cea3) ships the Looper pipeline AND the pre-processed GSE47122 human-monocyte matrix; re-running it on «our HPC» (Python 2.7 + pandas 0.24.2, SLURM 2176461) reproduced every primary headline number exactly: 18859 genes, top-0.5% = 95 genes, 102 phase-shifted pairs of 4465 possible, the deterministic 26/34 train-test split, and the IL1A-TNIP3 94% (=32/34) time-prediction accuracy; the angle-vs-time Pearson rho=0.988 matches the reported 0.98 (within-tol). The split is index-based (no random seed), so these are exactly reproducible. NOTE: the BRIEF's listed code repo (github.com/bytorres/PlosBio2015) is the predecessor MATLAB methods repo, NOT this paper's code; the real code is the cited GitLab repo, which we used. NOT attempted (optional ~20%): YF17D cohorts GSE13699/GSE13485 (Fig 5, 83/73/65% - shipped but not run), Ayasdi TDA (Fig 4, proprietary software - out of scope), Tableau rendering (visualization only), and the separate predicted-vs-actual-time R2=0.99 regression. No possible-fabrication flags: all reproduced values derive from shipped data+code.

💻 Code ↗ 🗄 Data: GSE47122

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 98
    assessed: 2026-06-14 ⛓ 0f65dbbf7eed
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether an automated computational method (Looper) can identify gene pairs whose expression traces looping trajectories in phase space, and whether these loops can be used to predict an individual's position (time/stage of perturbation) along a self-resolving immune response trajectory.

Core claims
  • Looper is a computational method that automatically identifies phase-shifted gene pairs (using SAX symbolic representation) that trace loops in longitudinal expression data. method
  • IL1A and TNIP3 expression are phase-shifted and trace a loop in human monocyte data, consistent with a feedback mechanism where IL1A induction precedes and is later suppressed by TNIP3-mediated inhibition of NF-kB. finding
  • The IL1A-TNIP3 loop predicts time of perturbation in withheld human monocyte test samples with 94% accuracy. finding
  • The IL1A-TNIP3 phase-shift and loop was experimentally validated by qRT-PCR in monocytes from independent human donors. finding
  • CDC20 and IFI44L expression are phase-shifted and trace a loop in YF17D-vaccinated individuals (Montreal cohort), with IFI44L peaking before CDC20. finding
  • The CDC20-IFI44L loop derived from the Montreal cohort predicts perturbation stage with 83% accuracy in withheld Montreal samples, 73% in the independent Lausanne cohort, and 65% in the independent Emory cohort. finding
  • Randomly sampled gene pairs yield poor prediction accuracy compared to Looper-identified looping gene pairs. finding
  • Looper's predictions remain robust (65% to 63% accuracy) when Gaussian noise is added to out-of-sample data. finding
Experimental setups
Assay System Perturbation Readout Platform
microarray gene expression profiling human monocytes from 12 donors (in vitro) sequential treatment with CCL2, LPS, TNFα, IFNγ (inflammatory) then IL10, TGFβ (deactivating) gene expression across 9 time points (0–48h)
qRT-PCR purified human monocytes from 3 independent donors (in vitro) same inflammatory/resolution stimuli, including LPS dose variants (5, 50, 500 ng/ml) IL1A and TNIP3 gene expression over time
microarray gene expression profiling whole blood, YF17D-vaccinated individuals, Montreal cohort (n=15) and Lausanne cohort (n=13) Yellow Fever Vaccine 17D vaccination gene expression at days 0, 3, 7, 10, 14, 28, 60 post-vaccination
microarray gene expression profiling PBMCs, YF17D-vaccinated individuals, Emory cohort (n=25, two trials) Yellow Fever Vaccine 17D vaccination gene expression at days 0, 3, 7 post-vaccination
topological data analysis (TDA) Montreal YF17D cohort gene expression subset (91 genes) Yellow Fever Vaccine 17D vaccination network structure/looping return to baseline expression
in silico sensitivity analysis (Gaussian noise simulation) Emory cohort CDC20-IFI44L expression data computationally added noise change in prediction accuracy
Key results
  • IL1A-TNIP3 loop angle (polar transform) correlates linearly with time in human monocyte training data Pearson's ρ = 0.98
  • Perturbation time predicted in 34 withheld human monocyte test samples using IL1A-TNIP3 loop and KNN (K=3) 94% accuracy, R² = 0.99
  • Phase-shift and loop of IL1A-TNIP3 confirmed experimentally across all 3 independent donors by qRT-PCR
  • CDC20-IFI44L loop angle correlates linearly with time in Montreal cohort training data Pearson's ρ = 0.91
  • Perturbation stage predicted in withheld Montreal cohort test samples (4 individuals, 24 data points) using CDC20-IFI44L loop 83% accuracy
  • Perturbation stage predicted in independent Lausanne cohort using same CDC20-IFI44L loop 73% accuracy
  • Perturbation stage predicted in independent Emory cohort using same CDC20-IFI44L loop 65% accuracy
  • Adding Gaussian noise to Emory cohort data reduced but did not eliminate prediction accuracy 65% to 63%
Key statistics
  • correlation Pearson's ρ = 0.98 (IL1A-TNIP3 loop angle vs time, human monocyte training data)
  • other 94% prediction accuracy, R² = 0.99 (withheld human monocyte test samples, IL1A-TNIP3 loop, KNN prediction)
  • correlation Pearson's ρ = 0.91 (CDC20-IFI44L loop angle vs time, Montreal cohort training data)
  • other 83% prediction accuracy (withheld Montreal cohort test samples (4 individuals), CDC20-IFI44L loop)
  • other 73% prediction accuracy (Lausanne cohort validation using Montreal-derived CDC20-IFI44L loop)
  • other 65% prediction accuracy (Emory cohort (25 individuals) validation using Montreal-derived CDC20-IFI44L loop)
  • other shift from 65% to 63% accuracy (in silico Gaussian noise sensitivity analysis on Emory cohort)
  • count top 0.5% (genes selected by expression range/standard deviation for loop identification and TDA)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper introduces Looper, a computational pipeline applied to publicly available longitudinal microarray datasets (GSE47122, GSE13699, GSE13485). Candidate phase-shifted gene pairs were identified by filtering for top-0.5% expression-range genes and applying Symbolic Aggregate Approximation (SAX) pattern matching; a K-nearest neighbor (K=3) classifier then predicted perturbation stage in withheld samples. Results were reported primarily as classification accuracy (%) and Pearson correlation coefficients (ρ) between polar-coordinate loop angles and time, with R² for linear fits of predicted vs. actual stage.

Replicationmixed Sample sizePer-cohort sample sizes stated: monocyte data 12 donors (9 time points, ~50% random train/test split); Montreal YF17D 15 individuals (11 train / 4 withheld test); Lausanne 13 individuals; Emory 25 individuals across 2 trials. Experimental qRT-PCR validation used 3 independent donors. GroupsPerturbation stages (early/middle/late inflammation-resolution) vs. baseline (pre- and post-perturbation time points), evaluated within and across independent cohorts Pairingmixed Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Pearson correlation (ρ) Angle derived from IL1A-TNIP3 polar transformation vs. ordinal time (ρ=0.98, training data); CDC20-IFI44L angle vs. time in Montreal YF17D cohort (ρ=0.91) 9 time-point means (monocyte training data); 6 time-point means across n=11 Montreal training individuals not stated
K-nearest neighbor classification (K=3) Predicting perturbation stage from IL1A-TNIP3 loop (94% accuracy, 34 monocyte test data points); CDC20-IFI44L loop: Montreal withheld (83%, 24 data points), Lausanne cohort (73%, 33 data points), Emory cohort (65%, 75 data points) 34 (monocyte test); 24 (Montreal withheld); 33 (Lausanne); 75 (Emory) not stated
Symbolic Aggregate Approximation (SAX) pattern matching with T/4 phase shift Identifying phase-shifted gene pair candidates among the top 0.5% expression-range genes in training data top 0.5% of all assayed genes by expression range not stated
Topological Data Analysis (TDA) Assessing overall temporal network structure of Montreal YF17D cohort (GSE13699) on 91 genes (top 0.5% by SD from day-0 baseline) 15 individuals, 91 genes not stated
Leave-one-individual-out cross-validation (LOOCV) Evaluating robustness of CDC20-IFI44L perturbation-stage predictions across all 15 Montreal cohort individuals 15 individuals na
Gaussian noise sensitivity analysis (in silico) Testing robustness of CDC20-IFI44L loop predictions on Emory cohort under noise sampled from gene-pair mean/SD distribution 75 data points (Emory cohort) not stated
Approaches that could also have been used
  • Classification accuracy (%) was the sole reported performance metric for the KNN perturbation-stage predictor across all cohorts
    Could also: Per-class sensitivity/recall, macro-averaged F1 score, Matthews correlation coefficient, or bootstrap confidence intervals around accuracy could also be reported — Accuracy alone can be misleading when stage class sizes differ; F1/MCC capture per-class error balance, and a bootstrap CI would convey statistical uncertainty in the accuracy estimate given the modest sample sizes
  • Multiple candidate gene pairs were evaluated and the highest-accuracy pair (IL1A-TNIP3; CDC20-IFI44L) was selected and reported, without adjustment for the number of pairs examined
    Could also: A permutation-based approach — comparing observed best-pair accuracy to the null distribution of best-pair accuracy from randomly shuffled labels or random gene pairs — could also contextualize the reported accuracy — Selecting the top performer from many candidates inflates apparent performance; a permutation test would estimate how much of the accuracy gain is attributable to the selection process itself
  • A K-nearest neighbor classifier (K=3) in the 2-D loop space was used to assign perturbation stage
    Could also: Ordinal logistic regression, a support vector machine with RBF kernel, or a random forest could also map 2-D loop coordinates to ordered stages — Model-based classifiers make decision boundaries explicit and, in the case of ordinal logistic regression, respect the ordering of perturbation stages; they also allow formal inference on classifier performance and calibrated probability outputs
  • Pearson's ρ was used to summarize the linear relationship between polar-coordinate angle and time using time-point means
    Could also: Spearman's rank correlation or a simple linear regression with a 95% confidence interval for the slope could also be reported — Pearson's ρ assumes linearity and normality; Spearman's ρ is robust to nonlinearity and outliers, and a regression CI would convey uncertainty in the angle–time relationship given the small number of time-point summary values used
  • Genes were pre-filtered by retaining those in the top 0.5% of expression range, using a fixed arbitrary threshold stated as such in the paper
    Could also: A data-driven threshold such as a permutation-derived cutoff, an elbow criterion on the ranked range distribution, or variance-stabilizing normalization followed by FDR-controlled selection could also be applied — A data-driven cutoff would generalize more transparently across datasets with different dynamic ranges and reduce sensitivity to the arbitrary 0.5% choice
  • Phase shifts between gene pairs were detected using SAX discretization with a fixed T/4 offset as the target pattern
    Could also: Pairwise cross-correlation functions or dynamic time warping distances could also quantify the lag between gene expression time series — Cross-correlation provides a continuous lag estimate at all offsets and a natural test statistic for significance; dynamic time warping handles irregular or unequal time intervals; both complement the binary SAX match with graded measures of phase similarity
Software: Metageo (Python module, used for GEO dataset download) · Looper (custom Python computational pipeline, described in paper)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE13485 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE13699 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE13845 GEO in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet
GSE47122 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-30296270

Paper: Rath P, Allen JA, Schneider DS. Predicting position along a looping immune response trajectory. PLoS One 2018;13(10):e0200147. DOI 10.1371/journal.pone.0200147.

Code artifact (corrected)

The BRIEF listed github.com/bytorres/PlosBio2015 — that is the predecessor methods repo (Torres BY et al., "Tracking resilience to infections by mapping disease space", PLOS Biology 2016; MATLAB+R polar transform). It is NOT the code for this paper.

The actual code for this paper is the GitLab repo cited in the article's "Code availability": https://gitlab.com/prath/resilience2018 (Poonam Rath, created 2018-01-10, commit a46cea3664e4284cec6833ee2b6e34dd6c947318). It ships the Python Looper library (looper.py), the SAX implementation (saxpy.py, N. Hoffman, MIT), the metageo GEO-parsing module, the analysis Jupyter notebooks, AND the pre-processed input CSVs. Per BRIEF rule 2 (P16), applying a third-party/own tool to the paper's own data is equally valid — here we run the authors' own shipped pipeline on their shipped data.

Pipeline (in scope)

Method = SAX (Symbolic Aggregate approXimation) phase-shift loop discovery + a K=3 nearest-neighbour time predictor, in Python 2.7 (pandas/numpy/scipy).

Primary dataset: GSE47122 (human monocytes, 12 donors, 9 time points 0–48 h, sequential immune elicitors). The processed matrix is shipped as code/human_mono_gse47122.csv (18859 genes × 60 samples). Reproduction notebook: code/script_for_human_monocyte_FINAL.ipynb.

In-scope reproducible results (deterministic — split is index-based, not random)

id result paper location
C1 18859 input genes; train 26 / test 34 samples Methods; Fig 3
C2 top 0.5% by range → 95 genes (of 18859) Results / Methods
C3 102 phase-shifted gene pairs of 4465 possible (=C(95,2)) Results
C4 IL1A–TNIP3 loop predicts perturbation time at 94% over 34 test samples Fig 3F; abstract
C5 IL1A–TNIP3 angle vs time Pearson ρ≈0.98, R²≈0.99 Fig 3E / S2B

We reproduce C1–C5 by running the shipped looper.py pipeline on the shipped human_mono_gse47122.csv (and shipped FigS2B polar CSV for C5), in a rebuilt Python-2.7 conda environment on «our HPC».

Out of scope (the optional hard ~20%)

  • YF17D vaccination cohorts (GSE13699 Montreal/Lausanne, GSE13485 Emory): 83% / 73% / 65% accuracies, ρ=0.91 (Fig 5). Reproducible in principle from the shipped *.csv + script_for_yellow_fever_FINAL.ipynb, but secondary; attempt only if primary lands cleanly with budget left.
  • Ayasdi 3.0 Topological Data Analysis (Fig 4): out of scope — proprietary commercial software (Ayasdi), not reproducible.
  • Tableau v9.0 figure rendering: out of scope (visualization only, no new number).
  • Wet-lab / GEO raw normalization upstream of the shipped matrix: not attempted (we start from the authors' shipped processed matrix, as the notebook does).

Honesty notes

  • The train/test split uses a deterministic index-order 50% cut per timepoint (index[:split_pt] / index[split_pt:]), so results are reproducible without a random seed — good for auditing.
  • create_composite_profile keeps the Time column among "genes" due to a set('Time') bug in the original; we replicate the original behaviour verbatim rather than fix it, so counts match the paper's own code path.
Figures / tables: Fig 3FFig 3E
C1
Reported
18859 input genes
Reproduced
18859
exact
C1b
Reported
34 test samples
Reproduced
26 train / 34 test
exact
C2
Reported
95 genes (top 0.5% by range)
Reproduced
95
exact
C3
Reported
102 phase-shifted gene pairs of 4465
Reproduced
102 of 4465
exact
C4
Reported
IL1A-TNIP3 predicts time at 94% (34 test)
Reproduced
0.9412 (32/34)
exact
C5
Reported
IL1A-TNIP3 angle-vs-time Pearson rho=0.98
Reproduced
rho=0.9884
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 98/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

A clean 1:1 reproduction: ran the authors' own shipped Looper pipeline (Python 2.7) on their own shipped GSE47122 human-monocyte matrix. Every headline number is exact — 18859 input genes, 95 top-0.5%-range genes, 102 phase-shifted pairs of 4465 (=C(95,2)), the deterministic 26-train/34-test split, and the IL1A-TNIP3 94% (=32/34) time-prediction accuracy — with the angle-vs-time Pearson rho=0.9884 vs reported 0.98 (rounding). The split is index-based with no random seed, so results are exactly reproducible. The registry code link is a false-positive (predecessor MATLAB repo); the agent correctly used the cited GitLab repo. No fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

141.9 k
tokens (I/O) · 8.8 M incl. cache
15 min
runtime · 0.03 CPU-h
1.3 GB
peak RAM
1
HPC jobs
hummel
machine