Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Integrative network modeling reveals mechanisms underlying T cell exhaustion.

Sci Rep · 2020
L1 59/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
59/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 19% of all assessed papers rank 925 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH + 1:1 on the primary self-contained result. Code-artifact correction: the paper's actual analysis code is github.com/hamid-bolouri/TCE (authors' R scripts), NOT ialbert/booleannet (a text-mining false positive: booleannet is only a logic simulator the authors cite as explored-and-abandoned). Reproduced C1 (logic-simulation robustness) on «our HPC» (SLURM 2176721, R 4.5.2, «infra») by running the authors' exact Boolean update loop (simRandTex.R + TexSimCmds.R, 22 rules, randomized update order) for 20 RNG seeds x 10000 runs: the reported single stochastic value 28/10000 lies inside the reproduced distribution [13,30] (mean 22.45); one seed reproduced exactly 28. Honest 1:1 at the distributional level (within-tol). NOT attempted (80/20): C2 concordance FDR 14.8% (in-scope and feasible but needs multi-GEO download + limma per dataset), C3 49 Mfuzz clusters (count is human-tuned, not deterministic), C4 exact 64-node count (repo edge table shows 57 raw labels; '64' depends on composite-node expansion not specified in the table). Out of scope entirely: wet-lab qRT-PCR, manual network curation, manual motif curation (4 GUI tools), Berkeley-Madonna ODE model (proprietary, parameters 'arbitrary for illustration'). No fabrication signal on C1.

💻 Code ↗ 🗄 Data: GSE41867

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 59
    assessed: 2026-06-15 ⛓ 9d67224af613
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The authors hypothesize that CD8+ T cell exhaustion (TCE) arises not from dysregulation of a single gene but from changes in the relative timing of largely shared gene regulatory interactions, such that differences between acute and chronic responses lie in timing rather than gene identity.

Core claims
  • An integrative, literature-curated and data-driven gene regulatory network underlies CD8+ T cell exhaustion and accurately captures expression states in chronic infection and tumor settings. resource
  • CD8+ T cells undergo two major state transitions after stimulation: an early pro-memory/proliferative (PP) state and a later effector/irreversibly exhausted (EE) state. finding
  • The duration cells spend in the early PP state is a fixed, inherent property of the network structure (set by overlapping incoherent feed-forward loops), independent of stimulation duration, whereas differentiation is prolonged with stimulation. mechanism
  • Overlapping incoherent feed-forward loops (iFFLs) fix the duration of FOXO1/TCF-1 expression and time the activation of late EE genes (TBET, ZEB2, BLIMP-1) and FAS/FASL signaling. mechanism
  • Widespread mutual inhibition (18/23 inhibitory interactions, 78%) between early PP and late EE genes enables mutually exclusive activation states. mechanism
  • Multiple positive feedback loops reinforce and maintain the late effector/irreversible exhaustion (EE) state, while inhibitory receptors implement overlapping negative feedback on TCR signaling. mechanism
  • Topology and simulation modeling predict the extent to which each node drives cells toward exhaustion, enabling prediction of drug effects. method
  • Drug-induced interference with EZH2 function increases the proportion of pro-memory/proliferative cells in the early days post-activation. finding
Experimental setups
Assay System Perturbation Readout Platform
Manual literature curation / network reconstruction CD8+ T cells (literature-derived) none regulatory interactions (nodes and edges) underlying TCE
Gene expression meta-analysis (published transcriptomic datasets) CD8+ T cells, acute and chronic stimulation (incl. LCMV infection, tumor-infiltrating lymphocytes) Ag stimulation (acute vs chronic) fold change in gene expression vs naive; consistency with interaction sense
Principal component analysis / time-course expression clustering CD8+ T cells, acute vs chronic stimulation duration expression trajectories and cluster membership switching
Logic / simulation network modeling in silico TCE network (64 nodes, 120 interactions) node activation delays / in silico perturbation recapitulation of gene expression state sequences; predicted exhaustion drive per node
Single-cell RNA-seq (re-analysis of published data) CD8+ T cells none mutually exclusive BCL6 vs BLIMP-1 expression
Experimental EZH2 drug-inhibition assay activated CD8+ T cells drug-induced EZH2 inhibition proportion of pro-memory/proliferative cells in early days post-activation
Key results
  • Only 17 interactions (~3.6%) in the network lacked supportive expression data at ~15% FDR. 17 interactions (~3.6%)
  • Acute and exhausted/chronic expression profiles showed similar PCA trajectories for TCE network genes.
  • 18 of 23 inhibitory interactions (excluding inhibitory-receptor feedback) are between early PP and late EE genes. 18/23 (78%)
  • Reduced TCE network comprises 64 nodes and 120 interactions. 64 nodes, 120 interactions
  • EZH2/PRC2 activity peaks ~4 days post-infection and AKT activation peaks ~day 5 post-infection, supporting timed iFFL repression of FOXO1/TCF-1. ~4 days / ~5 days
  • EZH2 drug inhibition increased the proportion of pro-memory/proliferative cells early post-activation, as predicted by the model.
  • CD200-R1 was the only TP53-related regulatory gene showing consistently large fold-change differences between chronic and acute settings across datasets.
Key statistics
  • count 17 interactions (~3.6%) lacked supportive expression data (network interactions without supportive expression data at ~15% FDR)
  • other false discovery rate ~15% (estimated via permutation testing for interaction support)
  • count 18 of 23 (78%) (inhibitory interactions between early PP and late EE genes)
  • count 64 nodes and 120 interactions (size of reduced literature-based TCE network)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is primarily a computational systems biology study that constructs a literature-curated gene regulatory network for CD8+ T cell exhaustion and validates it against multiple published gene expression datasets by assessing whether fold changes in source and target genes are directionally consistent with each reported interaction, using permutation-based FDR estimation. Principal component analysis and time-course clustering were used to compare expression trajectories across acute and chronic stimulation datasets. Network topology analysis identified functional motifs, and a mathematical simulation model was derived to predict cellular state transitions; experimental validation of one EZH2 inhibition prediction was also performed, though statistical details for those experiments are not present in the provided text excerpt.

Replicationunclear Sample sizeThe study reanalyzes multiple published gene expression datasets; individual dataset sample sizes are not described in the provided excerpt. Sample sizes for the experimental EZH2 inhibition validation are not present in the provided text. GroupsAcute vs. chronic/exhausted CD8+ T cell stimulation across multiple published datasets; naïve vs. activated T cell expression states Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionPermutation-based FDR estimation (~15% threshold)
Statistical tests used
Test Applied to n Assumptions
Permutation testing for false discovery rate (FDR) estimation Assessment of whether fold changes in source and target genes were directionally consistent with the reported activating or inhibitory sense of each network interaction across published datasets 17 interactions (~3.6% of total network interactions) lacked supportive expression data at FDR ~15%; exact total interaction count in full network not stated in excerpt not stated
Principal component analysis (PCA) Comparison of gene expression trajectories of TCE network genes between acute and chronic/exhausted stimulation published datasets null na
Time-course gene expression clustering Identification of transcription factors and signaling genes that switch cluster membership (i.e., timing of expression) between acute and chronic stimulation settings null not stated
Fold change calculation relative to naïve state Quantification of expression-level support for each regulatory interaction across multiple published datasets null not stated
Approaches that could also have been used
  • Network interaction support was summarized by a single aggregate FDR (~15%) estimated via permutation, characterizing how many interactions lacked directionally consistent fold changes across datasets
    Could also: A binomial or Fisher's exact test could also be applied per interaction to test whether directional consistency across datasets exceeds chance expectation, with Benjamini-Hochberg correction across all interactions — Per-interaction tests would provide individual uncertainty estimates rather than a single aggregate FDR, allowing the reader to distinguish interactions with strong multi-dataset support from those narrowly passing the threshold
  • PCA was used to compare gene expression trajectories between acute and chronic stimulation datasets
    Could also: Pseudotime trajectory inference methods (e.g., Monocle, Slingshot) or non-linear dimensionality reduction (UMAP, t-SNE) could also be applied to characterize and quantify state transitions along the activation-to-exhaustion continuum — Non-linear methods and pseudotime tools can capture curved or branching trajectories in high-dimensional expression space that PCA's linear projection may compress; they also provide a continuous ordering of cells along a differentiation axis
  • Time-course gene expression clustering was used to assign genes to early (PP) or late (EE) activity classes based on when expression changes occur
    Could also: Soft or probabilistic time-series clustering (e.g., Mfuzz, Gaussian process regression-based approaches) could also partition time-course profiles while providing graded membership scores — Soft clustering retains information about genes with intermediate or ambiguous timing, reflecting genuine biological uncertainty in a binary PP/EE assignment
  • Multiple published datasets were each analyzed separately to validate network interactions, with cross-dataset consistency described qualitatively
    Could also: A formal meta-analytic framework (e.g., random-effects meta-analysis of standardized effect sizes across datasets) could also synthesize quantitative evidence across studies — Meta-analysis yields a pooled estimate with a confidence interval and an explicit heterogeneity statistic (e.g., I²), making cross-dataset consistency or inconsistency transparent and quantifiable
  • Fold changes relative to the naïve state were used as the primary measure to assess expression-level support for each interaction direction
    Could also: Moderated differential expression methods (e.g., limma with empirical Bayes shrinkage, or DESeq2 for count data) could also be applied to the same datasets to generate shrunken fold-change estimates and adjusted p-values per gene per comparison — Shrinkage-based methods stabilize fold-change estimates for genes with high variance or low expression, which is especially relevant when pooling heterogeneous published datasets of varying size
  • The mathematical simulation model was used to predict qualitative drug effects, and one prediction (EZH2 inhibition) was experimentally tested
    Could also: Sensitivity analysis (e.g., parameter sweeps, Latin hypercube sampling, or Morris screening) could also be reported alongside simulation results to characterize how conclusions depend on specific parameter values — Sensitivity analysis quantifies the robustness of model predictions to parameter uncertainty, which is informative when network parameters are estimated from heterogeneous literature sources rather than directly measured in a single system
Software: Not stated in provided text

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
23
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32024856

Paper: Bolouri H, Young M, Beilke J, Johnson R, Fox B, Huang L, Santini CC, Hill CM. "Integrative network modeling reveals mechanisms underlying T cell exhaustion." Sci Rep 10:1915 (2020). DOI 10.1038/s41598-020-58600-8 · PMID 32024856 · PMCID PMC7002445.

Code artifact (corrected)

  • Authors' own code: https://github.com/hamid-bolouri/TCE (R scripts, public, master branch, license file present). This is the repo named in the paper's Code availability statement: "All R scripts used to carry out the analysis in this manuscript are freely available at GitHub (https://github.com/hamid-bolouri/TCE)."
  • The brief's code_url (ialbert/booleannet) is a text-mining false attribution: booleannet is only cited (ref 55) as one of two logic simulators the authors initially explored and then abandoned ("logic simulators are designed to explore network steady states ... To enable more flexible exploration ... we implemented our network models as a series of logic statements"). It is not the analysis code. Reproduction targets the real repo hamid-bolouri/TCE.

Data

  • All input data are previously-published public GEO microarray/RNA-seq sets (Data availability: "All data analyzed ... previously published and available publicly, as described in the Methods"): chronic infection — GSE41867 (Doering), GSE9650 (Wherry), GSE76279 (Leong), GSE74148 (He), GSE84105 (Im); cancer — GSE24536, GSE84072, GSE89307, GSE98638, …
  • The repo additionally bundles Scheitinger2017_exp.RData (Schietinger lab data) so the Schietinger-related scripts are self-contained.
  • The logic-simulation scripts need no external data (the model is a hard-coded set of Boolean logic rules).

Pipeline-derived results — IN SCOPE

# Reported result Pipeline / script Reproducibility
C1 Logic-sim robustness: model reaches the Tex steady state in "28 randomized update runs out of 10,000" (Methods §Logic simulation) simRandTex.R + TexSimCmds.R (base R; randomized update-order Boolean sim) HIGH — self-contained, no data needed. Stochastic → compare distribution/order-of-magnitude, not exact integer. PRIMARY TARGET.
C2 Concordance FDR: "Concordant edges occurred by chance in 14.8% of randomized controls (FDR ~15%)" from 500,000 randomized edges (Methods §Network evaluation by concordance matching) concordanceMatching3_25May2018.R (needs network edges + per-replicate fold changes from GEO; limma normalization) MEDIUM — needs GEO download + limma per dataset; randomization part is self-contained given concordance values. Stretch goal.
C3 "49 clusters met these requirements across all datasets" (Mfuzz soft clustering) mfuzz_clustering.R LOW — criterion is human-in-the-loop ("varied to find the number ... at least 1 cluster with few members ..."); not a deterministic output. Not attempted as 1:1.
C4 Network size: "64-node reduced/simplified TCE network" twoStateNetworkInteractionsForR.txt (structural count) LOW-as-1:1 — repo edge table = 120 edges / 57 raw node-labels; "64" depends on how composite/complex nodes are expanded. Reported as a structural note, not a clean match.

Out of scope (wet-lab / manual / external tool — NOT attempted)

  • In vitro T cell cultures, RNA isolation, qRT-PCR (Methods) — wet-lab.
  • Literature-based network construction — manual curation by co-authors.
  • Network motif analysis — manual curation combining 4 external GUI tools (mfinder, Cytoscape Motif-Discovery, FANMOD, MAVisto); not a scripted pipeline.
  • ODE model (Fig. 5) — run in Berkeley Madonna (proprietary GUI), parameters "selected arbitrarily for purely illustrative purposes" → no quantitative claim to match.
  • Network visualization (Cytoscape), PCA state-tracking plots — qualitative figures.

Plan

Primary 1:1 = C1 on «our HPC» (clone repo on «infra», run author

C1
Reported
28 / 10000 randomized update-order runs reach the Tex steady state (0.28%)
Reproduced
20 seeds x 10000 runs: counts 13-30 (mean 22.45, median 22, sd 4.24); reported 28 inside range; seed=14 = exactly 28
within tolerance
C2
Reported
14.8% concordance FDR from 500000 randomized control edges
Reproduced
not attempted (needs GEO microarray download + limma normalization; deferred under 80/20)
partial
C3
Reported
49 Mfuzz soft-clusters across all datasets
Reproduced
not attempted (cluster count is human-tuned, not a deterministic pipeline output)
partial
C4
Reported
64-node reduced/simplified TCE network
Reproduced
repo edge table has 57 distinct labels / 120 edges (90 promotes, 30 inhibits)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 59/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

The single fully self-contained quantitative result (C1, logic-sim robustness 28/10000) reproduces at the distributional level from the authors' own code — [13,30], mean 22.45, with seed 14 giving exactly 28 — so there is no fabrication or derivability concern on the tested claim. The remaining claims are either deferred under our 80/20 scope (C2 concordance FDR, C3 Mfuzz clusters) or only structurally checkable (C4: shipped table gives 57 raw labels vs the reported 64, explainable by unspecified composite-node expansion) — these are on our methodology/data-availability side, not authors' defects. A notable artifact correction: the registry code_url (booleannet) was a text-mining false positive, fixed to hamid-bolouri/TCE. Overall a solid but partial reproduction: 1:1 where self-contained, limited confirmation of the paper's broader conclusions.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

111.5 k
tokens (I/O) · 8.6 M incl. cache
15 min
runtime · 0.05 CPU-h
0.1 GB
peak RAM
1
HPC jobs
hummel
machine