Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Systems spatiotemporal dynamics of traumatic brain injury at single-cell resolution reveals humanin as a therapeutic target.

Cell Mol Life Sci · 2022
L1 98/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
98/100
Reproducibility score
1.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 93% of all assessed papers rank 65 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

1:1 REPRODUCED. Arneson et al 2022 (Cell Mol Life Sci), mouse mild-TBI Drop-seq scRNA-seq, GEO GSE180862. The paper is described well enough to reproduce its downstream pipeline-derived results directly from the authors' deposited processed matrices (no FASTQ re-alignment needed; valid own-data downstream reproduction). All four in-scope claims reproduce: total QC-pass cells = 78,895 EXACT (sum of the 3 main TBI-vs-Sham tissue matrices 22800+27082+29013); per-tissue cluster/cell-type counts Blood 8 / Cortex 13 / Hippocampus 17 EXACT (unique CellType labels); QC thresholds confirmed applied in the deposit (min 201 genes/cell, 0 cells <200; max mito 0.1499, 0 cells >=15%); mt-Rnr2 upregulated post-TBI in >=6 cell types in every tissue (Wilcoxon TBI vs Sham at 24h on Seurat LogNormalized expression -> 7/7/7 Bonferroni-significant, >=6 satisfied). Computed on «our HPC» SLURM «job» (numpy 2.4.3 / scipy 1.17.1 / pandas 3.0.2). Dataset profiled (grade A, delivers-promised yes, N exact). NOT attempted: wet-lab assays (behavior, RNAscope, in-vivo humanin) which are out of scope, and the GWAS/MSEA stretch item (needs external sumstats). All grades provisional and human-auditable.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 98
    assessed: 2026-06-22 ⛓ 60c4ca2f6cde
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The regulatory mechanisms underlying the spatiotemporal (regional and temporal) changes in mild traumatic brain injury (mTBI) pathology remain unclear at the cellular level, and the paper tests whether single-cell profiling across brain regions, timepoints, and blood can reveal cell types, genes, pathways, and cell-cell circuits driving mTBI pathophysiology and amenable to therapeutic targeting.

Core claims
  • Coordinated gene expression patterns across cell types are disrupted and re-organized by mTBI with distinct regional, cellular, and temporal (24-h vs 7-day) specificity. finding
  • Astrocytes act as a key regulator of cell-cell coordination following mTBI in both hippocampus and frontal cortex across timepoints. finding
  • mt-Rnr2 (encoding the mitochondrial peptide humanin) is a broadly and dynamically dysregulated gene following mTBI, prioritizing it as a therapeutic target. finding
  • Humanin treatment in a murine mTBI model reverses TBI-induced cognitive impairment by restoring metabolic pathways within astrocytes. finding
  • This is the first multi-tissue (hippocampus, frontal cortex, blood), multi-timepoint (acute/subacute), systems-level single-cell study of mTBI pathophysiology. resource
  • Astrocytes and activated microglia are consistently perturbed across brain regions and timepoints, whereas monocytes, T cells, B cells, neurons, endothelial cells, oligodendrocytes, and choroid plexus epithelial cells show spatiotemporally specific sensitivity to mTBI. finding
  • Cell-type-specific DEGs induced by mTBI show enrichment for human neurological disease GWAS signals, linking mouse cellular responses to human disease. finding
  • Ligand-receptor-based cell-cell communication (CellPhoneDB) increases consistently at the acute phase of mTBI across blood, hippocampus, and frontal cortex and mostly subsides by 7 days. finding
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing (scRNAseq, Drop-seq) mouse hippocampus mild fluid percussion injury (FPI) TBI vs sham cell-type identification, transcriptomic shifts, DEGs Drop-seq Tools/dropEst
single-cell RNA sequencing (scRNAseq, Drop-seq) mouse frontal cortex mild TBI (FPI) vs sham cell-type identification, transcriptomic shifts, DEGs Drop-seq Tools/dropEst
single-cell RNA sequencing (scRNAseq, Drop-seq) mouse peripheral blood leukocytes mild TBI (FPI) vs sham cell-type identification, transcriptomic shifts, DEGs Drop-seq Tools/dropEst
SVM-based classifier (machine learning on scRNAseq data) mouse hippocampus, frontal cortex, blood (all cell types) mild TBI vs sham classification accuracy of TBI vs control per cell type across 1000 bootstraps
ligand-receptor co-expression analysis (CellPhoneDB) mouse hippocampus, frontal cortex, blood (scRNAseq data) mild TBI vs sham number of significant inter-cell-type ligand-receptor interactions CellPhoneDB
pathway enrichment analysis of DEGs mouse hippocampus, frontal cortex, blood cell types mild TBI vs sham enriched biological pathways (e.g., apoptosis, mTOR signaling, immune response)
marker set enrichment analysis (MSEA) against human GWAS mouse cell-type DEG sets mapped to human disease genes mild TBI vs sham enrichment (-log10 FDR) of human neurological disease GWAS genes among cell-type DEGs Mergeomics
in vivo humanin treatment and cognitive behavioral testing murine mTBI model humanin administration post-TBI cognitive impairment reversal and restoration of metabolic pathways in astrocytes
Key results
  • Astrocytes were the top-ranked cell type for global transcriptional sensitivity to mTBI in both hippocampus and frontal cortex at 24-h and 7-day timepoints. avg rank 1.00-4.67 (Table 1)
  • CD8+ T cells and Ly6c+ monocytes were top sensitive cell types in peripheral blood at both acute and subacute phases. avg rank 1.67-3.67
  • Coordinated ligand-receptor gene expression across cell types increased at the acute (24-h) phase of mTBI in blood, hippocampus, and frontal cortex, and mostly subsided by 7 days.
  • Humanin treatment reversed cognitive impairment caused by mTBI in mice via restoration of astrocyte metabolic pathways.
  • Apoptosis and mTOR signaling pathways were consistently enriched among DEGs in glial cells (and immune cells) at the acute phase across hippocampus, frontal cortex, and blood.
  • 78,895 single cells were profiled and clustered into 24 distinct cell clusters across three tissues. n=78,895 cells; 24 clusters
  • General cellular transcriptomic response was stronger at 24-h than 7-day post-TBI, particularly in leukocyte populations.
Key statistics
  • count 78,895 single cells passed QC (total sequenced across blood, hippocampus, frontal cortex, both timepoints)
  • count n = 3 animals per group (per tissue and timepoint group)
  • count 8 blood leukocyte clusters, 13 frontal cortex clusters, 17 hippocampal clusters (cell-type clustering per tissue)
  • count 7 and 13 neuronal subtypes (subclustering of neurons in hippocampus and frontal cortex respectively)
  • other average rank 1.00 (hippocampus, 24-h, astrocytes) (top consensus sensitivity rank across ED, SVM, subset DEG methods (Table 1))
  • pvalue adjusted p-value < 0.05 (significance threshold for Euclidean distance transcriptome shift)
  • other median classification accuracy across 1000 bootstraps (SVM-based classifier sensitivity ranking method)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used single-cell RNA sequencing across three tissues (hippocampus, frontal cortex, peripheral blood) from mice with TBI versus sham controls at two timepoints (24-h, 7-day), with n = 3 animals per group per tissue/timepoint. Cell-type sensitivity to TBI was quantified using three complementary computational metrics (Euclidean distance from a null distribution, count of differentially expressed genes on subsampled cell numbers, and SVM classifier accuracy over 1000 bootstraps), averaged into a consensus rank. Cell-cell communication was inferred with CellPhoneDB, pathway enrichment was annotated for DEGs, and enrichment of human GWAS signals in DEG sets was assessed via MSEA (Mergeomics), with results reported primarily as fold-changes, adjusted p-values, and FDR-based significance.

Replicationmixed Sample sizen = 3 animals per group for each tissue and timepoint combination; single-cell data then provides many cells per animal GroupsTBI vs. sham control, across 3 tissues (hippocampus, frontal cortex, blood) and 2 timepoints (24-h, 7-day) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnot explicitly named in the provided text; referred to generically as 'adjusted p-value' and FDR (false discovery rate)
Statistical tests used
Test Applied to n Assumptions
Empirical vs. null-distribution comparison of Euclidean distance (logFC of empirical vs. null), yielding an adjusted p-value Cell-type transcriptomic shift ranking (Fig. 2b, Table 1) n=3 animals per group per tissue/timepoint; cell-level distances not stated
Differential gene expression testing on subsampled, equalized cell numbers per cluster (specific test not named in this excerpt) DEG counts per cell type used for sensitivity ranking (Fig. 2, Fig. 4a-c; Supplementary Tables 5, 6) subsampled to equal cell numbers across cell types per tissue/timepoint; n=3 animals per group not stated
Support vector machine (SVM) classifier accuracy with 1000 bootstraps Cell-type sensitivity ranking based on classification accuracy of TBI vs. sham cells not stated (bootstrap resampling of cells within each cell type/tissue/timepoint) not stated
CellPhoneDB ligand-receptor co-expression significance testing Cell-cell communication analysis (Fig. 3) not stated not stated
Pathway enrichment analysis with FDR-based significance (−log10(FDR) reported) Pathway enrichment among DEGs by tissue/timepoint/cell type (Fig. 4d; Supplementary Table 7) not stated not stated
Marker Set Enrichment Analysis (MSEA) in Mergeomics Enrichment of human GWAS disease genes in cell-type DEG sets (Fig. 4e) not stated not stated
Approaches that could also have been used
  • Cell-type sensitivity to TBI was quantified using cell-level Euclidean distance compared to an empirical null distribution, DEG counts, and SVM classification accuracy, averaged into a consensus rank across three methods.
    Could also: Pseudobulk aggregation per animal (summing/averaging counts per biological replicate) followed by standard differential expression tools such as DESeq2 or edgeR — Because animals, not individual cells, are the biological unit of replication (n=3 per group), pseudobulk approaches are a widely used alternative that directly models the animal-level sample size, which can be a helpful complementary check alongside cell-level metrics.
  • DEGs per cell type were identified using subsampled, count-equalized cell numbers to allow comparable statistical power across cell types.
    Could also: Mixed-effects or generalized linear mixed models (e.g., as implemented in MAST or glmmSeq) that explicitly include animal identity as a random effect — This approach can account for correlation among cells from the same animal (pseudoreplication) while still using all available cells rather than subsampling, and is a standard option in single-cell differential expression analysis.
  • SVM classifier performance for distinguishing TBI vs. sham cells was assessed using the median classification accuracy across 1000 bootstraps.
    Could also: A permutation test comparing observed classification accuracy to accuracy obtained after shuffling condition labels, or reporting accuracy with an AUC and associated 95% CI — Label-permutation null distributions are a common way to formally test whether classifier accuracy exceeds chance, and reporting an AUC with confidence interval is a standard way to convey both performance and its uncertainty.
  • Significance across many comparisons (cell types, genes, pathways, GWAS gene sets) was summarized using FDR or 'adjusted p-values' without naming the specific correction method in this excerpt.
    Could also: Explicitly specifying and applying a named method such as Benjamini-Hochberg FDR or Storey's q-value — Naming the specific multiple-testing procedure (and any assumptions such as independence or positive dependence among tests) is a standard practice that helps readers evaluate how conservatively the family-wise or false discovery rate was controlled.
  • Enrichment of human GWAS disease signals in TBI-associated DEG sets was assessed using MSEA in Mergeomics.
    Could also: Alternative gene-set/GWAS enrichment tools such as MAGMA or partitioned LD score regression (LDSC) — These are widely used alternative frameworks for linking GWAS summary statistics to gene sets or cell-type expression signatures, and can offer complementary ways to account for linkage disequilibrium structure when testing enrichment.
  • Effect magnitude across analyses was primarily conveyed via log fold change (logFC) values.
    Could also: Reporting logFC alongside a standard error or confidence interval for the fold-change estimate — Pairing an effect size with an interval estimate is a standard way to convey the precision of the estimated change, complementing the point estimate of logFC.
Software: Drop-seq Tools · dropEst · Snakemake · UMAP · Louvain clustering · CellPhoneDB

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35951114

Title: Systems spatiotemporal dynamics of traumatic brain injury at single-cell resolution reveals humanin as a therapeutic target. (Arneson et al., Cell Mol Life Sci 2022) PMID: 35951114 · PMCID: PMC9372016 · DOI: 10.1007/s00018-022-04495-9 Data: GEO GSE180862 (Drop-seq scRNA-seq, mouse; SRP329932 / PRJNA749786) Authors' code: github.com/darneson/DropSeq (Snakemake: FASTQ→DGE pipeline only); downstream R-markdown analysis referenced but not the in-repo content. Brief-named code: github.com/broadinstitute/Drop-seq (the upstream Drop-seq tools).

Study design (as reported)

  • Mild traumatic brain injury via fluid percussion injury (FPI) vs sham.
  • 3 tissues: hippocampus, frontal cortex, peripheral blood leukocytes.
  • 2 timepoints: acute (24 h) and subacute (7 day). n = 3 mice/group.
  • Drop-seq → STAR-2.5.0c on mm10 → Drop-seq tools v1.13 / dropEst → DGE matrices → Seurat 3.0.2 (Louvain clustering, UMAP, CCA integration) → Wilcoxon DE.
  • Separate humanin-treatment cohort (n=6/group) for behavior + reversal analysis.

Key resource for reproduction

GEO ships processed, human-auditable artifacts: per-tissue digital expression matrices (.mtx), barcodes.tsv, features.tsv, and metaData.tsv (cell-level annotations: tissue, condition, timepoint, cluster/cell-type labels). This lets us reproduce the downstream pipeline-derived results directly from the authors' own deposited matrices, without re-running alignment from FASTQ.

IN SCOPE (pipeline-derived, attempted)

id reported result location pipeline how reproduced
C1 78,895 single cells passed QC Methods / Results Seurat QC count cells across all metaData / mtx columns
C2 clusters per tissue: Blood 8, Cortex 13, Hippocampus 17 Fig 1b–d Seurat Louvain count unique cell-type labels per tissue in metaData
C3 QC thresholds ≥200 genes/cell, <15% mito fraction Methods Seurat filter verify min genes/cell and max mito fraction in matrices
C4 mt-Rnr2 upregulated in ≥6 cell types per tissue, acute (24h) post-TBI Results / Fig 5 Wilcoxon DE recompute Wilcoxon TBI vs sham per cell type; count up-regulated cell types

OUT OF SCOPE (not pipeline-derived → not attempted)

  • Barnes-maze cognitive behavior (Fig 6b) — in-vivo behavioral assay.
  • RNAscope spatial validation of mt-Rnr1/2, mt-Cytb (Fig 7) — wet-lab imaging.
  • Humanin in-vivo treatment / cognitive reversal — wet-lab intervention.

DEEPER / STRETCH (attempt after core floor)

  • GWAS enrichment (MSEA via Mergeomics, Fig 4e) — pipeline-derived but needs external GWAS sumstats + Mergeomics; attempt if feasible.
  • Re-run FASTQ→DGE alignment for one sample to spot-check the authors' matrix — heavy «our HPC» compute; only if core reproduces and time permits.

Compute plan

  • Downloads (GEO matrices) → «our HPC» front node → «infra» work dir.
  • Analysis (load mtx, count, Wilcoxon) → «our HPC» SLURM (or front for light steps), R/Seurat or Python/scanpy+scipy. Results (small values) → «host» dataset folder.
Figures / tables: Fig 1bFig 1cFig 1dFig 5
C1
Reported
78,895 cells passed QC
Reproduced
78895 (Blood 22800 + Cortex 27082 + Hippocampus 29013)
exact
C2a
Reported
Blood: 8 clusters/cell types
Reproduced
8
exact
C2b
Reported
Cortex: 13 clusters/cell types
Reproduced
13
exact
C2c
Reported
Hippocampus: 17 clusters/cell types
Reproduced
17
exact
C3a
Reported
QC: >=200 genes/cell
Reproduced
min 201, 0 cells <200 (all tissues)
within tolerance
C3b
Reported
QC: <15% mito
Reproduced
max 0.1499, 0 cells >=15% (all tissues)
exact
C4
Reported
mt-Rnr2 up in >=6 cell types/tissue (acute 24h, TBI vs Sham)
Reproduced
Blood 7, Cortex 7, Hippocampus 7 (Bonferroni adj p<0.05 & log2FC>0); raw p<0.05: 7/11/13
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 98/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

1:1 reproduction. All four in-scope claims for Arneson et al. 2022 recompute exactly from the authors' own GSE180862 Drop-seq matrices: 78,895 QC-pass cells (22800+27082+29013), Blood/Cortex/Hippocampus clusters 8/13/17, applied QC thresholds (>=200 genes, <15% mito), and mt-Rnr2 upregulation in 7/7/7 cell types post-TBI (>=6 met). Deviation is nil — the sole non-exact item (min 201 genes vs the >=200 cutoff) merely confirms the deposited matrices were already filtered. The central molecular claim holds; the wet-lab therapeutic arm is out of scope by design, not a derivability concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

155.1 k
tokens (I/O) · 9.9 M incl. cache
70 min
runtime · 0 CPU-h
1.2 GB
peak RAM
1
HPC jobs
hummel
machine