Systems spatiotemporal dynamics of traumatic brain injury at single-cell resolution reveals humanin as a therapeutic target.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
1:1 REPRODUCED. Arneson et al 2022 (Cell Mol Life Sci), mouse mild-TBI Drop-seq scRNA-seq, GEO GSE180862. The paper is described well enough to reproduce its downstream pipeline-derived results directly from the authors' deposited processed matrices (no FASTQ re-alignment needed; valid own-data downstream reproduction). All four in-scope claims reproduce: total QC-pass cells = 78,895 EXACT (sum of the 3 main TBI-vs-Sham tissue matrices 22800+27082+29013); per-tissue cluster/cell-type counts Blood 8 / Cortex 13 / Hippocampus 17 EXACT (unique CellType labels); QC thresholds confirmed applied in the deposit (min 201 genes/cell, 0 cells <200; max mito 0.1499, 0 cells >=15%); mt-Rnr2 upregulated post-TBI in >=6 cell types in every tissue (Wilcoxon TBI vs Sham at 24h on Seurat LogNormalized expression -> 7/7/7 Bonferroni-significant, >=6 satisfied). Computed on «our HPC» SLURM «job» (numpy 2.4.3 / scipy 1.17.1 / pandas 3.0.2). Dataset profiled (grade A, delivers-promised yes, N exact). NOT attempted: wet-lab assays (behavior, RNAscope, in-vivo humanin) which are out of scope, and the GWAS/MSEA stretch item (needs external sumstats). All grades provisional and human-auditable.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 98assessed: 2026-06-22 ⛓ 60c4ca2f6cde
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe regulatory mechanisms underlying the spatiotemporal (regional and temporal) changes in mild traumatic brain injury (mTBI) pathology remain unclear at the cellular level, and the paper tests whether single-cell profiling across brain regions, timepoints, and blood can reveal cell types, genes, pathways, and cell-cell circuits driving mTBI pathophysiology and amenable to therapeutic targeting.
- ★ Coordinated gene expression patterns across cell types are disrupted and re-organized by mTBI with distinct regional, cellular, and temporal (24-h vs 7-day) specificity. finding
- ★ Astrocytes act as a key regulator of cell-cell coordination following mTBI in both hippocampus and frontal cortex across timepoints. finding
- ★ mt-Rnr2 (encoding the mitochondrial peptide humanin) is a broadly and dynamically dysregulated gene following mTBI, prioritizing it as a therapeutic target. finding
- ★ Humanin treatment in a murine mTBI model reverses TBI-induced cognitive impairment by restoring metabolic pathways within astrocytes. finding
- ★ This is the first multi-tissue (hippocampus, frontal cortex, blood), multi-timepoint (acute/subacute), systems-level single-cell study of mTBI pathophysiology. resource
- ★ Astrocytes and activated microglia are consistently perturbed across brain regions and timepoints, whereas monocytes, T cells, B cells, neurons, endothelial cells, oligodendrocytes, and choroid plexus epithelial cells show spatiotemporally specific sensitivity to mTBI. finding
- Cell-type-specific DEGs induced by mTBI show enrichment for human neurological disease GWAS signals, linking mouse cellular responses to human disease. finding
- ★ Ligand-receptor-based cell-cell communication (CellPhoneDB) increases consistently at the acute phase of mTBI across blood, hippocampus, and frontal cortex and mostly subsides by 7 days. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA sequencing (scRNAseq, Drop-seq) | mouse hippocampus | mild fluid percussion injury (FPI) TBI vs sham | cell-type identification, transcriptomic shifts, DEGs | Drop-seq Tools/dropEst |
| single-cell RNA sequencing (scRNAseq, Drop-seq) | mouse frontal cortex | mild TBI (FPI) vs sham | cell-type identification, transcriptomic shifts, DEGs | Drop-seq Tools/dropEst |
| single-cell RNA sequencing (scRNAseq, Drop-seq) | mouse peripheral blood leukocytes | mild TBI (FPI) vs sham | cell-type identification, transcriptomic shifts, DEGs | Drop-seq Tools/dropEst |
| SVM-based classifier (machine learning on scRNAseq data) | mouse hippocampus, frontal cortex, blood (all cell types) | mild TBI vs sham | classification accuracy of TBI vs control per cell type across 1000 bootstraps | — |
| ligand-receptor co-expression analysis (CellPhoneDB) | mouse hippocampus, frontal cortex, blood (scRNAseq data) | mild TBI vs sham | number of significant inter-cell-type ligand-receptor interactions | CellPhoneDB |
| pathway enrichment analysis of DEGs | mouse hippocampus, frontal cortex, blood cell types | mild TBI vs sham | enriched biological pathways (e.g., apoptosis, mTOR signaling, immune response) | — |
| marker set enrichment analysis (MSEA) against human GWAS | mouse cell-type DEG sets mapped to human disease genes | mild TBI vs sham | enrichment (-log10 FDR) of human neurological disease GWAS genes among cell-type DEGs | Mergeomics |
| in vivo humanin treatment and cognitive behavioral testing | murine mTBI model | humanin administration post-TBI | cognitive impairment reversal and restoration of metabolic pathways in astrocytes | — |
- – Astrocytes were the top-ranked cell type for global transcriptional sensitivity to mTBI in both hippocampus and frontal cortex at 24-h and 7-day timepoints. avg rank 1.00-4.67 (Table 1)
- – CD8+ T cells and Ly6c+ monocytes were top sensitive cell types in peripheral blood at both acute and subacute phases. avg rank 1.67-3.67
- ▲ Coordinated ligand-receptor gene expression across cell types increased at the acute (24-h) phase of mTBI in blood, hippocampus, and frontal cortex, and mostly subsided by 7 days.
- ▲ Humanin treatment reversed cognitive impairment caused by mTBI in mice via restoration of astrocyte metabolic pathways.
- – Apoptosis and mTOR signaling pathways were consistently enriched among DEGs in glial cells (and immune cells) at the acute phase across hippocampus, frontal cortex, and blood.
- – 78,895 single cells were profiled and clustered into 24 distinct cell clusters across three tissues. n=78,895 cells; 24 clusters
- ▼ General cellular transcriptomic response was stronger at 24-h than 7-day post-TBI, particularly in leukocyte populations.
- count 78,895 single cells passed QC (total sequenced across blood, hippocampus, frontal cortex, both timepoints)
- count n = 3 animals per group (per tissue and timepoint group)
- count 8 blood leukocyte clusters, 13 frontal cortex clusters, 17 hippocampal clusters (cell-type clustering per tissue)
- count 7 and 13 neuronal subtypes (subclustering of neurons in hippocampus and frontal cortex respectively)
- other average rank 1.00 (hippocampus, 24-h, astrocytes) (top consensus sensitivity rank across ED, SVM, subset DEG methods (Table 1))
- pvalue adjusted p-value < 0.05 (significance threshold for Euclidean distance transcriptome shift)
- other median classification accuracy across 1000 bootstraps (SVM-based classifier sensitivity ranking method)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used single-cell RNA sequencing across three tissues (hippocampus, frontal cortex, peripheral blood) from mice with TBI versus sham controls at two timepoints (24-h, 7-day), with n = 3 animals per group per tissue/timepoint. Cell-type sensitivity to TBI was quantified using three complementary computational metrics (Euclidean distance from a null distribution, count of differentially expressed genes on subsampled cell numbers, and SVM classifier accuracy over 1000 bootstraps), averaged into a consensus rank. Cell-cell communication was inferred with CellPhoneDB, pathway enrichment was annotated for DEGs, and enrichment of human GWAS signals in DEG sets was assessed via MSEA (Mergeomics), with results reported primarily as fold-changes, adjusted p-values, and FDR-based significance.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Empirical vs. null-distribution comparison of Euclidean distance (logFC of empirical vs. null), yielding an adjusted p-value | Cell-type transcriptomic shift ranking (Fig. 2b, Table 1) | n=3 animals per group per tissue/timepoint; cell-level distances | not stated |
| Differential gene expression testing on subsampled, equalized cell numbers per cluster (specific test not named in this excerpt) | DEG counts per cell type used for sensitivity ranking (Fig. 2, Fig. 4a-c; Supplementary Tables 5, 6) | subsampled to equal cell numbers across cell types per tissue/timepoint; n=3 animals per group | not stated |
| Support vector machine (SVM) classifier accuracy with 1000 bootstraps | Cell-type sensitivity ranking based on classification accuracy of TBI vs. sham cells | not stated (bootstrap resampling of cells within each cell type/tissue/timepoint) | not stated |
| CellPhoneDB ligand-receptor co-expression significance testing | Cell-cell communication analysis (Fig. 3) | not stated | not stated |
| Pathway enrichment analysis with FDR-based significance (−log10(FDR) reported) | Pathway enrichment among DEGs by tissue/timepoint/cell type (Fig. 4d; Supplementary Table 7) | not stated | not stated |
| Marker Set Enrichment Analysis (MSEA) in Mergeomics | Enrichment of human GWAS disease genes in cell-type DEG sets (Fig. 4e) | not stated | not stated |
-
Cell-type sensitivity to TBI was quantified using cell-level Euclidean distance compared to an empirical null distribution, DEG counts, and SVM classification accuracy, averaged into a consensus rank across three methods.↳ Could also: Pseudobulk aggregation per animal (summing/averaging counts per biological replicate) followed by standard differential expression tools such as DESeq2 or edgeR — Because animals, not individual cells, are the biological unit of replication (n=3 per group), pseudobulk approaches are a widely used alternative that directly models the animal-level sample size, which can be a helpful complementary check alongside cell-level metrics.
-
DEGs per cell type were identified using subsampled, count-equalized cell numbers to allow comparable statistical power across cell types.↳ Could also: Mixed-effects or generalized linear mixed models (e.g., as implemented in MAST or glmmSeq) that explicitly include animal identity as a random effect — This approach can account for correlation among cells from the same animal (pseudoreplication) while still using all available cells rather than subsampling, and is a standard option in single-cell differential expression analysis.
-
SVM classifier performance for distinguishing TBI vs. sham cells was assessed using the median classification accuracy across 1000 bootstraps.↳ Could also: A permutation test comparing observed classification accuracy to accuracy obtained after shuffling condition labels, or reporting accuracy with an AUC and associated 95% CI — Label-permutation null distributions are a common way to formally test whether classifier accuracy exceeds chance, and reporting an AUC with confidence interval is a standard way to convey both performance and its uncertainty.
-
Significance across many comparisons (cell types, genes, pathways, GWAS gene sets) was summarized using FDR or 'adjusted p-values' without naming the specific correction method in this excerpt.↳ Could also: Explicitly specifying and applying a named method such as Benjamini-Hochberg FDR or Storey's q-value — Naming the specific multiple-testing procedure (and any assumptions such as independence or positive dependence among tests) is a standard practice that helps readers evaluate how conservatively the family-wise or false discovery rate was controlled.
-
Enrichment of human GWAS disease signals in TBI-associated DEG sets was assessed using MSEA in Mergeomics.↳ Could also: Alternative gene-set/GWAS enrichment tools such as MAGMA or partitioned LD score regression (LDSC) — These are widely used alternative frameworks for linking GWAS summary statistics to gene sets or cell-type expression signatures, and can offer complementary ways to account for linkage disequilibrium structure when testing enrichment.
-
Effect magnitude across analyses was primarily conveyed via log fold change (logFC) values.↳ Could also: Reporting logFC alongside a standard error or confidence interval for the fold-change estimate — Pairing an effect size with an interval estimate is a standard way to convey the precision of the estimated change, complementing the point estimate of logFC.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35951114
Title: Systems spatiotemporal dynamics of traumatic brain injury at single-cell resolution reveals humanin as a therapeutic target. (Arneson et al., Cell Mol Life Sci 2022) PMID: 35951114 · PMCID: PMC9372016 · DOI: 10.1007/s00018-022-04495-9 Data: GEO GSE180862 (Drop-seq scRNA-seq, mouse; SRP329932 / PRJNA749786) Authors' code: github.com/darneson/DropSeq (Snakemake: FASTQ→DGE pipeline only); downstream R-markdown analysis referenced but not the in-repo content. Brief-named code: github.com/broadinstitute/Drop-seq (the upstream Drop-seq tools).
Study design (as reported)
- Mild traumatic brain injury via fluid percussion injury (FPI) vs sham.
- 3 tissues: hippocampus, frontal cortex, peripheral blood leukocytes.
- 2 timepoints: acute (24 h) and subacute (7 day). n = 3 mice/group.
- Drop-seq → STAR-2.5.0c on mm10 → Drop-seq tools v1.13 / dropEst → DGE matrices → Seurat 3.0.2 (Louvain clustering, UMAP, CCA integration) → Wilcoxon DE.
- Separate humanin-treatment cohort (n=6/group) for behavior + reversal analysis.
Key resource for reproduction
GEO ships processed, human-auditable artifacts: per-tissue digital expression
matrices (.mtx), barcodes.tsv, features.tsv, and metaData.tsv (cell-level
annotations: tissue, condition, timepoint, cluster/cell-type labels). This lets us
reproduce the downstream pipeline-derived results directly from the authors'
own deposited matrices, without re-running alignment from FASTQ.
IN SCOPE (pipeline-derived, attempted)
| id | reported result | location | pipeline | how reproduced |
|---|---|---|---|---|
| C1 | 78,895 single cells passed QC | Methods / Results | Seurat QC | count cells across all metaData / mtx columns |
| C2 | clusters per tissue: Blood 8, Cortex 13, Hippocampus 17 | Fig 1b–d | Seurat Louvain | count unique cell-type labels per tissue in metaData |
| C3 | QC thresholds ≥200 genes/cell, <15% mito fraction | Methods | Seurat filter | verify min genes/cell and max mito fraction in matrices |
| C4 | mt-Rnr2 upregulated in ≥6 cell types per tissue, acute (24h) post-TBI | Results / Fig 5 | Wilcoxon DE | recompute Wilcoxon TBI vs sham per cell type; count up-regulated cell types |
OUT OF SCOPE (not pipeline-derived → not attempted)
- Barnes-maze cognitive behavior (Fig 6b) — in-vivo behavioral assay.
- RNAscope spatial validation of mt-Rnr1/2, mt-Cytb (Fig 7) — wet-lab imaging.
- Humanin in-vivo treatment / cognitive reversal — wet-lab intervention.
DEEPER / STRETCH (attempt after core floor)
- GWAS enrichment (MSEA via Mergeomics, Fig 4e) — pipeline-derived but needs external GWAS sumstats + Mergeomics; attempt if feasible.
- Re-run FASTQ→DGE alignment for one sample to spot-check the authors' matrix — heavy «our HPC» compute; only if core reproduces and time permits.
Compute plan
- Downloads (GEO matrices) → «our HPC» front node → «infra» work dir.
- Analysis (load mtx, count, Wilcoxon) → «our HPC» SLURM (or front for light steps), R/Seurat or Python/scanpy+scipy. Results (small values) → «host» dataset folder.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
1:1 reproduction. All four in-scope claims for Arneson et al. 2022 recompute exactly from the authors' own GSE180862 Drop-seq matrices: 78,895 QC-pass cells (22800+27082+29013), Blood/Cortex/Hippocampus clusters 8/13/17, applied QC thresholds (>=200 genes, <15% mito), and mt-Rnr2 upregulation in 7/7/7 cell types post-TBI (>=6 met). Deviation is nil — the sole non-exact item (min 201 genes vs the >=200 cutoff) merely confirms the deposited matrices were already filtered. The central molecular claim holds; the wet-lab therapeutic arm is out of scope by design, not a derivability concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.