Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Protocol for the generation of single-nuclei RNA-seq libraries and quantification of heterogeneous cell types activated during social interaction.

STAR Protoc · 2024
L1 78/100 PQI 84
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH + reproduced ~1:1 on the documented core. This STAR Protocols methods paper (Walker & Frost) documents a snRNA-seq pipeline: CellRanger v7.0.1/GRCm39 -> Seurat 4.4.0 (QC MT<5 & nFeature 800-6000, SCTransform regressing MT, 30 PCs, FindClusters res 0.8, UMAP) -> marker-based cell typing. GEO GSE269499 ships the authors' own CellRanger count matrices, so we started the Seurat pipeline from the exact input (no CellRanger re-run). On «our HPC» (Seurat 4.4.0, R 4.3.3) we merged the 14 deposited WT mPFC samples (111,679 nuclei), applied the documented QC (103,153 nuclei retained, 92.4%), and ran SCTransform+clustering verbatim. Result: 47 clusters at res 0.8 (vs 44 code-implied) and recovery of all 22 published PFC cell types' canonical markers as cluster-specific signals (excitatory layer subtypes via Slc17a7/Cux2/Etv1/Syt6/Oprk1/Car3; all inhibitory subtypes via Gad2/Pvalb/Sst/Vip/Lamp5/Chodl; Meis2/Pbx3; glia via Aqp4/Gfap/Mbp/Pdgfra/Cspg4/C1qa/Tmem119; vascular via Flt1/Ogn). Pipeline PARAMETERS reproduced exactly; cluster COUNT within tolerance (off by 3) and cell-type taxonomy fully recovered. DIFFERENCE vs original: the exact 47-vs-44 cluster index differs because GSE269499 deposits WT samples ONLY -- the repo's Fig-1 clustering merged 18 samples including 4 Shank3-KO that are not in GEO; we clustered the 14 WT samples (= the WT subset the published figure shows). NOT ATTEMPTED (the 80/20 tail): re-running CellRanger from FASTQ (unnecessary; matrices deposited), the full WT+KO 18-sample merge (KO data absent), the exact Enrichr-based 22-label assignment, the cerebellum (Supp Fig 1) and the IEG/social-ensemble differential analyses (Figs 2-6). No fabrication concern: every reproduced value is derivable from the shipped data+code. Grades are provisional; a human auditor confirms.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-14 ⛓ 95aba1fd4989
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Heterogeneous neuronal populations activated during social interaction can be identified by quantifying immediate early gene (IEG) expression coupled with transcriptional cell-type identification via single-nuclei RNA sequencing.

Core claims
  • A protocol for dissecting mouse PFC, hippocampus, and cerebellum and generating snRNA-seq libraries can identify neuronal populations activated during social interaction. method
  • Addition of Actinomycin-D to dissection PBS and lysis buffer prevents artificial transcriptional perturbation during tissue handling, which is critical for accurate IEG detection. method
  • Heterogeneous cell populations recruited during social interaction can be identified through quantification of immediate early genes coupled with transcriptional identification of different cell types following behavior. finding
  • Extending the interval between social interaction and dissection from 10 min to 35 min still allows IEG detection but decreases the number of nuclear IEG transcripts recovered. finding
  • A sucrose gradient (adapted from Ayhan et al.) effectively removes myelin and debris to isolate clean nuclei from brain tissue. method
  • The nuclei isolation protocol can also be used to obtain RNA for bulk RNA-sequencing. resource
  • This dissection and nuclei isolation protocol can be modified for other brain regions or behaviors of interest. method
Experimental setups
Assay System Perturbation Readout Platform
single-nuclei RNA-seq mouse medial prefrontal cortex (mPFC) social interaction with novel juveniles IEG expression and cell-type identification 10x Genomics Chromium Next GEM Single Cell 3' Reagent Kit v3.1
single-nuclei RNA-seq mouse hippocampus (dorsal/ventral optional) social interaction with novel juveniles IEG expression and cell-type identification 10x Genomics Chromium Next GEM Single Cell 3' Reagent Kit v3.1
single-nuclei RNA-seq mouse cerebellum social interaction (comparison region) / control cage-lid opening IEG expression and cell-type identification 10x Genomics Chromium Next GEM Single Cell 3' Reagent Kit v3.1
nuclei counting/viability (AO/PI staining) isolated brain nuclei suspensions none nuclei concentration and integrity Nexcelom Cellometer slide (alternative: trypan blue with Countess 3)
bulk RNA-seq (alternative downstream use) mouse brain tissue none stated RNA extraction/quantification
Key results
  • Post-social-interaction dissection delay extended to 35 min still permits IEG detection by sequencing, but nuclear IEG transcript numbers decrease relative to the standard 10 min interval.
  • Up to 8 samples from 4 mice can be routinely collected in one morning while maintaining high sample quality through nuclei isolation and library prep.
  • Samples must be diluted to 700-1200 nuclei/μL to achieve optimal recovery when targeting 10,000 cells on the 10x platform.
  • Sucrose gradient centrifugation separates a floating myelin layer from the nuclei pellet, enabling debris removal.
Key statistics
  • other 700-1200 nuclei/μL (recommended nuclei dilution concentration for 10x Genomics snRNA-seq)
  • count 10,000 cells (target cell recovery per 10x Genomics Chromium v3.1 kit)
  • count 8 samples from up to 4 mice (typical sample throughput collected in one morning)
  • other 500 × g for 5 min at 4°C (centrifugation speed for pelleting nuclei during lysis steps)
  • other 13000 × g for 45 min at 4°C (centrifugation for sucrose gradient debris removal)
  • other 15 μM (Actinomycin-D concentration in dissection PBS and lysis buffer)
  • other 0.5% (final BSA concentration in nuclei resuspension buffer)
  • other 40 μm (Flowmi cell strainer pore size used for filtering nuclei)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods protocol paper describing nuclei isolation and snRNA-seq library preparation from mouse brain regions to identify cell populations active during social interaction. The experimental design compares socially exposed mice against non-social controls, with IEG expression in snRNA-seq data used as a proxy for neuronal activity; downstream computational analysis uses Seurat in R. The protocol text is truncated before the analysis and statistical reporting sections, and primary statistical details are explicitly deferred to the companion study (Walker and Frost). No inferential statistical tests are described within the provided protocol text.

Replicationbiological Groupssocial interaction (sequential novel-juvenile exposures) vs. control (cage lid opening only, no social contact) Pairingunpaired Randomization/blindingnot stated Dispersionnone
Approaches that could also have been used
  • Nuclear IEG transcript counts in snRNA-seq are used as an indirect proxy for neuronal activity at the time of social behavior
    Could also: Activity-dependent genetic labeling systems (e.g., TRAP2, CaMPARI2) or in-vivo calcium imaging (fiber photometry, miniscope) could also be used to tag or record behaviorally active cell populations — These approaches provide temporally precise or real-time activity readouts during behavior; snRNA-seq with IEGs offers simultaneous transcriptomic identity across many heterogeneous cell types, a trade-off each approach handles differently
  • Seurat is specified for single-nuclei clustering and cell-type identification
    Could also: Alternative frameworks such as Scanpy/AnnData (Python), scran/scater (R/Bioconductor), or SnapATAC2 could also be applied for dimensionality reduction, clustering, and differential-abundance testing — These tools implement comparable graph-based clustering but differ in normalization strategies, batch-correction approaches, and ecosystem (R vs. Python); choice may also influence how IEG-positive cells are defined and compared across conditions
  • Actinomycin-D is added to both the dissection solution and lysis buffer to arrest transcription and preserve the in-vivo IEG expression state
    Could also: Immediate flash-freezing of tissue post-behavior followed by rapid cold-lysis could also limit post-mortem transcriptional changes — Flash-freezing arrests transcription mechanically and avoids handling a toxic compound; actinomycin-D actively inhibits RNA polymerase and may better preserve the IEG signal during the dissection window, particularly when processing multiple samples sequentially
  • Both male and female mice are included in the same experimental cohort
    Could also: Explicitly powered sex-stratified analyses or including sex as a covariate in differential expression models (e.g., in DESeq2 or Seurat's FindMarkers) could also be applied — Incorporating sex as a variable in downstream statistical models would allow detection of sex differences in behaviorally activated populations and is increasingly recommended by funding agencies; whether the companion study addresses this is not stated in this protocol text
  • Nuclei concentration is determined by averaging duplicate AO/PI counts on a Nexcelom Cellometer
    Could also: Flow cytometry with DAPI staining or a hemocytometer-based manual count could also be used to assess nuclei yield and integrity — Flow cytometry adds size and granularity metrics useful for detecting debris or doublets before sequencing; the protocol itself notes trypan blue on a Countess3 as a slightly more variable but valid counting alternative, indicating flexibility at this step
  • A fixed 10-minute post-behavior interval before dissection is used across all social mice
    Could also: A time-course design with multiple post-behavior intervals (e.g., 10, 20, 35 min) profiled in parallel could also characterize the temporal dynamics of IEG expression — The protocol notes that IEG nuclear transcript counts decrease by 35 min post-behavior; a systematic time-course would allow researchers to empirically select the optimal interval for their specific IEG targets and brain regions of interest
Software: R 4.3.1 · RStudio 2023.03.0 · Seurat 4.4.0

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
7
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE269499 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39423127

Paper: Walker H, Frost NA. Protocol for the generation of single-nuclei RNA-seq libraries and quantification of heterogeneous cell types activated during social interaction. STAR Protoc 2024. PMID 39423127 · PMC11532268 · DOI 10.1016/j.xpro.2024.103395 Code: https://github.com/nickfrostneuro/single-cell (commit c5f8c5f476e0, single commit, 2024-06-05) Data: GEO GSE269499 (Mus musculus, snRNA-seq, mPFC + cerebellum, WT only)

Nature of the publication

STAR Protocols methods paper. It describes a wet-lab + computational protocol and demonstrates it on the authors' own WT snRNA-seq data. The GitHub repo holds the R/Seurat analysis scripts from the associated research (Figures 1–6). The protocol paper itself prints very few hard pipeline-derived numbers; the concrete, checkable computational outputs live in (a) the documented pipeline parameters and (b) the repo's Figure 1 PFC clustering, whose cell-type taxonomy is fixed in code.

The documented computational pipeline

  1. Alignment/quantification: CellRanger v7.0.1, reference GRCm39 (Step 32) → per-sample MEX count matrices (barcodes/features/matrix). These are what GEO deposits (GSE269499_RAW.tar, MTX+TSV). ⇒ we start from the authors' own count matrices; CellRanger re-run is NOT required.
  2. QC (Seurat, Step 37): percent.mt via ^mt-; keep MT < 5 & nFeature_RNA > 800 & nFeature_RNA < 6000.
  3. Normalization (Step 38): SCTransform(vars.to.regress = "MT").
  4. Dim-reduction/clustering (Steps 38–39): RunPCA; FindNeighbors(dims=1:30); FindClusters (default res → SCT_snn_res.0.8); RunUMAP(dims=1:30).
  5. Cluster curation (Step 39): drop clusters present in <50% of samples or

    =75% from a single sample.

  6. Cell-type ID: canonical markers + Allen Brain Atlas (Enrichr) → named types.

IN SCOPE (pipeline-derived, attempted)

Reproduce the WT mPFC clustering (repo WT_Social_Ensemble_Scripts/Figure1_PFC clustering.R) starting from the GEO-deposited CellRanger count matrices:

  • C1 QC thresholds as documented (MT<5, nFeature 800–6000) — documentation/code match.
  • C2 Number of nuclei passing QC (reproduced value; no paper number to compare).
  • C3 Number of clusters at SCT_snn_res.0.8.
  • C4 Recovery of the 22 named PFC cell types (repo levels() list, Fig 1B/1D) via canonical markers (Slc17a7, Gad2, layer markers, Pvalb/Sst/Vip/Lamp5, Aqp4/Gfap, Mbp, Pdgfra, C1qa, Flt1, …).

Sample set — important honest caveat

GSE269499 contains WT samples only (14 PFC + 14 cerebellum). The repo's Figure-1 clustering merged 18 PFC samples including 4 Shank3-KO samples that are NOT deposited in GEO (20251X4, 20375X1, 20521X5, 20523X5). The published WT figure (Fig 1B/1D) subsets that 18-sample clustering down to the 14 WT PFC samples — which are exactly the 14 WT PFC GSMs in GEO. We therefore cluster the 14 deposited WT PFC samples directly. Consequence: the absolute cluster count at res 0.8 will not match the original 44 exactly (different merge → different Louvain partition); the robust, comparable claim is the cell-type taxonomy (C4), not the raw cluster index.

OUT OF SCOPE (not attempted) — and why

  • Wet-lab steps (nuclei isolation, library prep, sequencing) — non-computational.
  • CellRanger re-run from FASTQ — unnecessary (count matrices deposited) and the authors' exact CellRanger output is what GEO ships.
  • The full 18-sample (WT+KO) merge and exact cluster-index reproduction — KO data not in GEO; this is the hard last ~20% and is explicitly skipped.
  • Cerebellum (Supp Fig 1) and the social-ensemble / IEG differential analyses (Figs 2–6) — downstream of clustering; out of the 80/20 core. May spot-check if time permits.

Reproduction strategy

All compute on «our HPC» (SLURM, «infra»). Conda env pinned to Seurat 4.4.0 + sctransform inside the compute job. Start from GEO count matrices → documented pipeline → re

Figures / tables: Figure1_PFCFig 1B
C1
Reported
QC: MT<5 & nFeature_RNA 800-6000 (Step 37)
Reproduced
applied identically
exact
C2
Reported
CellRanger v7.0.1 / GRCm39 (Step 32)
Reproduced
not re-run; started from GEO-deposited CellRanger MEX matrices
partial
C3
Reported
SCTransform(regress MT), dims 1:30, FindClusters res 0.8 (Steps 38-39)
Reproduced
applied identically (Seurat 4.4.0, sctransform 0.4.2)
exact
C4
Reported
44 clusters at SCT_snn_res.0.8 (code-implied, 18-sample merge)
Reproduced
47 clusters (14-sample WT-only merge)
within tolerance
C5
Reported
22 named PFC cell types (Fig 1B/1D; repo levels())
Reproduced
all 22 types' canonical markers recovered with cluster-specific expression
within tolerance
C6
Reported
nuclei passing QC not numerically reported in paper
Reproduced
103,153 / 111,679 (92.4%) across 14 WT mPFC samples
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

The documented snRNA-seq pipeline reproduced faithfully on the authors' own GEO-deposited CellRanger matrices: QC thresholds and clustering parameters (SCTransform regress-MT, 30 PCs, res 0.8) are exact, and all 22 published PFC cell types' canonical markers were recovered as cluster-specific signals. The single numeric deviation (47 vs 44 clusters, off by 3) is fully explained on the data-availability side — GSE269499 deposits only the 14 WT samples, while the published Fig-1 merge included 4 Shank3-KO samples not in GEO — and is therefore an explainable cohort/input difference, not a computation defect. Every reproduced value is derivable from the shipped data+code with no fabrication concern; overall a solid (q8 yellow) reproduction whose core methods-paper claim holds.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

165.2 k
tokens (I/O) · 11.9 M incl. cache
36 min
runtime · 1.52 CPU-h
83 GB
peak RAM
2
HPC jobs
hummel
machine