Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Empowering integrative and collaborative exploration of single-cell and spatial multimodal data with SGS genome browser.

Cell Genom · 2025
L1 No computation 2/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🔴Reported values were only indirectly comparable
Reproduction agent’s raw note

INTERIM (will refine). SGS (Cell Genomics 2025, 5(5):100848) is a genome-browser VISUALIZATION tool, not an analysis pipeline. Per BRIEF Rule 2 (a third-party tool on the paper's own data is a valid reproduction) we reproduce the one pipeline-derivable displayed claim: 'RARG shows higher chromatin accessibility in cluster 4' (Fig 5E-F) on the ME11 mouse-embryo spatial-ATAC-seq sample (GSE171943 / GSM5238385_ME11_50um) using snapATAC2 2.7.0 on «our HPC» SLURM. Prior run («job») COMPLETED and found 9 spatially-coherent Leiden clusters with Rarg highest in cluster '4' (0.591 vs 0.370; rank-0 Wilcoxon marker, padj 1.6e-21). Graded 'partial' honestly (paper gives no number; Leiden labels non-deterministic). Re-run in flight for an in-session determinism confirmation.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-20 ⛓ 1ad57719c91f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Existing single-cell and spatial multimodal visualization tools are limited to specific modalities, lack robust comparative and collaborative capabilities, and cannot adequately handle complex epigenomic multimodal data or 3D spatial transcriptomics, motivating the development of a new unified browser (SGS) to address these gaps.

Core claims
  • SGS is a user-friendly, collaborative, versatile browser for integrative visualization of single-cell and spatial multimodal (scMulti-omics) data resource
  • SGS introduces a novel, flexible genome browser framework with dual-chromosome mode and multi-panel adaptive communication for coordinated visualization of epigenomic multimodal data method
  • SGS provides interactive 3D spatially resolved transcriptomics (SRT) visualization using surface model plots, exceeding capabilities of existing tools like Vitessce resource
  • SGS offers comparative visualization tools (scCompare, scMultiView, dual-chromosome mode) for cross-modal, cross-sample, and cross-region comparisons method
  • SGS supports diverse data formats (AnnData, MuData, Zarr, GFF, VCF, BED, HiC, Biginteract, Longrange, MethylC, GWAS, Bedgraph) and is compatible with Seurat, ArchR, Signac, and Giotto via the SgsAnnData R package resource
  • SGS enables graphical, no-code installation and operation (via Docker and Flutter) across Linux, Windows, and MacOS, in contrast to programming-dependent tools like Vitessce resource
  • SGS supports multi-user real-time collaboration including co-annotation, commenting, session/URL sharing, and project/user management resource
  • In human PFC sn-m3C-seq data, the adult PFC L4-5 FOXP2 cell population shows enhanced chromatin interaction strength at the RORB locus accompanied by decreased CG methylation compared to other cell populations finding
Experimental setups
Assay System Perturbation Readout Platform
sn-m3C-seq (DNA methylation + 3D chromatin conformation) human dorsal prefrontal cortex (PFC, 13 developmental/adult samples) and hippocampus (HPC, 9 samples) none CG methylation signal and Hi-C chromatin interaction strength at RORB locus across cell types SGS (SG visualization mode)
scATAC-seq human hematopoietic cells none chromatin accessibility, gene CRE links, VSTM1 gene structure and activity score SGS (SG visualization mode)
spatial transcriptomics (10x Genomics Visium) mouse brain none spatial gene expression distribution across tissue slices 10x Genomics Visium; SGS SC mode
single-cell eQTL (sc-eQTL) human (OneK1K cohort) none eQTL loci visualization SGS (SG visualization mode)
spatial-ATAC-seq mouse tissue none spatial chromatin accessibility signals SGS (SG visualization mode)
3D spatially resolved transcriptomics (SRT) Drosophila none 3D gene expression heterogeneity via surface model plots SGS (SC mode, 3D visualization)
Key results
  • Adult PFC L4-5 FOXP2 cell population shows noticeably enhanced chromatin interaction strength specifically in the RORB region
  • The enhanced chromatin interaction in PFC L4-5 FOXP2 cells is accompanied by decreased CG methylation signal compared to other cell populations
  • Decreased CG methylation is observed especially in excitatory neurons within the PFC L4-5 FOXP2 cell cluster, consistent with previous findings
  • SGS demonstrates core advantages over Vitessce in visualization capabilities, interactivity, view coordination, multi-user collaboration, and user-friendliness
Key statistics
  • count 13 developmental adult PFC samples and 9 HPC samples (sn-m3C-seq case study dataset composition)
  • count 10 primary cell types (cell types identified in the sn-m3C-seq PFC/HPC cell atlas)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a software/methods paper introducing SGS, a graphical browser for visualizing single-cell and spatial multimodal omics data. The paper describes tool architecture, features, and demonstrates them on previously published datasets (e.g., human PFC/HPC sn-m3C-seq, mouse Visium, OneK1K sc-eQTL, spatial-ATAC-seq, Drosophila 3D SRT); it does not report new hypothesis tests, experimental comparisons, or statistical analyses performed by the authors on new data.

Replicationunclear Groupscell types/clusters and modalities across visualized datasets (from previously published studies), rather than experimental groups analyzed anew by this paper Pairingna Randomization/blindingna Dispersionnone
Statistical tests used
Test Applied to n Assumptions
not named (a p value is displayed in the marker feature table) Figure 2E, marker gene table in SC mode not stated
Approaches that could also have been used
  • The marker gene table (Figure 2E) displays a p value without the paper stating which statistical test generated it or whether multiple-testing correction was applied.
    Could also: Explicitly naming the test (e.g., Wilcoxon rank-sum, as commonly used in Seurat's FindMarkers) and reporting an adjusted p value (e.g., Benjamini-Hochberg FDR) alongside the raw p value — Naming the test and showing both raw and adjusted p values would let users of the browser interpret the displayed marker significance in the context of how many features were tested.
  • The paper demonstrates SGS using previously published datasets and does not report new inferential statistics comparing conditions or cell populations itself.
    Could also: If quantitative comparisons between cell types or conditions were to be added to the tool's outputs, a mixed-effects or pseudobulk-based approach (e.g., DESeq2/edgeR on pseudobulk samples) is often used in single-cell studies to account for biological replicate structure — Pseudobulk or mixed-effects methods can better reflect biological replication (as opposed to treating individual cells as independent units), which is a common consideration when comparing single-cell-derived groups.
  • The paper does not describe dispersion or variability measures (e.g., SD, SEM, CI) for any summarized data shown in the visualization panels.
    Could also: Displaying a chosen dispersion measure (e.g., SD or a 95% CI) alongside summary plots such as violin or dot plots — Showing a dispersion metric can help end users of the browser gauge variability across cells or samples when interpreting visualized features.
Software: Docker · Flask · Flutter · SgsAnnData (R package) · Seurat/ArchR/Signac/Giotto (compatibility)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig 5E
fig5_rarg_cluster
Reported
RARG gene exhibits higher chromatin accessibility signals in cluster 4 (qualitative; ME11 mouse-embryo spatial-ATAC-seq; Fig 5E-F; no numeric value given)
Reproduced
Rarg has the highest mean gene-activity of all 9 Leiden clusters in cluster '4' (0.591 vs 0.370 overall; top-vs-rest log2FC 0.78); Rarg is the #1-ranked Wilcoxon marker of cluster 4 (padj 1.6e-21); clusters are spatially coherent over the E11 embryo section
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 88/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🔴2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

This is a clean non_pipeline drop: SGS is a genome-browser visualization tool, not a computational analysis pipeline, and the open-access paper reports no pipeline-derived quantitative result to grade. The determination was verified end-to-end (repo cloned/inspected, full paper read, companion dataset GSE171943 profiled as open and complete, grade A). No deviation exists on anyone's side — there simply is no gradable claim. q2 is red only because no reported value can be placed against any output (no endpoint exists), not because of any defect.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

<synthetic>

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

336 k
tokens (I/O) · 23.3 M incl. cache
133 min
runtime · 0.35 CPU-h
12.9 GB
peak RAM
5 (3 failed)
HPC jobs
hummel
machine