Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Inferring a spatial code of cell-cell interactions across a whole animal body.

PLoS Comput Biol · 2022
L1 90/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH -- clean 1:1 reproduction, re-verified from scratch. The Celegans-cell2cell repo (commit 25f03fc, 2022-11-22) is the authors' own code and ships every processed input plus the 100 precomputed GA runs, so the pipeline is fully self-contained. This run rebuilt everything from zero on «our HPC» «infra» (the prior workdir had been janitor-reclaimed): cloned the repo at the pinned commit, built a pinned cell2cell 0.6.8 / numpy<2 conda+pip env, and ran the pipeline via SLURM «job» (compute node n149, 2m15s, exit 0). The two headline deterministic results reproduce EXACTLY to the authors' own printed precision: full-list Bray-Curtis CCI vs spatial-distance Spearman rho = -0.2066 (paper -0.21, P=0.0016) and the GA-optimized 37-pair subset rho = -0.6334, P=2.629e-27 (paper -0.63) -- byte-identical floats to the prior independent run. The GA consensus selection (245-pair co-occurrence Jaccard matrix -> ward clustering into 2 groups {89,37} -> smaller cluster) reproduces all 37 pairs identically to the shipped list (37/37 overlap, zero set differences). Counts C1/C2/C6/C7/C8 exact. The distance-range classification AUCs (C12-C14) reproduce qualitatively (Bray-Curtis 0.62 > ICELLNET 0.61 > LR-Count 0.59, matching the paper's ranking and ~0.6 magnitude; CellChat/Smillie near chance) but differ ~0.02-0.03 in absolute value due to xgboost version drift; the released code uses XGBClassifier although the paper text says 'Random Forest' (naming inconsistency flagged, not fabrication). NOT attempted: C15 enrichment (NB09), wet-lab smFISH validation (out of scope), and the from-scratch GA optimization (1-2 days; authors provide the 100-run output, which we used). No fabrication detected -- every reproduced value is derivable from the shipped data+code. All grades PROVISIONAL pending human audit.

💻 Code ↗ 🗄 Data: GSE98561

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-22 ⛓ 8458cd7ac818
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

There exists a spatial code embedded in ligand-receptor interactions across the body of a multicellular organism that encodes spatial information and defines the organization/distribution of cells, and this code can be computationally inferred from single-cell transcriptomes of C. elegans.

Core claims
  • cell2cell computes cell-cell interaction (CCI) potential using a novel modified Bray-Curtis score based on complementary coexpression of ligand-receptor pairs between cells method
  • The Bray-Curtis CCI score captures spatial properties and separates short-, mid-, and long-range intercellular distances better than other CCI scoring approaches (LR Count, ICELLNET) finding
  • A genetic algorithm identifies the ligand-receptor pairs most informative of the spatial organization of cells (the spatial code) across the whole C. elegans body method
  • Inferred intercellular CCIs are negatively correlated with actual intercellular distances finding
  • Experimentally confirmed communicatory behavior for selected cell-cell and ligand-receptor pairs computationally inferred to contribute to the spatial code finding
  • Ligand production is cell-type specific (sender cells cluster together), whereas receptor production is promiscuous (receiver cells do not cluster) finding
  • Neurons show the highest CCI potential (with themselves and muscle cells), while germline cells show the lowest CCI potential, consistent with known tissue biology finding
  • Curated the most comprehensive list of 245 ligand-receptor interactions in C. elegans as a community resource resource
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq-based CCI inference (cell2cell, Bray-Curtis score) C. elegans, 22 of 27 annotated cell types across whole body none CCI score representing ligand-receptor coexpression complementarity between cell pairs cell2cell (Python/Jupyter tool)
Random forest classification computational: CCI score matrices derived from C. elegans scRNA-seq none classification of intercellular distance range (short/mid/long)
Genetic algorithm-based LR pair selection computational: C. elegans LR interaction list combined with 3D atlas intercellular distances none subset of LR pairs most informative of spatial cell organization
UMAP with Jaccard/Rand index similarity analysis C. elegans predicted cell-cell interaction network none clustering behavior of sender vs. receiver cells
Agglomerative hierarchical clustering C. elegans CCI score heatmap across 22 cell types none grouping of cell types by interaction potential and lineage
Experimental validation (in situ coexpression confirmation) C. elegans, selected cell-cell and ligand-receptor pairs none confirmation of predicted communicatory behavior/adjacent-cell LR coexpression
Key results
  • Intercellular distances are negatively correlated with inferred cell-cell interaction scores
  • Pairs of interacting cells cluster by sender cell identity in UMAP space, but not by receiver cell identity
  • Neurons have the largest CCI potential with other cell types (especially themselves and muscle cells); germline cells have the lowest CCI potential
  • Random forest classifiers using CCI scores predict short/mid/long-range intercellular distance categories, evaluated by ROC/AUC with 3-fold stratified cross-validation
  • Curated list of ligand-receptor interactions compiled for C. elegans 245 interactions
  • 22 of 27 identified C. elegans cell types were assigned spatial locations from a 3D atlas of 357 individual cells 22/27 cell types; 357 cells
  • Genetic algorithm selection produced a consensus list of LR interactions defining the spatial code 37 interactions
Key statistics
  • count 245 ligand-receptor interactions (manually curated LR interaction list for C. elegans)
  • count 37 interactions (consensus list from genetic algorithm selection defining spatial code)
  • count 357 individual cells (cells reported with locations in the 3D atlas of C. elegans)
  • count 22 of 27 cell types used (cell types assignable to spatial locations from the 3D atlas, used in CCI analysis)
  • other 3-fold stratified cross-validation (evaluation scheme for random forest distance-range classifiers)
  • correlation negative correlation (exact coefficient not specified in provided text) (relationship between intercellular distance and inferred CCI score)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational biology paper describing a novel cell-cell interaction (CCI) scoring method (a modified Bray-Curtis score) applied to single-cell transcriptomic and 3D spatial atlas data from C. elegans. Analyses were primarily computational/machine-learning in nature: unsupervised clustering and dimensionality reduction to explore CCI patterns, and supervised random forest classifiers (evaluated by ROC/AUC with cross-validation) to test whether CCI scores could distinguish short-, mid-, and long-range intercellular distances. A genetic algorithm was also used to select the ligand-receptor pairs most informative of spatial organization, and correlation was used to relate CCI scores to intercellular distance.

Replicationunclear Sample size22 of 27 identified C. elegans cell types were used (those with assigned locations in a published 3D atlas of 357 individual cells); 245 curated ligand-receptor interactions were used as input; a 3-fold stratified cross-validation scheme was used for classifier evaluation GroupsPairs of C. elegans cell types, and short- vs mid- vs long-range intercellular distance categories Pairingna Randomization/blindingnot stated DispersionSD Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Random forest classifier evaluated via ROC curves and AUC, with 3-fold stratified cross-validation Classifying cell-cell pairs into short-, mid-, or long-range intercellular distance categories from CCI scores (Fig 2C, S3 Fig) not stated
Agglomerative hierarchical clustering on a dissimilarity metric (1 − CCI score) Grouping cell types by their CCI score profiles, excluding autocrine interactions (Fig 2A) 22 cell types with assigned spatial locations (of 27 total cell types identified) not stated
UMAP dimensionality reduction based on Jaccard distance (and separately, Rand index) between directed LR-interaction profiles Visualizing similarity of cell-cell interaction pairs by sender/receiver identity (Fig 2B, S1 Fig, S2 Fig) not stated
Correlation between CCI scores and intercellular distance (type not specified in text) Assessing whether inferred CCI scores relate to physical distance between cells across the C. elegans body not stated
Genetic algorithm-based feature selection Identifying the ligand-receptor pairs most informative of spatial organization (consensus list of 37 interactions, S3 Table) 245 curated ligand-receptor interactions (S1 Table) as the candidate pool na
Approaches that could also have been used
  • Random forest classifier performance was summarized as mean ± standard deviation of ROC/AUC across a 3-fold stratified cross-validation.
    Could also: A repeated or higher-fold (e.g., 5- or 10-fold) cross-validation, or a nested cross-validation with hyperparameter tuning, could also be used — More folds/repeats or nested validation can give a more stable estimate of classifier performance and its variability, particularly useful when comparing several CCI scoring methods.
  • Variability in AUC across folds was expressed with standard deviation bands around the mean ROC curve.
    Could also: A 95% confidence interval (e.g., via bootstrapping) for the AUC could also be reported — Confidence intervals directly convey the precision of the estimated classification performance and are commonly used alongside or instead of SD when comparing AUCs between methods.
  • Cell types were grouped using agglomerative hierarchical clustering on a dissimilarity derived from the CCI score matrix.
    Could also: Other clustering approaches, such as k-means, spectral clustering, or community detection on a network graph, could also be applied — Different clustering algorithms make different assumptions about cluster shape and can be used to cross-validate whether the observed cell-type groupings are robust to the choice of method.
  • Similarity between interacting cell pairs was visualized with UMAP based on Jaccard distance, with the Rand index also used as an alternative similarity metric.
    Could also: Other dimensionality-reduction methods such as t-SNE or PCA, or other similarity metrics such as cosine similarity, could also be used — Comparing results across multiple embedding and distance-metric choices can help confirm that observed groupings (e.g., clustering by sender cell) are not an artifact of a single specific method.
  • The relationship between CCI scores and intercellular distance was described as a negative correlation, without specifying the correlation coefficient used.
    Could also: Reporting a specific coefficient (e.g., Pearson's r for linear relationships or Spearman's rho for monotonic, potentially nonlinear ones) along with its confidence interval could also be used — Specifying and reporting the correlation coefficient and its precision makes the strength and reliability of the CCI score-distance relationship more directly interpretable and comparable across studies.
  • A genetic algorithm was used to select the ligand-receptor pairs most informative of spatial organization, yielding a consensus list of 37 interactions.
    Could also: Regularized regression (e.g., LASSO) or exhaustive/greedy stepwise feature selection with cross-validated performance could also be used — These alternative feature-selection strategies offer different trade-offs between computational cost and guarantees of finding a globally optimal feature subset, and comparing results across methods can support the robustness of the selected LR pairs.
Software: cell2cell (custom open-source Python tool) · Python / Jupyter Notebooks

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36395331

Paper: Armingol et al. 2022, "Inferring a spatial code of cell-cell interactions across a whole animal body." PLoS Comput Biol. PMID 36395331 / PMC9714814. Code: https://github.com/LewisLabUCSD/Celegans-cell2cell Data: GEO GSE98561 (C. elegans L2 scRNA-seq, Cao et al. 2017 "sci-RNA-seq").

What the paper does (pipeline overview)

  1. Take a C. elegans single-cell transcriptome (GSE98561), aggregated to cell-type TPM profiles (27 cell types; 22 with a spatial location).
  2. Build a binary presence/absence call per gene per cell type (>10 TPM).
  3. Score every cell-type pair with a novel Bray-Curtis CCI score over a curated ligand-receptor (LR) list (245 interactions; 127 ligands, 66 receptors).
  4. Take a 3D spatial atlas (357 nuclei L1; 322 labelled) → cell-type Euclidean distances.
  5. Correlate CCI scores vs spatial distance (Spearman); show anticorrelation.
  6. Genetic algorithm (100 runs) selects an LR subset (37 consensus pairs) that maximizes anticorrelation → ρ improves from -0.21 to -0.63.
  7. Random-Forest classification of distance ranges from CCI scores (avg AUC 0.65 vs ICELLNET 0.63, LR-count 0.57).
  8. Pathway/enrichment analysis of selected pairs; smFISH validation (wet-lab).

In scope (pipeline-derived, reproducible)

  • C5–C8 LR-list sizes & threshold — directly checkable from repo data files.
  • C1–C4 cell-type / nuclei counts — checkable from repo/GEO inputs.
  • C10 full-list Spearman ρ = -0.21 (Bray-Curtis CCI vs distance) — CORE result, deterministic given inputs; primary reproduction target.
  • C11 GA-optimized subset ρ = -0.63 — depends on the GA (stochastic, 100 runs); reproducible in distribution, exact 37-pair set may differ.
  • C12–C14 RF AUCs — stochastic (RF seed, CV split) but reproducible in range.
  • C9 37 consensus LR pairs — GA output, stochastic.
  • C15 5 spatially enriched/depleted pairs — enrichment test, reproducible.

Out of scope (wet-lab / manual / external)

  • smFISH validation of 3 uncharacterized interactions (wet-lab imaging).
  • Manual curation of the LR list itself (literature curation; we USE the shipped list).
  • 3D atlas construction (from Long et al. external atlas; we USE shipped coordinates).

Reproduction strategy

The repo is the authors' own code (Python, cell2cell library + notebooks). Plan: clone repo on «infra», install cell2cell + deps in conda on «our HPC» front node, run the notebooks/scripts that compute the Bray-Curtis CCI matrix and the Spearman correlation vs distance (C10 is the deterministic anchor). Then attempt the GA (C11/C9) and RF (C12) which are stochastic. Primary anchor = C10 (-0.21) and the LR-list/threshold sanity counts (C5–C8), which are deterministic.

C1
Reported
27 cell types
Reproduced
27
exact
C2
Reported
22 cell types with spatial location
Reproduced
22
exact
C3
Reported
357 nuclei in 3D atlas
Reproduced
360 rows in shipped atlas csv
partial
C5
Reported
>10 TPM presence threshold
Reproduced
10 TPM constant_value
exact
C6
Reported
245 curated LR interactions
Reproduced
245
exact
C7
Reported
127 ligands
Reproduced
127
exact
C8
Reported
66 receptors
Reproduced
66
exact
C9
Reported
37 GA-consensus LR pairs (100 runs)
Reproduced
37 (smaller ward cluster; 37/37 identical to shipped list, 0 set differences)
exact
C10
Reported
Spearman rho=-0.21 (P=0.0016), full list
Reproduced
-0.2066 (P=0.00160)
exact
C11
Reported
Spearman rho=-0.63 (P=2.629e-27), 37-pair subset
Reproduced
-0.6334 (P=2.6295e-27)
exact
C12
Reported
avg AUC 0.65 (Bray-Curtis)
Reproduced
0.6212 (std 0.1124)
within tolerance
C13
Reported
avg AUC 0.63 (ICELLNET)
Reproduced
0.607 (std 0.0583)
within tolerance
C14
Reported
avg AUC 0.57 (LR Count)
Reproduced
0.5866 (std 0.0895)
within tolerance
C15
Reported
5 spatially enriched/depleted LR pairs
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

This is a clean, near-1:1 reproduction. The two headline results — full-list Bray-Curtis CCI vs spatial-distance Spearman rho=-0.21 (P=0.0016) and the GA-optimized 37-pair subset rho=-0.63 (P=2.629e-27) — reproduce exactly to the authors' own printed precision from their self-contained repo, and all hard counts (27/22/245/127/66, 37/37 GA pairs) match. The few deviations are all minor and explainable on the technical side: secondary classifier AUCs differ ~0.02-0.03 from xgboost version drift (with a paper-text 'Random Forest' vs code XGBClassifier naming inconsistency), the 3D atlas count is 357 vs 360 rows, and the C15 enrichment analysis was deferred. No fabrication signature — every value is derivable from shipped data+code — so this rates green on derivability and core claim, yellow overall only because the reproduction is not literally flawless.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

196.9 k
tokens (I/O) · 17.2 M incl. cache
53 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.