Celline: a flexible tool for one-step retrieval and integrative analysis of public single-cell RNA sequencing data.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED-WELL-ENOUGH: yes - Celline (tool paper) ships open code (GitHub @ ae00f57, v1.1.0) with a fully-specified, deterministic preprocess/QC algorithm, and the assigned dataset GSE153162/GSM4635075 deposits a ready Cell Ranger filtered count matrix, so we skipped FASTQ->Cell Ranger alignment (hard 20%) and ran Celline's EXACT QC code (verbatim _dynamic_cutoff + identical scanpy/Scrublet calls) on the paper's own E14.5 matrix on «our HPC» (SLURM 2176243). It runs cleanly and is DETERMINISTIC (identical across reruns): 3485 cells -> at n_mad=3 (paper text), nFeature 1195-3787, UMI cap 11491, mito<=5%, Scrublet auto-threshold 0.327, 83.3% retained, 19 Leiden clusters. OUTCOME = PARTIAL / faithful-method, NOT exact 1:1, because the paper publishes NO standalone numbers for our assigned dataset (GSE153162/E14.5) - verified against full text. The paper's only pinnable per-sample QC numbers (nFeature 536-3446, UMI 9654, doublet 0.366, 70% retained, 3xMAD) are for a DIFFERENT sample, E18/GSM2453043/GSE93421, which deposits no per-sample matrix (only an aggregated 133-sample hdf5 + BAMs) and would need Cell Ranger alignment to match exactly. NOT ATTEMPTED (and why): exact E18 QC numbers (different sample, no matrix, needs alignment = hard 20%); scVI integration +0.22/30-clusters/scIB-0.57 (stochastic VAE); cell-type proportions & 11 cell types (need E18 alignment + annotation); LLM metadata accuracy 90.8/97.0% (OpenAI GPT, non-deterministic, paid) and timing benchmarks (hardware-bound). FLAG: paper text says 3xMAD but code default is n_mad=2.5 (text/code inconsistency; not fabrication).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 64assessed: 2026-06-14 ⛓ 9efb8827afb9
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusExisting scRNA-seq tools address only isolated steps and require manual curation, leaving no end-to-end solution; the paper tests whether a single Python package (Celline) can automatically retrieve, preprocess, integrate, and analyze public scRNA-seq data via one-line commands while standardizing metadata and reducing batch effects.
- ★ Celline is a Python package executing an entire scRNA-seq workflow (retrieval, preprocessing, integration, analysis) using single-line commands per step resource
- ★ Celline automatically gathers raw scRNA-seq data from multiple public repositories (GEO, SRA, CNCB, ArrayExpress) and extracts/standardizes metadata using large language models method
- ★ Celline wraps established tools (Scrublet, Seurat, Scanpy, Harmony, scVI, Slingshot, Cell Ranger, STARsolo, scPred) into one-line commands for a unified workflow method
- ★ Applied to two mouse brain cortex datasets (E14.5 and E18), Celline retrieved data, standardized metadata, removed low-quality cells, annotated 11 major cell types, improved integration, and completed trajectory analysis finding
- ★ Integration with Celline improved integration quality by a scIB score of +0.22 finding
- Celline's modular, extensible architecture allows users to add custom functions inheriting the CellineFunction base class without modifying core code method
- Preprocessing applies consistent QC: removing cells with mitochondrial fraction >5%, nFeature outlier filtering by 3x median absolute deviation, and Scrublet doublet removal method
- Celline uniquely integrates automated multi-repository retrieval, LLM-based metadata standardization, and downstream analysis in a unified command-line workflow compared to existing tools finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| droplet-based scRNA-seq (10x Chromium) | Mus musculus E14.5 C57BL/6 brain cortex (GSE153162, GSM4635075) | none | single-cell gene expression / gene-barcode matrix | Illumina HiSeq 2000; Cell Ranger count |
| droplet-based scRNA-seq (10x Chromium) | Mus musculus E18 C57BL/6 whole brain (pooled cortex, hippocampus, subventricular zone) (GSE93421, GSM2453043) | none | single-cell gene expression / gene-barcode matrix | Illumina HiSeq 4000; Cell Ranger count |
| scRNA-seq quality control / preprocessing | mouse brain scRNA-seq cells | none | mitochondrial fraction, nFeatures, doublet status; flag for high-quality cells | Scrublet; Seurat/Scanpy |
| cell-type annotation | mouse brain scRNA-seq cells | none | assigned cell identities (11 major cell types) | canonical markers or scPred |
| batch correction / integration | two integrated mouse brain datasets (E14.5 + E18) | none | corrected latent embeddings, UMAP, scIB integration score | scVI or Harmony |
| trajectory inference | integrated mouse brain scVI latent space | none | pseudotime, lineages, minimum spanning tree | Slingshot |
| droplet-based scRNA-seq (10x) | Homo sapiens peripheral blood mononuclear cells (PBMCs) (GSE115189) | none | single-cell gene expression | Illumina HiSeq 2500 |
- – Celline annotated 11 major cell types in the mouse brain datasets 11 cell types
- ▲ Integration with Celline improved integration quality as measured by scIB score +0.22
- – Preprocessing removed low-quality cells (high mitochondrial content, nFeature outliers, doublets)
- – Celline successfully retrieved data and standardized metadata across repositories
- – Trajectory analysis was completed using Slingshot in the corrected latent space
- other scIB score +0.22 (improvement in integration quality after Celline integration)
- count 11 major cell types (cell types annotated in mouse brain datasets)
- count 78,655 samples (GEO samples returned for 'scRNA-seq' search, illustrating data growth)
- other mitochondrial gene fraction > 5% (QC threshold for removing stressed/dying cells)
- other 3x median absolute deviation (nFeature outlier filtering threshold from the median)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a software tool paper describing Celline, a Python package for end-to-end retrieval and analysis of public scRNA-seq data. Technical validation was performed on two publicly archived mouse brain developmental datasets (E14.5 and E18), demonstrating automated data acquisition, QC, cell-type annotation, batch correction, and trajectory inference. Outcomes were reported descriptively (cell-type counts, scIB integration score improvement), with no inferential statistical comparisons between biological conditions. The paper's primary contribution is the pipeline itself, not a biological discovery.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Median absolute deviation (MAD)-based outlier filtering | QC step: filtering cells by nFeatures (unique genes per cell); cells outside 3 MADs from the median were excluded | — | not stated |
| Fixed threshold filtering | QC step: mitochondrial gene fraction >5% used to remove stressed/dying cells | — | not stated |
| Scrublet doublet scoring | QC step: probabilistic doublet detection applied to each sample | — | not stated |
| scIB (single-cell integration benchmarking) composite score | Evaluation of batch correction quality after Harmony/scVI integration (reported as +0.22 improvement) | — | not stated |
| Slingshot minimum spanning tree / pseudotime assignment | Trajectory inference in scVI latent space for developmental lineage reconstruction | — | not stated |
| scPred supervised machine learning classifier | Reference-based cell-type annotation (optional second annotation strategy) | — | not stated |
-
Integration quality was summarized using a single scIB composite score (+0.22), reported without uncertainty↳ Could also: Individual scIB sub-metrics (ASW, graph connectivity, kBET, LISI, etc.) could be reported separately, along with bootstrap confidence intervals around each score — Decomposing the composite score reveals which aspects of integration improved (batch mixing vs. biological conservation), and uncertainty estimates indicate whether the observed change is likely stable across random seeds or initializations
-
The MAD-based nFeatures threshold was fixed at 3 MADs from the median for all samples↳ Could also: Adaptive per-sample thresholds, or sensitivity analyses across a range of MAD cutoffs (e.g., 2–4 MADs), could also be applied — Optimal QC thresholds can vary substantially across tissue types and sequencing depths; reporting sensitivity to the chosen cutoff helps users calibrate the pipeline for their own datasets
-
The mitochondrial-fraction cutoff was fixed at 5% across all samples↳ Could also: A MAD-based adaptive threshold for mitochondrial fraction (analogous to the nFeatures filter) could also be used — Appropriate mitochondrial thresholds differ by tissue (e.g., brain vs. heart); an adaptive approach applies the same statistical logic as the nFeatures filter and may generalize better across Celline's intended multi-dataset use cases
-
Batch correction was evaluated by comparing Harmony vs. scVI only via the scIB score on one pair of datasets↳ Could also: A broader benchmarking across multiple dataset pairs or a statistical comparison (e.g., permutation test on scIB scores across resampled cell subsets) could also be performed — A single dataset pair limits generalizability; repeated sampling or multiple dataset pairs would allow uncertainty around the scIB improvement to be quantified
-
Cell-type annotation accuracy was reported by counting 11 annotated cell types, without a quantitative accuracy metric↳ Could also: Annotation accuracy could also be measured using F1 score, adjusted Rand index, or confusion matrices against a held-out reference with known labels — Quantitative annotation metrics allow direct comparison with other tools and reveal which cell types are harder to classify, supporting users in choosing between the marker-gene and scPred strategies
-
Trajectory inference was validated qualitatively (visual inspection of UMAP and pseudotime plots) with the root cluster specified by user expertise↳ Could also: Quantitative trajectory evaluation metrics such as correlation of pseudotime with known developmental markers, or comparison with alternative trajectory methods (PAGA, Monocle 3), could also be reported — Quantitative benchmarking of trajectory outputs against orthogonal developmental time information (e.g., embryonic day) would provide an objective measure of biological plausibility beyond visual inspection
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41458999 (Celline)
Paper: Sato Y, Asahi T, Kataoka K. Celline: a flexible tool for one-step
retrieval and integrative analysis of public single-cell RNA sequencing data.
Front Bioinform 2025. PMID 41458999 · PMCID PMC12738925 · DOI 10.3389/fbinf.2025.1684227.
Code: https://github.com/Kataoka-K-Lab/Celline (cloned @ HEAD ae00f57688f2900733576f1d00b8c748c8ec17ea).
Assigned data: GEO GSE153162 (E14.5 mouse cortex; sample GSM4635075, "RNA-Seq E14_5").
What kind of paper
This is a tool paper. Celline is a Python CLI that wraps a standard scRNA-seq pipeline: retrieve public data (SRA/GEO) → align (Cell Ranger / STARsolo) → QC/preprocess (scanpy + Scrublet, MAD-based dynamic thresholds) → cell-type prediction → integration (Harmony/scVI) → trajectory (Slingshot). The paper's case study applies Celline to two public samples:
- GSM4635075 — E14.5 cortex, from GSE153162 (our assigned RU data).
- GSM2453043 — E18 whole brain, from GSE93421 (the paper's other sample).
In scope (pipeline-derived, attempted)
The deterministic preprocess / QC stage of Celline applied to the paper's own
data. Celline's Preprocess.call() (src/celline/functions/preprocess.py) is fully
specified and deterministic given a count matrix:
- Scrublet doublet detection (
scr.Scrublet(adata.X).scrub_doublets(), default seed=0). - mito% via genes prefixed
mt-(case-insensitive);sc.pp.calculate_qc_metrics. - Dynamic MAD thresholds
_dynamic_cutoff(vec, n_mad):median ± n_mad·MAD(scipymedian_abs_deviation, scale=1) for nFeature (both sides) and total counts (upper only). Paper text states n_mad = 3 ("three times the MAD"); code default is 2.5 — we run both and flag the discrepancy. - Keep cells with
lower ≤ nGenes ≤ upper,counts ≤ upper_counts,mito% ≤ 5,not predicted_doublet. - Downstream (deterministic, seeded): normalize→log1p→HVG(2000,seurat_v3)→scale→ PCA(arpack)→neighbors(40pc/15nn)→UMAP→Leiden(res=1.0). Reports n_clusters.
Cheap because alignment is skipped: GSM4635075 ships a ready Cell Ranger
filtered count matrix (GSM4635075_E14_5_filtered_gene_bc_matrices_h5.h5), so the
heavy FASTQ→Cell Ranger step is not needed. We feed this matrix into Celline's
exact preprocess code.
Out of scope / NOT attempted (with reason)
- Exact E18 QC numbers (paper's only pinnable per-sample values: nFeature 536–3446, UMI cap 9654, doublet 0.366, 70% retained, 3×MAD). These belong to GSM2453043 / GSE93421, which deposits no per-sample count matrix (only an aggregated 133-sample hdf5 + BAMs). Matching them would require FASTQ→Cell Ranger per-sample alignment = the heavy ~20% we deliberately skip (HARD RULE 3).
- scVI integration (+0.22 scIB, 30 clusters, scIB 0.57): stochastic VAE over both samples; not 1:1 reproducible; hard 20%. Not attempted.
- Cell-type proportions (EN 36.5% … for E18): depend on alignment + marker annotation of the E18 sample we cannot cheaply rebuild. Not attempted.
- LLM metadata accuracy (90.8%/97.0%) and timing benchmarks (6.9–39.2 s, R²=0.995): depend on OpenAI GPT API (non-deterministic, paid) and on hardware; not reproducible 1:1. Not attempted.
Honest framing
The paper reports zero standalone numeric results for our assigned dataset (GSE153162 / E14.5) — verified against the full text. So there is no exact published per-sample value for GSE153162 to grade 1:1 against. What we CAN do, and do, is run Celline's exact, deterministic preprocess algorithm on the paper's own GSE153162 matrix and report the documented pipeline outputs + confirm run-to-run determinism. This is a faithful method reproduction / partial: it demonstrates the tool installs, runs on the paper's data, and is internally reproducible — but it is not an exact match to a published number, because the paper published none for this dataset.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a tool-paper reproduction: Celline's exact, open, deterministic QC algorithm ran cleanly on the authors' own GSE153162/GSM4635075 (E14.5) matrix and was bit-identical across reruns, confirming the central methodological claim. No exact 1:1 numeric match is possible because the paper's only per-sample QC numbers (doublet 0.366, UMI 9654, nFeature 536–3446, 70% retained) are for a different sample (E18/GSM2453043/GSE93421) that deposits no per-sample matrix — a data-availability constraint, not an authors' computational defect or fabrication. Two benign flags: the cross-sample comparison (C3–C6) is method-corroboration not reproduction, and a genuine text/code inconsistency (paper '3×MAD' vs code default n_mad=2.5). Overall yellow: faithful, deterministic, well-specified, with explainable deviations on our/data side.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.