Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Mapping the Development of Human Spermatogenesis Using Transcriptomics-Based Data: A Scoping Review.

Int J Mol Sci · 2024
L1 No computation 2/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
Reproduction agent’s raw note

DROP / non_pipeline. The publication (Kwaspen, Kanbar, Wyns 2024, IJMS) is a PRISMA SCOPING REVIEW of human-spermatogenesis transcriptomics: a narrative literature mapping, not a bioinformatic study. Methods state verbatim 'The synthesis of the results is provided in a narrative, descriptive format'; the work was PubMed/Embase/GEO screening + Endnote/Excel extraction, not data analysis. It produces NO original pipeline-derived computational result: Figure 1 is a PRISMA flowchart, Figure 2 a hand-drawn developmental timeline, Table 1 a catalogue of reviewed primary studies, and all reported numbers are tallies quoted from those studies. The code link on file (github.com/eisascience/HISTA) and the data accession (GSE63818) are two UNRELATED cited resources sitting in Table 1: HISTA is the 'Other Database' link for Mahyari et al. 2021 (a third-party Shiny atlas built on that group's own data, not on GSE63818), and GSE63818 is the GEO accession for Guo et al. 2015 (prenatal PGC dataset). They belong to two different reviewed papers and were never combined into any analysis by these authors. So there is no claim of THIS paper to regenerate; pairing the two links (e.g. running HISTA on GSE63818) would fabricate a pipeline the paper never describes and reproduce zero of its claims. Correct, honest outcome is a drop with no «our HPC» compute spent. NOT attempted (out of scope): re-running HISTA or the Guo-2015 pipeline on their respective data — that would reproduce those cited primary studies, a different RU, not this review. No possible-fabrication concern about the paper itself: it is transparently a review and does not present computed values as its own.

💻 Code ↗ 🗄 Data: GSE63818

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-14 ⛓ fe971580b695
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This scoping review asks how the human testis develops and matures transcriptomically from the fetal period to late adulthood, using single-cell RNA-seq data, in order to identify cell signaling pathways and conditions that could optimize in vitro maturation (IVM) of human testicular tissue toward in vitro spermatogenesis.

Core claims
  • scRNA-seq consistently shows major transcriptional modifications of the testis around 11 years of age before reaching the adult state finding
  • scRNA-seq data favor a paradigm shift in which Adark and Apale spermatogonia cannot be distinctly identified among the different SSC substates (States 0-4) finding
  • The mitotically arrested fetal germ cell (State f0 SSC) loses pluripotency hallmarks but expresses SSC markers (PIWIL4, EGR4, MSL3, TSPAN33), acting as precursor of the neonatal/adult undifferentiated SSC pool (State 0) mechanism
  • Undifferentiated SSCs are favored in culture in the presence of an AKT-signaling pathway inhibitor finding
  • Involvement of the oxidative phosphorylation (OXPHOS) pathway depends on the maturational state of the cells finding
  • Multiple conserved signaling pathways (BMP, NODAL, KIT, NOTCH, WNT, Hedgehog/DHH, TGF-β) govern germ-soma and soma-soma communication across developmental stages mechanism
  • Data on the somatic cell lineage, especially Sertoli cells, are limited due to technical issues related to cell size finding
  • A scoping review synthesizing scRNA-seq studies of native and cultured human testicular cells across developmental stages to inform IVM resource
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-sequencing (scRNA-seq) human prenatal/fetal testicular tissue (4-25 WPC) none germ and somatic cell transcriptional states and gene expression
single-cell RNA-sequencing (scRNA-seq) human neonatal testis (2-7 days, 5 months) none germ cell states (PGC-like, PreSPG-1/2) and somatic cell transcriptomes
single-cell RNA-sequencing (scRNA-seq) human prepubertal/peripubertal testis (1-14 years) none SSC substates, metabolic/signaling pathway gene expression
single-cell RNA-sequencing (scRNA-seq) human adult testis (17-55 years) none germ and somatic cell transcriptomes across spermatogenesis
single-cell RNA-sequencing (scRNA-seq) human elderly testis (60-90 years) none testicular cell transcriptomes
single-cell RNA-sequencing (scRNA-seq) cultured human fetal testicular cells / adult SSCs / adult PTMCs in vitro cell culture (incl. AKT-signaling pathway inhibitor) transcriptional dysregulation and SSC maintenance markers
DNA methylome analysis human fetal germ cells (7-19 WPC) none DNA methylation levels over time
immunofluorescent staining human prepubertal testis (~11-year-old) none undifferentiated SSC morphology (round to flat)
Key results
  • Mitotic arrested GC fraction rose from <10% to >60%, becoming the most abundant GC state at 23 WPC <10% to >60%
  • At 8-10 WPC mainly mitotic quiescent GCs present, with slightly more than 25% of mitotic GCs in proliferative state >25%
  • 17 cell-cycle-arrest-related genes upregulated in mitotic arrest GC state, 8 of them (NANOS2, CDKN2B, TCF7L2, CDK6, PCBP4, NFATC1, SESN3, PDK1) specific to that state 17 genes
  • Differentiating spermatogonia (States 2-4), spermatocytes and rare spermatids first identifiable at age 11, while 1-7 year olds have only undifferentiated SSCs (States 0-1)
  • OXPHOS and ATPase genes upregulated in first year of life, stable for ~10 years, then downregulated around age 11 with simultaneous transient hypoxia pathway re-upregulation
  • Hypoxia-related genes upregulated in SSCs only during the first year of life then declined
  • Fetal GC DNA methylation decreased over time (7-19 WPC), consistent with base-excision-repair pathway upregulation
  • Garcia-Alonso et al. identified two resident testis-specific macrophage populations (osteoclast-like interstitial and microglia-like intratubular) in 6-21 WPC testis
Key statistics
  • count 1091 articles identified (records identified from PubMed, Embase, and GEO DataSet)
  • count 26 studies included (studies fulfilling eligibility criteria)
  • count 69 prenatal samples (4-25 WPC) (prenatal scRNA-seq samples)
  • count 53 adult testicular samples (17-55 years) (adult scRNA-seq samples)
  • count 14 testis samples (60-90 years) (elderly scRNA-seq samples)
  • other 94,200 to 279,300 copies of mRNA per germ cell (average gene expression per GC in 4-19 WPC testis)
  • other 50,900 to 147,000 copies of mRNA per somatic cell (average gene expression per somatic cell in 4-19 WPC testis)
  • count 5 GC clusters (Chitiashvili); 7 clusters on reanalyzed Li et al. data (germ cell clustering of 4-16 WPC samples)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper is a scoping review following PRISMA guidelines; the authors performed no original statistical tests. The analytic approach consists entirely of narrative, descriptive synthesis of scRNA-seq findings extracted from 26 eligible studies spanning fetal to elderly human testicular tissue. Quantitative data reported by the authors are limited to counts of included studies and samples per developmental stage; all biological findings (cell states, GO terms, pathway enrichments) are summarized as reported in the primary sources.

Replicationunclear Sample size26 studies met eligibility criteria; sample counts per developmental stage are reported narratively (e.g., 69 prenatal samples, 53 adult samples across included studies); no power calculation or formal sample-size justification is described for the review itself GroupsDevelopmental stages compared narratively: fetal (4–25 WPC), neonatal, prepubertal (1–11 years), peripubertal (13–14 years), adult (17–55 years), elderly (60–90 years) Pairingna Randomization/blindingstated Dispersionnone
Approaches that could also have been used
  • The review used narrative synthesis without any quantitative pooling of findings across the 26 included scRNA-seq studies
    Could also: A semi-quantitative vote-counting or frequency analysis—tabulating how many independent studies reported each cell state, pathway, or marker gene—could also have been used alongside the narrative — Vote-counting provides a transparent, reproducible signal for which findings are consistently replicated across studies versus reported in only one or two, without requiring raw-data access or harmonization of heterogeneous scRNA-seq pipelines
  • The review did not formally assess the methodological quality or risk of bias of included studies
    Could also: Adapted quality-appraisal tools (e.g., a scRNA-seq-specific checklist covering sequencing depth, clustering resolution, cell-type validation, and sample provenance) could also have been applied to each included study — Quality assessment, though not mandatory in scoping reviews, allows readers to gauge how much weight to place on individual findings and is increasingly recommended even in scoping contexts to improve transparency
  • Inter-rater agreement during dual independent screening was not quantified, only described as resolved by consensus
    Could also: Reporting a kappa coefficient or percent agreement at both the title/abstract and full-text screening stages could also have been included — A quantitative agreement statistic gives readers a standardized, reproducible measure of how consistently the eligibility criteria were applied across screeners, which is a common reporting standard for PRISMA-based reviews
  • Cell-type labels and marker gene sets from included studies were harmonized narratively across different clustering schemes and naming conventions
    Could also: Computational cross-study integration (e.g., Harmony, Seurat label transfer, or scANVI applied to publicly available raw count matrices) could also have been used to build a unified reference atlas for direct cell-state comparison — Computational integration would enable quantitative comparison of cell proportions and transcriptional distances across studies, reducing ambiguity introduced by different authors using different resolutions, reference datasets, and terminology for equivalent cell populations
  • The search covered three databases (PubMed, Embase, GEO DataSet) with a lower publication-date bound of 2009
    Could also: Adding preprint servers (bioRxiv, Research Square) and screening reference lists of included systematic reviews in parallel could also have been used to increase recall — ScRNA-seq is a rapidly evolving field where relevant findings often appear first as preprints; supplementing database searches with preprint servers and citation chasing is increasingly recommended for scoping reviews in fast-moving domains
Software: Endnote 21.0.1 · Microsoft Excel v2311 Build 16.0.17029.20178

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
5
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope analysis — pmid-39000031

Title: Mapping the Development of Human Spermatogenesis Using Transcriptomics-Based Data: A Scoping Review. Authors: Kwaspen L, Kanbar M, Wyns C. (UCLouvain) · Int J Mol Sci 2024 · DOI 10.3390/ijms25136925 · PMID 39000031 · PMCID PMC11241379.

Verdict: DROP — non_pipeline (text-mining false positive)

This publication is a PRISMA scoping review of the literature on human spermatogenesis transcriptomics. It produces no original pipeline-derived computational result. The code/data links recorded for this RU are cited resources inside the review's summary table, not an analysis the authors ran. There is therefore nothing pipeline-derived to reproduce.

Evidence (verbatim, from PMC full text)

  • Study type / synthesis method (Methods):

    "This is a scoping review with a published protocol, written via the PRISMA guidelines for protocols..." "The synthesis of the results is provided in a narrative, descriptive format."

  • The work performed was literature screening + extraction, not data analysis:

    "The databases PubMed, Embase, and the GEO DataSet were screened for articles from 2009 to 31 March 2024." "Full-text screening was performed separately, independently, and blinded by K.L. and K.M." using "Endnote (v.21.0.1) and an Excel file".

  • The Results report values quoted from the reviewed primary studies, not computed by the authors:

    "ScRNA-seq data were identified for 69 prenatal samples..." (a tally of what the reviewed studies reported, not a recomputed value).

  • Figures are not computational outputs:
    • Figure 1 = "PRISMA flowchart of screened papers..." (attrition counts).
    • Figure 2 = hand-drawn "Overview of the germ cell lineage development..." (a developmental timeline schematic).
  • Table 1 = a catalogue of the included primary studies, with columns Author, Year, #Participants, Age, GSE Number, Other Database, Reference.

Why the recorded code + data links are not in scope

RU field Value What it actually is in the paper Reproducible result?
code_url github.com/eisascience/HISTA A single link in Table 1's "Other Database" column for Mahyari et al. 2021 (a different primary study reviewed here). HISTA is a third-party Shiny atlas built by the Mahyari/OHSU group on their own infertility-testis data — not on GSE63818, and not run by these review authors. No — not the authors' code, and not applied to this dataset.
data_accession GSE63818 The GEO accession listed in Table 1 for F. Guo et al. 2015 (prenatal PGC dataset). Cited as a catalogued dataset; the review authors did not reprocess it. No — never reprocessed by the authors.

The two links come from two different reviewed papers and were never combined into any analysis. Pairing them (e.g. "run HISTA on GSE63818") would fabricate a pipeline the publication never describes and would reproduce zero claim of this paper. Per HARD RULE 6 / 80-20, the honest outcome is a drop, not a manufactured result.

In scope (pipeline-derived results the authors produced): NONE

There is no bioinformatic pipeline output (no UMAP/clustering/DE/marker table) generated by these authors. The only author-computed numbers are PRISMA screening counts (identified / excluded / included study counts), which are literature-mapping tallies, explicitly out of scope for a pipeline reproduction (HARD RULE 2).

Out of scope (not attempted)

  • Re-running other groups' tools (HISTA / Guo-2015 pipeline) on their respective data: that reproduces those cited papers, not this scoping review, and is a different RU's job.

Decision

drop · drop_reason = non_pipeline · no «our HPC» compute spent (correctly — there is no pipeline result to regenerate).

No individual results have been recorded for this entry yet.

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

This publication is a PRISMA scoping review (Kwaspen, Kanbar, Wyns 2024) with no original pipeline-derived result — Methods explicitly use 'narrative, descriptive' synthesis, and every reported number is a screening count or a tally quoted from reviewed primary studies. The recorded code_url (HISTA) and data_accession (GSE63818) are two unrelated resources cited in Table 1 belonging to different reviewed studies, not an analysis these authors ran, so there is nothing to regenerate. The drop / non_pipeline verdict is correct and honest: q1/q2 are red because no comparable data/endpoint of this paper exists, but there is no fabrication concern and no authors'-side defect (q5/q7/q8 green). Severity is negligible — the deviation flags simply reflect that the study is out of scope for reproduction, not a discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

37.6 k
tokens (I/O) · 1.8 M incl. cache
4 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.