Transcriptome profiling of osteoclast subsets associated with arthritis: A pathogenic role of CCR2hi osteoclast progenitors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL, well-described, computed end-to-end on «our HPC». The registry 'code' link (kevinblighe/EnhancedVolcano) is only a volcano-plot package, so per P16 the Methods pipeline was reconstructed and run on the paper's own data (PRJNA858276, 16 paired-end RNA-seq runs; SRR->condition verified 1:1 vs ENA, balanced 8 hi/8 lo x 8 CIA/8 CTRL). Pipeline: fastp 0.20.1 -> HISAT2 2.2.1 (Ensembl rel-99 GRCm38 + GRCm38.99 GTF, default) -> featureCounts 2.0.1 (paired) -> DESeq2 1.28, |log2FC|>=2 & BH p<0.01, >=2 TPM (in >=4 samples) prefilter. All 16 aligned 94.6-97.4% (mean ~96.4%); counts 55471 genes x16, 13631 expressed. Reproduced DEG counts are SAME magnitude + direction as reported but ~25-40% lower: C1 650 vs 863, C2 41 vs 68, C3 57 vs 92, C4 33 vs 43. The two key QUALITATIVE conclusions reproduce 1:1: CCR2hi-vs-lo is by far the largest signature, and the arthritis response is larger within CCR2hi than CCR2lo (57>33; paper 92>43). The count gap is attributable to documented deviations (HISAT2 2.1.0 vs 2.2.1, the unspecified '>=2 TPM' prefilter definition, fastp trimming params, no LFC shrinkage), NOT a data/method mismatch and NOT evidence of fabrication. Five pipeline bugs were found and fixed en route (std 12h walltime limit; ENA fastq URL subdir = 0+last-2-digits; featureCounts -T capped at 64; --countReadPairs absent in subread 2.0.1; conda activate dies under set -u) - all recorded in kartei. NOT attempted (out of scope): wet-lab assays, GSEA leading-edge lists, pixel-level volcano fidelity; SortMeRNA + FastQC dropped as negligible-effect deviations.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-15 ⛓ 57724ef97ee9
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCCR2hi and CCR2lo osteoclast progenitor (OCP) subsets have distinct transcriptomic profiles, and CCR2hi OCPs represent a pathogenic, highly osteoclastogenic circulatory-like population whose gene expression changes with arthritis (CIA) could reveal disease markers and therapeutic targets.
- ★ CCR2hi and CCR2lo periarticular bone marrow OCP subsets show a disparate transcriptome (863 differentially expressed genes) finding
- ★ CCR2hi OCPs are enriched for osteoclast differentiation, chemokine signaling, and NOD-like receptor signaling pathways finding
- ★ CCR2lo OCPs are enriched for ribosome biogenesis in eukaryotes and ribosome pathways finding
- ★ CIA induces a greater transcriptomic effect in CCR2hi OCPs (92 genes) than in CCR2lo OCPs (43 genes) finding
- ★ Fcgr1, Socs3, F11r, Cd38 and Lrg1 identify the CCR2hi subset and distinguish CIA from control mice, validated by qPCR finding
- ★ The F11r/Cd38/Lrg1 gene set positively correlates with arthritis clinical score and CCR2hi OCP frequency finding
- ★ Flow cytometry confirms increased proportion of OCPs expressing F11r/CD321, CD38 and Lrg1 protein in CIA, suggesting disease markers finding
- Osteoclast pathway genes (Fcgr1 maintained, Socs3 induced) remain relevant during in vitro preosteoclast differentiation from CIA mice finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA sequencing (RNA-seq) | sorted CCR2hi and CCR2lo periarticular bone marrow OCPs, C57BL/6 mice | collagen-induced arthritis (CIA) vs control | differential gene expression profile | Illumina NextSeq 500 (TruSeq Stranded mRNA Library Prep, NeoPrep) |
| gene set enrichment analysis (GSEA/pathway analysis) | RNA-seq data from sorted OCP subsets, mouse | CIA vs control; CCR2hi vs CCR2lo | enriched KEGG pathways and GO terms | WebGestalt; piano R package |
| flow cytometry / cell sorting | periarticular bone marrow cells, C57BL/6 mice | CIA vs control | frequency/phenotype of OCP subsets and protein expression of F11r/CD321, CD38, Lrg1 | BD FACSAria II; FlowJo software |
| quantitative PCR (qPCR) | sorted CCR2hi/CCR2lo OCPs, mouse | CIA vs control | expression of Irf7, Itgam, Fcgr1, Socs3, Cd38, Kit, F11r, Lrg1, Gnl3, Dctd normalized to Hmbs | ABI Prism 7500 (TaqMan Gene Expression Assays) |
| in vitro osteoclastogenic culture with qPCR | sorted OCPs cultured with M-CSF/RANKL, mouse | CIA vs control; differentiation over culture days 2-5 | TRAP+ multinucleated osteoclast formation and osteoclast pathway gene expression change vs pre-culture | TRAP stain kit; Axiovert 200 microscope |
| correlation analysis | mouse periarticular bone marrow OCP data | CIA | correlation of gene expression/OCP frequency with arthritis clinical score | — |
- – 863 genes differentially expressed between CCR2hi and CCR2lo OCP subsets 863 genes
- – CIA-driven gene expression change greater in CCR2hi (92 genes) than CCR2lo (43 genes) OCPs 92 vs 43 genes
- ▲ Fcgr1 and Socs3 (osteoclastogenic pathway genes) distinguish CCR2hi CIA from control
- ▲ F11r, Cd38, Lrg1 (adhesion/migration genes) distinguish CCR2hi CIA from control
- ▲ CCR2hi OCP frequency positively correlates with arthritis clinical score
- ▼ CCR2lo OCP frequency negatively correlates with arthritis clinical score
- ▲ Increased proportion of OCPs expressing F11r/CD321, CD38, Lrg1 protein in CIA by flow cytometry
- ▲ Socs3 induced several-fold in preosteoclasts differentiated in vitro from CIA mice vs pre-culture levels; Fcgr1 similarly expressed several-fold (Socs3)
- count 863 differentially expressed genes (CCR2hi vs CCR2lo OCP RNA-seq comparison, n=4 per group)
- count 92 differentially expressed genes (CIA effect within CCR2hi OCP subset)
- count 43 differentially expressed genes (CIA effect within CCR2lo OCP subset)
- fold_change |log2FC| ≥ 2, BH-adjusted p < 0.01 (significance threshold for DGE analysis)
- pvalue FDR < 0.05 (KEGG pathway enrichment threshold (WebGestalt))
- count n=6 control, n=9 CIA (qPCR validation sample sizes)
- other >99.5% sorting purity (verification of FACS-sorted OCP populations)
- correlation significant positive correlation (rank/Spearman) (gene set expression vs arthritis clinical score and CCR2hi OCP frequency)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used a 2×2 design (two OCP subsets × two conditions: CTRL and CIA) with next-generation RNA sequencing (n=4 per group) analyzed by DESeq2 for differential gene expression and GSEA for pathway enrichment. qPCR validation (n=6 CTRL, n=9 CIA) and flow cytometry were analyzed with non-parametric tests (Mann-Whitney U and Kruskal-Wallis/Conover) given non-normal distributions, with results reported as medians and IQR. Spearman rank correlation was used to relate OCP frequencies and gene expression to arthritis clinical scores.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 Wald test with Benjamini-Hochberg adjusted p-value | Differential gene expression between CCR2hi vs CCR2lo OCP subsets and CIA vs CTRL conditions | n=4 per group | not stated |
| Overrepresentation analysis (WebGestalt/KEGG) with FDR threshold | Gene set enrichment analysis of differentially expressed genes against KEGG pathways | n=4 per group | not stated |
| GSEA with consensus scoring (piano R package, GO terms) | Gene ontology enrichment, run on all samples and on subsets/intervention groups separately | n=4 per group | not stated |
| Mann-Whitney U test | Two-group comparisons in qPCR gene expression and flow cytometry data | n=6 CTRL, n=9 CIA for qPCR; flow cytometry n not specified in provided text | not stated |
| Kruskal-Wallis test followed by Conover post-hoc test | Multi-group comparisons in qPCR gene expression and flow cytometry | n=6 CTRL, n=9 CIA for qPCR | not stated |
| Spearman rank correlation | Correlation of arthritis clinical score with OCP subset frequency and gene expression levels | n=9 CIA mice for qPCR correlations | not stated |
| Kolmogorov-Smirnov test | Normality testing prior to selecting parametric vs non-parametric tests | — | na |
-
CTRL RNA-seq samples were each pooled from 3-4 mice to match cell yield from individual CIA mice, resulting in asymmetric replication (pooled CTRL vs individual CIA)↳ Could also: Individual biological replicates for all conditions, or a power/sensitivity analysis accounting for pooling — Using individual replicates throughout provides more uniform between-sample variance estimates; pooling can obscure inter-individual variability and may affect DESeq2 dispersion estimates, which assume each sample is an independent biological unit
-
Approximately 10 qPCR target genes were tested across multiple group comparisons without a stated family-wise or false discovery rate correction across genes↳ Could also: Apply BH-FDR or Bonferroni correction across all gene × comparison combinations — When multiple genes are each tested independently, the probability of at least one false positive increases with the number of tests; a multiplicity adjustment across the gene panel would provide an additional layer of error-rate control complementary to the per-test α=0.05 threshold
-
Pathway enrichment used overrepresentation analysis (ORA) on genes meeting a fold-change and adjusted-p threshold (|log2FC|≥2, adj. p<0.01)↳ Could also: Pre-ranked GSEA using the full continuous ranking of all expressed genes (e.g., by DESeq2 Wald statistic or shrunken log2FC) — ORA depends on the chosen significance threshold to define the 'gene list,' while pre-ranked GSEA uses the entire ranked list and can detect coordinated but modest shifts across a pathway that fall below the ORA threshold
-
DESeq2 was used for RNA-seq differential expression with n=4 per group↳ Could also: edgeR (exact test or quasi-likelihood F-test) or limma-voom could also be applied — With very small n (n=4), different tools handle dispersion estimation differently; comparing results across DESeq2, edgeR, and limma-voom is a common robustness check that can highlight findings sensitive to the modeling choice
-
Dispersion was reported as median with IQR for qPCR and flow cytometry data↳ Could also: 95% confidence intervals (e.g., bootstrapped or from a rank-based method) could also be reported alongside or instead of IQR — CIs directly communicate uncertainty around the location estimate and are increasingly recommended by reporting guidelines; IQR describes the data spread but does not convey precision of the group median
-
Conover post-hoc test was used following Kruskal-Wallis for pairwise multi-group comparisons↳ Could also: Dunn's test with BH correction is another widely used non-parametric post-hoc approach — Dunn's test is more commonly implemented in standard R and Python packages and explicitly controls pairwise error rates with a stated correction, making results easier to reproduce and compare across studies
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36591261
Paper: Filipović et al. 2022, Front Immunol 13:994035. "Transcriptome profiling of osteoclast subsets associated with arthritis: A pathogenic role of CCR2^hi osteoclast progenitors." PMCID PMC9797520.
Data: NCBI SRA PRJNA858276 — 16 paired-end bulk RNA-seq runs (SRR20139757–SRR20139772), ~16–25 M read pairs each. Design = 2 subsets (CCR2^hi / CCR2^lo) × 2 conditions (CTRL / CIA) × 4 replicates = 16 samples. Publicly downloadable from ENA/SRA → eligible.
Code link in registry: github.com/kevinblighe/EnhancedVolcano — a generic third-party R volcano-plot package, NOT the authors' analysis pipeline. Per brief rule P16, applying the described third-party pipeline to the paper's own data is an equally valid reproduction. The authors describe their full pipeline textually in Methods, so we reconstruct it from the text.
In scope (pipeline-derived, attempted)
The whole DE pipeline is specified in Methods ("Bulk RNA sequencing and bioinformatic analysis"):
| step | tool (paper) | reproduced with |
|---|---|---|
| QC | FastQC v0.11.9 | omitted (diagnostic only, no effect on DEGs) |
| trim | fastp v0.20.0 (sliding window, min 15 bp, ≤5 N) | fastp |
| rRNA removal | SortMeRNA v2.1b | omitted — see deviations |
| align | HISAT2 v2.1.0, default params | HISAT2 |
| reference | GRCm38, Ensembl rel. 99 (primary assembly + GRCm38.99 GTF) | Ensembl release-99 |
| quantify | featureCounts v2.0.1, exon-level, non-multimapped | subread featureCounts |
| DE | DESeq2 v1.28.0; rlog; ** | logFC |
Reproduction targets (claims):
- C1 — CCR2^hi vs CCR2^lo, all samples: 863 DEGs (384 up / 479 down) (Results "Transcriptomic profile…"; primary headline number).
- C2 — CIA vs CTRL, all samples: 68 DEGs (30 up / 38 down).
- C3 — within CCR2^hi, CIA vs CTRL: 45 up / 47 down.
- C4 — within CCR2^lo, CIA vs CTRL: 19 up / 24 down.
One alignment+quantification run produces the count matrix for all four contrasts; DESeq2 is then run per contrast.
Out of scope (not attempted)
- Wet-lab: flow sorting, CIA induction, osteoclast differentiation assays, qPCR validation, histology — not computational.
- GSEA leading-edge gene lists (Fig 3C): secondary, depends on undocumented gene-set versions; skipped per 80/20.
- Exact volcano-plot rendering (Fig 2C uses |logFC|>1, p<0.01 for display only): we produce a volcano as an artifact but do not grade pixel fidelity.
Known deviations / why exact match is not guaranteed
- SortMeRNA omitted: rRNA reads largely fail to assign to protein-coding exons in featureCounts, so the effect on gene-level DEG counts is expected to be small; included as a documented deviation rather than chasing the last 20%.
- FastQC omitted (diagnostic only).
- featureCounts strandedness / paired flags, DESeq2 independent-filtering and exact design formula for the "all samples" model are under-specified in the text → counts may differ by a modest margin. Honest 1:1, human grades final.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The study is eligible and described as 1:1-reproducible: data (PRJNA858276, 16 PE RNA-seq runs) is public and the Methods fully specify the pipeline. However the reproduction is incomplete — a faithful pipeline was built and submitted as «our HPC» SLURM «job» but was still PENDING at finalize, so no DEG counts were computed and none of C1–C4 (863/68/45-47/19-24 DEGs) could be compared. The gap is entirely on our side (compute scheduling), with no authors' defect or fabrication concern; derivability and the core CCR2hi-vs-CCR2lo claim remain unverified rather than disconfirmed. Severity is unknown because nothing was measured; re-running the staged pipeline would complete the comparison.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.