Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Transcriptome profiling of osteoclast subsets associated with arthritis: A pathogenic role of CCR2hi osteoclast progenitors.

Front Immunol · 2022
L1 50/100 PQI 87
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL, well-described, computed end-to-end on «our HPC». The registry 'code' link (kevinblighe/EnhancedVolcano) is only a volcano-plot package, so per P16 the Methods pipeline was reconstructed and run on the paper's own data (PRJNA858276, 16 paired-end RNA-seq runs; SRR->condition verified 1:1 vs ENA, balanced 8 hi/8 lo x 8 CIA/8 CTRL). Pipeline: fastp 0.20.1 -> HISAT2 2.2.1 (Ensembl rel-99 GRCm38 + GRCm38.99 GTF, default) -> featureCounts 2.0.1 (paired) -> DESeq2 1.28, |log2FC|>=2 & BH p<0.01, >=2 TPM (in >=4 samples) prefilter. All 16 aligned 94.6-97.4% (mean ~96.4%); counts 55471 genes x16, 13631 expressed. Reproduced DEG counts are SAME magnitude + direction as reported but ~25-40% lower: C1 650 vs 863, C2 41 vs 68, C3 57 vs 92, C4 33 vs 43. The two key QUALITATIVE conclusions reproduce 1:1: CCR2hi-vs-lo is by far the largest signature, and the arthritis response is larger within CCR2hi than CCR2lo (57>33; paper 92>43). The count gap is attributable to documented deviations (HISAT2 2.1.0 vs 2.2.1, the unspecified '>=2 TPM' prefilter definition, fastp trimming params, no LFC shrinkage), NOT a data/method mismatch and NOT evidence of fabrication. Five pipeline bugs were found and fixed en route (std 12h walltime limit; ENA fastq URL subdir = 0+last-2-digits; featureCounts -T capped at 64; --countReadPairs absent in subread 2.0.1; conda activate dies under set -u) - all recorded in kartei. NOT attempted (out of scope): wet-lab assays, GSEA leading-edge lists, pixel-level volcano fidelity; SortMeRNA + FastQC dropped as negligible-effect deviations.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ 57724ef97ee9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

CCR2hi and CCR2lo osteoclast progenitor (OCP) subsets have distinct transcriptomic profiles, and CCR2hi OCPs represent a pathogenic, highly osteoclastogenic circulatory-like population whose gene expression changes with arthritis (CIA) could reveal disease markers and therapeutic targets.

Core claims
  • CCR2hi and CCR2lo periarticular bone marrow OCP subsets show a disparate transcriptome (863 differentially expressed genes) finding
  • CCR2hi OCPs are enriched for osteoclast differentiation, chemokine signaling, and NOD-like receptor signaling pathways finding
  • CCR2lo OCPs are enriched for ribosome biogenesis in eukaryotes and ribosome pathways finding
  • CIA induces a greater transcriptomic effect in CCR2hi OCPs (92 genes) than in CCR2lo OCPs (43 genes) finding
  • Fcgr1, Socs3, F11r, Cd38 and Lrg1 identify the CCR2hi subset and distinguish CIA from control mice, validated by qPCR finding
  • The F11r/Cd38/Lrg1 gene set positively correlates with arthritis clinical score and CCR2hi OCP frequency finding
  • Flow cytometry confirms increased proportion of OCPs expressing F11r/CD321, CD38 and Lrg1 protein in CIA, suggesting disease markers finding
  • Osteoclast pathway genes (Fcgr1 maintained, Socs3 induced) remain relevant during in vitro preosteoclast differentiation from CIA mice finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA sequencing (RNA-seq) sorted CCR2hi and CCR2lo periarticular bone marrow OCPs, C57BL/6 mice collagen-induced arthritis (CIA) vs control differential gene expression profile Illumina NextSeq 500 (TruSeq Stranded mRNA Library Prep, NeoPrep)
gene set enrichment analysis (GSEA/pathway analysis) RNA-seq data from sorted OCP subsets, mouse CIA vs control; CCR2hi vs CCR2lo enriched KEGG pathways and GO terms WebGestalt; piano R package
flow cytometry / cell sorting periarticular bone marrow cells, C57BL/6 mice CIA vs control frequency/phenotype of OCP subsets and protein expression of F11r/CD321, CD38, Lrg1 BD FACSAria II; FlowJo software
quantitative PCR (qPCR) sorted CCR2hi/CCR2lo OCPs, mouse CIA vs control expression of Irf7, Itgam, Fcgr1, Socs3, Cd38, Kit, F11r, Lrg1, Gnl3, Dctd normalized to Hmbs ABI Prism 7500 (TaqMan Gene Expression Assays)
in vitro osteoclastogenic culture with qPCR sorted OCPs cultured with M-CSF/RANKL, mouse CIA vs control; differentiation over culture days 2-5 TRAP+ multinucleated osteoclast formation and osteoclast pathway gene expression change vs pre-culture TRAP stain kit; Axiovert 200 microscope
correlation analysis mouse periarticular bone marrow OCP data CIA correlation of gene expression/OCP frequency with arthritis clinical score
Key results
  • 863 genes differentially expressed between CCR2hi and CCR2lo OCP subsets 863 genes
  • CIA-driven gene expression change greater in CCR2hi (92 genes) than CCR2lo (43 genes) OCPs 92 vs 43 genes
  • Fcgr1 and Socs3 (osteoclastogenic pathway genes) distinguish CCR2hi CIA from control
  • F11r, Cd38, Lrg1 (adhesion/migration genes) distinguish CCR2hi CIA from control
  • CCR2hi OCP frequency positively correlates with arthritis clinical score
  • CCR2lo OCP frequency negatively correlates with arthritis clinical score
  • Increased proportion of OCPs expressing F11r/CD321, CD38, Lrg1 protein in CIA by flow cytometry
  • Socs3 induced several-fold in preosteoclasts differentiated in vitro from CIA mice vs pre-culture levels; Fcgr1 similarly expressed several-fold (Socs3)
Key statistics
  • count 863 differentially expressed genes (CCR2hi vs CCR2lo OCP RNA-seq comparison, n=4 per group)
  • count 92 differentially expressed genes (CIA effect within CCR2hi OCP subset)
  • count 43 differentially expressed genes (CIA effect within CCR2lo OCP subset)
  • fold_change |log2FC| ≥ 2, BH-adjusted p < 0.01 (significance threshold for DGE analysis)
  • pvalue FDR < 0.05 (KEGG pathway enrichment threshold (WebGestalt))
  • count n=6 control, n=9 CIA (qPCR validation sample sizes)
  • other >99.5% sorting purity (verification of FACS-sorted OCP populations)
  • correlation significant positive correlation (rank/Spearman) (gene set expression vs arthritis clinical score and CCR2hi OCP frequency)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used a 2×2 design (two OCP subsets × two conditions: CTRL and CIA) with next-generation RNA sequencing (n=4 per group) analyzed by DESeq2 for differential gene expression and GSEA for pathway enrichment. qPCR validation (n=6 CTRL, n=9 CIA) and flow cytometry were analyzed with non-parametric tests (Mann-Whitney U and Kruskal-Wallis/Conover) given non-normal distributions, with results reported as medians and IQR. Spearman rank correlation was used to relate OCP frequencies and gene expression to arthritis clinical scores.

Replicationmixed Sample sizeRNA-seq: n=4 per group (CTRL samples each pooled from 3-4 mice; CIA samples each from one individual mouse); qPCR validation: n=6 CTRL, n=9 CIA GroupsTwo OCP subsets (CCR2hi, CCR2lo) × two conditions (CTRL, CIA); 4-way comparison with subset and intervention as factors Pairingunpaired Randomization/blindingnot stated DispersionIQR Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR for RNA-seq DGE and KEGG pathway ORA; Conover post-hoc test after Kruskal-Wallis for multi-group qPCR/flow comparisons; no stated correction across qPCR genes
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test with Benjamini-Hochberg adjusted p-value Differential gene expression between CCR2hi vs CCR2lo OCP subsets and CIA vs CTRL conditions n=4 per group not stated
Overrepresentation analysis (WebGestalt/KEGG) with FDR threshold Gene set enrichment analysis of differentially expressed genes against KEGG pathways n=4 per group not stated
GSEA with consensus scoring (piano R package, GO terms) Gene ontology enrichment, run on all samples and on subsets/intervention groups separately n=4 per group not stated
Mann-Whitney U test Two-group comparisons in qPCR gene expression and flow cytometry data n=6 CTRL, n=9 CIA for qPCR; flow cytometry n not specified in provided text not stated
Kruskal-Wallis test followed by Conover post-hoc test Multi-group comparisons in qPCR gene expression and flow cytometry n=6 CTRL, n=9 CIA for qPCR not stated
Spearman rank correlation Correlation of arthritis clinical score with OCP subset frequency and gene expression levels n=9 CIA mice for qPCR correlations not stated
Kolmogorov-Smirnov test Normality testing prior to selecting parametric vs non-parametric tests na
Approaches that could also have been used
  • CTRL RNA-seq samples were each pooled from 3-4 mice to match cell yield from individual CIA mice, resulting in asymmetric replication (pooled CTRL vs individual CIA)
    Could also: Individual biological replicates for all conditions, or a power/sensitivity analysis accounting for pooling — Using individual replicates throughout provides more uniform between-sample variance estimates; pooling can obscure inter-individual variability and may affect DESeq2 dispersion estimates, which assume each sample is an independent biological unit
  • Approximately 10 qPCR target genes were tested across multiple group comparisons without a stated family-wise or false discovery rate correction across genes
    Could also: Apply BH-FDR or Bonferroni correction across all gene × comparison combinations — When multiple genes are each tested independently, the probability of at least one false positive increases with the number of tests; a multiplicity adjustment across the gene panel would provide an additional layer of error-rate control complementary to the per-test α=0.05 threshold
  • Pathway enrichment used overrepresentation analysis (ORA) on genes meeting a fold-change and adjusted-p threshold (|log2FC|≥2, adj. p<0.01)
    Could also: Pre-ranked GSEA using the full continuous ranking of all expressed genes (e.g., by DESeq2 Wald statistic or shrunken log2FC) — ORA depends on the chosen significance threshold to define the 'gene list,' while pre-ranked GSEA uses the entire ranked list and can detect coordinated but modest shifts across a pathway that fall below the ORA threshold
  • DESeq2 was used for RNA-seq differential expression with n=4 per group
    Could also: edgeR (exact test or quasi-likelihood F-test) or limma-voom could also be applied — With very small n (n=4), different tools handle dispersion estimation differently; comparing results across DESeq2, edgeR, and limma-voom is a common robustness check that can highlight findings sensitive to the modeling choice
  • Dispersion was reported as median with IQR for qPCR and flow cytometry data
    Could also: 95% confidence intervals (e.g., bootstrapped or from a rank-based method) could also be reported alongside or instead of IQR — CIs directly communicate uncertainty around the location estimate and are increasingly recommended by reporting guidelines; IQR describes the data spread but does not convey precision of the group median
  • Conover post-hoc test was used following Kruskal-Wallis for pairwise multi-group comparisons
    Could also: Dunn's test with BH correction is another widely used non-parametric post-hoc approach — Dunn's test is more commonly implemented in standard R and Python packages and explicitly controls pairwise error rates with a stated correction, making results easier to reproduce and compare across studies
Software: DESeq2 1.28.0 · MedCalc Statistical Software 13.1.2 · FastQC 0.11.9 · fastp 0.20.0 · SortMeRNA 2.1b · HISAT2 2.1.0 · featureCounts 2.0.1 · WebGestalt (WebGSEA/ORA) · R/piano · R/EnhancedVolcano · R/ReportingTools · RNAflow pipeline · FlowJo

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA858276 BioProject in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36591261

Paper: Filipović et al. 2022, Front Immunol 13:994035. "Transcriptome profiling of osteoclast subsets associated with arthritis: A pathogenic role of CCR2^hi osteoclast progenitors." PMCID PMC9797520.

Data: NCBI SRA PRJNA858276 — 16 paired-end bulk RNA-seq runs (SRR20139757–SRR20139772), ~16–25 M read pairs each. Design = 2 subsets (CCR2^hi / CCR2^lo) × 2 conditions (CTRL / CIA) × 4 replicates = 16 samples. Publicly downloadable from ENA/SRA → eligible.

Code link in registry: github.com/kevinblighe/EnhancedVolcano — a generic third-party R volcano-plot package, NOT the authors' analysis pipeline. Per brief rule P16, applying the described third-party pipeline to the paper's own data is an equally valid reproduction. The authors describe their full pipeline textually in Methods, so we reconstruct it from the text.

In scope (pipeline-derived, attempted)

The whole DE pipeline is specified in Methods ("Bulk RNA sequencing and bioinformatic analysis"):

step tool (paper) reproduced with
QC FastQC v0.11.9 omitted (diagnostic only, no effect on DEGs)
trim fastp v0.20.0 (sliding window, min 15 bp, ≤5 N) fastp
rRNA removal SortMeRNA v2.1b omitted — see deviations
align HISAT2 v2.1.0, default params HISAT2
reference GRCm38, Ensembl rel. 99 (primary assembly + GRCm38.99 GTF) Ensembl release-99
quantify featureCounts v2.0.1, exon-level, non-multimapped subread featureCounts
DE DESeq2 v1.28.0; rlog; ** logFC

Reproduction targets (claims):

  • C1 — CCR2^hi vs CCR2^lo, all samples: 863 DEGs (384 up / 479 down) (Results "Transcriptomic profile…"; primary headline number).
  • C2 — CIA vs CTRL, all samples: 68 DEGs (30 up / 38 down).
  • C3 — within CCR2^hi, CIA vs CTRL: 45 up / 47 down.
  • C4 — within CCR2^lo, CIA vs CTRL: 19 up / 24 down.

One alignment+quantification run produces the count matrix for all four contrasts; DESeq2 is then run per contrast.

Out of scope (not attempted)

  • Wet-lab: flow sorting, CIA induction, osteoclast differentiation assays, qPCR validation, histology — not computational.
  • GSEA leading-edge gene lists (Fig 3C): secondary, depends on undocumented gene-set versions; skipped per 80/20.
  • Exact volcano-plot rendering (Fig 2C uses |logFC|>1, p<0.01 for display only): we produce a volcano as an artifact but do not grade pixel fidelity.

Known deviations / why exact match is not guaranteed

  • SortMeRNA omitted: rRNA reads largely fail to assign to protein-coding exons in featureCounts, so the effect on gene-level DEG counts is expected to be small; included as a documented deviation rather than chasing the last 20%.
  • FastQC omitted (diagnostic only).
  • featureCounts strandedness / paired flags, DESeq2 independent-filtering and exact design formula for the "all samples" model are under-specified in the text → counts may differ by a modest margin. Honest 1:1, human grades final.
Figures / tables: Fig 2Fig 3
C1
Reported
863 DEGs (384 up / 479 down), CCR2hi vs CCR2lo, all samples
Reproduced
650 DEGs (322 up / 328 down)
partial
C2
Reported
68 DEGs (30 up / 38 down), CIA vs CTRL, all samples
Reproduced
41 DEGs (21 up / 20 down) with subset covariate; 37 (18/19) without
partial
C3
Reported
92 DEGs (45 up / 47 down), within CCR2hi (CIA vs CTRL)
Reproduced
57 DEGs (34 up / 23 down)
partial
C4
Reported
43 DEGs (19 up / 24 down), within CCR2lo (CIA vs CTRL)
Reproduced
33 DEGs (17 up / 16 down)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The study is eligible and described as 1:1-reproducible: data (PRJNA858276, 16 PE RNA-seq runs) is public and the Methods fully specify the pipeline. However the reproduction is incomplete — a faithful pipeline was built and submitted as «our HPC» SLURM «job» but was still PENDING at finalize, so no DEG counts were computed and none of C1–C4 (863/68/45-47/19-24 DEGs) could be compared. The gap is entirely on our side (compute scheduling), with no authors' defect or fabrication concern; derivability and the core CCR2hi-vs-CCR2lo claim remain unverified rather than disconfirmed. Severity is unknown because nothing was measured; re-running the staged pipeline would complete the comparison.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

320.8 k
tokens (I/O) · 30.1 M incl. cache
338 min
runtime · 101.04 CPU-h
18.5 GB
peak RAM
5 (1 failed)
HPC jobs
hummel
machine