Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Loss of mutual protection between human osteoclasts and chondrocytes in damaged joints initiates osteoclast-mediated cartilage degradation by MMPs.

Sci Rep · 2021
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

IN PROGRESS. Paper reproducible result is DESeq2 differential expression across 3 osteoclast substrate groups, primarily Table 1 (top-15 over-expressed genes cartilage-vs-dentine) + MMP8 8.89-fold/p=0.0133. The featureCounts count matrix (GSE166535_final_counts.tsv.gz) is deposited so the DESeq2->Table 1 step is directly reproducible; running on «our HPC». Wet-lab results (zymography, siRNA, GAG release, RT-qPCR) are out of scope. Note discrepancy: paper says n=7/group, GEO has 8 donors/group (24 samples). Awaiting «our HPC» tunnel (currently unreachable) to download matrix + run DESeq2.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 6aab26d9fdd9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This study tests whether osteoclasts degrade cartilage using the same cellular machinery as bone resorption, and whether chondrocytes regulate this process, aiming to understand the osteoclast–cartilage interaction in joint degradation.

Core claims
  • Human osteoclasts can differentiate on acellular cartilage, express osteoclast markers, and degrade cartilage matrix in a contact-dependent manner without forming F-actin rings or resorption pits. finding
  • MMP8 and MMP9 are contributory MMPs driving osteoclast-mediated cartilage degradation (GAG release), with degradation being MMP-dependent rather than acidification/cathepsin-K dependent. mechanism
  • Bone-resident osteoclasts and chondrocytes exert mutually protective effects on their native tissue; chondrocytes stimulate osteoclast formation at non-bone sites but inhibit it on dentine. finding
  • MMP8 is the most highly upregulated gene in osteoclasts on cartilage versus dentine, while MMP9 is the most highly expressed MMP. finding
  • Direct contact between osteoclasts and cartilage is required for proteoglycan (GAG) degradation. finding
  • Comparing the osteoclast gene-expression profile on its two primary natural substrates (bone/dentine vs cartilage) by RNA-seq. method
Experimental setups
Assay System Perturbation Readout Platform
Histology (H&E, TRAP staining) Human OA knee, RA, and GCTB tissue none Multinucleated osteoclast presence and cartilage/bone erosion
Osteoclast differentiation / light microscopy & SEM Human CD14+ monocyte-derived osteoclasts on dentine vs acellular articular cartilage substrate (dentine vs cartilage vs plastic) Resorption pits, F-actin rings, cell morphology, surface degradation SEM
GAG and collagen release assay Human osteoclasts on dentine, acellular cartilage, cellular cartilage substrate; Transwell (indirect contact) Collagen release from dentine, GAG release from cartilage
TRAP staining / osteoclast quantification (co-culture) Human CD14+ monocyte-derived osteoclasts co-cultured with chondrocytes on plastic or dentine chondrocyte co-culture + M-CSF/RANKL Number of TRAP+ cells (≥10 nuclei / total)
Inhibitor treatment + GAG release Human osteoclasts on acellular/cellular cartilage bafilomycin, E64, GM6001 (pan-MMP), recombinant TIMP1 (100 nM) GAG release
RNA-seq and RT-qPCR Human osteoclasts differentiated on plastic, acellular cartilage, dentine substrate Differential gene/MMP expression (PCA, fold change)
Gelatin zymography Conditioned media of human osteoclasts on acellular cartilage none vs recombinant protein control Active MMP8 (~65 kDa) and MMP9 (~82 kDa) production
siRNA knockdown + GAG release; Immunohistochemistry Human osteoclasts on cartilage explants; human OA tissue siRNA targeting MMP8 or MMP9 GAG release reduction; MMP8/MMP9 protein localization
Key results
  • MMP8 showed greatest fold upregulation in osteoclasts on cartilage versus dentine 8.89-fold
  • Direct co-culture with chondrocytes increased large osteoclast (>10 nuclei) formation on plastic p < 0.0001
  • Chondrocyte co-culture inhibited osteoclast formation on dentine p = 0.0002
  • Osteoclasts cultured on dentine inhibited basal cartilage degradation p = 0.012
  • Pan-MMP inhibitor GM6001 significantly reduced osteoclast-mediated GAG release from acellular cartilage (bafilomycin and E64 had no effect)
  • siRNA knockdown of MMP8 and MMP9 reduced osteoclast-mediated GAG release from cartilage (non-significant trend) 39% (MMP8), 28% (MMP9)
  • Recombinant TIMP1 inhibited osteoclast-mediated GAG release from acellular cartilage
  • Osteoclasts on dentine clustered separately from cartilage/plastic by PCA, and on cartilage failed to form F-actin rings or resorption pits
Key statistics
  • fold_change 8.89-fold (LFC 8.89993669) (MMP8 overexpression in osteoclasts on cartilage vs dentine)
  • pvalue p = 0.0133 (padj 1.33E−02) (MMP8 upregulation significance)
  • pvalue p < 0.0001 (Increased large osteoclast formation by chondrocyte co-culture on plastic)
  • pvalue p = 0.0002 (Inhibition of osteoclast formation on dentine by chondrocyte co-culture)
  • pvalue p = 0.012 (Dentine-cultured osteoclasts inhibit basal cartilage degradation)
  • other 39% and 28% reduction (GAG release reduction after MMP8/MMP9 siRNA knockdown)
  • fold_change PER2 LFC 5.33 (padj 7.07E−08) (Second top over-expressed gene in osteoclasts on cartilage)
  • other 100 nM (Recombinant human TIMP1 concentration in GAG release assay)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This in vitro study examined human osteoclast behaviour on multiple substrates (dentine, acellular cartilage, cellular cartilage, cell culture plastic) using functional assays, RT-qPCR, and bulk RNA-seq. Two-group comparisons employed Student's t-tests or Mann-Whitney U tests, while multi-group inhibitor and co-culture experiments used one-way ANOVA or Kruskal-Wallis ANOVA; test choice appeared to vary by experiment without a stated selection criterion. RNA-seq transcriptomic profiling compared osteoclasts across the three substrates, reporting adjusted p-values and log-fold changes. Results were presented as individual data points with exact or threshold p-values; dispersion measures and post-hoc procedures were not explicitly named.

Replicationbiological Sample sizeSample sizes stated per experiment (ranging n=7 to n=20 across figures); no formal a priori power analysis or sample size justification reported GroupsOsteoclasts on dentine vs acellular cartilage vs cellular cartilage vs cell culture plastic; with vs without chondrocyte direct or indirect co-culture; with vs without enzyme inhibitors (bafilomycin, E64, GM6001, TIMP1) or siRNA knockdown Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionAdjusted p-values (padj) reported for RNA-seq family of tests (method not explicitly named; likely Benjamini-Hochberg FDR); no correction stated for multiple Mann-Whitney comparisons across genes or for multiple ANOVA comparisons across experiments
Statistical tests used
Test Applied to n Assumptions
Student's t-test (unpaired assumed; not stated) Collagen release from dentine (Fig 2a); GAG release from acellular cartilage (Fig 2a); large osteoclast formation on plastic with chondrocytes (Fig 2d); RANKL:OPG ratio (Fig 2f); TIMP1 effect on GAG from acellular cartilage (Fig 3c); MMP8 RT-qPCR (Fig 4b) n=16 (dentine collagen), n=20 (acellular GAG), n=10 (Fig 2d), n=11 (Fig 2f), n=18 (Fig 3c), n=8 (Fig 4b) not stated
Mann-Whitney U test GAG release from cellular cartilage (Fig 2a) n=17 not stated
Multiple Mann-Whitney U tests (no correction stated) Classical osteoclast marker gene expression on dentine vs acellular cartilage (Fig 1d); MMP1, 2, 3, 7, 9, 12, 13, 14 and ADAMTS1 RT-qPCR (Fig 4c) n=8 not stated
One-way ANOVA (post-hoc test not named) GAG release from indirect co-culture with acellular cartilage (Fig 2b) and cellular cartilage (Fig 2c); total osteoclast number on dentine with chondrocytes (Fig 2e); bafilomycin and E64 inhibitor effects on GAG from acellular cartilage (Fig 3a); bafilomycin and E64 effects on GAG from cellular cartilage (Fig 3b) n=10 (Figs 2b, 2c), n=8 (Fig 2e); N varies by group for Fig 3 not stated
Kruskal-Wallis ANOVA (post-hoc test not named) Bafilomycin effect on GAG from acellular cartilage (Fig 3a); GM6001 effect on GAG from cellular cartilage (Fig 3b); siRNA MMP8 and MMP9 knockdown vs control on GAG release from acellular cartilage (Fig 4e) N varies for Fig 3; n=7 (Fig 4e) not stated
RNA-seq differential expression analysis with adjusted p-values (padj; pipeline not named; DESeq2 Wald test with Benjamini-Hochberg FDR inferred from padj notation and log fold change reporting in Table 1) Genome-wide transcriptomic comparison of osteoclasts on acellular cartilage vs dentine vs cell culture plastic (Fig 4a, Table 1) n=7 per substrate group (inferred from PCA plot legend) not stated
Approaches that could also have been used
  • Multiple individual Mann-Whitney U tests were applied across several gene comparisons within each figure (6 genes in Fig 1d; 9 genes in Fig 4c) without a stated multiplicity correction
    Could also: Apply a family-wise error rate or FDR correction (e.g. Benjamini-Hochberg) across all gene comparisons within each figure, or use a linear mixed model that treats gene as a within-subject factor — Controlling the false discovery rate across the family of gene-level tests would explicitly bound the expected proportion of false positives across the panel, which is standard practice in multi-gene RT-qPCR analysis
  • One-way ANOVAs and Kruskal-Wallis tests were reported for multi-group comparisons without naming a post-hoc pairwise test
    Could also: Follow significant omnibus tests with a named post-hoc procedure — Tukey HSD or Dunnett's test (vs a control) for ANOVA, and Dunn's test with Bonferroni or BH correction for Kruskal-Wallis — A significant omnibus test indicates that at least one pair differs but does not identify which pairs; a named post-hoc procedure with multiplicity control makes specific pairwise conclusions explicit and reproducible
  • The choice between parametric (t-test, one-way ANOVA) and non-parametric (Mann-Whitney, Kruskal-Wallis) tests appeared to vary by experiment without a stated selection criterion
    Could also: Pre-specify a decision rule for test selection (e.g. Shapiro-Wilk normality test on residuals, or visual Q-Q inspection) and apply it consistently across all experiments — A transparent, pre-specified criterion for distributional testing aids reproducibility and helps readers understand why particular tests were chosen for particular experiments
  • The siRNA knockdown experiment reported biologically meaningful reductions in GAG release (39% for MMP8, 28% for MMP9) that did not reach statistical significance at n=7 (Kruskal-Wallis)
    Could also: Report an effect size with 95% confidence interval (e.g. rank-biserial correlation for non-parametric data) to characterize the magnitude of the observed difference independently of the p-value; alternatively a paired Wilcoxon signed-rank test could be used if knockdown and control conditions were matched within each biological replicate — Effect size with CI communicates the precision and plausible range of the true effect, which is informative when sample size limits power; a paired design can reduce within-experiment variability and increase sensitivity when measurements share a common biological source
  • No a priori power analysis or sample size justification was reported; n per experiment ranged from 7 to 20
    Could also: Report a power calculation based on effect sizes from pilot experiments or the literature, stating the minimum detectable effect at the chosen alpha and power — Pre-specified power calculations help readers assess whether individual experiments were adequately powered to detect biologically relevant effect sizes, and are particularly informative for experiments yielding trends that did not reach significance
  • Dispersion in figures is not described in the text; it is not stated whether error bars represent SD, SEM, or another measure
    Could also: Explicitly label the dispersion measure in each figure legend and state it in the Methods; for small n (7–20), SD or 95% CI is often recommended over SEM because SEM shrinks with n and can visually understate variability — SD describes sample-level variability, SEM describes the precision of the mean estimate, and 95% CI conveys inferential range; stating which was used — and overlaying individual data points — allows readers to correctly interpret the spread and assess distributional assumptions
Software: RNA-seq analysis pipeline (not explicitly named; DESeq2 inferred from padj and LFC reporting convention in Table 1) · Statistical software for non-RNA-seq tests (not stated)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34811438

Paper: Larrouture et al. 2021, Sci Rep 11:22841. "Loss of mutual protection between human osteoclasts and chondrocytes in damaged joints initiates osteoclast-mediated cartilage degradation by MMPs." DOI 10.1038/s41598-021-02246-7 · PMCID PMC8608887.

Code: https://github.com/cgat-developers/cgat-flow (CGAT toolkit; ruffus-based NGS pipelines — the framework named in Methods. Authors used "a computational pipeline calling scripts from the CGAT toolkit"; co-author A.P. Cribbs is a cgat-developers maintainer). Applying this framework + DESeq2 to the paper's own deposited data is a valid reproduction (BRIEF P16).

Data: GEO GSE166535 (SRA SRP305654 / BioProject PRJNA701237). Human osteoclasts cultured 10 days on three substrates: plastic ("no-sub"), dentine ("dent"), and OA cartilage ("cart"). Illumina NextSeq 500, paired-end. Series-level processed file: GSE166535_final_counts.tsv.gz (featureCounts matrix).

Reported pipeline (Methods)

  • Alignment: STAR → human genome GRCh38.
  • Quantification: featureCounts v1.4.6 (subread package), uniquely-mapped reads only.
  • QC: "at least 14 million aligned reads per sample".
  • Differential expression: DESeq2, three groups (plastic / dentine / cartilage).
  • Orchestration: CGAT toolkit pipeline.

IN SCOPE (pipeline-derived → attempt to reproduce)

# Result Paper location Reproduction route
C1 Table 1 — top-15 over-expressed genes in osteoclasts on cartilage vs dentine, with LFC + p-adj Table 1 DESeq2 on GSE166535_final_counts.tsv (deposited), contrast cart vs dent
C2 MMP8 greatest upregulation: 8.89-fold, p = 0.0133 (cartilage vs dentine) Results text + Table 1 row 1 same DESeq2 run
C3 PCA plot separating the three substrate groups Fig 4 (PCA panel) DESeq2/vst + prcomp on count matrix
C4 (stretch) Full pipeline from raw FASTQ: STAR→GRCh38 + featureCounts → counts matching deposited matrix Methods Download SRP305654 (24 runs) on «infra», run STAR+featureCounts on «our HPC», compare to deposited final_counts.tsv

The most faithful + feasible target is C1/C2: the count matrix is deposited, so the DESeq2 step that produces Table 1 can be reproduced directly. C4 (re-deriving counts from FASTQ) is the full-pipeline stretch goal.

Audit flag to test

Table 1 column "LFC" = 8.89993669 for MMP8 while the text says "8.89-fold". DESeq2 reports log2FoldChange. If DESeq2 log2FC(MMP8) ≈ 8.9, then "8.89-fold" in the text is a mislabel (a log2FC of 8.9 = ~477-fold). If instead the table's LFC column is linear fold (= 2^log2FC), DESeq2 log2FC(MMP8) should be ≈ 3.15. Reproducing the DESeq2 output resolves which, and whether 8.89 is the log2FC or the linear fold change. Record this in AUDIT.md.

OUT OF SCOPE (wet-lab / manual — not attempted)

  • RT-qPCR validation of MMP expression (Fig 4).
  • Gelatin zymography (Fig 4).
  • siRNA knockdown of MMP8/MMP9 and GAG-release assays from cartilage explants (39% / 28% reduction) — wet-lab functional assays, no pipeline.
  • Histology / immunostaining, resorption pit assays.

Known discrepancy to verify

Paper Methods state n = 7 biological replicates per group; GEO GSE166535 lists n = 8 donors per group (24 samples: BA, BB, BF, BH, BI, BJ, BK, BE). Confirm N actually present in the count-matrix columns and reconcile (possibly one donor/sample excluded from the published DE analysis).

Figures / tables: TableFig 4
C1
Reported
Table 1: top-15 over-expressed genes (cartilage vs dentine) with LFC + p-adj; rank1 MMP8 LFC=8.89993669 padj=1.33E-02
Reproduced
partial
C2
Reported
MMP8 greatest upregulation cartilage vs dentine: 8.89-fold, p=0.0133
Reproduced
partial
C3
Reported
PCA separates the 3 substrate groups (Fig 4)
Reproduced
partial
D1
Reported
n=7 biological replicates per group (Methods)
Reproduced
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

42.7 k
tokens (I/O) · 2.2 M incl. cache
7 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.