Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Partial correlation network analysis identifies coordinated gene expression within a regional cluster of COPD genome-wide association signals.

PLoS Comput Biol · 2024
L1 23/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
23/100
Reproducibility score
2.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 1% of all assessed papers rank 1160 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described-well-enough? Partially. Code is public (MIT, P16) and the algorithm is shipped and runnable; nreg_0 PCORR provably reduces to Pearson correlation (verified by source inspection). 1:1 numeric reproduction NOT achieved. Decisive finding: the repo's shipped data/expression.csv is a 200-sample MIXED UBC+GRNG test fixture and reproduces NONE of the three deposited supplementary networks (obs_GSE23352/23529/23545) at nreg_0 (mean|diff| 0.22-0.34, ~0% of 9453 edges within 1e-6 for every cohort). The deposited networks were each generated from a full single-cohort GEO matrix (Laval 499 / UBC 405 / GRNG 445 samples) with the authors' unspecified probe->gene preprocessing, none of which is shipped. So a faithful reproduction requires downloading the full GEO series + reverse-engineering the preprocessing then running PCORR = the hard, under-specified 20% -- NOT attempted. NOT attempted also: the LTRC manuscript main networks (restricted-access cohort, out of scope), the nreg>0 ridge configs (need prior.csv, would run on «our HPC»), and downstream significant-edges/module-clustering. Compute-access caveat: «our HPC» SLURM job was fully prepared (run.sbatch + pcorr_repro.py staged) but NOT executed -- the «our HPC» VPN 2FA was not completed in-session across 5 attempts; the small nreg_0 check was done transiently in-memory (stdlib, no data written to «host» disk). NO fabrication signal: the mismatch is benign, fully explained by the repo shipping a generic test fixture instead of the real per-cohort inputs; deposited matrices are internally plausible. All grades provisional; human review required.

💻 Code ↗ 🗄 Data: GSE23352

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 23
    assessed: 2026-06-15 ⛓ 3130ff2652b7
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Does the dense cluster of COPD GWAS loci on chromosome 4q (within the ~70 Mb region bounded by HHIP and BTC) represent statistically significant regional enrichment, and do these regional genes share coordinated/co-regulated gene expression relevant to COPD pathogenesis?

Core claims
  • COPD GWAS loci are more spatially clustered across the genome than expected by chance, with significant enrichment on chromosome 4q. finding
  • Genes in the 4q COPD risk region form a dense co-expression network containing multiple partial-correlation links between COPD candidate genes (e.g., NPNT-HHIP, BTC-NPNT, FAM13A-TET2). finding
  • Network clustering places four previously COPD-associated genes (BTC, HHIP, NPNT, PPM1K) in the same network community/module. finding
  • A differential co-expression sub-network distinguishes COPD cases from controls, including FAM13A, PPA2, PPM1K and TET2, with SPP1 as a potential local key regulatory gene. finding
  • A gene-specific ridge partial correlation method that uses protein-protein interaction networks (personalized PageRank) as a regularization prior to estimate direct co-expression relationships. method
  • Key partial correlations (HHIP-NPNT, HHIP-PPA2, BTC-NPNT, FAM13A-TET2) replicate in independent lung tissue cohorts. finding
  • A publicly available Python package implementing the Partial Correlation algorithm. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-Seq co-expression / partial correlation network analysis human lung tissue (LTRC cohort: 458 COPD cases, 329 controls) none (observational case/control) gene expression levels / partial correlation coefficients among 224 genes in the 4q region
microarray gene expression (replication co-expression analysis) human lung tissue (GSE23352 Laval, GSE23545 Groningen, GSE23529 University of British Columbia) none replication of partial correlations between gene pairs
GWAS / regional genetic association and clustering simulation human (Sakornsakolpat et al. and Hobbs et al.; 35,735 cases and 222,076 controls) none genome-wide significant COPD loci and spatial clustering p-values
protein-protein interaction network integration STRING PPI database none PPI-derived regularization prior (mediating genes) for partial correlation STRING PPI
Key results
  • Regions with 3 COPD GWAS hits within 34.28 Mb were more frequent than expected by chance (n=26 regions). p=0.006
  • Five COPD GWAS signals co-occur within a 70 Mb region on chromosome 4q bounded by BTC and HHIP.
  • Partial correlations HHIP-NPNT, HHIP-PPA2, BTC-NPNT, and FAM13A-TET2 were consistently strong and replicated in at least two of three independent lung cohorts.
  • BTC, HHIP, NPNT and PPM1K cluster together in the same co-expression network community.
  • Differential co-expression sub-network between COPD and controls includes FAM13A, PPA2, PPM1K and TET2, notably the CXCL10-CXCL11 edge.
Key statistics
  • pvalue p = 0.006 (enrichment of regions with 3 COPD GWAS hits within 34.28 Mb vs 100,000 simulations)
  • count 458 COPD cases (LTRC lung tissue RNA-Seq COPD cases)
  • count 329 control subjects (LTRC lung tissue RNA-Seq controls)
  • count 224 genes (genes expressed within the 4q COPD risk region analyzed)
  • count 83 GWAS loci (COPD GWAS loci identified by Sakornsakolpat et al. and Hobbs et al.)
  • pvalue p = 5 x 10^-8 (genome-wide significance threshold in regional association plot)
  • mean FEV1 %predicted 42.0 (20.2) cases vs 96.3 (12.6) controls (lung function difference between COPD cases and controls (p<10^-3))
  • count 35,735 cases and 222,076 controls (sample size of source COPD GWAS (Sakornsakolpat et al. 2019))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper constructed partial correlation gene co-expression networks for a 70 Mb chromosome 4q COPD risk region using RNA-seq lung tissue data from 458 COPD cases and 329 controls (LTRC cohort). A novel gene-specific ridge partial correlation algorithm, informed by STRING protein-protein interaction distances via personalized PageRank, was used to estimate direct gene-gene relationships; edge significance was assessed via t-tests with permutation and bootstrap procedures. GWAS locus clustering enrichment was evaluated against 100,000 genomic permutation simulations, and key co-expression findings were replicated in three independent public lung tissue cohorts (GSE23352, GSE23545, GSE23529).

Replicationbiological Sample sizePrimary cohort: 458 COPD cases and 329 controls (LTRC). Three independent replication cohorts: GSE23352 (Laval), GSE23545 (Groningen), GSE23529 (University of British Columbia). No formal power calculation described. GroupsCOPD cases vs non-COPD controls (lung tissue RNA-seq); a COPD-only partial correlation network was also constructed separately from the differential network Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
permutation test (100,000 genome simulations) enrichment of COPD GWAS loci clustering within small genomic windows (Fig 1 / S5 Fig) 26 regions with ≥3 GWAS hits within 34.28 Mb not stated
gene-specific ridge partial correlation with t-test (permutation and bootstrap procedures) significance of pairwise co-expression edges in the 4q partial correlation network (COPD cases) and differential co-expression network (COPD vs controls) 458 COPD cases; 329 controls (LTRC) not stated
two-sided Wald statistic GWAS association p-values shown in Fig 1 regional association plot (from Sakornsakolpat et al. 2019; not computed by this paper) 35,735 cases and 222,076 controls not stated
statistical test not named (threshold p<10⁻³ noted) comparison of clinical characteristics between COPD cases and controls in Table 1 329 controls, 458 cases not stated
Approaches that could also have been used
  • Pairwise partial correlations were computed across all 224 genes (up to ~24,976 pairs) and 29 parameter combinations without a stated multiple testing correction
    Could also: Apply a false discovery rate correction (e.g., Benjamini-Hochberg) across all tested gene pairs to control the expected proportion of false positives in the network — With thousands of simultaneous hypothesis tests, an FDR procedure provides a principled bound on spurious edges and is a standard complement to permutation-based significance in large-scale co-expression network analyses
  • Gene-specific ridge regularization informed by PPI-based personalized PageRank distances was used to estimate partial correlations
    Could also: Use graphical lasso (GLASSO) or related sparse precision-matrix methods (e.g., HUGE, SPACE) to estimate the full partial correlation network jointly — GLASSO and related methods impose sparsity across the entire precision matrix simultaneously, have well-characterized high-dimensional statistical properties, and provide a single coherent regularization framework without requiring parameter grid search
  • Differential co-expression between COPD cases and controls was assessed via the authors' custom differential partial correlation approach
    Could also: Apply dedicated differential co-expression tools such as DINGO, DiffCoEx, or a z-test on Fisher-transformed pairwise correlation differences with permutation-based FDR — These methods are explicitly designed for the two-group differential network setting and provide standardized significance frameworks that account for within-group variability in correlation estimates, facilitating comparison with other published differential network studies
  • Continuous clinical variables in Table 1 are summarized as mean (SD) and group differences are flagged at p<10⁻³ without naming the specific test
    Could also: Name the test used (e.g., two-sample t-test, Mann-Whitney U for non-normal variables) and report exact p-values alongside a standardized effect size (e.g., Cohen's d, rank-biserial r) — Naming the test aids reproducibility, and effect sizes allow readers to assess the magnitude of group differences independently of sample size, which is particularly informative given the unequal group sizes here (n=329 vs n=458)
  • Replication of co-expression edges in three independent cohorts was assessed by directional concordance of partial correlations across parameter settings
    Could also: Apply a formal meta-analytic framework (e.g., fixed- or random-effects meta-analysis of Fisher z-transformed correlation coefficients) across discovery and replication cohorts — Meta-analysis yields a pooled estimate with confidence intervals and a formal test for heterogeneity (I²), providing a more quantitative and portable summary of replicability than visual concordance across cohorts
  • Network community structure was identified by clustering the partial correlation network (specific algorithm not named in the provided text)
    Could also: Compare multiple community detection algorithms (e.g., Louvain, Leiden, spectral clustering) and assess module stability via bootstrap resampling of the gene expression data — Community assignments in co-expression networks can be sensitive to the choice of algorithm and to sampling variability; reporting stability metrics (e.g., proportion of bootstrap replicates recovering each module) would indicate how robustly the observed communities emerge from the data
Software: Python (custom partial correlation package, github.com/michelegentili93/Partial_Correlation) · STRING PPI database

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

rs13140176 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs2047409 RefSNP in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
rs34712979 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs4585380 RefSNP in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
rs62375246 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 39418301

Title: Partial correlation network analysis identifies coordinated gene expression within a regional cluster of COPD genome-wide association signals. Gentili, Platig, Silverman, et al. PLoS Comput Biol 2024. DOI 10.1371/journal.pcbi.1011079. PMCID PMC11521246.

Code: https://github.com/michelegentili93/Partial_Correlation (MIT, main, last push 2024-04-29). Algorithm: PCORR = Gene-Specific Ridge Partial Correlation (src/PartialCorrelation.py).

Pipeline-derived results (candidate in-scope)

The core pipeline output is a gene–gene partial-correlation network over the genes in a regional GWAS cluster (chr4 region around FAM13A/HHIP and chr5 region), computed by PCORR, then thresholded into significant edges and clustered into modules.

Result Pipeline Deposited artifact In scope?
Manuscript main networks (LTRC cohort, case/control) PCORR data/LTRC/, computed on /proj/regeps/.../LTRC/ OUT — LTRC = Lung Tissue Research Consortium, restricted-access clinical cohort (not public)
Supplementary networks (public GEO cohorts) PCORR data/GSE/nreg_*/obs_GSE2335{2,29,45}.csv (184×184 chr4) IN (targeted) — public GEO lung-eQTL data
Significant edges / module clustering threshold + community detection Pcorr_significant_edges_*.csv, *_Clustering_MaxMod.csv downstream of obs (not reached)
GWAS-loci clustering simulation notebook Gwas_hits/ Results.csv not attempted

Public data: the three supplementary cohorts are the Lung eQTL study sets (superseries GSE23546): GSE23352 = Laval (499 samples), GSE23529 = UBC (405), GSE23545 = GRNG/Groningen (445), all on platform GPL10379.

What was attempted

Targeted the cleanest deterministic output: obs_GSE23352.csv at parameter nreg_0_minL_0 (the network with n_controlling_genes==0). By source inspection of PartialCorrelation._residual(), at n_controlling_genes==0 the residual returns the raw expression unchanged, so PCORR reduces exactly to the Pearson correlation of the two genes (diagonal forced to 0). This makes nreg_0 fully deterministic and reproducible from the input expression alone — no prior, no Ray, no regression.

Out of scope / not attempted (and why)

  • LTRC manuscript networks — restricted-access cohort; data not public.
  • nreg>0 ridge-PCORR configs — require the shipped prior.csv (PageRank prior, ~100 MB) and the full controlling-gene pool; deferred (would run on «our HPC»).
  • Full GEO re-derivation — see AUDIT: the deposited obs were generated from the full per-cohort GEO matrices with the authors' (unspecified) probe→gene preprocessing, which are not shipped; re-deriving them is the hard, under-specified 20%.
C1
Reported
PCORR (Gene-Specific Ridge Partial Correlation); at n_controlling_genes==0 reduces to Pearson correlation
Reproduced
Confirmed by source inspection of src/PartialCorrelation._residual (returns raw expr) + np.corrcoef; deposited obs diagonals verified all 0
partial
C2
Reported
obs_GSE23352.csv, nreg_0/minL_0, 184x184 chr4-region partial-correlation matrix
Reproduced
Pearson from shipped data/expression.csv does NOT match (138-gene overlap, 9453 edges; mean|diff|=0.299, max=0.971, 0% within 1e-6)
did not match
C4
Reported
ReadMe: data/expression.csv is 'subset from the GSE dataset, used to test the PCORR'
Reproduced
expression.csv = 200 MIXED UBC+GRNG samples (titles hLung_UBC_*/hLung_GRNG_*); 0 sample overlap with GSE23352 Laval(499); reproduces NONE of the 3 deposited cohorts (mean|diff| 0.224-0.338, 0% within 1e-6)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 23/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The PCORR algorithm is public and verifiably reduces to Pearson correlation at nreg_0 (confirmed by source inspection), but no 1:1 numeric reproduction of any deposited value was achieved. The decisive issue is input non-correspondence: the repo ships a 200-sample mixed UBC+GRNG expression.csv that the ReadMe mislabels as the PCORR input, yet it reproduces none of the three deposited cohorts (mean|diff| 0.224–0.338, 0% of edges within 1e-6). This is an authors'/data-availability defect (real per-cohort matrices and probe→gene preprocessing not shipped), not a method error or fabrication — the deposited matrices are internally plausible. The central conclusion is therefore untested rather than refuted (main LTRC cohort is restricted-access, out of scope).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

241.1 k
tokens (I/O) · 21.8 M incl. cache
52 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.