Similarities and Differences in Gene Expression Networks Between the Breast Cancer Cell Line Michigan Cancer Foundation-7 and Invasive Human Breast Cancer Tissu
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH -> reproduced 1:1 on structure, partial on one exact count. Reproduced the GSE50705 WGCNA network (the cleanest deterministic pipeline output) by running the authors' own repo script #3 (commit 9b1935c) verbatim on GEO GSE50705 on «our HPC» (SLURM 2176994, R 4.3.3, WGCNA 1.73). EXACT: 88 estradiol samples (cols 174:261), platform GPL570, top-10000-MAD nodes, soft-threshold beta=7, minModuleSize=80, and the module STRUCTURE (turquoise is the largest non-grey module, followed by blue/brown). PARTIAL: the turquoise gene count is 2497 vs the paper's 2244 (+11%), consistent with biomaRt Ensembl release-116 vs ~Aug-2020 probe->symbol drift (the unique-symbol filter runs before the top-10000 selection, so a changed annotation shifts the node set). Faithfully reproduced subtlety: paper's beta=7 sits below WGCNA's usual R^2>=0.9 cutoff (R^2 hits 0.9 only at power 18; powerEstimate=14). FIXED a real WGCNA gotcha to run it: blockwiseModules resolved stats::cor instead of WGCNA::cor ('unused arguments weights.x') -> set cor<-WGCNA::cor. FLAG (HARD RULE 5): canSAR ligand-druggability scores (Fig 6, Table 1) are NOT in-repo derivable -- script reads an unshipped object from the external canSAR web tool; not a fabrication signal but not independently auditable. NOT ATTEMPTED (the ~20%): ARCHS4 (1032) + TCGA-BRCA (1212) networks and hence the cross-dataset overlap Venn (5252/2203/681/6440), BRCA-brown<->GSE-turquoise overlap (1013/354), and scaled-connectivity difference (mean abs 3223) -- need 2-3 more large downloads + fundamentalNetworkConcepts (>30 min/net); the single GSE50705 network already gives a clear honest data point.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 92assessed: 2026-06-15 ⛓ 2a5e26500a05
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDoes the breast cancer cell line MCF-7 adequately recapitulate the gene expression networks of human invasive breast cancer tissue, and what are the key similarities and differences in their co-expression networks that affect MCF-7's relevance as a model for human breast cancer?
- ★ MCF-7 cell lines and human breast cancer tissues share only minimal similarity in biological processes, though fundamental functions such as cell cycle are conserved finding
- ★ Network topology (scaled connectivity) differs drastically between MCF-7 and BRCA, with some genes acting as hubs in MCF-7 but not in tissue finding
- ★ Ligand-based druggability analysis suggests using MCF-7 to study breast cancer may lead to missing important gene targets finding
- ★ WGCNA combined with functional annotation can be used to compare an immortalized cancer cell line to human cancer tissues method
- A reusable bioinformatics pipeline comparing cell lines to corresponding human tissues was made available resource
- ★ Only 681 genes were conserved among the top 10,000 most variant genes across all three datasets, indicating minimal conservation of gene expression signatures finding
- Ribosomal subunit genes showing patient-to-patient variation are present in BRCA tissue but absent from MCF-7 datasets, suggesting MCF-7 fails to capture ribosomal protein variation linked to poor outcomes finding
- BRCA tissue uniquely shows immune-related gene enrichment due to immune infiltration not present in cell lines finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (gene-level expression, reanalyzed) | MCF-7 human breast adenocarcinoma cell line | none/various (mixed biological conditions across GEO series) | gene expression levels / co-expression network modules | ARCHS4 database (1032 samples, 107 GEO series); aligned and gene-mapped per ARCHS4 pipeline |
| RNA microarray gene expression | MCF-7 human breast cancer cell line | drug — 17β-estradiol treatment for 48 h (dose-response) | gene expression / co-expression network modules | GEO GSE50705 microarray (n=88) |
| bulk RNA-seq (RSEM-normalized Level 3) | human breast invasive ductal carcinoma (BRCA) tissue, 1212 samples | none (clinical biopsies) | gene expression / co-expression network modules | TCGA, downloaded from FireBrowse |
| Weighted Gene Co-expression Network Analysis (WGCNA) | MCF-7 (ARCHS4, GSE50705) and BRCA datasets | none (computational) | network modules, adjacency matrix, scaled connectivity | WGCNA R package (blockwiseModules) |
| functional enrichment / protein-protein interaction annotation | WGCNA modules from MCF-7 and BRCA datasets | none (computational) | GO biological process enrichment, PPI enrichment p-values | STRINGdb (STRING database), Enrichr, Bioplanet |
| network topology analysis (scaled connectivity) | MCF-7 (GSE50705) and BRCA networks | none (computational) | scaled connectivity K = Connectivity/max(Connectivity) and ranking differences | fundamentalNetworkConcepts function, WGCNA |
| ligand-based druggability scoring | top 10,000 variant genes in MCF-7 (GSE50705) and BRCA datasets | none (computational) | ligand druggability scores (ligand efficiency, med-chem friendliness, molecular weight) | canSAR knowledgebase Protein Annotation Tool |
| batch effect correction and outlier/QC check | MCF-7 ARCHS4, GSE50705, BRCA datasets | none (computational) | batch-corrected expression matrix, outlier/missing-value detection | ComBat (sva); goodSamplesGenes (WGCNA); hierarchical clustering |
- – Only 681 genes were conserved among the top 10,000 most variant genes across all three datasets 681 genes
- – 6,440 of the top 10,000 ARCHS4 genes were unique to that dataset 6,440 genes
- ▲ Genes common to all three datasets were enriched for mitotic cell cycle adjusted p = 1.08E-21
- – Higher gene overlap between GSE50705 and BRCA than between ARCHS4 and BRCA 5252 vs 2203 overlapping genes
- – Largest cell-cycle module in BRCA (brown, 1013 genes) shares only 354 genes with its GSE50705 counterpart (turquoise, 2244 genes) 354/1013 genes overlap
- – Substantial scaled-connectivity difference between MCF-7 and BRCA, with a cluster of genes differing by >9,000 in ranking, ranking higher in MCF-7 mean absolute difference 3,223; max >9,000
- – SUSD2 had the largest scaled-connectivity ranking difference (rank 68 in MCF-7 vs 9995 in BRCA) absolute ranking difference 9927
- – Among top 30 druggability genes, only 17 overlap between GSE50705 and BRCA 17/30 overlapping
- count 681 conserved genes across three datasets (genes shared among top 10,000 most variant genes)
- count 6,440 genes unique to ARCHS4 (of top 10,000 ARCHS4 genes)
- pvalue adjusted p = 1.08E-21 (mitotic cell cycle enrichment in common genes)
- pvalue adjusted p = 2.23E-59 (cell-cell signaling enrichment in ARCHS4-unique genes)
- count mean absolute difference in scaled connectivity = 3,223 (between MCF-7 and BRCA networks)
- count 5252 vs 2203 (overlapping genes GSE50705-BRCA vs ARCHS4-BRCA)
- count 354 of 1013 / 2244 (gene overlap between BRCA brown and GSE50705 turquoise cell-cycle modules)
- count 1032; 88; 1212 samples (sample sizes for ARCHS4, GSE50705, and BRCA datasets)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study applied Weighted Gene Correlation Network Analysis (WGCNA) to three large gene expression datasets (MCF-7 ARCHS4 n=1032, MCF-7 GSE50705 n=88, BRCA TCGA n=1212) to construct co-expression modules and compare gene network topology between a cell line and human breast tissue. Gene modules were annotated via hypergeometric over-representation tests against a Homo sapiens background in STRINGdb, with Benjamini-Hochberg FDR correction. Differences in network topology (scaled connectivity) and ligand-based druggability were summarized descriptively rather than through formal inferential testing.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Hypergeometric over-representation analysis | Functional enrichment of GO Biological Processes and STRING pathway annotations for WGCNA modules across all three datasets | Top 10,000 most variant genes per dataset as query; full Homo sapiens gene universe as background; sample n=1032 / 88 / 1212 depending on dataset | not stated |
| Hierarchical clustering | Sample outlier detection in all three datasets; gene module detection via blockwiseModules in WGCNA | 1032 (ARCHS4), 88 (GSE50705), 1212 (BRCA) | not stated |
| ComBat empirical Bayes batch correction | Batch effect adjustment of MCF-7 ARCHS4 data aggregated from 107 GEO series | 1032 samples from 107 GEO series | not stated |
| Scale-free topology criterion (soft-thresholding power selection) | Selection of WGCNA soft-threshold β per dataset (β=5 ARCHS4, β=7 GSE50705, β=6 BRCA) | — | not stated |
-
Module conservation across datasets was assessed by counting overlapping gene members between modules annotated to the same biological process↳ Could also: WGCNA's built-in modulePreservation function (Zsummary and medianRank statistics) could also have been used — Module preservation statistics provide a quantitative, statistically calibrated measure of how well the co-expression structure of one dataset's modules is reproduced in another, going beyond gene-membership overlap counts and yielding a p-value or Z-score for preservation
-
Network topology differences between datasets were summarized descriptively using scaled connectivity ranking differences↳ Could also: Permutation-based or bootstrap statistical tests of network topology metrics (e.g., comparing hub gene connectivity distributions across datasets) could also have been applied — Formal testing would attach a probability statement to the observed ranking differences, allowing readers to assess whether the magnitude of difference exceeds what would be expected by chance given the different sample sizes and data platforms
-
GO and pathway enrichment was assessed using hypergeometric over-representation analysis in STRINGdb↳ Could also: Rank-based enrichment methods such as Gene Set Enrichment Analysis (GSEA) or fgsea could also have been used — Rank-based methods use the full ranked gene list rather than a binary module-membership threshold, which can be more sensitive to coordinated but modest expression shifts and avoids dependence on the arbitrary module boundary produced by hierarchical clustering
-
Batch effects in the ARCHS4 dataset were corrected using ComBat prior to WGCNA↳ Could also: Surrogate variable analysis (SVA) or limma's removeBatchEffect could also have been applied — SVA can estimate and remove latent sources of variation beyond explicitly known batch labels, which may be relevant when aggregating 107 GEO series with diverse experimental conditions whose batch structure may not be fully captured by series ID alone
-
Genes were filtered to the top 10,000 by median absolute deviation (MAD) prior to network construction↳ Could also: A sensitivity analysis varying the gene-number threshold (e.g., top 5,000 or 15,000 genes) could also have been reported alongside the primary analysis — Reporting how module membership and the overlap conclusions change across a range of filtering thresholds would allow readers to assess the stability of the network structure and the robustness of the comparison results to this analysis choice
-
Co-expression networks were constructed with WGCNA using a Pearson-correlation-based soft-thresholding approach↳ Could also: Mutual information-based network methods (e.g., ARACNE) or sparse partial correlation approaches could also have been used — Correlation-based methods capture linear co-expression; mutual information methods can capture nonlinear dependencies, and partial correlation methods can distinguish direct from indirect gene associations, which may yield a different view of which regulatory connections are preserved or lost between the cell line and tissue context
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34056582 (Tran et al. 2021, Front. Artif. Intell.)
"Similarities and Differences in Gene Expression Networks Between the Breast Cancer Cell Line MCF-7 and Invasive Human Breast Cancer Tissue." DOI 10.3389/frai.2021.674370 · PMCID PMC8155268.
Code: github.com/vy-p-tran/Similarities-and-differences-in-gene-expression-networks-between-MCF7-and-BRCA (7 R-Markdown scripts, no shipped data, no license, last push 2021-03-05).
Datasets used by the paper
- MCF-7 ARCHS4 — 1032 RNA-seq samples (107 GEO series) from the ARCHS4 db.
- MCF-7 GSE50705 — 88 estradiol-treated Affymetrix U133 Plus 2.0 microarray samples (GEO). [es = series-matrix columns 174:261 in script #3]
- BRCA — TCGA breast invasive carcinoma, 1212 samples (FireBrowse/FireHose).
Pipeline(s)
WGCNA (Langfelder & Horvath) on the top-10,000-MAD genes of each dataset:
pickSoftThreshold → blockwiseModules (TOMType unsigned) →
fundamentalNetworkConcepts for scaled connectivity. Probe→symbol mapping via
biomaRt (Ensembl). Downstream: STRINGdb enrichment, RRHO, gene-overlap Venn.
IN SCOPE (pipeline-derived, deterministic, attempted)
Primary 80/20 target = GSE50705 WGCNA network (one GEO download, fully
deterministic — blockwiseModules has no RNG once the input matrix is fixed):
- C1 88 estradiol-treated samples (data dimension).
- C2 soft-threshold power β = 7 for GSE50705 (Methods + Suppl. Fig S1).
- C3 minModuleSize = 80; module table from
table(net$colors). - C4 turquoise module = 2244 genes (Results: "turquoize module, 2244 genes"). This is the cleanest pinnable number for the GSE50705 network.
- C5 10,000 most variant genes used as network nodes.
IN SCOPE but HARDER (the ~20%, attempted only if primary is clean)
- Gene-overlap Venn (Fig 6 / script #1): GSE50705∩BRCA = 5252; ARCHS4∩BRCA = 2203; 681 conserved across all three; 6440 unique to ARCHS4. Needs all 3 datasets.
- BRCA brown module = 1013 genes (cell cycle), 354 overlap with GSE50705 turquoise. Needs the BRCA (TCGA) WGCNA network too.
- Scaled-connectivity difference GSE50705 vs BRCA (Fig 3/5, Table 1; mean abs
diff 3223). Needs both networks +
fundamentalNetworkConcepts(>30 min/network).
OUT OF SCOPE (not pipeline-derived / not attempted)
- canSAR ligand druggability scores (Fig 6, Table-1 druggability column):
computed by the external canSAR web tool, not reproducible from the repo
(
GEO_top_connectivity.gene_CANSAR.scoreis an unshipped manual input). Flagged. - STRINGdb enrichment p-values, GO Biological-Process annotations (script #5): depend on STRING/GO db versions; qualitative, no pinned numeric claim.
- Cytoscape figure layouts (Figs 2,3,5) — manual visualization.
- Wet-lab / interpretive statements.
Reproduction risk to flag (env drift)
biomaRt's Ensembl affy_hg_u133_plus_2 → hgnc_symbol map changes between releases.
The repo collapses probes to symbols, drops NA-symbol / multi-mapping probes, then
takes the top-10,000-MAD probes. A newer Ensembl release shifts which probes
survive → the exact module sizes (incl. turquoise 2244) can drift. We run with a
current pinned biomaRt and report the delta honestly rather than forcing a match.
The script comment says "power of 6 was chosen" but the code calls
blockwiseModules(power = 7) and the paper text says β=7 — we use the code/paper
value 7.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The deterministic GSE50705 WGCNA network reproduces cleanly on its own repo+data: identical input (88 estradiol samples), all stated parameters (beta=7, minModuleSize=80, 10000 nodes), and the module structure (turquoise = largest cell-cycle module). The one quantitative gap is the turquoise gene count (2497 vs 2244, ~11%), squarely on the technical/version side — biomaRt annotation-DB drift acting before the top-MAD node selection, not an authors' or methodology defect. A real non-derivability caveat is the canSAR druggability scores (Fig 6/Table 1), which read an unshipped external object and cannot be audited from the deposited artifacts. Core claim holds; overall a solid reproduction with explainable deviations rather than a perfect 1:1.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.