Reusable building blocks in biological systems.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce DIRECTIONALLY. Ran the authors' own ModuleReusability pipeline (commit 50e5c15, ported Py2->Py3) on the paper's own miRNA data: shipped GSE33045 fluid+plasma (a paper system) and the RU-target GSE47652 (raw Ct from GSE47652_non_normalized.txt.gz; the GEO series matrix is normalized and unusable). All four core directional claims reproduce 3/3 systems: real systems have smaller mean PBB size (C2), larger max PBB size vs DP-Rand (C4), reusability NOT characteristically high (C1, real << DP-Rand), and mean size closer to RSS-Rand than DP-Rand (C3); GSE47652 reusability ~= its RSS-Rand equivalent, mirroring the paper's 'one system close to RSS-Rand'. C5 (reusability entropy) reproduces as ratio<1 but the paper's prose vs figure are ambiguous -> partial. Overall PARTIAL because grading is directional only: the paper reports NO per-system numeric values and the decomposition is a stochastic heuristic, so exact numeric match is neither reported nor possible. NOT attempted (80/20): other 5 miRNA systems, protein-EST GO-enrichment validation, exact figure regen, production-scale 50-surrogate run. Run is internally deterministic (fixed seed; identical across 2 submissions). Repo needed mechanical Py3 fixes (range/reload/hashlib/scipy.misc + a real weightVector ordering bug) that do not change the algorithm.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 82assessed: 2026-06-14 ⛓ 614eb027a80d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusAre the modular building blocks of biological systems more reusable—or reusable in a distinctive way—than those found in randomized versions of the same systems? The paper tests how the reusability and size distribution of phenotypic building blocks in real biological systems compare to random equivalents.
- ★ Biological systems can be decomposed into phenotypic building blocks (PBBs) via k-maximally reusable decompositions (k-MRD) that maximize average reusability across conditions. method
- ★ k-MRDs of real biological systems are composed of smaller average-size PBBs than their random equivalents, while simultaneously having a larger maximum PBB size. finding
- ★ Real biological systems exhibit PBBs with a wider, more uniformly distributed range of reusabilities (both condition-specific and constitutive PBBs) than random systems. finding
- ★ Smaller average PBB size implies less overlapping, more independent building blocks, corroborating the near-decomposability property of natural systems. mechanism
- ★ The element-usage (expression breadth) distribution of real systems differs from the binomial distribution of a density-matched random matrix and partly drives the existence of large, highly reusable PBBs. finding
- ★ Average reusability is not characteristically higher in biological systems than in random equivalents; several systems matched their random counterparts. finding
- PBBs found in k-MRDs of human tissue protein data are more often significantly enriched for gene ontology terms than modules from agglomerative Jaccard-distance clustering. finding
- The bimodal distribution of PBB sizes is not exclusive to maximally reusable decompositions, but the high reusability of large modules is. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| EST-based protein presence/absence profiling (computational analysis of expression data) | 21 human tissues | none | presence/absence of proteins (EST equated with protein presence) decomposed into PBBs; PBB size, reusability, GO-term enrichment | — |
| miRNA expression profiling by quantitative RT-PCR | human conditions/samples (GEO datasets, e.g. GSE47652) | none/various conditions | miRNA presence/absence across conditions (threshold 35 PCR cycles); PBB mean size, max size, reusability entropy | GEO platform GPK13987 (RT-PCR) |
| In silico randomization (density-preserving, DP-Rand) | randomized binary matrices of real datasets | randomization preserving per-condition element count | PBB mean size, max size, reusability entropy compared to real | — |
| In silico randomization (row-sum-sequence-preserving, RSS-Rand) | randomized binary matrices of real datasets | randomization preserving element-usage distribution | PBB mean size, max size, reusability entropy compared to real | — |
- ▼ Average PBB size of k-MRDs is smaller in real systems than randomized versions across all k (AUC ratios all >1). AUC ratio >1
- ▲ Maximum PBB size is larger in real systems than DP-Rand equivalents, a feature recovered in RSS-Rand equivalents.
- ▲ Reusability entropy is higher for real systems than random equivalents (AUC ratio below 1). AUC ratio <1
- – Element-usage distribution in miRNA datasets deviates from the binomial expectation, favoring many constitutive elements and large reusable PBBs.
- – Of nine systems studied, four had average reusabilities within one s.d. of their DP-Rand equivalents and one within range of its RSS-Rand equivalents. 4 of 9 (plus 1)
- ▲ k-MRDs of 21-human-tissue protein data yielded more PBBs significantly enriched for GO terms than agglomerative Jaccard clustering. p<0.01 after Bonferroni
- – In 1529 decompositions of GSE47652 at k=26, only near-optimal decompositions exhibited large, highly reusable PBBs. 1529 decompositions, k=26
- count 9 systems studied; 4 within one s.d. of DP-Rand, 1 within RSS-Rand reusability (average reusability not characteristically high in biological systems)
- count 21 human tissues (protein presence/absence dataset from Souiai et al.)
- pvalue p < 0.01 after Bonferroni correction (GO term enrichment significance for PBBs vs agglomerative clustering)
- count 35 PCR cycles threshold (tested 25–35) (miRNA non-detection threshold; no difference in results across range)
- count 100 randomized versions (50 DP-Rand + 50 RSS-Rand) (random equivalents per miRNA dataset GSE47652)
- count k-MRDs of between 22 and 4000 PBBs (decomposition range for human tissue protein data)
- count 1529 decompositions at k=26 (GSE47652 decompositions, most far from maximally reusable)
- count size 120 separation between two size modes (split between small and large PBBs in bimodal size distribution)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper develops a mathematical framework for decomposing biological systems into phenotypic building blocks (PBBs) and compares three summary properties of those decompositions (mean PBB size, maximum PBB size, reusability-entropy) between real biological datasets and 50 density-preserving (DP-Rand) and 50 row-sum-sequence-preserving (RSS-Rand) randomized equivalents per dataset. The primary quantitative comparison method is the ratio of areas under the curve (AUC) of each summary statistic plotted as a function of decomposition size k, while departure from the random baseline is assessed visually by whether the real curve falls outside the range or one-standard-deviation band of randomized curves. Gene ontology enrichment is tested separately with Bonferroni correction at p < 0.01.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| AUC ratio (ratio of area under summary-statistic-vs.-k curve for each randomized equivalent to that of the real system) | Main comparison of mean PBB size, maximum PBB size, and reusability entropy between real and randomized systems (figures 2, 4) | 9 biological datasets; 50 DP-Rand and 50 RSS-Rand per dataset for miRNA data; count for protein data not explicitly stated | not stated |
| One-standard-deviation threshold comparison (whether real system value falls within mean ± 1 SD of randomized distribution) | Cross-dataset summary of average reusability vs. DP-Rand and RSS-Rand equivalents (results text: 'four had average reusabilities close, within one standard deviation') | 9 datasets | not stated |
| Gene ontology term enrichment test (specific test, e.g. hypergeometric/Fisher's exact, not named) with Bonferroni correction | Functional relevance of PBBs in k-MRDs of human tissue protein expression data vs. agglomerative clustering (figure 7) | null | not stated |
| Visual/graphical comparison of empirical element-usage distribution to expected binomial distribution (no formal test named) | Element usage distributions in miRNA datasets vs. binomial expectation for a density-matched random matrix (figure 5) | null | not stated |
-
Departure of the real system from the random baseline is assessed by whether the real curve falls outside the one-SD band or full range of 50 randomized equivalents, without a formal p-value↳ Could also: A permutation-based p-value could be derived directly from the empirical rank of the real AUC among the 50 simulated AUCs (e.g., proportion of randomized AUCs exceeding the real value) — A formal permutation p-value would allow exact probabilistic statements and enable application of a multiple-testing correction across the nine datasets and three summary quantities, making claims about systematic deviation from random more precisely quantified
-
Results across nine datasets are summarized narratively (e.g., 'four had average reusabilities within one SD of DP-Rand') without a combined test across datasets↳ Could also: A binomial sign test or Fisher's combined probability method could aggregate whether each dataset's real AUC ratio exceeds 1 (or the median of randomized values), providing a single cross-dataset summary statistic — A combined-test approach would quantify how consistently biological systems depart from their random equivalents and reduce reliance on informal counting of datasets that cross a threshold
-
The GO term enrichment analysis uses Bonferroni correction, but the underlying enrichment test is not named↳ Could also: Explicitly stating the enrichment test (e.g., hypergeometric test or Fisher's exact test) and reporting adjusted p-values; alternatively, Benjamini-Hochberg FDR is widely used for GO enrichment as a less conservative alternative — Naming the test aids reproducibility; BH-FDR is often preferred over Bonferroni for GO term sets because GO terms are correlated, which makes the independence assumption underlying Bonferroni conservative and can reduce power to detect genuinely enriched terms
-
The departure of the empirical element-usage distribution from a binomial expectation is shown graphically without a formal goodness-of-fit test↳ Could also: A chi-squared goodness-of-fit test or Kolmogorov-Smirnov test could formally compare the observed usage distribution to the binomial null for each dataset — A formal test would provide a quantitative measure of the departure and allow comparison of the degree of non-randomness across datasets, which would complement the visual comparison in figure 5
-
The dispersion of randomized equivalents is summarized as one SD around the mean and as the full range↳ Could also: An empirical percentile interval (e.g., 2.5th–97.5th percentile across the 50 simulations) could also be reported alongside or instead of the SD band — Percentile-based intervals directly reflect the empirical simulation distribution without assuming normality of the randomized AUC values, and a 95% interval corresponds directly to an approximate two-tailed p = 0.05 threshold, making it easier for readers to interpret the visual comparison
-
The nine datasets vary in the number of conditions (shown in parentheses in figure 4) but are treated equally in the narrative cross-dataset summary↳ Could also: A weighted analysis or a mixed-effects model treating dataset as a random effect and system type (real vs. DP-Rand vs. RSS-Rand) as a fixed factor could account for differences in dataset size when pooling evidence across datasets — Datasets with more conditions may yield more stable AUC estimates; a model that weights by precision or accounts for dataset-level variance would separate within-dataset signal from between-dataset heterogeneity and yield a formal omnibus test of the cross-dataset pattern
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-30958230 (Mireles & Conrad 2018, "Reusable building blocks in biological systems")
- DOI 10.1098/rsif.2018.0595 · J R Soc Interface · code https://github.com/syats/ModuleReusability (commit 50e5c15, GPL-3, authors' own Python)
- Data: 7 miRNA qRT-PCR GEO series on platform GPL13987 (GSE37766, GSE48910, GSE48909, GSE48908, GSE47652 [RU target], GSE45387, GSE33045) + 1 protein-EST system (presence of proteins in 21 human tissues). "Nine systems studied" total.
In scope (pipeline-derived, attempted)
The whole paper IS a single computational pipeline applied to binary presence/absence matrices:
- Preprocessing (
aux/auxFunctionsGEO.load_TaqMan_csv, thr=35 PCR cycles): Ct matrix → binary C (present = detected before 35 cycles), drop empty rows →cleanInputMatrix. Deterministic. - Decomposition into k phenotypic building blocks (PBBs) via the authors' minimum-reusable-decomposition heuristic (
gradDescent/heuristic5DMandList), for k = n…m. Stochastic heuristic. - Per-k size & reusability measures (
sizeMeasures): module-size and reusability distributions per decomposition. - Comparison to random null models DP-Rand (preserve column density) and RSS-Rand (preserve row-sum / element-usage distribution); the paper's Fig 2/3 plot AUC ratios real/random of: mean PBB size, max PBB size, reusability Shannon entropy, mean reusability.
Reproduced quantities → see original/claims.tsv C1–C6:
- C1 reusability of real ≈ DP-Rand (not characteristically high) — central claim, directional.
- C2 mean PBB size real < random; C3 mean size closer to RSS-Rand than DP-Rand; C4 max PBB size real > DP-Rand; C5 reusability entropy/range real > random.
- C6 binary-matrix dimensions of shipped GSE33045 datasets (deterministic anchor).
Reproduction strategy (80/20)
- Run the authors' own pipeline on shipped test data (GSE33045 fluid + plasma — these ARE one of the paper's 7 systems) and on the RU target accession GSE47652 (downloaded from GEO, formatted to the repo's
.datCt-matrix format). - Compute the 4 AUC-ratio measures real-vs-DP-Rand and real-vs-RSS-Rand and test the directions of claims C1–C5; report deterministic dims for C6.
- The heuristic is stochastic and the paper reports no per-system numeric values (only directional/qualitative statements + figures), so grading is at the directional level; exact byte-reproduction is not expected and not attempted.
Out of scope / not attempted
- The other 5 miRNA systems and the protein-EST GO-enrichment validation (would add coverage, not new evidence about the method; 80/20 cut).
- Exact AUC ratio numeric matching — impossible: no reported numbers + stochastic heuristic + Py2
randomvs Py3. - The full multi-core, many-surrogate (50 per dataset) production run — we use 3+3 surrogates and the "fast" heuristic params shipped in
analysisModuleSizeAndReusabilities.py.
Env / porting notes
Python 2.7 origin; ported in-job (Py3.10): inject reload=importlib.reload; add repo + gradDescent/ + aux/ to sys.path (implicit relative imports); 2to3 -f print; scipy.misc→scipy.special (comb). Heuristic run single-core (nc=1).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Authors' own ModuleReusability pipeline (Py2→Py3) re-run on the paper's own miRNA data reproduces all four core directional claims 3/3 (smaller mean PBB, larger max PBB, reusability not high, mean size closer to RSS-Rand). Because the paper reports no per-system numeric values and the decomposition is a stochastic heuristic, grading is directional only — an inherent property of the source, not a defect on our side. The only real wrinkle is C5, where the reproduced reusability-entropy ratio (<1, real<random) matches the Fig 2 caption but contradicts the paper's prose — an authors'-side internal ambiguity that a human must resolve against the actual figure, hence partial. Overall yellow: methodologically solid with explainable, non-critical deviations; core conclusion fully holds.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.