A network-based model of Aspergillus fumigatus elucidates regulators of development and defensive natural products of an opportunistic pathogen.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for an internal-consistency reproduction. The paper (GRAsp; Carriel et al., NAR 2026, gkaf1439) infers an A. fumigatus regulatory network with MERLIN-P-TFA from 294 bulk RNA-seq samples (18 studies incl. PRJEB2987). The named code repo (github.com/Roy-lab/merlin-preprocess) is ONLY preprocessing (Trimmomatic/FastQC/MultiQC) and yields no quantitative paper output; the inference is MERLIN-P-TFA (Roy-lab/merlin-p C++) which would need re-aligning ~294 samples + multi-day stability-selection compute = the out-of-scope hard 20%. Instead I reproduced by an independent RECOUNT of the authors' shipped supplementary tables (gkaf1439_supplemental_files.zip via EuropePMC; all download+parse on «our HPC»/«infra»). 7 of 7 fully-shipped headline numbers reproduce EXACTLY from the per-item data: 9859 genes, 164 modules (>=5 genes), 3381 genes-in-modules, largest module 323, average 21 (20.616), 74 GO-enriched modules, and the 189292-edge prior network. The candidate-regulator count is a 95% near-miss (783 recounted vs 820 reported), most plausibly a named-only/_nca-variant ID or overlap-counting convention rather than a fabricated value. NO fabrication indicators. NOT attempted (data not shipped / hard 20%): the inferred-network headline counts 7422 edges / 669 regulators / 5274 targets and top-regulator out-degrees (MAT1-2=216, AtfA=155) — the inferred edge list lives only on the interactive grasp.wid.wisc.edu site (a bounded download probe hung the compute node and was cancelled); these remain unverified, not refuted. Verdict: 1:1 on every recountable shipped statistic, partial overall because the inferred-network numbers require a from-scratch pipeline run beyond 80/20 scope.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 78assessed: 2026-06-16 ⛓ ca980eebcf39
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a comprehensive genome-scale gene regulatory network of Aspergillus fumigatus, inferred from integrated public RNA-seq data, recapitulate known regulatory pathways and generate experimentally testable hypotheses about regulators of development, virulence, and defensive natural products?
- ★ A comprehensive genome-wide gene regulatory network resource for A. fumigatus (GRAsp) was constructed from 18 RNA-seq datasets using MERLIN-P-TFA. resource
- ★ GRAsp recapitulates known regulatory pathways including hypoxia response, iron and zinc homeostasis, ergosterol biosynthesis, and secondary metabolite synthesis. finding
- ★ GRAsp identified an uncharacterized transcription factor that negatively regulates production of the virulence factor gliotoxin. finding
- ★ GRAsp revealed the bZip protein AtfA as required for fungal responses to lipo-chitooligosaccharides (LCOs). finding
- ★ The MERLIN-P-TFA network was computationally validated against published ChIP-seq TF-target relationships and showed high precision. method
- Zero-mean/quantile-normalization batch correction outperformed Combat-seq for recovering known regulatory edges, justifying its use. method
- MERLIN-P-TFA estimates hidden transcription factor activity levels from a noisy prior network and infers regulator-target edges plus gene module assignments via regularized regression. method
- ★ GRAsp is provided as a user-friendly online resource (grasp.wid.wisc.edu) offering module analysis, Steiner tree estimation, and node diffusion. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (publicly available, reanalyzed) | Aspergillus fumigatus (strain Af293 reference) | various (diverse environmental conditions across 18 datasets) | transcript abundance (TPM) | Trimmomatic v0.32, RSEM, reference ASM265v1.49 |
| computational GRN inference | A. fumigatus integrated transcriptome (9859 genes, 294 measurements) | none | inferred TF-target edges, TFA matrix, gene modules | MERLIN-P-TFA |
| ChIP-seq-based network validation | A. fumigatus | none | recovery/precision of TF-target gold-standard edges | — |
| gliotoxin production assay (experimental validation of TF) | A. fumigatus | transcription factor manipulation | gliotoxin production | — |
| LCO response assay (experimental validation of AtfA) | A. fumigatus | AtfA gene perturbation | fungal response to lipo-chitooligosaccharides | — |
- – GRAsp recovered well-known regulatory relationships for ergosterol biosynthesis, iron homeostasis, and secondary metabolite regulation.
- ▼ An uncharacterized TF predicted by GRAsp was confirmed to negatively regulate gliotoxin production.
- – AtfA was confirmed as required for A. fumigatus responses to LCOs.
- – MERLIN-P-TFA network showed high precision against ChIP-seq gold-standard edges.
- ▲ Zero-mean batch correction offered a slight benefit over Combat-seq for recovering known edges.
- count 18 bulk RNA-seq studies/datasets (datasets curated for network inference)
- count 9859 genes across 294 measurements (final zero-mean transformed expression matrix)
- count 820 putative regulator genes (candidate regulators prepared)
- count 632 TF binding site sequence motifs (used to construct prior network)
- count 7 samples removed (S1-S4 and rep3 of PRJEB2987) (samples with unexpected correlation/expression removed)
- other ~50% mortality rate (COVID-19-associated pulmonary aspergillosis (CAPA))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This computational study integrates 18 publicly available A. fumigatus bulk RNA-seq datasets (294 samples, 9859 genes after QC) using zero-mean batch correction, then applies the MERLIN-P-TFA probabilistic graphical model—combining L1-regularized regression with network component analysis (NCA)-based TF activity (TFA) estimation—to infer a genome-wide gene regulatory network. Network quality was assessed by recovery of ChIP-seq-derived gold-standard edges and literature-supported regulatory relationships. Batch correction approaches (zero-mean vs. COMBAT-seq) were compared using PCA, variance explained, ANOVA of PC scores, and global pairwise correlation; the experimental validation sections (gliotoxin TF, AtfA/LCO) were not included in the provided text excerpt.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| PCA (principal component analysis) | QC and visualization of batch correction across 18 datasets | 294 samples × 9859 genes | not stated |
| ANOVA of PCA component scores | Comparison of dataset-driven variance before and after batch correction (zero-mean vs. COMBAT-seq) | 294 samples | not stated |
| L2-minimization / alternating least squares (NCA matrix factorization) | TFA estimation step within MERLIN-P-TFA | 294 samples × number of regulators with prior motif edges | not stated |
| L1-regularized (LASSO-type) regression | Regulator-to-target-gene edge inference step within MERLIN-P-TFA, iterated per gene | 294 samples per gene model | not stated |
| Precision / edge recovery rate (against ChIP-seq gold standard) | Computational validation of inferred GRN edges | Number of ChIP-seq-validated edges not stated | na |
-
Zero-mean (log + quantile normalization + within-dataset mean subtraction) was chosen over COMBAT-seq for batch correction based on ChIP-seq edge recovery↳ Could also: limma::removeBatchEffect, SVA (surrogate variable analysis), or Harmony could also be applied to integrated multi-study RNA-seq data — SVA and Harmony are commonly used for multi-cohort transcriptomic integration and can model latent batch variables without requiring explicit batch labels; comparing all three via the same gold-standard recovery metric would further substantiate the chosen approach
-
ANOVA on PCA component scores was used to quantify residual batch-driven variance after correction↳ Could also: PVCA (principal variance component analysis) or a mixed-model variance partitioning approach (e.g., variancePartition in R) could also decompose variance by batch and biological factors jointly — These methods explicitly partition variance into batch and biological components simultaneously, providing a single quantitative summary of how much residual batch variance remains relative to biologically meaningful variation
-
MERLIN-P-TFA (probabilistic graphical model with L1-regularized regression + NCA-based TFA) was used for GRN inference↳ Could also: GENIE3 (random-forest regression), ARACNE (mutual-information based), or SCENIC (GENIE3 + motif-based regulon pruning) could also infer directed TF–target networks from the same expression matrix — These methods represent well-benchmarked alternative paradigms (tree-based, information-theoretic, and combined expression+motif); including one as a comparison in the ChIP-seq gold-standard evaluation would contextualize MERLIN-P-TFA's precision within the landscape of GRN inference tools
-
TPM was used as the expression unit prior to batch correction and network inference↳ Could also: TMM-normalized counts (edgeR) or variance-stabilizing transformation (DESeq2 vst) could also be used as input to GRN inference — TMM and VST account for library-size differences and mean–variance relationships inherent to count data; some GRN benchmarks have found normalized count-based inputs perform comparably or better than TPM for certain inference algorithms
-
Network validation relied on recovery of ChIP-seq-derived edges (precision) and literature-supported edges↳ Could also: AUROC and AUPR (area under the precision-recall curve) computed across a range of network edge-weight thresholds could also summarize recovery performance — Threshold-free metrics like AUPR are less sensitive to the chosen cutoff and are standard in GRN benchmarking studies (e.g., DREAM challenges), allowing more direct comparison with published inference algorithms
-
Batch correction choice between methods was based on visual inspection of PCA/correlation plots and a single ChIP-seq precision comparison↳ Could also: A quantitative kBET (k-nearest-neighbour batch-effect test) or mixing score could also provide a statistical summary of residual batch structure — kBET and related mixing statistics give a single numeric score with an associated p-value for batch mixing, making the comparison between correction strategies less dependent on visual assessment
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41505094
Paper: Carriel et al. (2026) NAR 54(1):gkaf1439 — "A network-based model of Aspergillus fumigatus (GRAsp) elucidates regulators of development and defensive natural products." PMCID PMC12781895, DOI 10.1093/nar/gkaf1439.
Pipeline: 18 published bulk RNA-seq studies (incl. PRJEB2987) → Trimmomatic v0.32 trim → RSEM count vs Af293 ASM265v1.49 → zero-mean batch correction → expression matrix (9859 genes × 294 samples) → MERLIN-P-TFA GRN inference (820 candidate regulators, motif prior restricted to top 189,292 edges, λ=1.0, stability selection 100× subsamples of 147 samples, edges kept at ≥80% confidence) → inferred network (7422 edges, 669 regs, 5274 targets) → 164 consensus modules → GO enrichment → GRAsp web resource.
Code/data artifacts
- Authors' preprocessing repo: github.com/Roy-lab/merlin-preprocess (Trimmomatic/ FastQC/MultiQC only — no inference code, no quantitative output).
- Method repo: github.com/Roy-lab/MERLIN-P-TFA (companion methods paper biorxiv 2025.06.09.658650; ships mESC/yeast/hESC benchmark data, not Aspergillus).
- Inference tool: github.com/Roy-lab/merlin-p (C++).
- Data: SRA/ENA PRJEB2987 (+17 other GEO studies, Supp Table S1).
- Shipped results (EuropePMC supplementaryFiles → gkaf1439_supplemental_files.zip):
- Table S13 Prior_network.xlsx — prior network edge list (Regulator,Target,score).
- Table S14 Module_details.xlsx — per-gene consensus module assignment (9859) + per-module GO enrichment.
- Table S3 Regulators.xlsx — candidate regulator lists (5 source sheets).
- Table S2 — inferred-network targets (conf≥0.8) for 12 ChIP-validated regulators.
IN SCOPE (low-hanging, clearly specified) — what we attempt
Internal-consistency reproduction: recompute the paper's reported summary statistics directly from the shipped per-item supplementary tables (independent recount). This is the auditable, fabrication-detecting target.
- C1 expression-matrix gene count (9859) — from S14.
- C2–C6 module statistics: #modules≥5 (164), genes-in-modules (3381), largest (323), average (21), GO-enriched modules (74) — from S14.
- C7 prior-network size (189,292 edges) — from S13.
- C8 candidate-regulator count (820) — union over S3 source sheets.
OUT OF SCOPE (the hard ~20%, not attempted — why)
- Full MERLIN-P-TFA re-run producing the inferred network (7422 edges / 669 regs / 5274 targets) and top-regulator counts (MAT1-2=216, AtfA=155): requires downloading + RSEM-aligning ~294 RNA-seq samples across 18 studies (hundreds of GB) and multi-day C++ stability-selection inference (100 subsamples). The inferred network edge list is not shipped in the supplement (only interactively on grasp.wid.wisc.edu), so the headline edge/regulator/target counts cannot be recounted from shipped data. A bounded probe of the GRAsp site for a downloadable network was attempted (see AUDIT.md).
- Wet-lab validation (ChIP-seq generation, mutant phenotypes), AUPR/gold-standard benchmarking, OrthoFinder ortholog mapping — out of computational-recount scope.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is an internal-consistency recount of the authors' own shipped supplementary tables, and 7 of 7 fully-shipped headline numbers reproduce exactly (9859 genes, 164 modules, 3381 genes-in-modules, largest 323, 74 GO modules, 189292 prior edges) with the average-module value 20.616→21 a clean rounding match — no fabrication indicator. The only deviation, 820 vs 783 candidate regulators (~5%), sits on our side as an ID-format/counting-convention gap, not an authors' defect. The paper's central result — the GRAsp inferred regulatory network (7422 edges; MAT1-2=216/AtfA=155) — could not be verified because that network was never shipped (web-only), so core-claim support is limited, not refuted. Overall a solid partial reproduction with explainable deviations and a data-availability gap on the central claim.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.