Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A network-based model of Aspergillus fumigatus elucidates regulators of development and defensive natural products of an opportunistic pathogen.

Nucleic Acids Res · 2026
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for an internal-consistency reproduction. The paper (GRAsp; Carriel et al., NAR 2026, gkaf1439) infers an A. fumigatus regulatory network with MERLIN-P-TFA from 294 bulk RNA-seq samples (18 studies incl. PRJEB2987). The named code repo (github.com/Roy-lab/merlin-preprocess) is ONLY preprocessing (Trimmomatic/FastQC/MultiQC) and yields no quantitative paper output; the inference is MERLIN-P-TFA (Roy-lab/merlin-p C++) which would need re-aligning ~294 samples + multi-day stability-selection compute = the out-of-scope hard 20%. Instead I reproduced by an independent RECOUNT of the authors' shipped supplementary tables (gkaf1439_supplemental_files.zip via EuropePMC; all download+parse on «our HPC»/«infra»). 7 of 7 fully-shipped headline numbers reproduce EXACTLY from the per-item data: 9859 genes, 164 modules (>=5 genes), 3381 genes-in-modules, largest module 323, average 21 (20.616), 74 GO-enriched modules, and the 189292-edge prior network. The candidate-regulator count is a 95% near-miss (783 recounted vs 820 reported), most plausibly a named-only/_nca-variant ID or overlap-counting convention rather than a fabricated value. NO fabrication indicators. NOT attempted (data not shipped / hard 20%): the inferred-network headline counts 7422 edges / 669 regulators / 5274 targets and top-regulator out-degrees (MAT1-2=216, AtfA=155) — the inferred edge list lives only on the interactive grasp.wid.wisc.edu site (a bounded download probe hung the compute node and was cancelled); these remain unverified, not refuted. Verdict: 1:1 on every recountable shipped statistic, partial overall because the inferred-network numbers require a from-scratch pipeline run beyond 80/20 scope.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-16 ⛓ ca980eebcf39
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a comprehensive genome-scale gene regulatory network of Aspergillus fumigatus, inferred from integrated public RNA-seq data, recapitulate known regulatory pathways and generate experimentally testable hypotheses about regulators of development, virulence, and defensive natural products?

Core claims
  • A comprehensive genome-wide gene regulatory network resource for A. fumigatus (GRAsp) was constructed from 18 RNA-seq datasets using MERLIN-P-TFA. resource
  • GRAsp recapitulates known regulatory pathways including hypoxia response, iron and zinc homeostasis, ergosterol biosynthesis, and secondary metabolite synthesis. finding
  • GRAsp identified an uncharacterized transcription factor that negatively regulates production of the virulence factor gliotoxin. finding
  • GRAsp revealed the bZip protein AtfA as required for fungal responses to lipo-chitooligosaccharides (LCOs). finding
  • The MERLIN-P-TFA network was computationally validated against published ChIP-seq TF-target relationships and showed high precision. method
  • Zero-mean/quantile-normalization batch correction outperformed Combat-seq for recovering known regulatory edges, justifying its use. method
  • MERLIN-P-TFA estimates hidden transcription factor activity levels from a noisy prior network and infers regulator-target edges plus gene module assignments via regularized regression. method
  • GRAsp is provided as a user-friendly online resource (grasp.wid.wisc.edu) offering module analysis, Steiner tree estimation, and node diffusion. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (publicly available, reanalyzed) Aspergillus fumigatus (strain Af293 reference) various (diverse environmental conditions across 18 datasets) transcript abundance (TPM) Trimmomatic v0.32, RSEM, reference ASM265v1.49
computational GRN inference A. fumigatus integrated transcriptome (9859 genes, 294 measurements) none inferred TF-target edges, TFA matrix, gene modules MERLIN-P-TFA
ChIP-seq-based network validation A. fumigatus none recovery/precision of TF-target gold-standard edges
gliotoxin production assay (experimental validation of TF) A. fumigatus transcription factor manipulation gliotoxin production
LCO response assay (experimental validation of AtfA) A. fumigatus AtfA gene perturbation fungal response to lipo-chitooligosaccharides
Key results
  • GRAsp recovered well-known regulatory relationships for ergosterol biosynthesis, iron homeostasis, and secondary metabolite regulation.
  • An uncharacterized TF predicted by GRAsp was confirmed to negatively regulate gliotoxin production.
  • AtfA was confirmed as required for A. fumigatus responses to LCOs.
  • MERLIN-P-TFA network showed high precision against ChIP-seq gold-standard edges.
  • Zero-mean batch correction offered a slight benefit over Combat-seq for recovering known edges.
Key statistics
  • count 18 bulk RNA-seq studies/datasets (datasets curated for network inference)
  • count 9859 genes across 294 measurements (final zero-mean transformed expression matrix)
  • count 820 putative regulator genes (candidate regulators prepared)
  • count 632 TF binding site sequence motifs (used to construct prior network)
  • count 7 samples removed (S1-S4 and rep3 of PRJEB2987) (samples with unexpected correlation/expression removed)
  • other ~50% mortality rate (COVID-19-associated pulmonary aspergillosis (CAPA))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational study integrates 18 publicly available A. fumigatus bulk RNA-seq datasets (294 samples, 9859 genes after QC) using zero-mean batch correction, then applies the MERLIN-P-TFA probabilistic graphical model—combining L1-regularized regression with network component analysis (NCA)-based TF activity (TFA) estimation—to infer a genome-wide gene regulatory network. Network quality was assessed by recovery of ChIP-seq-derived gold-standard edges and literature-supported regulatory relationships. Batch correction approaches (zero-mean vs. COMBAT-seq) were compared using PCA, variance explained, ANOVA of PC scores, and global pairwise correlation; the experimental validation sections (gliotoxin TF, AtfA/LCO) were not included in the provided text excerpt.

Replicationmixed Sample size18 RNA-seq datasets aggregated; 7 QC-failing samples removed, yielding 294 samples; biological vs. technical replication varies by contributing study GroupsExperimental conditions pooled across 18 heterogeneous RNA-seq studies; no single controlled comparison; network inference is genome-wide regulator-to-target Pairingunclear Randomization/blindingnot stated Dispersionnone Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
PCA (principal component analysis) QC and visualization of batch correction across 18 datasets 294 samples × 9859 genes not stated
ANOVA of PCA component scores Comparison of dataset-driven variance before and after batch correction (zero-mean vs. COMBAT-seq) 294 samples not stated
L2-minimization / alternating least squares (NCA matrix factorization) TFA estimation step within MERLIN-P-TFA 294 samples × number of regulators with prior motif edges not stated
L1-regularized (LASSO-type) regression Regulator-to-target-gene edge inference step within MERLIN-P-TFA, iterated per gene 294 samples per gene model not stated
Precision / edge recovery rate (against ChIP-seq gold standard) Computational validation of inferred GRN edges Number of ChIP-seq-validated edges not stated na
Approaches that could also have been used
  • Zero-mean (log + quantile normalization + within-dataset mean subtraction) was chosen over COMBAT-seq for batch correction based on ChIP-seq edge recovery
    Could also: limma::removeBatchEffect, SVA (surrogate variable analysis), or Harmony could also be applied to integrated multi-study RNA-seq data — SVA and Harmony are commonly used for multi-cohort transcriptomic integration and can model latent batch variables without requiring explicit batch labels; comparing all three via the same gold-standard recovery metric would further substantiate the chosen approach
  • ANOVA on PCA component scores was used to quantify residual batch-driven variance after correction
    Could also: PVCA (principal variance component analysis) or a mixed-model variance partitioning approach (e.g., variancePartition in R) could also decompose variance by batch and biological factors jointly — These methods explicitly partition variance into batch and biological components simultaneously, providing a single quantitative summary of how much residual batch variance remains relative to biologically meaningful variation
  • MERLIN-P-TFA (probabilistic graphical model with L1-regularized regression + NCA-based TFA) was used for GRN inference
    Could also: GENIE3 (random-forest regression), ARACNE (mutual-information based), or SCENIC (GENIE3 + motif-based regulon pruning) could also infer directed TF–target networks from the same expression matrix — These methods represent well-benchmarked alternative paradigms (tree-based, information-theoretic, and combined expression+motif); including one as a comparison in the ChIP-seq gold-standard evaluation would contextualize MERLIN-P-TFA's precision within the landscape of GRN inference tools
  • TPM was used as the expression unit prior to batch correction and network inference
    Could also: TMM-normalized counts (edgeR) or variance-stabilizing transformation (DESeq2 vst) could also be used as input to GRN inference — TMM and VST account for library-size differences and mean–variance relationships inherent to count data; some GRN benchmarks have found normalized count-based inputs perform comparably or better than TPM for certain inference algorithms
  • Network validation relied on recovery of ChIP-seq-derived edges (precision) and literature-supported edges
    Could also: AUROC and AUPR (area under the precision-recall curve) computed across a range of network edge-weight thresholds could also summarize recovery performance — Threshold-free metrics like AUPR are less sensitive to the chosen cutoff and are standard in GRN benchmarking studies (e.g., DREAM challenges), allowing more direct comparison with published inference algorithms
  • Batch correction choice between methods was based on visual inspection of PCA/correlation plots and a single ChIP-seq precision comparison
    Could also: A quantitative kBET (k-nearest-neighbour batch-effect test) or mixing score could also provide a statistical summary of residual batch structure — kBET and related mixing statistics give a single numeric score with an associated p-value for batch mixing, making the comparison between correction strategies less dependent on visual assessment
Software: Trimmomatic 0.32 · RSEM · COMBAT-seq · MERLIN-P-TFA · BioRender

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41505094

Paper: Carriel et al. (2026) NAR 54(1):gkaf1439 — "A network-based model of Aspergillus fumigatus (GRAsp) elucidates regulators of development and defensive natural products." PMCID PMC12781895, DOI 10.1093/nar/gkaf1439.

Pipeline: 18 published bulk RNA-seq studies (incl. PRJEB2987) → Trimmomatic v0.32 trim → RSEM count vs Af293 ASM265v1.49 → zero-mean batch correction → expression matrix (9859 genes × 294 samples) → MERLIN-P-TFA GRN inference (820 candidate regulators, motif prior restricted to top 189,292 edges, λ=1.0, stability selection 100× subsamples of 147 samples, edges kept at ≥80% confidence) → inferred network (7422 edges, 669 regs, 5274 targets) → 164 consensus modules → GO enrichment → GRAsp web resource.

Code/data artifacts

  • Authors' preprocessing repo: github.com/Roy-lab/merlin-preprocess (Trimmomatic/ FastQC/MultiQC only — no inference code, no quantitative output).
  • Method repo: github.com/Roy-lab/MERLIN-P-TFA (companion methods paper biorxiv 2025.06.09.658650; ships mESC/yeast/hESC benchmark data, not Aspergillus).
  • Inference tool: github.com/Roy-lab/merlin-p (C++).
  • Data: SRA/ENA PRJEB2987 (+17 other GEO studies, Supp Table S1).
  • Shipped results (EuropePMC supplementaryFiles → gkaf1439_supplemental_files.zip):
    • Table S13 Prior_network.xlsx — prior network edge list (Regulator,Target,score).
    • Table S14 Module_details.xlsx — per-gene consensus module assignment (9859) + per-module GO enrichment.
    • Table S3 Regulators.xlsx — candidate regulator lists (5 source sheets).
    • Table S2 — inferred-network targets (conf≥0.8) for 12 ChIP-validated regulators.

IN SCOPE (low-hanging, clearly specified) — what we attempt

Internal-consistency reproduction: recompute the paper's reported summary statistics directly from the shipped per-item supplementary tables (independent recount). This is the auditable, fabrication-detecting target.

  • C1 expression-matrix gene count (9859) — from S14.
  • C2–C6 module statistics: #modules≥5 (164), genes-in-modules (3381), largest (323), average (21), GO-enriched modules (74) — from S14.
  • C7 prior-network size (189,292 edges) — from S13.
  • C8 candidate-regulator count (820) — union over S3 source sheets.

OUT OF SCOPE (the hard ~20%, not attempted — why)

  • Full MERLIN-P-TFA re-run producing the inferred network (7422 edges / 669 regs / 5274 targets) and top-regulator counts (MAT1-2=216, AtfA=155): requires downloading + RSEM-aligning ~294 RNA-seq samples across 18 studies (hundreds of GB) and multi-day C++ stability-selection inference (100 subsamples). The inferred network edge list is not shipped in the supplement (only interactively on grasp.wid.wisc.edu), so the headline edge/regulator/target counts cannot be recounted from shipped data. A bounded probe of the GRAsp site for a downloadable network was attempted (see AUDIT.md).
  • Wet-lab validation (ChIP-seq generation, mutant phenotypes), AUPR/gold-standard benchmarking, OrthoFinder ortholog mapping — out of computational-recount scope.
C1_expr_genes
Reported
9859 genes (x294 samples)
Reproduced
9859
exact
C2_modules_ge5
Reported
164 modules >=5 genes
Reproduced
164
exact
C3_genes_in_modules
Reported
3381
Reproduced
3381
exact
C4_largest_module
Reported
323
Reproduced
323
exact
C5_avg_module
Reported
21
Reproduced
20.616 (rounds to 21)
within tolerance
C6_go_modules
Reported
74 of 164
Reproduced
74
exact
C7_prior_edges
Reported
189292
Reproduced
189292
exact
C8_regulators
Reported
820 distinct regulators
Reproduced
783 (union of AFUA IDs over 5 S3 sheets)
partial
C9_inferred_edges
Reported
7422
Reproduced
not attempted (not shipped)
partial
C10_inferred_regulators
Reported
669
Reproduced
not attempted (not shipped)
partial
C11_inferred_targets
Reported
5274
Reproduced
not attempted (not shipped)
partial
C12_top_regulators
Reported
MAT1-2=216, AtfA=155 targets
Reproduced
not attempted (not shipped)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is an internal-consistency recount of the authors' own shipped supplementary tables, and 7 of 7 fully-shipped headline numbers reproduce exactly (9859 genes, 164 modules, 3381 genes-in-modules, largest 323, 74 GO modules, 189292 prior edges) with the average-module value 20.616→21 a clean rounding match — no fabrication indicator. The only deviation, 820 vs 783 candidate regulators (~5%), sits on our side as an ID-format/counting-convention gap, not an authors' defect. The paper's central result — the GRAsp inferred regulatory network (7422 edges; MAT1-2=216/AtfA=155) — could not be verified because that network was never shipped (web-only), so core-claim support is limited, not refuted. Overall a solid partial reproduction with explainable deviations and a data-availability gap on the central claim.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

132.3 k
tokens (I/O) · 7.3 M incl. cache
20 min
runtime · 0 CPU-h
0.1 GB
peak RAM
4 (2 failed)
HPC jobs
hummel
machine