Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The genome and development-dependent transcriptomes of Pyronema confluens: a window into fungal evolution.

PLoS Genet · 2013
L1 76/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
76/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 48% of all assessed papers rank 586 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Pipeline-derived, quantitative results from PMID 24068976 were reproduced with a modern substitute toolchain (HISAT2 in place of the deprecated Tophat + featureCounts) rather than 1:1 with the authors' original code. Genome assembly metrics (scaffolds, length, N50) match exactly; annotated gene counts are within tolerance of the paper's approximate figure. Independently re-aligned and re-quantified RNA-seq expression correlates strongly (Spearman 0.93-0.96) with the authors' own processed/normalized values across all 6 samples, and recomputed differential-expression flags agree at 85-93% raw agreement, weakest for the DD/vegmix comparison which also has the fewest reported positives. The paper's linked gene-prediction parameter repository does not actually contain species-specific parameters for this organism, so that specific pipeline claim could not be directly re-run and is recorded as a scope/data-availability drop for that claim only, not for the room as a whole. No completeness claim is made: NAT detection, functional enrichment, and phylogenomic claims were not attempted.

💻 Code ↗ 🗄 Data: GSE41631

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Because only the atypical, mycorrhizal, subterranean-fruiting black truffle (Tuber melanosporum) genome represents the early-diverging Pezizomycetes, it is unclear which of its features are ancestral to filamentous ascomycetes versus truffle-specific adaptations; the authors therefore sequenced the genome and development-dependent transcriptomes of the saprobic, apothecium-forming Pezizomycete Pyronema confluens to close this sequence gap and draw conclusions about the evolution of fungal genomes and sexual development.

Core claims
  • The 50 Mb P. confluens genome with 13,369 predicted protein-coding genes is more characteristic of higher filamentous ascomycetes than of the large, repeat-rich Tuber melanosporum genome, showing that the truffle's expanded genome is not typical of the Pezizales. finding
  • The P. confluens genome contains an unusually high number of predicted orphan genes, many of which are upregulated during sexual development, consistent with rapid evolution of sex-associated genes. finding
  • The transcription factor gene pro44 is upregulated during development in both P. confluens and Sordaria macrospora, and the P. confluens gene (PCON_06721) complements the S. macrospora pro44 deletion mutant, demonstrating functional conservation of this developmental regulator. finding
  • The genomic environment of the mating type genes that is conserved among higher filamentous ascomycetes is only partly conserved in P. confluens, which has two conserved mating type genes. finding
  • Analysis of spliced RNA-seq reads in antisense orientation identified natural antisense transcripts for 281 genes, indicating NATs are present in P. confluens at levels similar to other filamentous fungi but not as pervasive as in metazoans. finding
  • P. confluens has a full complement of fungal photoreceptors, and expression studies indicate light perception may resemble that of distantly related ascomycetes, thus representing a basic feature of filamentous ascomycetes. finding
  • P. confluens has a low, highly divergent repeat content and gene sets for chromatin modification/silencing (with slight expansion of putative RNAi gene families), while RIP does not appear to play a major role in its genome defense. finding
  • Phylogenomic analysis places the P. confluens lineage at the base of the filamentous ascomycetes, with the Pezizomycetes as sister group to the Orbiliomycetes. finding
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome shotgun sequencing and de novo assembly Pyronema confluens strain CBS100304 none Genome assembly (scaffolds/contigs, assembly size, N50, GC content) Roche/454 and Illumina/Solexa
k-mer analysis of sequence reads for genome size estimation Pyronema confluens (Illumina/Solexa reads) none Predicted total genome size and k-mer coverage peak (k = 31 and 41), ploidy inference algorithm described for the potato genome
RNA-seq (non-strand-specific), two biological replicates per condition P. confluens mycelia: sexual development (sex, minimal medium surface culture, constant light, pooled 3d/4d/5d), long-term dark submerged culture (DD), and a mixture of vegetative tissues (vegmix) growth condition/light (light-dependent fruiting body formation vs. conditions preventing fruiting body formation) Transcript levels and differential expression between conditions (sex/DD, sex/vegmix, DD/vegmix); read coverage for annotation
Gene model prediction and annotation (ab initio plus RNA-seq evidence-based, merged, ~10% manually curated; UTR modeling with custom Perl scripts) P. confluens genome assembly none Number of protein-coding genes, tRNAs, CDS/mRNA lengths and GC, 5′/3′ UTR lengths, rDNA unit (18S, 5.8S, 28S, ITS1/2) MAKER
Spliced-read mapping analysis for natural antisense transcript (NAT) detection, with manual curation P. confluens RNA-seq data mapped to the genome none Antisense splice sites (≥5 spliced reads, >10% of average sense-transcript coverage) and number of genes with NATs Tophat
Repeat annotation by similarity to known repeat classes plus de novo repeat finding; RIP signature analysis of genomic DNA P. confluens genome (compared with T. melanosporum and other fungal genomes) none Fraction of genome as transposable elements >200 bp, low-complexity/simple repeats, sequence identity to repeat consensus, fraction of genome with RIP signatures
Gene-space completeness assessment by BLASTP against a eukaryotic core gene set P. confluens predicted peptides none Number of single-copy core genes recovered BLASTP
Phylome reconstruction, species-tree building, super-tree reconstruction, and molecular dating 18 fungal species including P. confluens, T. melanosporum, A. oligospora and Schizosaccharomyces pombe none Tree topology from 426 single-copy widespread genes, bootstrap values, and estimated divergence times PhyML, DupTree, r8s
Key results
  • Final assembly of 1,588 scaffolds (1,898 contigs) totaling 50 Mb with N50 of 135 kb and 47.8% GC; k-mer analysis independently predicted ~50.1 Mb, close to the assembly length 50 Mb assembly; ~50.1 Mb by k-mer
  • Compared with its closest sequenced relative T. melanosporum, the P. confluens genome is much smaller yet contains nearly twice as many protein-coding genes 50 Mb vs 125 Mb; 13,369 vs 7,496 genes
  • Transposable elements >200 bp make up only 12% (~6 Mb) of the P. confluens genome versus ~58% (~71 Mb) in T. melanosporum; few repeats show high identity to consensus, indicating old, divergent repeats 12% (~6 Mb) vs 58% (~71 Mb)
  • Only a small fraction of the genome showed evidence of RIP, while the N. crassa rid homolog PCON_06255 is slightly upregulated during sexual development 0.46% of genome with RIP signature
  • 376 antisense splice sites were identified in 281 genes; this is likely an underestimate given the stringent criteria and inability to detect unspliced antisense transcripts 376 splice sites in 281 genes
  • All single-copy eukaryotic core genes were present in the predicted peptide set, suggesting the assembly covers the complete core gene space 248/248 core genes
  • The species tree based on 426 single-copy widespread genes placed the Pezizomycetes as sister group to the Orbiliomycetes with full bootstrap support, and the same topology was obtained by super-tree reconstruction from all 6,949 phylome trees bootstrap 100% for all branches
  • Divergence time of the Pyronema and Tuber lineages was estimated by molecular dating under two calibrations 260 or 413 Mya
Key statistics
  • count 13,369 predicted protein-coding genes; 605 predicted tRNAs (P. confluens genome annotation (Table 1))
  • other assembly size 50 Mb; 1,588 scaffolds; N50 135 kb; GC content 47.8%; coding regions 29.2% of genome (Main assembly features)
  • mean average CDS length 1,093 nt (GC 51.1%); average mRNA length 1,483 nt (GC 49.9%); mean/median 5′ UTR 260/156 nt; mean/median 3′ UTR 281/200 nt (Gene/transcript structure statistics)
  • count 376 antisense splice sites in 281 genes (NAT detection from spliced RNA-seq reads; compared with only 33 NATs in T. melanosporum)
  • other transposable elements >200 bp = 12% (~6 Mb); low complexity regions 0.21%; simple repeats 1.03%; RIP-affected regions 0.46% (Repeat and genome-defense analysis of P. confluens)
  • other 58% (~71 Mb) of the T. melanosporum genome consists of repeats larger than 200 bp; genome 125 Mb with 7,496 protein-coding genes (Comparison values for the closest sequenced relative)
  • count all 248 single-copy core genes present (Eukaryotic core gene set BLASTP completeness check)
  • count 426 single-copy widespread genes for species tree; 6,949 phylome trees; 100 bootstrap repetitions with 100% support; divergence estimates 260/413 Mya calibrated at 723.86 and 1147.78 Mya (Phylogenomic analysis across 18 fungal species)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study reports a de novo genome assembly (Roche/454 + Illumina/Solexa) of Pyronema confluens, validated by k-mer-based genome size estimation and BLASTP-based core-gene completeness checks, together with RNA-seq transcriptome profiling of three growth/developmental conditions (sexual development, dark-grown, vegetative mix) using two biological replicates per condition. Comparative genomics and phylogenomics were performed using maximum-likelihood tree reconstruction (PhyML) with bootstrap support and a super-tree consistency check (DupTree), plus molecular-clock divergence-time estimation (r8s) calibrated with fixed reference dates. Other analyses (natural antisense transcript detection, repeat content characterization) were reported using coverage- and similarity-based thresholds with manual curation rather than classical inferential hypothesis tests, and results in this excerpt are presented mainly as counts, percentages, and point estimates.

Replicationbiological Sample sizetwo biological replicates per condition (sex, DD, vegmix) for RNA-seq; no formal power/sample-size calculation described Groupsthree growth/developmental transcriptome conditions (sexual development, dark-grown submerged culture, vegetative tissue mix); genome/gene-content comparisons across fungal species Pairingunclear Randomization/blindingnot stated Dispersionnone
Statistical tests used
Test Applied to n Assumptions
bootstrap resampling (100 replicates) for phylogenetic node support species tree reconstruction (Figure 2), PhyML analysis of 426 single-copy genes 426 single-copy widespread genes; 100 bootstrap replicates not stated
not stated (fixed coverage/read-count thresholds with manual curation, not a formal statistical test) identification of natural antisense transcripts (NATs) from spliced RNA-seq reads antisense splice sites with ≥5 spliced reads and >10% of average sense-transcript coverage; two biological replicates per condition not stated
not stated in provided excerpt calling genes as differentially regulated between sex, DD, and vegmix conditions two biological replicates per condition not stated
Approaches that could also have been used
  • Differential expression across the sex/DD/vegmix conditions is described qualitatively (genes 'differentially regulated' in certain comparisons) without a named statistical test in this excerpt.
    Could also: a replicate-aware count-based model such as DESeq2 or edgeR with per-gene dispersion estimation across the two biological replicates — such models explicitly model biological variability between replicates and yield p-values/FDR-adjusted significance for each condition contrast, which can complement a threshold- or fold-change-based call
  • Natural antisense transcripts were called using fixed coverage thresholds (≥5 spliced reads, >10% of average sense coverage) followed by manual review.
    Could also: a model-based statistical comparison of antisense versus sense read counts (e.g., using a count model with FDR control) — a formal statistical test would allow an explicit false-discovery-rate estimate for the NAT calls, complementing the threshold-and-curation approach
  • Phylogenetic support for the species tree was assessed with 100 bootstrap replicates plus one alternative super-tree method (DupTree).
    Could also: additional or complementary support measures such as SH-aLRT or Bayesian posterior probabilities (e.g., via RAxML or MrBayes/BEAST) — combining multiple support metrics can help cross-validate node confidence, particularly for short internal branches near the base of the tree
  • Divergence times for the Pyronema/Tuber split were reported as single point estimates (260 or 413 Mya) depending on the calibration used.
    Could also: a relaxed molecular-clock approach (e.g., BEAST) reporting posterior credibility intervals around divergence estimates — credibility intervals convey the uncertainty from calibration choice and substitution-rate variation alongside the point estimate
  • Genome completeness was assessed as a binary presence/absence BLASTP check against a 248-gene core set.
    Could also: a graduated completeness assessment such as BUSCO, which categorizes genes as complete single-copy, duplicated, fragmented, or missing — a graduated metric can convey partial completeness (e.g., fragmented hits) that a simple presence/absence check does not capture
  • Lower repeat-sequence identity to consensus repeats in P. confluens versus other species was illustrated visually (Figure S3) rather than via a formal statistical comparison.
    Could also: a non-parametric comparison (e.g., Mann-Whitney U) of repeat-identity distributions between species — a formal test would quantify whether the apparent difference in repeat divergence between species is unlikely to arise by chance, complementing the visual comparison
Software: PhyML · DupTree · r8s · MAKER · Tophat · BLASTP

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

genome_assembly_stats
Reported
Genome assembly ~50 Mb, 1,588 scaffolds, N50 ~135 kb (Pyronema confluens Pcon_v1 / GCA_000981605.1); paper Results/genome section
Reproduced
total-length 50,026,054 bp; scaffold-count 1,588; scaffold-N50 135,140 bp (NCBI assembly_stats.txt, GCA_000981605.1)
exact
protein_coding_gene_count
Reported
~13,400 protein-coding genes (paper abstract/results, approximate figure)
Reproduced
13,978 gene features / 13,369 mRNA features in NCBI feature_table.txt; 13,368 gene rows in GEO supplementary processed table (GSE41631)
within tolerance
rnaseq_alignment_pipeline
Reported
RNA-seq reads (12 SRA runs; 6 conditions x 2 biological replicates: sexual mycelium, vegetative dark-grown [DD], vegetative mixed-light [vegmix]) aligned to the genome (originally via Tophat) and quantified per gene
Reproduced
All 12 ENA FASTQ run-pairs re-downloaded; HISAT2 (splice-aware, GTF-guided known-splice-sites) re-alignment per merged biological sample; overall alignment rates: sex_1 86.61%, sex_2 92.77%, veg_1 89.10%, veg_2 91.88%, vegmix_1 93.15%, vegmix_2 91.46%. featureCounts read-assignment rates: 82.0-88.2% across the 6 samples.
within tolerance
expression_level_concordance_vs_geo_table
Reported
Author-processed, normalized per-gene RNA-seq expression values for 6 samples (GEO supplementary file GSE41631_Pyronema_confluens_expression.txt)
Reproduced
Independently re-aligned + re-quantified + RPKM-computed values correlate with the paper's normalized counts: log2 Pearson 0.908-0.926, Spearman 0.934-0.957 (n=13,368 genes common to both, all 6 samples)
within tolerance
differential_expression_calls
Reported
Per-gene lenient (>2x/<0.5x) and stringent (>4x/<0.25x) differential-expression flags for 3 pairwise comparisons: sex/DD, sex/vegmix, DD/vegmix (GEO supplementary table flag columns)
Reproduced
Recomputed fold-change flags from independently re-quantified RPKM replicate means. Raw agreement: sex/DD lenient 86.72%/stringent 91.68%; sex/vegmix lenient 87.25%/stringent 92.82%; DD/vegmix lenient 84.49%/stringent 92.77%. Jaccard-on-positives (more informative given class imbalance): sex/DD 0.643/0.661; sex/vegmix 0.649/0.694; DD/vegmix 0.269/0.178 (weakest, and also the comparison with fewest paper-reported positives: n=774 lenient, 232 stringent).
partial
gene_prediction_pipeline_reproduction
Reported
Gene models predicted via an AUGUSTUS/SNAP-based pipeline using organism-trained parameters; code repo linked as hyphaltip/fungi-gene-prediction-params
Reproduced
Not run: repository search found no Pyronema/Pcon-specific parameter set in the linked repo. JGI MycoCosm's Pyrco1 entry states its annotation copy was sourced from the paper's own deposited data, not an independent JGI gene-prediction run. Only indirect check available: deposited annotation gene count (see protein_coding_gene_count claim).
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 76/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

Data availability is exemplary: all 12 raw ENA run-pairs, the assembly (GCA_000981605.1) and the authors' processed matrix were retrieved 1:1, and the headline numbers reproduce - 50,026,054 bp / 1,588 scaffolds / N50 135,140 bp match the paper exactly, and 13,369 mRNA features match the '~13,400' gene figure to within 0.2%. The deviations sit on our side of the pipeline: with Tophat deprecated we substituted HISAT2 + featureCounts + our own RPKM normalization, which still regenerates the authors' expression matrix at Spearman 0.934-0.957 across all six samples - strong evidence the deposited processed values genuinely derive from the deposited reads - but shifts enough genes across the fixed 2x/4x thresholds that binary DE-flag set overlap drops, worst for the low-signal DD/vegmix comparison (Jaccard 0.269 lenient / 0.178 stringent, only 774/232 reported positives). One genuine authors'-side gap: the linked code repo contains no Pyronema-specific AUGUSTUS/SNAP parameters, so the gene-prediction step is not re-runnable from the cited resource. Overall this is a solid, non-suspicious reproduction whose residual disagreements are attributable to toolchain and normalization substitution rather than to any defect in the reported science - hence yellow rather than green on the overall judgement.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.