The genome and development-dependent transcriptomes of Pyronema confluens: a window into fungal evolution.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Pipeline-derived, quantitative results from PMID 24068976 were reproduced with a modern substitute toolchain (HISAT2 in place of the deprecated Tophat + featureCounts) rather than 1:1 with the authors' original code. Genome assembly metrics (scaffolds, length, N50) match exactly; annotated gene counts are within tolerance of the paper's approximate figure. Independently re-aligned and re-quantified RNA-seq expression correlates strongly (Spearman 0.93-0.96) with the authors' own processed/normalized values across all 6 samples, and recomputed differential-expression flags agree at 85-93% raw agreement, weakest for the DD/vegmix comparison which also has the fewest reported positives. The paper's linked gene-prediction parameter repository does not actually contain species-specific parameters for this organism, so that specific pipeline claim could not be directly re-run and is recorded as a scope/data-availability drop for that claim only, not for the room as a whole. No completeness claim is made: NAT detection, functional enrichment, and phylogenomic claims were not attempted.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBecause only the atypical, mycorrhizal, subterranean-fruiting black truffle (Tuber melanosporum) genome represents the early-diverging Pezizomycetes, it is unclear which of its features are ancestral to filamentous ascomycetes versus truffle-specific adaptations; the authors therefore sequenced the genome and development-dependent transcriptomes of the saprobic, apothecium-forming Pezizomycete Pyronema confluens to close this sequence gap and draw conclusions about the evolution of fungal genomes and sexual development.
- ★ The 50 Mb P. confluens genome with 13,369 predicted protein-coding genes is more characteristic of higher filamentous ascomycetes than of the large, repeat-rich Tuber melanosporum genome, showing that the truffle's expanded genome is not typical of the Pezizales. finding
- ★ The P. confluens genome contains an unusually high number of predicted orphan genes, many of which are upregulated during sexual development, consistent with rapid evolution of sex-associated genes. finding
- ★ The transcription factor gene pro44 is upregulated during development in both P. confluens and Sordaria macrospora, and the P. confluens gene (PCON_06721) complements the S. macrospora pro44 deletion mutant, demonstrating functional conservation of this developmental regulator. finding
- ★ The genomic environment of the mating type genes that is conserved among higher filamentous ascomycetes is only partly conserved in P. confluens, which has two conserved mating type genes. finding
- ★ Analysis of spliced RNA-seq reads in antisense orientation identified natural antisense transcripts for 281 genes, indicating NATs are present in P. confluens at levels similar to other filamentous fungi but not as pervasive as in metazoans. finding
- ★ P. confluens has a full complement of fungal photoreceptors, and expression studies indicate light perception may resemble that of distantly related ascomycetes, thus representing a basic feature of filamentous ascomycetes. finding
- P. confluens has a low, highly divergent repeat content and gene sets for chromatin modification/silencing (with slight expansion of putative RNAi gene families), while RIP does not appear to play a major role in its genome defense. finding
- Phylogenomic analysis places the P. confluens lineage at the base of the filamentous ascomycetes, with the Pezizomycetes as sister group to the Orbiliomycetes. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome shotgun sequencing and de novo assembly | Pyronema confluens strain CBS100304 | none | Genome assembly (scaffolds/contigs, assembly size, N50, GC content) | Roche/454 and Illumina/Solexa |
| k-mer analysis of sequence reads for genome size estimation | Pyronema confluens (Illumina/Solexa reads) | none | Predicted total genome size and k-mer coverage peak (k = 31 and 41), ploidy inference | algorithm described for the potato genome |
| RNA-seq (non-strand-specific), two biological replicates per condition | P. confluens mycelia: sexual development (sex, minimal medium surface culture, constant light, pooled 3d/4d/5d), long-term dark submerged culture (DD), and a mixture of vegetative tissues (vegmix) | growth condition/light (light-dependent fruiting body formation vs. conditions preventing fruiting body formation) | Transcript levels and differential expression between conditions (sex/DD, sex/vegmix, DD/vegmix); read coverage for annotation | — |
| Gene model prediction and annotation (ab initio plus RNA-seq evidence-based, merged, ~10% manually curated; UTR modeling with custom Perl scripts) | P. confluens genome assembly | none | Number of protein-coding genes, tRNAs, CDS/mRNA lengths and GC, 5′/3′ UTR lengths, rDNA unit (18S, 5.8S, 28S, ITS1/2) | MAKER |
| Spliced-read mapping analysis for natural antisense transcript (NAT) detection, with manual curation | P. confluens RNA-seq data mapped to the genome | none | Antisense splice sites (≥5 spliced reads, >10% of average sense-transcript coverage) and number of genes with NATs | Tophat |
| Repeat annotation by similarity to known repeat classes plus de novo repeat finding; RIP signature analysis of genomic DNA | P. confluens genome (compared with T. melanosporum and other fungal genomes) | none | Fraction of genome as transposable elements >200 bp, low-complexity/simple repeats, sequence identity to repeat consensus, fraction of genome with RIP signatures | — |
| Gene-space completeness assessment by BLASTP against a eukaryotic core gene set | P. confluens predicted peptides | none | Number of single-copy core genes recovered | BLASTP |
| Phylome reconstruction, species-tree building, super-tree reconstruction, and molecular dating | 18 fungal species including P. confluens, T. melanosporum, A. oligospora and Schizosaccharomyces pombe | none | Tree topology from 426 single-copy widespread genes, bootstrap values, and estimated divergence times | PhyML, DupTree, r8s |
- – Final assembly of 1,588 scaffolds (1,898 contigs) totaling 50 Mb with N50 of 135 kb and 47.8% GC; k-mer analysis independently predicted ~50.1 Mb, close to the assembly length 50 Mb assembly; ~50.1 Mb by k-mer
- – Compared with its closest sequenced relative T. melanosporum, the P. confluens genome is much smaller yet contains nearly twice as many protein-coding genes 50 Mb vs 125 Mb; 13,369 vs 7,496 genes
- ▼ Transposable elements >200 bp make up only 12% (~6 Mb) of the P. confluens genome versus ~58% (~71 Mb) in T. melanosporum; few repeats show high identity to consensus, indicating old, divergent repeats 12% (~6 Mb) vs 58% (~71 Mb)
- ▲ Only a small fraction of the genome showed evidence of RIP, while the N. crassa rid homolog PCON_06255 is slightly upregulated during sexual development 0.46% of genome with RIP signature
- – 376 antisense splice sites were identified in 281 genes; this is likely an underestimate given the stringent criteria and inability to detect unspliced antisense transcripts 376 splice sites in 281 genes
- – All single-copy eukaryotic core genes were present in the predicted peptide set, suggesting the assembly covers the complete core gene space 248/248 core genes
- – The species tree based on 426 single-copy widespread genes placed the Pezizomycetes as sister group to the Orbiliomycetes with full bootstrap support, and the same topology was obtained by super-tree reconstruction from all 6,949 phylome trees bootstrap 100% for all branches
- – Divergence time of the Pyronema and Tuber lineages was estimated by molecular dating under two calibrations 260 or 413 Mya
- count 13,369 predicted protein-coding genes; 605 predicted tRNAs (P. confluens genome annotation (Table 1))
- other assembly size 50 Mb; 1,588 scaffolds; N50 135 kb; GC content 47.8%; coding regions 29.2% of genome (Main assembly features)
- mean average CDS length 1,093 nt (GC 51.1%); average mRNA length 1,483 nt (GC 49.9%); mean/median 5′ UTR 260/156 nt; mean/median 3′ UTR 281/200 nt (Gene/transcript structure statistics)
- count 376 antisense splice sites in 281 genes (NAT detection from spliced RNA-seq reads; compared with only 33 NATs in T. melanosporum)
- other transposable elements >200 bp = 12% (~6 Mb); low complexity regions 0.21%; simple repeats 1.03%; RIP-affected regions 0.46% (Repeat and genome-defense analysis of P. confluens)
- other 58% (~71 Mb) of the T. melanosporum genome consists of repeats larger than 200 bp; genome 125 Mb with 7,496 protein-coding genes (Comparison values for the closest sequenced relative)
- count all 248 single-copy core genes present (Eukaryotic core gene set BLASTP completeness check)
- count 426 single-copy widespread genes for species tree; 6,949 phylome trees; 100 bootstrap repetitions with 100% support; divergence estimates 260/413 Mya calibrated at 723.86 and 1147.78 Mya (Phylogenomic analysis across 18 fungal species)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study reports a de novo genome assembly (Roche/454 + Illumina/Solexa) of Pyronema confluens, validated by k-mer-based genome size estimation and BLASTP-based core-gene completeness checks, together with RNA-seq transcriptome profiling of three growth/developmental conditions (sexual development, dark-grown, vegetative mix) using two biological replicates per condition. Comparative genomics and phylogenomics were performed using maximum-likelihood tree reconstruction (PhyML) with bootstrap support and a super-tree consistency check (DupTree), plus molecular-clock divergence-time estimation (r8s) calibrated with fixed reference dates. Other analyses (natural antisense transcript detection, repeat content characterization) were reported using coverage- and similarity-based thresholds with manual curation rather than classical inferential hypothesis tests, and results in this excerpt are presented mainly as counts, percentages, and point estimates.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| bootstrap resampling (100 replicates) for phylogenetic node support | species tree reconstruction (Figure 2), PhyML analysis of 426 single-copy genes | 426 single-copy widespread genes; 100 bootstrap replicates | not stated |
| not stated (fixed coverage/read-count thresholds with manual curation, not a formal statistical test) | identification of natural antisense transcripts (NATs) from spliced RNA-seq reads | antisense splice sites with ≥5 spliced reads and >10% of average sense-transcript coverage; two biological replicates per condition | not stated |
| not stated in provided excerpt | calling genes as differentially regulated between sex, DD, and vegmix conditions | two biological replicates per condition | not stated |
-
Differential expression across the sex/DD/vegmix conditions is described qualitatively (genes 'differentially regulated' in certain comparisons) without a named statistical test in this excerpt.↳ Could also: a replicate-aware count-based model such as DESeq2 or edgeR with per-gene dispersion estimation across the two biological replicates — such models explicitly model biological variability between replicates and yield p-values/FDR-adjusted significance for each condition contrast, which can complement a threshold- or fold-change-based call
-
Natural antisense transcripts were called using fixed coverage thresholds (≥5 spliced reads, >10% of average sense coverage) followed by manual review.↳ Could also: a model-based statistical comparison of antisense versus sense read counts (e.g., using a count model with FDR control) — a formal statistical test would allow an explicit false-discovery-rate estimate for the NAT calls, complementing the threshold-and-curation approach
-
Phylogenetic support for the species tree was assessed with 100 bootstrap replicates plus one alternative super-tree method (DupTree).↳ Could also: additional or complementary support measures such as SH-aLRT or Bayesian posterior probabilities (e.g., via RAxML or MrBayes/BEAST) — combining multiple support metrics can help cross-validate node confidence, particularly for short internal branches near the base of the tree
-
Divergence times for the Pyronema/Tuber split were reported as single point estimates (260 or 413 Mya) depending on the calibration used.↳ Could also: a relaxed molecular-clock approach (e.g., BEAST) reporting posterior credibility intervals around divergence estimates — credibility intervals convey the uncertainty from calibration choice and substitution-rate variation alongside the point estimate
-
Genome completeness was assessed as a binary presence/absence BLASTP check against a 248-gene core set.↳ Could also: a graduated completeness assessment such as BUSCO, which categorizes genes as complete single-copy, duplicated, fragmented, or missing — a graduated metric can convey partial completeness (e.g., fragmented hits) that a simple presence/absence check does not capture
-
Lower repeat-sequence identity to consensus repeats in P. confluens versus other species was illustrated visually (Figure S3) rather than via a formal statistical comparison.↳ Could also: a non-parametric comparison (e.g., Mann-Whitney U) of repeat-identity distributions between species — a formal test would quantify whether the apparent difference in repeat divergence between species is unlikely to arise by chance, complementing the visual comparison
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Data availability is exemplary: all 12 raw ENA run-pairs, the assembly (GCA_000981605.1) and the authors' processed matrix were retrieved 1:1, and the headline numbers reproduce - 50,026,054 bp / 1,588 scaffolds / N50 135,140 bp match the paper exactly, and 13,369 mRNA features match the '~13,400' gene figure to within 0.2%. The deviations sit on our side of the pipeline: with Tophat deprecated we substituted HISAT2 + featureCounts + our own RPKM normalization, which still regenerates the authors' expression matrix at Spearman 0.934-0.957 across all six samples - strong evidence the deposited processed values genuinely derive from the deposited reads - but shifts enough genes across the fixed 2x/4x thresholds that binary DE-flag set overlap drops, worst for the low-signal DD/vegmix comparison (Jaccard 0.269 lenient / 0.178 stringent, only 774/232 reported positives). One genuine authors'-side gap: the linked code repo contains no Pyronema-specific AUGUSTUS/SNAP parameters, so the gene-prediction step is not re-runnable from the cited resource. Overall this is a solid, non-suspicious reproduction whose residual disagreements are attributable to toolchain and normalization substitution rather than to any defect in the reported science - hence yellow rather than green on the overall judgement.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.