Corpus 1,273 assessed · 1,174 scored · 643 reproduced ≥75 · 169 flagged ·∅ 74.1/100
← New search

A comprehensive resource of genomic, epigenomic and transcriptomic sequencing data for the black truffle Tuber melanosporum.

Gigascience · 2014
61/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
61/100
Reproducibility score
0.7 SD below mean
vs. all fields · 1174 studies
🎯 Scores higher than 22% of all assessed papers rank 906 of 1174 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the three named pipelines (RNA-seq TopHat+Cufflinks+HTSeq+DESeq2; WGBS BS-Seeker2 methylation calling; CRI/RIP repeat analysis via critool+RepeatMasker) from PMID 25392735 end-to-end on all 9 GSE49700 samples (5 WGBS, 4 RNA-seq). RNA-seq raw read counts match the paper exactly for all 4 samples; unique-mapping rate matches closely for FLM/FLM_aza/FLM_untreated but is notably lower than reported for FB. WGBS genome-wide %methylation reproduces well for CG context in most samples and very well for FLM_aza across all 3 contexts, but shows a systematic upward inflation in CHG/CHH for FB/FLM/FLM_untreated and a large, consistent mismatch for ECM across all 3 contexts. CRI/RIP results are qualitatively consistent with the paper's claim that T. melanosporum shows a weaker RIP-like signature than N. crassa, though the paper gives no numeric CRI value to match exactly. Cufflinks novel-gene counts are consistently ~2.3-3x higher than reported for all 4 samples (same direction, different magnitude). DESeq2 differential expression (5-aza treated vs untreated, via an approximate no-replicate workaround) found 139 significant genes vs the paper's reported 115 -- same order of magnitude. WG-seq (whole-genome resequencing) was correctly excluded as out of scope. All large reproducible intermediates (~208GB: raw/trimmed reads, BAM/CGmap/ATCGmap, conda pkg caches, RepeatMasker scratch output) were deleted after result extraction; curated small artifacts and logs are retained under scratch/pmid-25392735/curated/ and logs/.

💻 Code ↗ 🗄 Data: GSE49700

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-05
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper asks if and how the black truffle Tuber melanosporum, which has one of the largest and most transposable-element-rich fungal genomes (125 Mb, >58% repetitive DNA, >45,000 repeated elements), exploits DNA methylation to silence transposable elements and maintain genome integrity, and whether this methylation is reversible and linked to transcription.

Core claims
  • T. melanosporum shows a high rate of cytosine methylation (>44%) that selectively targets transposable elements rather than genes, with a strong preference for CpG sites. finding
  • Silenced TEs are highly methylated whereas expressed TEs show an anti-correlation between methylation and transcription, indicating TE expression is regulated by DNA methylation; gene-body methylation shows no relationship with gene expression level. mechanism
  • Whole-genome sequencing uncovered multiple TE-enriched copy number variant regions containing a significant fraction of hypomethylated and expressed TEs, almost exclusively in in vitro-propagated free-living mycelium. finding
  • Treatment of mycelia with the demethylating agent 5-azacytidine partially reduced DNA methylation and increased TE transcription, indicating a non-exhaustive and partly reversible methylation process. finding
  • A comprehensive BS-seq, WG-seq and RNA-seq dataset across fruitbody, free-living mycelium and ectomycorrhiza developmental stages is provided as a community resource for evolutionary (epi)genomics of Pezizomycotina. resource
  • Transcriptome assembly identified 614 novel genes not present in the original truffle v1.0 annotation, including 68 novel genes/transcripts present only in 5-aza treated mycelium. finding
  • Genes differentially expressed between untreated and 5-aza treated mycelia are enriched for oxidation-reduction processes and transmembrane transport, supporting an epigenetic regulatory system responsive to environmental stress. finding
  • An implementation of the Composite Repeat Induced Point mutation index (CRI), a dinucleotide frequency distribution tool for assessing RIP-induced methylation, is provided. method
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome bisulfite sequencing (BS-seq) Tuber melanosporum Mel28 strain free-living mycelium (grown 5 weeks on 1% malt agar) and field-collected mature fruitbody none Per-cytosine methylation level #C/(#C+#T) in CG/CHG/CHH contexts; genome-wide methylation profiles and read-depth-based CNV Illumina HiSeq 2000; EpiTect bisulfite kit (QIAGEN); DNeasy Plant Mini kit; BS Seeker 2 alignment to Tuber_melanosporum_v1.0
Whole-genome bisulfite sequencing (BS-seq) Ectomycorrhizal (ECM) root tips of common hazel (Corylus avellana L.) plantlets inoculated with T. melanosporum mycelium slurry symbiotic association (inoculation) Low-resolution global CG/CHG/CHH methylation profiles of the symbiotic stage (1.18X per-strand coverage) Illumina HiSeq 2000; BS Seeker 2
Whole-genome bisulfite sequencing (BS-seq) T. melanosporum mycelia grown in the dark at 23 °C in synthetic liquid medium 5-azacytidine (drug) added every 5 days at 10, 40 and 100 μM final concentration for 45 days, versus water-added control Genome, gene and TE methylation levels in CG/CHG/CHH contexts in treated vs untreated mycelia Illumina HiSeq 2000; EpiTect kit (QIAGEN)
Whole-genome (non-bisulfite) DNA sequencing (WG-seq) T. melanosporum free-living mycelium and fruitbody genomic DNA none Read-depth coverage used to independently confirm copy number variant regions called from BS-seq Standard Illumina sequencing (51-bp reads)
mRNA sequencing (RNA-seq) T. melanosporum fruitbody and free-living mycelium none Read counts per gene and per TE; expression levels as RPKM; coverage depth per gene/TE Illumina TruSeq RNA Sample Preparation kit, Illumina HiSeq 2000, 50-bp single-end reads; RNeasy Plant Mini kit; Bioanalyzer (Agilent); Qubit RNA BR Assay kit; TopHat/HTSeq v0.5.4p3/FastQC
mRNA sequencing (RNA-seq) T. melanosporum free-living mycelium in synthetic liquid medium 5-azacytidine treatment (10, 40, 100 μM over 45 days) vs untreated control Differential gene expression (RPKM, fold change, adjusted p value), TE transcription, functional enrichment via Blast2GO, and treatment-specific novel transcripts Illumina, 51-bp reads; Blast2GO
Guided transcriptome assembly / novel gene prediction (computational) RNA-seq data from FB, FLM, 5-aza treated and untreated FLM mapped to truffle v1.0 reference assembly (7,496 genes) none Number of novel intergenic transcripts (Cuffcompare class code 'u', FPKM ≥4 in the 95% confidence interval) clustered into novel genes with BlastX protein homology TopHat, Cufflinks, Cuffcompare, Blast2GO (BlastX)
Composite Repeat Induced Point mutation index (CRI) dinucleotide frequency analysis T. melanosporum genome sequence none Likelihood that genome DNA methylation is induced by Repeat-Induced Point mutation (RIP)
Key results
  • Cytosine methylation rate across the T. melanosporum genome is high (>44%), selectively targeting TEs rather than genes with a strong preference for CpG sites. >44%
  • TEs are heavily CG-methylated (70.51% FLM, 69.71% FB, 59.16% ECM) while genes are almost devoid of methylation (CG 0.64% FLM, 0.87% FB, 0.84% ECM). TE CG 70.51% vs gene CG 0.64% (FLM)
  • Genome-wide CG methylation was 30.3% (FLM), 28.9% (FB) and 6.4% (ECM); CHG and CHH levels were much lower (8.1-10.3% in FLM/FB). CG 30.3%/28.9%/6.4%
  • Genes are methylation-depleted in all contexts with no difference between high, medium and low expression classes, whereas silenced TEs are 70-80% methylated and expressed TEs show clear methylation/transcription anti-correlation. silenced TEs 70-80% methylated
  • 107 genomic regions with significant CNV between FLM and FB (100 kb windows with |log coverage ratio FLM/FB| ≥ 0.3) were identified, covering 7.3% of the genome; 102 (95%) were independently confirmed by non-bisulfite Illumina WGS of FLM DNA. 107 regions, 7.3% of genome, 95% confirmed
  • A total of 614 novel genes were identified after merging assemblies (282 in FB, 328 in FLM, 466 in 5-aza treated, 440 in 5-aza untreated), following removal of 11 extremely long genes (≥36,573 bp). 614 novel genes
  • 68 novel genes/transcripts were detected only in 5-aza treated mycelium and not in untreated mycelium. 68 novel genes/transcripts
  • Oxidation-reduction and transmembrane transport genes were strongly induced by 5-aza, e.g. flavin-binding monooxygenase-like protein (RPKM 8 → 258) and extracellular dioxygenase (RPKM 18 → 196), while others such as MFS drug efflux were repressed (RPKM 663 → 84). 33.80-fold up; 10.71-fold up; 0.13-fold (down)
Key statistics
  • fold_change 33.80 (GSTUMT00004482001 flavin-binding monooxygenase-like protein, 5-aza treated vs untreated mycelium (RPKM 8 → 258), adj. p = 3.51E-09)
  • pvalue 4.71E-13 (Lowest reported adjusted p value; GSTUMT00001850001 endo-beta-glucanase eng1, 16.02-fold up in 5-aza treated mycelium)
  • fold_change 0.10 (GSTUMT00007117001 D-isomer specific 2-hydroxyacid dehydrogenase, most strongly down-regulated by 5-aza (RPKM 4 → 0.4), adj. p = 4.66E-02)
  • count 107 CNV regions corresponding to 7.3% of the genome; 102 (95%) confirmed by WGS (Copy number variant regions between FLM and FB called from BS-seq read depth (100 kb windows, |log ratio| ≥ 0.3))
  • count 614 (Novel genes identified from the merged RNA-seq transcript assembly against truffle v1.0 (7,496 annotated genes))
  • other 31.61X (FLM), 34.73X (FB), 1.18X (ECM), 8.44X (5-aza treated), 7.21X (5-aza untreated) coverage per strand (BS-seq library coverage; ~90% of cytosines retained for FB/FLM analysis and ~75% for 5-aza samples (coverage ≥4 reads))
  • count 63,742,213 (FB) and 74,109,065 (FLM) uniquely mapped RNA-seq reads, 93.34% and 84.88% overall mapping (RNA-seq alignment statistics for fruitbody and free-living mycelium)
  • other RIN 7.0 (FB) and 6.5 (FLM); ~75% of genes covered at ≥10 reads in FB/FLM, ~90% in 5-aza samples; 55-60% of TEs covered in both samples per comparison (RNA quality and RNA-seq coverage completeness metrics)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a data-note resource paper describing whole-genome bisulfite sequencing (BS-seq), whole-genome sequencing (WGS) and RNA-seq of Tuber melanosporum across developmental stages (FLM, FB, ECM) and 5-azacytidine treated/untreated mycelium, apparently one sample per condition. Differential gene expression between untreated and 5-aza treated mycelia is reported as RPKM fold-change with an adjusted p-value per gene (Table 5), and novel transcripts were called from Cufflinks/Cuffcompare assemblies using an FPKM threshold with a 95% confidence interval. Copy-number variant regions were defined using a fixed log-coverage-ratio cutoff, and results are otherwise reported as summary percentages/coverage statistics rather than through classical group-comparison tests.

Replicationunclear Groupsdevelopmental stages (FLM, FB, ECM) and 5-aza treated vs untreated mycelium Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionnot stated (Table 5 reports 'adj. p value' without naming the correction method)
Statistical tests used
Test Applied to n Assumptions
not stated (differential expression reported as fold change with an adjusted p-value) Table 5, genes differentially expressed between untreated and 5-aza treated mycelia not stated
FPKM 95% confidence interval threshold (Cufflinks/Cuffcompare) used to call novel transcripts novel gene/transcript prediction from RNA-seq assemblies not stated
Approaches that could also have been used
  • Differential expression between untreated and 5-aza treated mycelia was summarized as fold change with an adjusted p-value, without the underlying statistical/count model being named in the text.
    Could also: A count-based RNA-seq differential expression tool such as DESeq2 or edgeR could also be used — These explicitly model read-count overdispersion and provide a documented statistical framework (e.g., negative-binomial Wald or likelihood-ratio tests) alongside the fold-change and adjusted p-value already reported.
  • Comparisons across developmental stages and treatment appear to be based on one sample per condition rather than multiple biological replicates.
    Could also: Designs with multiple biological replicates per condition, analyzed with replicate-aware methods (e.g., DESeq2/edgeR or ANOVA), could also be used — Replicates allow estimation of biological variability and would let dispersion (e.g., SD, CI) be reported alongside the point estimates already given.
  • Multiple testing correction was applied to the differential expression results (an 'adjusted p value' is reported) without naming the specific method.
    Could also: Explicitly stating and applying a named method such as Benjamini-Hochberg FDR or Bonferroni could also be used — Naming the method makes the stringency of the correction and the assumptions behind it transparent and reproducible.
  • Novel transcripts were defined using a fixed FPKM threshold (>=4) within a 95% FPKM confidence interval from Cufflinks/Cuffcompare.
    Could also: A permutation-based or model-based transcript-confidence approach (e.g., as implemented in Cuffdiff or a Bayesian expression caller) could also be used — Such approaches can provide an alternative, model-based estimate of calling confidence that complements a fixed FPKM cutoff.
  • Copy-number variant regions were defined using a fixed threshold (|log2 coverage ratio FLM/FB| >= 0.3 over 100 kb windows) rather than a formal statistical test.
    Could also: A dedicated CNV-calling method with statistical significance testing (e.g., circular binary segmentation, as in DNAcopy) could also be used — Such methods provide segment-level significance or confidence estimates in addition to a coverage-ratio threshold, which can complement the concordance check already performed against standard Illumina sequencing.
Software: BS Seeker2 · TopHat · Cufflinks/Cuffcompare · HTSeq 0.5.4p3 · Blast2GO · FastQC

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

rnaseq_raw_reads_FB
Reported
87,310,639 raw reads (Table 1, FB)
Reproduced
87,310,639 (TopHat align_summary.txt 'Input' count)
exact
rnaseq_raw_reads_FLM
Reported
68,287,049 raw reads (Table 1, FLM)
Reproduced
68,287,049 (TopHat align_summary.txt 'Input' count)
exact
rnaseq_raw_reads_FLM_aza
Reported
72,050,137 raw reads (Table 1, FLM_aza)
Reproduced
72,050,137 (TopHat align_summary.txt 'Input' count)
exact
rnaseq_raw_reads_FLM_untreated
Reported
75,043,437 raw reads (Table 1, FLM_untreated)
Reproduced
75,043,437 (TopHat align_summary.txt 'Input' count)
exact
rnaseq_unique_map_FB
Reported
84.86% uniquely mapped (74,090,987/87,310,639, Table 1)
Reproduced
72.95% ((Mapped 80,251,496 - multi-alignments 16,555,713) / Input 87,310,639 = 63,695,783/87,310,639)
did not match
rnaseq_unique_map_FLM
Reported
93.36% uniquely mapped (63,751,666/68,287,049, Table 1)
Reproduced
93.08% ((Mapped 64,479,091 - multi-alignments 920,059) / Input 68,287,049 = 63,559,032/68,287,049)
exact
rnaseq_unique_map_FLM_aza
Reported
84.17% uniquely mapped (60,644,746/72,050,137, Table 1)
Reproduced
83.80% ((Mapped 61,263,197 - multi-alignments 885,592) / Input 72,050,137 = 60,377,605/72,050,137)
exact
rnaseq_unique_map_FLM_untreated
Reported
82.76% uniquely mapped (62,107,389/75,043,437, Table 1)
Reproduced
81.70% ((Mapped 63,656,167 - multi-alignments 2,346,693) / Input 75,043,437 = 61,309,474/75,043,437)
within tolerance
wgbs_meth_CG_ECM
Reported
6.4% CG methylation, genome-wide (Table 3, ECM)
Reproduced
22.08% (weighted mC/total from BS-Seeker2 CGmap, CG context)
did not match
wgbs_meth_CHG_ECM
Reported
3.4% CHG methylation, genome-wide (Table 3, ECM)
Reproduced
11.93% (weighted mC/total from BS-Seeker2 CGmap, CHG context)
did not match
wgbs_meth_CHH_ECM
Reported
3.3% CHH methylation, genome-wide (Table 3, ECM)
Reproduced
9.46% (weighted mC/total from BS-Seeker2 CGmap, CHH context)
did not match
wgbs_meth_CG_FB
Reported
28.9% CG methylation, genome-wide (Table 3, FB)
Reproduced
32.83% (weighted mC/total from BS-Seeker2 CGmap, CG context)
within tolerance
wgbs_meth_CHG_FB
Reported
8.9% CHG methylation, genome-wide (Table 3, FB)
Reproduced
19.85% (weighted mC/total from BS-Seeker2 CGmap, CHG context)
did not match
wgbs_meth_CHH_FB
Reported
10.3% CHH methylation, genome-wide (Table 3, FB)
Reproduced
23.21% (weighted mC/total from BS-Seeker2 CGmap, CHH context)
did not match
wgbs_meth_CG_FLM
Reported
30.3% CG methylation, genome-wide (Table 3, FLM)
Reproduced
26.53% (weighted mC/total from BS-Seeker2 CGmap, CG context)
within tolerance
wgbs_meth_CHG_FLM
Reported
8.1% CHG methylation, genome-wide (Table 3, FLM)
Reproduced
10.75% (weighted mC/total from BS-Seeker2 CGmap, CHG context)
partial
wgbs_meth_CHH_FLM
Reported
10.1% CHH methylation, genome-wide (Table 3, FLM)
Reproduced
13.91% (weighted mC/total from BS-Seeker2 CGmap, CHH context)
partial
wgbs_meth_CG_FLM_aza
Reported
26.4% CG methylation, genome-wide (Table 3, FLM_aza)
Reproduced
23.74% (weighted mC/total from BS-Seeker2 CGmap, CG context)
within tolerance
wgbs_meth_CHG_FLM_aza
Reported
7.3% CHG methylation, genome-wide (Table 3, FLM_aza)
Reproduced
8.09% (weighted mC/total from BS-Seeker2 CGmap, CHG context)
within tolerance
wgbs_meth_CHH_FLM_aza
Reported
8.8% CHH methylation, genome-wide (Table 3, FLM_aza)
Reproduced
9.39% (weightedmC/total from BS-Seeker2 CGmap, CHH context)
within tolerance
wgbs_meth_CG_FLM_untreated
Reported
25.3% CG methylation, genome-wide (Table 3, FLM_untreated)
Reproduced
24.99% (weighted mC/total from BS-Seeker2 CGmap, CG context)
exact
wgbs_meth_CHG_FLM_untreated
Reported
8.3% CHG methylation, genome-wide (Table 3, FLM_untreated)
Reproduced
10.59% (weighted mC/total from BS-Seeker2 CGmap, CHG context)
partial
wgbs_meth_CHH_FLM_untreated
Reported
9.9% CHH methylation, genome-wide (Table 3, FLM_untreated)
Reproduced
12.55% (weighted mC/total from BS-Seeker2 CGmap, CHH context)
partial
cri_rip_CG_TG
Reported
Qualitative only (paper reports no numeric CRI table in main text/available excerpt): T. melanosporum relies on reversible DNA-methylation-based silencing 'rather than an irreversible mechanism (such as RIP)' (Conclusions), implicitly contrasted with N. crassa's strong RIP.
Reproduced
critool CG->TG CRI: T.melanosporum_v1 mean=-0.452 median=-0.383 (n=6634), vs N.crassa mean=-1.156 median=-0.974 (n=626); U.reesii mean=0.829, A.nidulans mean=0.407, S.cerevisiae mean=0.123
partial
cri_rip_CA_TA
Reported
Same qualitative claim as above (no numeric CA->TA CRI value given in paper text).
Reproduced
critool CA->TA CRI: T.melanosporum_v1 mean=0.312 median=0.519 (n=6936), vs N.crassa mean=0.991 median=1.064 (n=655); U.reesii mean=-0.710, A.nidulans mean=-0.161, S.cerevisiae mean=-0.275
partial
cufflinks_novel_genes_FB
Reported
282 novel genes (Table 4, FB)
Reproduced
673 novel 'u'-class loci passing FPKM_conf_lo>=4 filter (4453 raw novel loci before FPKM filtering)
partial
cufflinks_novel_genes_FLM
Reported
328 novel genes (Table 4, FLM)
Reproduced
1001 novel 'u'-class loci passing FPKM_conf_lo>=4 filter (5044 raw novel loci before FPKM filtering)
partial
cufflinks_novel_genes_FLM_aza
Reported
466 novel genes (Table 4, FLM_aza)
Reproduced
1090 novel 'u'-class loci passing FPKM_conf_lo>=4 filter (6129 raw novel loci before FPKM filtering)
partial
cufflinks_novel_genes_FLM_untreated
Reported
440 novel genes (Table 4, FLM_untreated)
Reproduced
1074 novel 'u'-class loci passing FPKM_conf_lo>=4 filter (5759 raw novel loci before FPKM filtering)
partial
deseq2_sig_genes_aza_vs_untreated
Reported
115 significant genes (adjusted p<0.05) upon 5-aza treatment (Table 5); 53 hypothetical, 62 functionally classified
Reproduced
139 significant genes (padj<0.05) out of 7353 genes tested; DESeq2 v1.50.2, design=~1 dispersion estimation (geneEst+fit) then design switch to ~condition + Wald test -- workaround for modern DESeq2's hard block on naive no-replicate designs, approximating the paper's original no-replicate DE method (exact original tool/version not fully specified)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.