A comparative gene co-expression analysis using self-organizing maps on two congener filmy ferns identifies specific desiccation tolerance mechanisms associated
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction, honest 1:1 where checkable. The paper's own method scripts (eostria/Gavel @ 866a3d0; a PCA-SOM(kohonen)+WGCNA/igraph recipe carried over from an earlier Gevuina study) were transcribed and applied to the paper's own public GEO matrices (GSE140234 H. dentatum, GSE140238 H. caudiculatum). DESCRIBED-WELL-ENOUGH for the upstream/structural results, which reproduce cleanly: assembly sizes match EXACTLY (HC 34,726; HD 69,599 = matrix rows), SwissProt annotation rate ~50% (HC 54.3%, HD 49.1%), the upper-50% CV filter retains ~half the genes, and the 2x3 Kohonen SOM yields 6 nodes/species with the reported FH-dominant vs DH/RH-enriched pattern. The downstream co-expression network is NOT reproducible 1:1: (C1) the reported 67/183-gene network is a small CURATED subset with no deposited list/script — the described selection yields ~10,773/~5,993 annotated accumulation genes; (C2) with only n=3 hydration states (no replicates in the deposited matrix) WGCNA's scale-free fit is degenerate and the reported beta=9 is not data-justifiable; (C3) fastgreedy does give 2 modules at matched scale; (C4) the reported '12 hub genes with >200 connections' in a 183-gene network is MATHEMATICALLY IMPOSSIBLE (max degree 182) and our reproduction finds 0 such nodes — flagged as a probable metric/reporting error for human review. NOT ATTEMPTED (out of scope): wet-lab physiology (RWC, Fv/Fm, gas exchange), de-novo Trinity re-assembly from raw reads (we verified against the deposited assembly/matrix instead), GO/functional-enrichment narrative, and DE gene lists. Code is a valid P16 third-party/own-method reproduction. Verdict provisional — human auditor decides.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 62assessed: 2026-06-16 ⛓ ce37ba14e3be
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusAlthough desiccation tolerance in filmy ferns has been proposed to rely mainly on constitutive features rather than dehydration-induced responses, the authors hypothesize that inter-specific differences in vertical microhabitat distribution between two Hymenophyllum species are associated with different dynamics of gene expression during the dehydration and rehydration phases.
- ★ H. caudiculatum (lower canopy) exhibits about twice as many differentially expressed genes as H. dentatum (upper canopy), with a higher proportion of both increased and decreased gene abundance during dehydration. finding
- ★ In H. dentatum, gene abundance decreases significantly when transitioning from dehydration to rehydration. finding
- ★ H. caudiculatum enhances osmotic responses and phenylpropanoid-related pathways, whereas H. dentatum enhances defense system responses and protection against high light stress, reflecting their microhabitat preferences. mechanism
- ★ Desiccation tolerance in these two filmy ferns broadly involves three constitutively highly abundant processes: translation, photosynthesis, and antioxidant activity. finding
- ★ A comparative transcriptomic approach combining WGCNA with self-organizing maps (artificial neural networks) identifies species-specific desiccation tolerance mechanisms. method
- ★ De novo transcriptome assemblies provide reference resources of 34,726 transcripts for H. caudiculatum and 69,599 transcripts for H. dentatum. resource
- Top BLAST hits of both transcriptomes are enriched for the moss Physcomitrella patens and the lycophyte Selaginella moellendorffii, consistent with poikilohydry and the regressive evolution hypothesis. finding
- At full hydration, H. caudiculatum co-expression network contains twelve hub genes (>200 connections) involved in oxidative stress protection, photosystem light harvesting, and lipid metabolism/transport, while H. dentatum networks contain no hub genes. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (de novo transcriptome assembly and differential expression) | Hymenophyllum caudiculatum and Hymenophyllum dentatum fronds | experimental desiccation-rehydration cycle (cessation/reestablishment of irrigation) | transcript abundance (FPKM), differential gene expression across full-hydrated, dehydrated, and rehydrated states | Illumina HiSeq2000; Trinity assembler; SwissProt annotation |
| relative water content measurement | Hymenophyllum caudiculatum and Hymenophyllum dentatum fronds | desiccation-rehydration cycle | relative water content (RWC, %) | — |
| chlorophyll fluorescence (PAM) and imaging fluorescence | dark-adapted detached fronds of H. caudiculatum and H. dentatum | desiccation-rehydration cycle | maximum quantum efficiency Fv/Fm and Y(II) fluorescence images | — |
| Gene Ontology enrichment analysis | annotated differentially expressed genes of both species | none | GO category distribution (BP, CC, MF) | — |
| self-organizing maps (SOM) clustering / artificial neural network | differentially expressed genes of both species | none | gene clustering into six nodes by accumulation pattern per hydration state | — |
| weighted gene co-expression network analysis (WGCNA) | SOM-selected node genes of both species | none | network modules, gene connectivity, hub genes (Fast Greedy modularity) | — |
- ▲ H. caudiculatum had the highest number of DE genes during dehydration, 265 total, with 139 increased in abundance 265 DE genes (139 up)
- ▼ During 3-25 h without irrigation, H. dentatum lost water faster, reaching 18% RWC vs 30% RWC in H. caudiculatum 18% vs 30% RWC
- ▼ Fv/Fm dropped from ~0.7 to ~0.2 in H. dentatum during dehydration but remained near 0.78 in H. caudiculatum 0.7 to 0.2 (H. dentatum); ~0.78 maintained (H. caudiculatum)
- – Constitutively highly abundant transcripts (fold change <2): 102 in H. caudiculatum and 128 in H. dentatum, spanning translation, photosynthesis, and antioxidant activity 102 and 128 transcripts
- – H. caudiculatum full-hydrated network had highest connectivity with 6216 connections; H. dentatum dehydrated network had 866 connections 6216 vs 866 connections
- ▼ Both species reached ca. 60% RWC during first three hours after cessation of irrigation ~60% RWC
- – Approximately 80% (H. caudiculatum) and 70% (H. dentatum) of high-quality reads mapped to the respective transcriptomes ~80% and ~70%
- – H. caudiculatum desiccation network had six hub genes (>100 connections); rehydration network had eighteen equally connected genes 6 hub genes; 18 genes
- count 34,726 contigs (H. caudiculatum) and 69,599 contigs (H. dentatum) final transcriptomes (refined de novo transcriptome assemblies)
- count 111,495,169 (H. caudiculatum) and 110,988,488 (H. dentatum) paired-end reads (101 bp) (sequencing output)
- count 161,689 contigs (H. caudiculatum) and 332,003 contigs (H. dentatum) (initial assemblies before refinement)
- fold_change fold change >=2 and FDR <0.05 (cutoff for differential expression analysis)
- count 265 DE genes, 139 up (H. caudiculatum during dehydration)
- count 6216 and 866 connections (network connectivity FH H. caudiculatum vs DH H. dentatum)
- count 102 and 128 transcripts (constitutive highly abundant genes per species)
- other ~50% alignment rate to SwissProt (annotation rate of transcripts)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study performed de novo transcriptome assembly (Trinity) from RNA-seq data collected from two filmy fern species across three hydration states (full hydration, dehydration, rehydration). Differential expression was identified by pairwise comparisons using fold-change ≥2 and FDR < 0.05 thresholds, with the specific DE software not named. Co-expression structure was then explored by Self-Organizing Maps (SOM) for clustering, followed by Weighted Gene Co-expression Network Analysis (WGCNA) with Fast Greedy modularity optimization, and GO enrichment was applied to resulting node transcripts.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression — pairwise comparisons with fold-change ≥2 and FDR < 0.05 cutoff; specific software/test not named | All pairwise hydration-state comparisons (FH vs DH, FH vs RH, DH vs RH) in both H. caudiculatum and H. dentatum | — | not stated |
| GO enrichment analysis; method not specified | Annotated DE transcripts and transcripts within selected SOM nodes for both species | — | not stated |
| Self-organizing map (SOM) clustering — 2×3 hexagonal topology, Euclidean distance | DE gene expression patterns partitioned into six nodes across hydration states for both species | — | na |
| Weighted Gene Co-expression Network Analysis (WGCNA) with Fast Greedy modularity optimization for community structure | Co-expression network construction from genes in selected SOM nodes per hydration state per species | — | not stated |
-
The differential expression software is not named and the specific FDR procedure is unspecified↳ Could also: Naming the DE tool (e.g., DESeq2, edgeR, limma-voom) and the FDR method (e.g., Benjamini-Hochberg) would be an alternative reporting practice — Different DE tools make different distributional assumptions (negative binomial, empirical Bayes, etc.); explicit identification enables readers to assess those assumptions, assess reproducibility, and compare results across studies
-
The number of biological replicates is not reported in the text; the design involves fronds sampled at multiple hydration states↳ Could also: Explicitly reporting the number of independent biological replicates (individual plants) and distinguishing them from technical replicates is also standard practice in transcriptomic studies — Biological replication is the primary basis for variance estimation in DE inference; its ambiguity limits the reader's ability to assess the statistical power underlying FDR-controlled calls
-
A hard fold-change threshold (≥2) was applied alongside FDR < 0.05 to select DE genes↳ Could also: Shrinkage-based log fold-change estimation (e.g., DESeq2 lfcShrink or the ashr prior) without a hard FC cutoff would also be used to rank and filter genes — Hard FC thresholds can exclude biologically relevant genes with smaller but precisely estimated changes, particularly for low-abundance transcripts; shrinkage stabilizes FC estimates across the abundance range
-
SOM clustering was used to group DE genes by expression pattern across hydration states↳ Could also: Hierarchical clustering with a correlation- or Euclidean-based distance metric, or fuzzy c-means clustering (as in the Mfuzz R package, common in time-course transcriptomics), would also partition genes by expression trajectory — SOM and hierarchical/fuzzy clustering differ in their sensitivity to topology and initialization; comparing results across methods can help distinguish robust clusters from algorithm-specific groupings
-
Physiological data (RWC, Fv/Fm) across the desiccation-rehydration time course were presented as curves without stated dispersion metrics or formal between-species statistical tests↳ Could also: Reporting mean ± SD or SEM at each time point with a mixed-effects model or repeated-measures ANOVA comparing species trajectories would also characterize the physiological response — Formal dispersion metrics and between-species tests allow readers to assess whether observed differences in water-loss rates and Fv/Fm recovery are within expected biological variation for each species
-
Fast Greedy modularity optimization was used to identify community structure in the WGCNA co-expression networks↳ Could also: The native WGCNA topological overlap matrix (TOM)-based module detection, or alternative community detection algorithms such as Louvain or label propagation, would also identify co-expression modules — Different algorithms optimize different objective functions and can yield different module memberships; reporting sensitivity to algorithm choice, or comparing multiple algorithms, can support claims about the robustness of identified hub genes
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32019526
Paper: Ostria-Gallardo et al. 2020, BMC Plant Biol. "A comparative gene co-expression analysis using self-organizing maps on two congener filmy ferns identifies specific desiccation tolerance mechanisms." DOI 10.1186/s12870-019-2182-3.
Code: https://github.com/eostria/Gavel (R scripts, public, not archived,
last push 2018-05-03, commit resolved at run time). NOTE: the repo ships the
authors' method scripts from an earlier study (Gevuina avellana, light
study; the figure numbers in the filenames — "Fig. 5", "Fig. 6c" — refer to that
paper, and the hard-coded input files Gavel.light.ORF.ave.csv /
Gave.interestORF.light.csv are the Gevuina data, NOT the ferns). They are the
exact PCA-SOM (kohonen) + WGCNA/igraph co-expression recipe the ferns paper says
it used. Per brief rule P16, applying this shipped recipe to the paper's own data
is a valid reproduction. The scripts are R-console transcripts (leading >), not
turnkey scripts, so they are transcribed into a runnable pipeline preserving the
same steps/parameters.
Data (both public, "Public on Nov 19 2019"):
- GSE140234 — H. dentatum (HD). suppl:
GSE140234_Hdentatum_2019.matrix.gz(FPKM-like expression matrix) +..._2019.fasta.gz(assembly). 69,599 genes × 3 hydration states (Hdent_FH, Hdent_DH, Hdent_RH). md5 c780933ecbd784b348e9b3a8b9eba8c1. - GSE140238 — H. caudiculatum (HC). suppl:
GSE140238_Hcaudiculatum_2019.matrix.gz..._2019.fasta.gz. 34,726 genes × 3 states (Hcau_FH, Hcau_DH, Hcau_RH). md5 e299e64a09d6790e557184f9933d7335.
- Hydration states: FH = fully hydrated, DH = dehydrated, RH = rehydrated.
- «infra»: «path»
IN SCOPE (pipeline-derived, attempted)
| # | Result | Pipeline | Difficulty |
|---|---|---|---|
| A1 | HC refined assembly = 34,726 contigs | Trinity assembly → matrix row count | verified from data |
| A2 | HD refined assembly = 69,599 contigs | Trinity assembly → matrix row count | verified from data |
| A3 | ~85M reads HC, ~87M reads HD used | SRA run metadata | check SRA |
| S1 | SOM = 2×3 hexagonal = 6 nodes/species | kohonen SOM on upper-50%-CV, scaled FPKM | runnable |
| S2 | SOM node expression patterns (one node FH-dominant; others DH/RH-enriched) | SOM codebook vectors | qualitative |
| C1 | Co-expression subset = 67 (HD) / 183 (HC) annotated genes | SOM node selection + SwissProt annotation | runnable (needs BLAST) |
| C2 | Soft-threshold β = 9 | WGCNA pickSoftThreshold | runnable |
| C3 | 2 modules per network (each species) | igraph fastgreedy on TOM | runnable |
| C4 | HC: 12 hub genes (>200 conn) at FH; 6 (>100) at desiccation; HD: no hub genes | igraph degree on TOM>0.05 | runnable |
OUT OF SCOPE (wet-lab / manual / external — not attempted)
- Physiological measurements (RWC, chlorophyll fluorescence Fv/Fm, gas exchange).
- GO/functional enrichment narrative and biological interpretation of specific genes.
- De-novo Trinity assembly from raw reads (we use the authors' deposited assembly + matrix; re-assembly is a separate, non-deterministic step — the deposited 34,726/69,599-contig matrices ARE the assembly output we verify against).
- Differential-expression gene lists (FC≥2, FDR<0.05) beyond counts (no DE table with exact values pinnable from text).
Notes / caveats
- Only n=3 columns (one value per hydration state, no replicates in the deposited matrix) → WGCNA correlation across 3 points is statistically degenerate; β/scale-free fit and exact edge counts are fragile. Reported honestly.
- SOM node labels are arbitrary (depend on init/seed); we compare node patterns, not node numbers.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The upstream/structural pipeline reproduces cleanly from the authors' own deposited GEO matrices — assembly sizes match exactly (34,726/69,599), ~50% SwissProt annotation, the upper-50% CV filter, and the 2×3 SOM with the FH-dominant vs DH/RH pattern all hold. The failures are concentrated in the downstream co-expression network and sit on the authors' side: the 67/183-gene subset is an undeposited curated set (our run yields ~5,993/~10,773), beta=9 is unjustifiable at n=3, and most seriously C4's '12 hub genes >200 connections' is mathematically impossible in a 183-node graph (max degree 182) where we find 0. Severity is high for the core network claims (impossible value + orders-of-magnitude gene-count gap), even though the comparative SOM story survives qualitatively. Overall: solid core, but fabrication-suspect and non-auditable on the key hub/network result.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.