Comparing time series transcriptome data between plants using a network module finding algorithm.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Methods/protocol paper (CompareTranscriptome): per-species co-expression networks (rcorr, PCC>=0.99, p<0.001) clustered cross-species by OrthoClust (kappa=3). DESCRIBED WELL ENOUGH to reproduce the pipeline; ran it unmodified on «our HPC» from the repo's shipped processed inputs (FPKM, edge lists, RBH) + shipped OrthoClust_1.0 package (patched only rBind/cBind->rbind/cbind, defunct in Matrix>=1.3, semantics identical). RESULT = mixed/honest: (C3) the paper's reported numbers are EXACTLY the authors' shipped output files -> no fabrication. (C2) module COUNT reproduces within tolerance (independent seeds 345-358; one hits 353 exactly) but the EXACT Table 3 module sizes do NOT reproduce because OrthoClust2's only randomness (node update order, no seed pinned in the paper) lands in different local optima with larger top modules -> a published-method reproducibility limit, not a number discrepancy. (C1) co-expression network sizes reproduce in method within ~5-8% but not byte-exact because the shipped Step5 script's gene pre-filter is commented out (var>0.5 reproduces the ATH edge count 17648 exactly). NOT ATTEMPTED (hard 20%): raw FASTQ->STAR->featureCounts->FPKM, BLASTp+RBH (registration-walled genomes/proteomes; both shipped processed), Cytoscape figures, GO interpretation, kappa/PCC sweeps.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 59assessed: 2026-06-14 ⛓ d06f97c6e697
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a one-to-one mapping of developmental stages between two plant species be avoided by converting time-series gene expression data into co-expression networks and applying a network module finding algorithm (OrthoClust) to reveal conserved versus divergent gene function across species?
- ★ Converting time-series expression data into co-expression networks and applying network module finding (OrthoClust) enables cross-species comparison without requiring one-to-one developmental stage mapping. method
- ★ The pipeline integrates RBH-based homology, per-species PCC co-expression networks, and OrthoClust simulated-annealing clustering to find orthologous co-expression modules. method
- ★ The method reveals functional conservation and divergence of genes between Arabidopsis and soybean during embryo development. finding
- ★ Many soybean genes in a co-expression cluster change their expression pattern in Arabidopsis, suggesting potential functional divergence, while RBH pairs tend to share similar expression patterns. finding
- ★ The coupling constant kappa controls the relative weight of homology vs co-expression edges; kappa>0 includes homologous edges and increases the number of cross-species modules. mechanism
- Higher PCC threshold yields fewer co-expression edges, a less connected network, and therefore more modules. mechanism
- ★ A publicly accessible pipeline with bash/Python/R scripts and example data is provided via GitHub. resource
- OrthoClust output can be visualized as networks in Cytoscape and as expression profile plots. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Reciprocal best hit (RBH) homology identification via protein BLAST | Soybean (Glycine max) and Arabidopsis thaliana proteomes | none | Number of RBH orthologous gene pairs | BLAST; custom Python script (CompareTranscriptome GitHub) |
| Time-series RNA-seq co-expression network construction (PCC-based) | Arabidopsis seed embryo (developmental time course) | none | Co-expression edges and genes passing PCC/p-value filtering | — |
| Time-series RNA-seq co-expression network construction (PCC-based) | Soybean seed embryo (developmental time course) | none | Co-expression edges and genes passing filtering | — |
| Cross-species module finding (OrthoClust, simulated annealing) | Combined Arabidopsis + soybean embryo co-expression networks with RBH orthologous pairs | none | Number and composition of cross-species modules | OrthoClust in R |
| Network visualization | Module 8 genes from Arabidopsis and soybean | none | Network layout of nodes/edges/homology pairs | Cytoscape |
| Parameter sensitivity analysis | Combined Arabidopsis + soybean OrthoClust networks | varying PCC threshold and kappa | Number of modules | OrthoClust |
- – 13,024 RBH ortholog pairs identified between soybean and Arabidopsis 13,024 pairs
- – Arabidopsis embryo co-expression network: 17,648 edges among 853 genes after filtering (from 1267 selected, 1092 retained, 595,686 initial edges) 17,648 edges / 853 genes
- – Soybean embryo co-expression network: 62,185 edges among 1401 genes 62,185 edges / 1401 genes
- – One OrthoClust trial yielded 353 cross-species modules; some balanced between species, others single-species 353 modules
- ▼ RBH gene AT5G52560 (raffinose biosynthesis) and soybean RBH Glyma.04G245100 share a similar decreasing expression pattern across embryo development
- ▲ Number of modules at kappa=1 is two to three times that at kappa=0; increasing kappa to 2 or 3 further increases modules but not dramatically 2-3x
- ▲ For the same kappa, higher PCC threshold always produces more modules
- – Top module (ID 2) contained 352 genes: 273 soybean (77.6%) and 79 Arabidopsis (22.4%) 352 genes
- count 13,024 (RBH ortholog pairs in each species)
- count 595,686 (Initial Arabidopsis co-expression edges from PCC step)
- count 17,648 edges among 853 genes (Arabidopsis embryo co-expression network after filtering)
- count 62,185 edges among 1401 genes (Soybean embryo co-expression network)
- count 353 (Modules from one OrthoClust trial)
- count 352 (273 soybean / 79 Arabidopsis) (Largest module (Module ID 2) gene composition)
- fold_change two to three times (Modules at kappa=1 vs kappa=0)
- count 48,375 proteins (56,044 gene models) soybean; 24,148 proteins (37,336 gene models) Arabidopsis (Proteome/gene model totals used for BLAST)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a methods and protocol paper presenting a computational pipeline for comparative transcriptome analysis between two plant species (Arabidopsis and soybean embryo time-series data). The primary statistical tool is the Pearson Correlation Coefficient (PCC), applied pairwise across genes to construct per-species co-expression networks; PCC values and their associated p-values serve as filters to retain high-confidence edges. Cross-species network modules are then identified using OrthoClust, a simulated annealing optimization algorithm that jointly minimizes a cost function over co-expression edges and orthologous gene-pair edges weighted by the coupling constant κ. Results are reported descriptively as counts of genes, edges, and modules under varied parameter settings, with no formal inferential statistics on biological outcomes.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson Correlation Coefficient (PCC) with associated p-value thresholding | Pairwise gene co-expression quantification for Arabidopsis and soybean networks; edges retained at PCC ≥ 0.99 (Table 3 caption) and passing a p-value cutoff | — | not stated |
| Simulated annealing optimization (OrthoClust module finding) | Cross-species network module identification integrating co-expression edges and orthologous gene pairs from both species | — | na |
-
Pairwise PCC p-values were used as filters across hundreds of thousands of gene pairs without a stated multiple-testing correction↳ Could also: A Benjamini-Hochberg FDR correction (or permutation-based FDR) could also be applied to the full family of pairwise correlation tests — With ~595,686 pairwise tests in the Arabidopsis network alone, the expected number of false-positive edges under a nominal p-value threshold can be substantial; FDR control is a standard approach in large-scale correlation screens to calibrate the proportion of spurious edges retained
-
Pearson Correlation Coefficient was used to quantify co-expression between gene pairs across time points↳ Could also: Spearman rank correlation or mutual information (MI) could also be used to construct the co-expression network — Spearman correlation is robust to non-normal expression distributions and outlier time points; MI captures non-linear dependencies; either is commonly used in co-expression network construction and may yield a different edge set, particularly in small time-series experiments where distributional assumptions of PCC may not hold
-
Reciprocal Best BLAST Hit (RBH) was used as the primary orthology criterion in the main demonstration↳ Could also: Graph-based orthology tools such as OrthoFinder, OMA, or PLAZA (noted in the Discussion) could also define the ortholog edge set — RBH is conservative and misses many-to-many relationships; graph-based methods model gene family evolution explicitly and may recover a broader or more biologically representative ortholog set, potentially altering module membership and cross-species connectivity
-
OrthoClust with simulated annealing was used to find cross-species co-expression modules↳ Could also: Weighted Gene Co-expression Network Analysis (WGCNA) applied per species followed by module-preservation statistics could also identify and assess conserved co-expression modules — WGCNA provides well-established module-preservation Z-scores and permutation-based p-values for formally quantifying whether a module found in one species is recapitulated in another, offering a complementary inferential layer beyond cluster enumeration
-
Parameter sensitivity (PCC threshold and κ) was explored by reporting module counts under different settings (Fig. 5)↳ Could also: A stability analysis such as bootstrap resampling of genes or time points, or a consensus clustering approach, could also characterize how robustly module memberships are recovered across perturbations — Reporting module counts under varying thresholds describes threshold sensitivity but not whether the same gene memberships recur; resampling-based stability indices would provide complementary evidence that identified modules are reproducible features of the data rather than artifacts of a specific parameter choice
-
Expression profile similarity between homologous gene pairs was assessed qualitatively by visual inspection of time-course plots (Fig. 4)↳ Could also: A quantitative similarity score such as dynamic time warping distance, Euclidean distance between normalized time-course vectors, or the correlation of RBH pair trajectories could also be computed across all homologous pairs — Quantitative trajectory similarity measures would provide an objective, distributional summary of expression conservation across the full set of homologous pairs, enabling ranked comparisons and threshold-based classification of conserved versus diverged pairs beyond individual illustrative examples
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-31164912
Paper: Lee J, Heath LS, Grene R, Li S. Comparing time series transcriptome data between plants using a network module finding algorithm. Plant Methods 2019. DOI 10.1186/s13007-019-0440-x · PMCID PMC6544932.
Repo: https://github.com/LiLabAtVT/CompareTranscriptome (commit
487b170a2dc2d6a8713d09d9ae0a4f5186945a62). This is a methods/protocol paper:
the contribution is a workflow (= MIMB book-chapter protocol) for cross-species
comparison of time-series transcriptomes, using the OrthoClust module-finding
algorithm. The repo ships the processed inputs (FPKM matrices, co-expression edge
lists, RBH ortholog pairs), the OrthoClust R package (software/OrthoClust_1.0.tar.gz),
all scripts, and the authors' own final OrthoClust output
(Orthoclust_Results_GmxAth_...kappa3.csv/.RData).
Data: Arabidopsis thaliana embryo time series (PRJNA301162 / GSE74692, 7 dpp timepoints) and Glycine max (soybean) embryo time series (PRJNA197379, 10 timepoints). FPKM matrices shipped: ATH 1267 genes × 7, GMX 2092 genes × 10.
In scope (pipeline-derived, reproduced here)
| ID | Result | Pipeline | Determinism |
|---|---|---|---|
| C1 | Co-expression network size (edges/nodes) per species | Step5 FPKM2NETWORK.R: Hmisc rcorr Pearson on FPKM → filter PCC≥0.99 & p<0.001 |
deterministic |
| C2 | OrthoClust modules: total count (353) + top-10 module sizes & species ratios (Table 3) | Step1 OrthoClust.R: OrthoClust2(GMX,ATH,RBH,kappa=3) |
stochastic (random node-update order sample.int(N); seed-controlled) |
| C3 | Shipped result vs paper Table 3 / shipped edge lists vs reported network sizes (anti-fabrication audit) | direct comparison of shipped files to the paper | n/a (file inspection) |
Out of scope (not attempted — 80/20, heavy raw-seq, or external)
- Raw FASTQ download (PRJNA301162/197379) + STAR genome index + STAR mapping + featureCounts → counts → FPKM. Heavy raw-seq processing; the paper's contribution is the comparison method, and FPKM matrices are shipped. Skipped (the hard 20%).
- BLASTp all-vs-all (24k×48k proteins) + RBH extraction. Needs Araport11 + Gmax_275 proteomes (registration-walled at araport.org / phytozome). RBH subset is shipped. Skipped (external data + compute).
- Cytoscape visualization (Figs 3–4), GO/pathway functional interpretation (raffinose pathway, AT5G52560↔Glyma.04G245100): manual/external. Out of scope.
- κ-sweep (Fig 5, κ=0,1,2,3) and PCC-threshold sweep: optional secondary analysis, not attempted beyond κ=3.
Notes on reproducibility
OrthoClust2is a greedy Potts/modularity spin optimizer; the only nondeterminism is the initialseq=sample.int(N)update order. Module count and exact membership therefore vary run-to-run; we report several seeds and treat C2 as a partial/within-tolerance reproduction, not byte-exact.- The package uses removed APIs (
rBind/cBind,graph.edgelist,get.adjacency), so it is run unmodified on R 3.6 + Matrix <1.3 + igraph <2.0.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All reported numbers are exactly present in the authors' shipped outputs (no fabrication, q5 green), and the headline result — ~353 cross-species modules at kappa=3 — reproduces independently (one seed hits 353 exactly; networks reproduce in-method within ~5–8%). Two explainable deviations remain, both benign and on the method/artifact side: OrthoClust2 pins no RNG seed so the exact Table 3 module sizes are one stochastic realization (352→465 etc.), and the shipped Step5 script ships its gene pre-filter commented out (var>0.5 reproduces 17648 ATH edges exactly), so byte-exact network sizes aren't recoverable. Severity is moderate (magnitude/direction and module count preserved), so overall a yellow: a faithful, well-documented reproduction with stochastic- and code-config deviations, not a substantive discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.