Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Analysis of a photosynthetic cyanobacterium rich in internal membrane systems via gradient profiling by sequencing (Grad-seq).

Plant Cell · 2021
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce: YES — 1:1 EXACT reproduction. The authors' GitHub repo (commit 8aedf73) is a self-contained, end-to-end reproducible bundle: it ships the raw UMI-deduplicated count tables + spike-in factors, the R normalization + WGCNA clustering scripts, the WGCNA input matrices, AND the expected clustering output. We rebuilt an R 4.3.3 + WGCNA 1.73 conda env on a «our HPC» compute node (SLURM «job», n093) and re-ran the full normalization + WGCNA pipeline via faithful driver scripts (computation byte-for-byte from the authors' code; only Windows setwd() dropped, interactive plotting removed, commented-out write.csv re-enabled, and one content-preserving fix: spikeIns.csv read with fileEncoding='latin1' because its header contains 'ng/uL'). Results: C1 2,394 proteins EXACT; C2 4,251 transcripts EXACT; C3 17 modules EXACT; C4 6,645-feature partition reproduces with 100% exact module-label agreement and Adjusted Rand Index = 1.000 (perfectly diagonal contingency), reported module sizes all present; C5 normalized matrices reproduce to floating-point precision (max abs diff 5e-08, corr 1.0) for both RNAseq and proteomics (proteomics required column-name alignment because dcast emits fraction columns alphabetically vs the shipped numeric order — per-cell values are identical, normalization is column-order invariant). NOT attempted (out of scope): re-alignment from raw SRA reads (PRJNA608723; count tables are the provided starting point, deposit profiled separately and grades A), MaxQuant MS database search (wet-lab/proprietary upstream), SVM RBP classification (no runnable training script shipped). Large outputs kept on «infra» with SHA256 manifest; small results pulled to reproduction/outputs/.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 42c55b8d4a9b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether gradient profiling by sequencing (Grad-seq), applied for the first time to a photosynthetic cyanobacterium (Synechocystis sp. PCC 6803), can resolve RNA-protein complexes and reveal candidate RNA chaperones and RNA-binding proteins despite the absence of known regulatory RNA chaperones (e.g., functional ProQ, CsrA, or Hfq homologs) in cyanobacteria.

Core claims
  • Grad-seq resolves complexes with overlapping subunits, such as CpcG1-type versus CpcL-type phycobilisomes or PsaK1 versus PsaK2 photosystem I (pre)complexes, validating the approach. finding
  • Clustering of in-gradient distribution profiles yielded a short list of candidate RNA chaperones including a YlxR homolog and a cyanobacterial homolog of the KhpA/B complex. finding
  • The data suggest previously undetected complexes between accessory proteins and CRISPR-Cas systems, such as a Csx1-Csm6 ribonucleolytic defense complex. finding
  • RpoZ or 6S RNA associate exclusively with the core RNA polymerase complex, and a reservoir of inactive sigma-antisigma complexes is suggested. finding
  • A publicly available Synechocystis Grad-seq database/resource (GradSeqExplorer) is provided for functional assignment of RNA-protein and multisubunit protein complexes. resource
  • Hierarchical clustering (WGCNA dynamic tree cut) splits the dataset into a protein-dominated branch (Clusters 1-9) and an RNA-dominated branch (Clusters 10-17). finding
  • Most Synechocystis transcripts show bimodal sedimentation profiles, suggesting they occupy either a large-complex-bound state or an unbound/chaperone-protected state. finding
  • RNA sedimentation profiles are largely independent of transcript length but depend on associated binding protein(s). mechanism
Experimental setups
Assay System Perturbation Readout Platform
Grad-seq (sucrose density gradient ultracentrifugation coupled to MS proteomics and RNA-seq) Synechocystis sp. PCC 6803, whole-cell lysates, triplicate cultures none sedimentation/in-gradient distribution profiles of proteins and RNAs mass spectrometry and RNA deep sequencing
Denaturing urea-polyacrylamide gel electrophoresis RNA from gradient fractions of Synechocystis lysate none size separation/visualization of fractionated RNA species 10% urea-PAGE with low range ssRNA ladder (NEB)
SDS-polyacrylamide gel electrophoresis Proteins from gradient fractions of Synechocystis lysate none size separation/visualization of fractionated proteins 15% SDS-PAGE with PageRuler ladder (Thermo Fisher)
Hierarchical clustering (WGCNA dynamic tree cut algorithm) Grad-seq protein and RNA abundance dataset from Synechocystis none cluster assignment of proteins/RNAs by sedimentation profile similarity WGCNA R package
Comparative phylogenomics/ortholog detection (domclust algorithm) 57-59 cyanobacterial genomes, Arabidopsis thaliana, E. coli, Salmonella enterica proteomes none presence/absence of protein homologs, degree of synteny Microbial Genome Database (domclust)
RNA-binding protein prediction Cosedimenting proteins from Grad-seq dataset none SVM-based RBP likelihood score RNApred
Key results
  • Strong correlation between triplicate gradient replicates for RNA-seq and MS profiles median RNA-seq R=0.76; median MS R=0.85
  • 2,394 proteins detected in all three replicates, corresponding to 67.3% of 3,559 annotated protein-coding genes 67.3%
  • 4,251 distinct transcripts detected, corresponding to 2,544 of 4,091 previously defined transcriptional units 62.2%
  • Dataset split into 17 clusters: protein-dominated branch (Clusters 1-9) contains ~80% of detected proteins; RNA-dominated branch (Clusters 10-17) contains ~90% of detected transcripts 80% / 90%
  • Only a minority of the proteome localizes to RNA-dominated (large complex) clusters 20%
  • 53 ribosomal or ribosome-associated proteins and their rRNAs detected, with the complete ribosomal component set concentrated mainly in one fraction Fraction 18
  • Several ribosomal proteins (Rpl11, Rpl12, Rpl16, Rps2, Rps10) and Ycf65 diverge from the main ribosome sedimentation pattern, suggesting alternative associations
Key statistics
  • correlation median R=0.76 (RNA-seq reproducibility between gradient replicates)
  • correlation median R=0.85 (MS proteome reproducibility between gradient replicates)
  • count 2,394 proteins (proteins detected in all three replicates)
  • count 4,251 transcripts (distinct transcripts detected in all three replicates)
  • other 67.3% (percentage of 3,559 annotated protein-coding genes detected)
  • other 62.2% (percentage of 4,091 transcriptional units detected)
  • count 17 clusters (clusters from WGCNA dynamic tree cut of proteins and RNAs)
  • count 53 ribosomal/ribosome-associated proteins (detected in highest-sedimentation-coefficient fractions)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a discovery-oriented profiling study in which triplicate Synechocystis 6803 cultures were lysed and fractionated by sucrose density gradient ultracentrifugation; each fraction was then characterized by both mass spectrometry (proteomics) and RNA-seq (transcriptomics). Reproducibility across biological triplicates was assessed with Spearman correlation coefficients. Co-sedimentation profiles of all detected proteins and transcripts were then clustered using the WGCNA dynamic tree cut algorithm, yielding 17 clusters. Downstream analyses included z-score standardization for heatmap visualization, SVM-based prediction of RNA-binding proteins via RNApred, and a phylogenetic conservation screen across 57+ genomes using the domclust algorithm; no formal null-hypothesis significance tests comparing experimental groups are reported in the available text.

Replicationbiological Sample sizeThree independent triplicate cultures prepared and fractionated identically; no formal power calculation mentioned in the available text GroupsSingle growth condition (moderate light 50 µE, BG11 medium); no between-group experimental comparison — primary aim is whole-complexome profiling Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Spearman correlation coefficient (profile-to-profile, across replicates) Reproducibility of in-gradient RNA-seq and MS sedimentation profiles between the three biological replicates (Supplemental Figure 1) 3 biological replicates; 4,251 transcripts and 2,394 proteins not stated
WGCNA dynamic tree cut hierarchical clustering Assignment of all detected proteins and transcripts to 17 sedimentation-profile clusters (Figure 2A) 4,251 transcripts + 2,394 proteins across 18 gradient fractions not stated
Support vector machine (SVM) classification score via RNApred Prediction of RNA-binding potential among co-sedimenting proteins not stated
Z-score standardization (not a significance test; normalization for visualization) Heatmap representation of relative protein and RNA abundances across gradient fractions (Figure 4) na
Ortholog presence/absence classification via domclust algorithm Phylogenetic conservation screen of detected proteins across 57 selected genomes (Figure 2C, Supplemental Table 1) 57 genomes (56 cyanobacterial + E. coli + S. enterica + Arabidopsis) not stated
Approaches that could also have been used
  • Reproducibility between replicates was summarized with Spearman correlation coefficients of the per-fraction abundance vectors
    Could also: Principal component analysis (PCA) or sample-to-sample Euclidean distance on the full profile matrix could also summarize replicate concordance globally, and intraclass correlation coefficient (ICC) would additionally quantify the proportion of variance attributable to biological versus technical sources — PCA provides a visual overview of all sources of variation simultaneously, while ICC gives a single interpretable reliability estimate useful when gradient profiles are later used as input to downstream models
  • Sedimentation profiles were grouped with the WGCNA dynamic tree cut algorithm applied to hierarchical clustering
    Could also: Gaussian mixture modelling, k-medoids (PAM), or self-organizing maps could also partition profile shapes; cluster stability could additionally be assessed with bootstrap resampling (e.g., via the clusterboot function in R) — These alternatives provide probabilistic cluster membership, allow formal comparison of candidate cluster numbers, or yield stability metrics that help judge how reproducibly a profile belongs to a given cluster
  • RNA-binding protein candidates were ranked using an SVM score from RNApred trained on known RBPs
    Could also: Random forest classifiers, gradient boosting (e.g., XGBoost), or logistic regression with regularization trained on the same feature sets could also score RBP likelihood; an ensemble of multiple classifiers could further improve calibration — Ensemble methods often outperform single SVMs on imbalanced binary classification tasks and provide feature-importance rankings that can help interpret which sequence or structural properties drive the prediction
  • Phylogenetic conservation was assessed as binary presence/absence of orthologs across 57 genomes
    Could also: Quantitative conservation scores (e.g., normalized BLAST bit-score, percent identity, or dN/dS ratio) or a continuous phylogenetic profiling distance metric could also be computed, and enrichment of conserved proteins within each cluster could be tested with a Fisher's exact test or hypergeometric test — Continuous scores preserve information about the degree of conservation and allow statistical comparison across clusters, whereas binary presence/absence conflates highly diverged orthologs with close homologs
  • Cluster-level functional enrichment was visualized with KEGG pie charts per cluster
    Could also: Formal overrepresentation analysis (ORA) with a hypergeometric test or gene set enrichment analysis (GSEA) using the gradient profile as a ranked metric could also be applied, with Benjamini–Hochberg FDR correction across GO/KEGG categories — Statistical enrichment testing provides a principled way to distinguish categories that are genuinely over-represented in a cluster from those appearing by chance in the pie-chart visualization
  • The study used a single growth condition with three biological replicates and no control group for cross-condition comparison
    Could also: A paired two-condition design (e.g., standard light vs. high light or iron-replete vs. iron-depleted), with differential sedimentation tested via linear models (limma) or edgeR on the per-fraction count data, could also quantify condition-dependent shifts in complex formation — Comparative Grad-seq across conditions allows formal statistical identification of proteins or RNAs that change their sedimentation — and hence complex association state — in response to a perturbation, extending the purely descriptive complexome map to mechanistic inference
Software: R / WGCNA (Weighted Correlation Network Analysis) · RNApred (SVM-based RBP predictor) · Microbial Genome Database / domclust (ortholog clustering)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33793824

Paper: Riediger M et al. (2021) Analysis of a photosynthetic cyanobacterium rich in internal membrane systems via gradient profiling by sequencing (Grad-seq). Plant Cell 33(2):248–269. DOI 10.1093/plcell/koaa017. PMCID PMC8136920.

Organism / assay: Synechocystis sp. PCC 6803. Grad-seq = glycerol-gradient fractionation of a cell lysate (16 analysed fractions, 3 biological replicates), each fraction profiled by RNA-seq (transcriptome sedimentation) and shotgun proteomics / MS (proteome sedimentation). Co-sedimenting RNAs and proteins are clustered to reveal complexes.

Code: https://github.com/MatthiasRiediger/Analysis-of-a-photosynthetic-cyanobacterium-rich-in-internal-membrane-systems-via-Grad-seq commit 8aedf73b1d7ce03882d7b5942e4ece3a28136b78 (cloned to «infra»). This is the authors' own analysis code (R), and it ships its own input + intermediate + output CSVs — a self-contained, end-to-end reproducible bundle.

Repo layout:

  • normalization/GradSeq_RNAseq_git.R, GradSeq_Proteomics_git.R + the raw input CSVs (GradSeqInput_features.csv, GradSeqInput_transcripts.csv, GradSeqInput_proteomics.csv, spikeIns.csv). Turn raw UMI-deduplicated read counts / MS intensities into spike-in / relative normalized fraction profiles.
  • clustering/GradSeq_Clustering_git.R + its WGCNA input CSVs (WGCNA_input_RNAseq_features.csv, WGCNA_input_proteomics.csv) and the expected output (WGCNA_full_out.csv). Runs WGCNA dynamic-tree-cut on the combined RNA+protein matrix → module assignment per feature.
  • GradSeqExplorer_v2_2/ — a Shiny app + precomputed tables (SVM RBP scores, conservation, KEGG). Visualization layer, not a batch pipeline.

IN SCOPE (pipeline-derived computational results)

ID Result Pipeline How verified
C1 2,394 proteins detected in all 3 replicates (Fig 1C / Results, "67.3% of 3,559") proteomics normalization (GradSeq_Proteomics_git.R) → filtered protein set row count of reproduced WGCNA_input_proteomics.csv
C2 4,251 transcripts/features detected in all 3 replicates (Fig 1C / Results, "62.2% of 4,091 TUs") RNAseq normalization (GradSeq_RNAseq_git.R) → filtered feature set row count of reproduced WGCNA_input_RNAseq.csv
C3 17 WGCNA clusters (Results / Fig 2) WGCNA dynamic tree cut, signed, power=16 (GradSeq_Clustering_git.R) n distinct modules in reproduced output
C4 Per-feature module assignment for 6,645 features (4,251 RNA + 2,394 protein) same row-by-row agreement vs shipped WGCNA_full_out.csv
C5 Intermediate normalization outputs reproduce the shipped WGCNA input matrices 1:1 normalization scripts numeric diff of reproduced vs shipped WGCNA_input_*.csv

Pipelines named: spike-in / relative normalization (custom R, uses lme4, reshape2, plyr); WGCNA (adjacency signed power=16 → TOMsimilarityhclust average → cutreeDynamic deepSplit=1, pamStage=TRUE, minClusterSize=10).

OUT OF SCOPE (not attempted — wet-lab / external / manual)

  • Gradient fractionation, RNA/protein extraction, the PhiX174-derived spike-in RNA synthesis (wet lab).
  • MS database search / protein quantification (MaxQuant) producing GradSeqInput_proteomics.csv — upstream proprietary/heavy MS step; the count table is the provided starting point.
  • Read processing from raw SRA reads (FastQC → Cutadapt → segemehl → UMI dedup → featureCounts) producing GradSeqInput_features.csv/_transcripts.csv. The GradSeqInput count tables are the provided pipeline starting point; the raw reads (PRJNA608723) are profiled (see data/dataset_profile.json) but the full re-alignment is a large secondary effort (recorded, attempted only if core done).
  • SVM training/classification of RNA-binding proteins (Table 1, "14 candidates") — the SVM scores are shipped as a precomputed table in the Shiny app; th
Figures / tables: Fig 1CFig 2
C1
Reported
2,394 proteins detected in all 3 replicates
Reproduced
2,394 proteins (regenerated WGCNA_input_proteomics.csv = 2394 rows)
exact
C2
Reported
4,251 transcripts/features detected in all 3 replicates
Reproduced
4,251 transcripts (regenerated WGCNA_input_RNAseq.csv = 4251 rows)
exact
C3
Reported
17 WGCNA clusters
Reproduced
17 clusters
exact
C4
Reported
6,645 features (4251 RNA + 2394 protein) partitioned into 17 modules; largest 4057; protein modules 242/241/196/191/168
Reproduced
6,645 features, 17 modules, 100% exact label agreement, Adjusted Rand Index = 1.000; module sizes match (4057,...,242,241,...,196,191,...,168,...)
exact
C5
Reported
normalized WGCNA input matrices from raw inputs
Reproduced
RNAseq max abs diff 5e-08 corr 1.0; Proteomics (name-aligned) max abs diff 5e-08 corr 1.0
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

405.1 k
tokens (I/O) · 32.4 M incl. cache
60 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
1
HPC jobs
hummel
machine