Grad-seq identifies KhpB as a global RNA-binding protein in Clostridioides difficile that regulates toxin production.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PRELIMINARY (in progress). Third-party tool reproduction (GRADitude, P16) on paper's own GEO data GSE165672. Key friction: GEO deposits only WIG coverage tracks + raw FASTQ (SRA SRP303528); NO processed count table, so GRADitude input must be regenerated from FASTQ via READemption 4.3. Pipeline (READemption->GRADitude / DESeq2) being run on «our HPC». In scope: C1-C8 (sedimentation clustering, RIP-seq enrichment, dRNA-seq DEGs). Out of scope: proteomics/MS (1867 proteins), wet-lab. «our HPC» VPN intermittently unreachable during setup.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-19
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-07-29
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether Grad-seq (gradient profiling of native RNA-protein complexes) can systematically uncover previously unidentified RNA-binding proteins (RBPs) in the Gram-positive pathogen Clostridioides difficile, and whether the discovered protein KhpB acts as a pervasive, Hfq-like global RBP that regulates toxin (TcdA) expression.
- ★ Grad-seq resolves in-gradient sedimentation profiles for ~87-88% of annotated C. difficile transcripts and ~50% of annotated proteins, providing a comprehensive RNA-protein complexome resource resource
- ★ KhpB (Jag), previously characterized in S. pneumoniae, is identified as an uncharacterized, pervasive RNA-binding protein in C. difficile via Grad-seq sedimentation profiling and pulldown approaches finding
- ★ Global RIP-seq shows KhpB binds a large suite of mRNA and sRNA targets in C. difficile, comparable in scope to the Hfq targetome finding
- ★ KhpB-bound transcripts include functionally related mRNAs encoding virulence-associated metabolic pathways and the toxin A transcript finding
- ★ Toxin A transcript levels are increased in a khpB deletion strain finding
- ★ Toxin A protein production is increased upon khpB deletion finding
- 6S RNA co-migrates with RNA polymerase subunits and RNAP-derived pRNAs are detected, providing evidence for a functional 6S RNA-RNAP complex in C. difficile finding
- The abundant ncRNA RaiA exists in a stable RNP complex, providing the first experimental evidence of its potential function finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Grad-seq (glycerol gradient ultracentrifugation of native lysate) | C. difficile strain 630 (WT), late-exponential phase, BHI broth | none | sedimentation/separation of RNA-protein complexes into 20 fractions + pellet | linear glycerol gradient |
| RNA-seq of gradient fractions | C. difficile 630 WT gradient fractions | none | transcript sedimentation profiles / abundance across fractions | — |
| LC-MS/MS (mass spectrometry) of gradient fractions | C. difficile 630 WT gradient fractions | none | protein sedimentation profiles / abundance across fractions | LC-MS/MS |
| A260 absorbance profiling | C. difficile 630 WT gradient fractions | none | bulk RNA/complex absorbance profile across fractions | — |
| Northern blot | C. difficile 630 WT gradient fractions (RaiA, 6S RNA, 5'UTRs of bglF, CD0426, CD1541, CD2512) | none | detection/size of specific RNA species and their gradient distribution | radioactively/DNA-probe labeled Northern blot |
| RIP-seq (RNA immunoprecipitation sequencing) | C. difficile, KhpB protein | none (pulldown of KhpB-associated RNA) | genome-wide identification of KhpB-bound mRNA and sRNA targets | — |
| Pulldown / co-purification approach | C. difficile lysate | none | identification of KhpB as RBP via associated complexes | — |
| t-SNE dimensionality reduction of RNA-seq gradient data | C. difficile Grad-seq RNA-seq dataset (computational) | none | clustering of similarly-behaving ncRNAs/transcripts | — |
| khpB deletion strain analysis (transcript and toxin protein quantification) | C. difficile khpB deletion mutant vs WT | khpB gene deletion | toxin A (TcdA) transcript levels and toxin protein production | — |
- – Sedimentation profiles obtained for 3541 transcripts (~87% of cellular transcripts) and 1867 proteins (~50% of annotated proteome) ~87% transcripts / ~50% proteome
- – KhpB identified as pervasive RBP via combined sedimentation profiling and pulldown data
- – RIP-seq establishes large suite of KhpB mRNA and sRNA targets, similar in scope to the Hfq targetome (Hfq binds at least 10% of C. difficile transcripts) at least 10% of transcripts (Hfq comparison)
- ▲ Toxin A (TcdA) transcript levels increased in khpB deletion strain
- ▲ Toxin protein production increased upon khpB deletion
- – 6S RNA species (~200 nt and ~175 nt) co-migrate with RNAP subunits and RNAP-derived pRNAs detected initiating within the central bulge 6S RNA ~200/175 nt
- – RaiA detected as two RNA species (~260 nt and ~220 nt) at levels comparable to rRNA, in a stable RNP complex ~260 nt / ~220 nt
- – UTR-CDS sedimentation correlation analysis identifies numerous 5'/3'UTRs with low correlation (R<0.2) to parental mRNA, often harboring riboswitches, PTT elements, or UTR-derived sRNAs R<0.2
- count 3541 transcripts (~87%) (transcripts with resolved Grad-seq sedimentation profiles)
- count 1867 proteins (~50%) (proteins with resolved Grad-seq sedimentation profiles)
- other 20 gradient fractions plus pellet (Grad-seq fractionation scheme)
- correlation R < 0.2 (low correlation between UTR and parental CDS sedimentation profiles, associated with riboregulatory elements)
- other at least 10% (proportion of C. difficile transcripts bound by Hfq, used as comparison for KhpB targetome scope)
- other ~180 annotated RBPs (number of annotated RNA-binding proteins in E. coli, cited for comparison)
- other ~260 nt and ~220 nt (sizes of two RaiA RNA species detected by Northern blot)
- other ~200 nt and ~175 nt (sizes of two 6S RNA species detected by Northern blot)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper presents a discovery-oriented resource study applying Grad-seq (glycerol gradient fractionation coupled to RNA-seq and mass spectrometry) to map RNA-protein complexes in Clostridioides difficile. Analytical approaches described in the available text include Spearman's rank correlation to compare sedimentation profiles between UTRs and their parental mRNAs, and t-SNE dimensionality reduction to cluster RNAs by similarity of in-gradient behavior, with candidate RNA-binding proteins and RNA partners inferred via a 'guilt-by-association' correlation logic. The provided text does not include the Methods section or the later results describing quantification of toxin transcript/protein levels in the khpB deletion strain, so inferential statistics for those comparisons are not visible here.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Spearman's rank correlation | correlation of 5'/3' UTR sedimentation profiles with their corresponding CDS/mRNA profiles (Fig. 2B) | — | not stated |
| t-SNE (t-distributed stochastic neighbor embedding) dimensionality reduction | clustering of RNA sedimentation profiles to find similarly behaving ncRNAs (Fig. 3A) | — | na |
-
UTR–CDS sedimentation profile relationships were assessed with Spearman's rank correlation.↳ Could also: Pearson correlation could also be used if the profiles are expected to have an approximately linear relationship, or a distance/similarity metric paired with hierarchical clustering. — Pearson correlation can be more sensitive to the magnitude of proportional co-variation when the underlying relationship is linear, while Spearman is often preferred when only monotonicity (not linearity) is expected, so the choice can be tailored to the anticipated shape of the relationship.
-
Candidate RNA-binding proteins and RNA partners were nominated via a correlation-based 'guilt-by-association' approach across gradient fractions.↳ Could also: A formal significance or false-discovery-rate threshold (e.g. permutation-based p-values with Benjamini-Hochberg correction) applied to the correlation scores could also be used. — Explicit FDR control over the large number of pairwise RNA-protein comparisons is a standard way to help readers gauge the expected proportion of chance associations among nominated candidates.
-
t-SNE was used to visualize clustering of RNA sedimentation profiles.↳ Could also: UMAP or PCA could also be used for dimensionality reduction and visualization of the same data. — UMAP often better preserves global structure alongside local clusters, and PCA offers a linear, more directly interpretable projection; comparing multiple reduction methods can help confirm that observed clusters are not an artifact of one particular technique's parameters.
-
Sedimentation profile similarity (e.g. RNA-protein co-migration) was interpreted primarily through visual/heatmap comparison of normalized profiles.↳ Could also: A quantitative co-migration or profile-similarity score with an associated null distribution (e.g. based on shuffled fractions) could also be reported alongside the heatmaps. — Pairing visual heatmap comparisons with a quantitative statistic and reference null distribution can help readers distinguish visually similar profiles that are statistically distinguishable from background sedimentation patterns.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37223250
Paper: Lamm-Schmidt et al., "Grad-seq identifies KhpB as a global RNA-binding protein in Clostridioides difficile that regulates toxin production." microLife (2023). PMID 37223250 · PMCID PMC10117727 · DOI 10.1093/femsml/uqab004.
Code: https://github.com/foerstner-lab/GRADitude (third-party tool, P16 — equally
valid). Default branch main, HEAD 98824aa18cb76dd4b495219d7634ce6a42a31ef3
(2026-06-04). Paper used GRADitude v0.1.0 + READemption v4.3 + cutadapt.
Will pin a released/dated commit at run time.
Data: GEO GSE165672 (SRA SRP303528), C. difficile 630, Illumina NextSeq 500. 40 GSM samples:
- 22 Grad-seq gradient fractions (GSM5048012–33, Grad_01_Fraction_00L..21P, Rep_1)
- 3 FLAG RIP-seq pulldowns + 3 WT controls (GSM5048034–39)
- 12 dRNA-seq (GSM5048040–51): E1–E6 (exponential) + S1–S6 (stationary), WT/KO
Critical data note: GEO ships ONLY
GSE165672_RAW.tar(355 MB, WIG coverage tracks) + raw FASTQ on SRA. No processed/normalized count table (gene_wise_quantifications_combined.csv) is deposited. The paper's Grad-seq quantification matrix that feeds GRADitude must therefore be regenerated from FASTQ via READemption — it is not directly downloadable. Supplementary Tables S1–S4 (paper) contain the reported result values (sedimentation data, RIP-seq enriched genes, DEGs) used here as the "reported" side of each claim.
Reference genome
C. difficile 630 — CP010905.2 with ncRNA + UTR annotations; ERCC92 spike-in (mix 1) added for normalization.
IN SCOPE (pipeline-derived — bioinformatic)
| id | result | reported | pipeline |
|---|---|---|---|
| C1 | transcripts with resolved sedimentation profile (pass GRADitude filter, ≥100 reads) | 3541 (~87% of cellular transcripts) | READemption→GRADitude min_row_sum |
| C2 | sRNAs grouped by k-means into clusters of similarity | 3 clusters | GRADitude clustering (k-means, Lloyd) |
| C3 | t-SNE map of in-gradient sedimentation profiles | Fig (perplexity=40) | GRADitude dimension-reduction t-SNE |
| C4 | RIP-seq transcripts significantly enriched by KhpB | ~1400 (FC ≥ 5) | READemption→DESeq2 (FLAG vs WT) |
| C5 | DEGs in ΔkhpB, late-exponential | 80 (log2FC>|1|) | READemption→DESeq2 (dRNA-seq E, KO vs WT) |
| C6 | DEGs in ΔkhpB, stationary | 122 (log2FC>|1|) | READemption→DESeq2 (dRNA-seq S, KO vs WT) |
| C7 | sRNAs enriched among detected | 10 of 42 | overlap of C4 with sRNA annotation |
| C8 | riboswitches enriched among detected | 9 of 80 | overlap of C4 with riboswitch annotation |
Quick-minimum (~80% floor): C1 + C5/C6 are the most discrete, pinnable numbers. Headline reproduction = the GRADitude analysis itself (C1–C3), since GRADitude is the tool under test.
OUT OF SCOPE (wet-lab / external / manual — not attempted)
- Proteomics: 1867 proteins (~50% proteome) — mass spectrometry, wet-lab.
- Northern blots, growth curves, toxin (TcdA/TcdB) ELISA/Western, microscopy.
- Strain/plasmid construction (Table S5).
- The online sedimentation browser (helmholtz-hiri.de/datasets/gradseqcd) — a visualization, not a recomputable result.
Reproduction strategy
- Front1: clone GRADitude on «infra»; download SRA FASTQs; fetch CP010905.2 + ERCC92.
- SLURM: READemption project → align → gene quantification (Grad-seq 22 frac; dRNA-seq 12). cutadapt trim.
- GRADitude: filter (min_row_sum 100) → ERCC normalize (robust regression) → scaling → t-SNE (perplexity 40) → k-means → compare C1–C3.
- DESeq2 on dRNA-seq + RIP-seq counts → C4–C8.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.