Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Grad-seq identifies KhpB as a global RNA-binding protein in Clostridioides difficile that regulates toxin production.

Microlife · 2021
50/100 3/4
Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PRELIMINARY (in progress). Third-party tool reproduction (GRADitude, P16) on paper's own GEO data GSE165672. Key friction: GEO deposits only WIG coverage tracks + raw FASTQ (SRA SRP303528); NO processed count table, so GRADitude input must be regenerated from FASTQ via READemption 4.3. Pipeline (READemption->GRADitude / DESeq2) being run on «our HPC». In scope: C1-C8 (sedimentation clustering, RIP-seq enrichment, dRNA-seq DEGs). Out of scope: proteomics/MS (1867 proteins), wet-lab. «our HPC» VPN intermittently unreachable during setup.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
not recorded
Assessed by
Last updated
2026-07-29

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether Grad-seq (gradient profiling of native RNA-protein complexes) can systematically uncover previously unidentified RNA-binding proteins (RBPs) in the Gram-positive pathogen Clostridioides difficile, and whether the discovered protein KhpB acts as a pervasive, Hfq-like global RBP that regulates toxin (TcdA) expression.

Core claims
  • Grad-seq resolves in-gradient sedimentation profiles for ~87-88% of annotated C. difficile transcripts and ~50% of annotated proteins, providing a comprehensive RNA-protein complexome resource resource
  • KhpB (Jag), previously characterized in S. pneumoniae, is identified as an uncharacterized, pervasive RNA-binding protein in C. difficile via Grad-seq sedimentation profiling and pulldown approaches finding
  • Global RIP-seq shows KhpB binds a large suite of mRNA and sRNA targets in C. difficile, comparable in scope to the Hfq targetome finding
  • KhpB-bound transcripts include functionally related mRNAs encoding virulence-associated metabolic pathways and the toxin A transcript finding
  • Toxin A transcript levels are increased in a khpB deletion strain finding
  • Toxin A protein production is increased upon khpB deletion finding
  • 6S RNA co-migrates with RNA polymerase subunits and RNAP-derived pRNAs are detected, providing evidence for a functional 6S RNA-RNAP complex in C. difficile finding
  • The abundant ncRNA RaiA exists in a stable RNP complex, providing the first experimental evidence of its potential function finding
Experimental setups
Assay System Perturbation Readout Platform
Grad-seq (glycerol gradient ultracentrifugation of native lysate) C. difficile strain 630 (WT), late-exponential phase, BHI broth none sedimentation/separation of RNA-protein complexes into 20 fractions + pellet linear glycerol gradient
RNA-seq of gradient fractions C. difficile 630 WT gradient fractions none transcript sedimentation profiles / abundance across fractions
LC-MS/MS (mass spectrometry) of gradient fractions C. difficile 630 WT gradient fractions none protein sedimentation profiles / abundance across fractions LC-MS/MS
A260 absorbance profiling C. difficile 630 WT gradient fractions none bulk RNA/complex absorbance profile across fractions
Northern blot C. difficile 630 WT gradient fractions (RaiA, 6S RNA, 5'UTRs of bglF, CD0426, CD1541, CD2512) none detection/size of specific RNA species and their gradient distribution radioactively/DNA-probe labeled Northern blot
RIP-seq (RNA immunoprecipitation sequencing) C. difficile, KhpB protein none (pulldown of KhpB-associated RNA) genome-wide identification of KhpB-bound mRNA and sRNA targets
Pulldown / co-purification approach C. difficile lysate none identification of KhpB as RBP via associated complexes
t-SNE dimensionality reduction of RNA-seq gradient data C. difficile Grad-seq RNA-seq dataset (computational) none clustering of similarly-behaving ncRNAs/transcripts
khpB deletion strain analysis (transcript and toxin protein quantification) C. difficile khpB deletion mutant vs WT khpB gene deletion toxin A (TcdA) transcript levels and toxin protein production
Key results
  • Sedimentation profiles obtained for 3541 transcripts (~87% of cellular transcripts) and 1867 proteins (~50% of annotated proteome) ~87% transcripts / ~50% proteome
  • KhpB identified as pervasive RBP via combined sedimentation profiling and pulldown data
  • RIP-seq establishes large suite of KhpB mRNA and sRNA targets, similar in scope to the Hfq targetome (Hfq binds at least 10% of C. difficile transcripts) at least 10% of transcripts (Hfq comparison)
  • Toxin A (TcdA) transcript levels increased in khpB deletion strain
  • Toxin protein production increased upon khpB deletion
  • 6S RNA species (~200 nt and ~175 nt) co-migrate with RNAP subunits and RNAP-derived pRNAs detected initiating within the central bulge 6S RNA ~200/175 nt
  • RaiA detected as two RNA species (~260 nt and ~220 nt) at levels comparable to rRNA, in a stable RNP complex ~260 nt / ~220 nt
  • UTR-CDS sedimentation correlation analysis identifies numerous 5'/3'UTRs with low correlation (R<0.2) to parental mRNA, often harboring riboswitches, PTT elements, or UTR-derived sRNAs R<0.2
Key statistics
  • count 3541 transcripts (~87%) (transcripts with resolved Grad-seq sedimentation profiles)
  • count 1867 proteins (~50%) (proteins with resolved Grad-seq sedimentation profiles)
  • other 20 gradient fractions plus pellet (Grad-seq fractionation scheme)
  • correlation R < 0.2 (low correlation between UTR and parental CDS sedimentation profiles, associated with riboregulatory elements)
  • other at least 10% (proportion of C. difficile transcripts bound by Hfq, used as comparison for KhpB targetome scope)
  • other ~180 annotated RBPs (number of annotated RNA-binding proteins in E. coli, cited for comparison)
  • other ~260 nt and ~220 nt (sizes of two RaiA RNA species detected by Northern blot)
  • other ~200 nt and ~175 nt (sizes of two 6S RNA species detected by Northern blot)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper presents a discovery-oriented resource study applying Grad-seq (glycerol gradient fractionation coupled to RNA-seq and mass spectrometry) to map RNA-protein complexes in Clostridioides difficile. Analytical approaches described in the available text include Spearman's rank correlation to compare sedimentation profiles between UTRs and their parental mRNAs, and t-SNE dimensionality reduction to cluster RNAs by similarity of in-gradient behavior, with candidate RNA-binding proteins and RNA partners inferred via a 'guilt-by-association' correlation logic. The provided text does not include the Methods section or the later results describing quantification of toxin transcript/protein levels in the khpB deletion strain, so inferential statistics for those comparisons are not visible here.

Replicationunclear GroupsWT C. difficile Grad-seq gradient fractions; later comparison of khpB deletion strain vs WT for toxin/transcript levels (details not in excerpt) Pairingunclear Randomization/blindingnot stated Dispersionunclear
Statistical tests used
Test Applied to n Assumptions
Spearman's rank correlation correlation of 5'/3' UTR sedimentation profiles with their corresponding CDS/mRNA profiles (Fig. 2B) not stated
t-SNE (t-distributed stochastic neighbor embedding) dimensionality reduction clustering of RNA sedimentation profiles to find similarly behaving ncRNAs (Fig. 3A) na
Approaches that could also have been used
  • UTR–CDS sedimentation profile relationships were assessed with Spearman's rank correlation.
    Could also: Pearson correlation could also be used if the profiles are expected to have an approximately linear relationship, or a distance/similarity metric paired with hierarchical clustering. — Pearson correlation can be more sensitive to the magnitude of proportional co-variation when the underlying relationship is linear, while Spearman is often preferred when only monotonicity (not linearity) is expected, so the choice can be tailored to the anticipated shape of the relationship.
  • Candidate RNA-binding proteins and RNA partners were nominated via a correlation-based 'guilt-by-association' approach across gradient fractions.
    Could also: A formal significance or false-discovery-rate threshold (e.g. permutation-based p-values with Benjamini-Hochberg correction) applied to the correlation scores could also be used. — Explicit FDR control over the large number of pairwise RNA-protein comparisons is a standard way to help readers gauge the expected proportion of chance associations among nominated candidates.
  • t-SNE was used to visualize clustering of RNA sedimentation profiles.
    Could also: UMAP or PCA could also be used for dimensionality reduction and visualization of the same data. — UMAP often better preserves global structure alongside local clusters, and PCA offers a linear, more directly interpretable projection; comparing multiple reduction methods can help confirm that observed clusters are not an artifact of one particular technique's parameters.
  • Sedimentation profile similarity (e.g. RNA-protein co-migration) was interpreted primarily through visual/heatmap comparison of normalized profiles.
    Could also: A quantitative co-migration or profile-similarity score with an associated null distribution (e.g. based on shuffled fractions) could also be reported alongside the heatmaps. — Pairing visual heatmap comparisons with a quantitative statistic and reference null distribution can help readers distinguish visually similar profiles that are statistically distinguishable from background sedimentation patterns.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37223250

Paper: Lamm-Schmidt et al., "Grad-seq identifies KhpB as a global RNA-binding protein in Clostridioides difficile that regulates toxin production." microLife (2023). PMID 37223250 · PMCID PMC10117727 · DOI 10.1093/femsml/uqab004.

Code: https://github.com/foerstner-lab/GRADitude (third-party tool, P16 — equally valid). Default branch main, HEAD 98824aa18cb76dd4b495219d7634ce6a42a31ef3 (2026-06-04). Paper used GRADitude v0.1.0 + READemption v4.3 + cutadapt. Will pin a released/dated commit at run time.

Data: GEO GSE165672 (SRA SRP303528), C. difficile 630, Illumina NextSeq 500. 40 GSM samples:

  • 22 Grad-seq gradient fractions (GSM5048012–33, Grad_01_Fraction_00L..21P, Rep_1)
  • 3 FLAG RIP-seq pulldowns + 3 WT controls (GSM5048034–39)
  • 12 dRNA-seq (GSM5048040–51): E1–E6 (exponential) + S1–S6 (stationary), WT/KO

Critical data note: GEO ships ONLY GSE165672_RAW.tar (355 MB, WIG coverage tracks) + raw FASTQ on SRA. No processed/normalized count table (gene_wise_quantifications_combined.csv) is deposited. The paper's Grad-seq quantification matrix that feeds GRADitude must therefore be regenerated from FASTQ via READemption — it is not directly downloadable. Supplementary Tables S1–S4 (paper) contain the reported result values (sedimentation data, RIP-seq enriched genes, DEGs) used here as the "reported" side of each claim.

Reference genome

C. difficile 630 — CP010905.2 with ncRNA + UTR annotations; ERCC92 spike-in (mix 1) added for normalization.

IN SCOPE (pipeline-derived — bioinformatic)

id result reported pipeline
C1 transcripts with resolved sedimentation profile (pass GRADitude filter, ≥100 reads) 3541 (~87% of cellular transcripts) READemption→GRADitude min_row_sum
C2 sRNAs grouped by k-means into clusters of similarity 3 clusters GRADitude clustering (k-means, Lloyd)
C3 t-SNE map of in-gradient sedimentation profiles Fig (perplexity=40) GRADitude dimension-reduction t-SNE
C4 RIP-seq transcripts significantly enriched by KhpB ~1400 (FC ≥ 5) READemption→DESeq2 (FLAG vs WT)
C5 DEGs in ΔkhpB, late-exponential 80 (log2FC>|1|) READemption→DESeq2 (dRNA-seq E, KO vs WT)
C6 DEGs in ΔkhpB, stationary 122 (log2FC>|1|) READemption→DESeq2 (dRNA-seq S, KO vs WT)
C7 sRNAs enriched among detected 10 of 42 overlap of C4 with sRNA annotation
C8 riboswitches enriched among detected 9 of 80 overlap of C4 with riboswitch annotation

Quick-minimum (~80% floor): C1 + C5/C6 are the most discrete, pinnable numbers. Headline reproduction = the GRADitude analysis itself (C1–C3), since GRADitude is the tool under test.

OUT OF SCOPE (wet-lab / external / manual — not attempted)

  • Proteomics: 1867 proteins (~50% proteome) — mass spectrometry, wet-lab.
  • Northern blots, growth curves, toxin (TcdA/TcdB) ELISA/Western, microscopy.
  • Strain/plasmid construction (Table S5).
  • The online sedimentation browser (helmholtz-hiri.de/datasets/gradseqcd) — a visualization, not a recomputable result.

Reproduction strategy

  1. Front1: clone GRADitude on «infra»; download SRA FASTQs; fetch CP010905.2 + ERCC92.
  2. SLURM: READemption project → align → gene quantification (Grad-seq 22 frac; dRNA-seq 12). cutadapt trim.
  3. GRADitude: filter (min_row_sum 100) → ERCC normalize (robust regression) → scaling → t-SNE (perplexity 40) → k-means → compare C1–C3.
  4. DESeq2 on dRNA-seq + RIP-seq counts → C4–C8.
Figures / tables: Fig 1Table
C1
Reported
3541 transcripts (~87%) with sedimentation profile
Reproduced
pending
partial
C5
Reported
80 DEGs delta-khpB late-exponential (log2FC>|1|)
Reproduced
pending
partial
C6
Reported
122 DEGs delta-khpB stationary (log2FC>|1|)
Reproduced
pending
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

55.9 k
tokens (I/O) · 2.8 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.