Phosphorylation of ribosomal protein S6 differentially affects mRNA translation based on ORF length.
The main results reproduced, with only marginal, non-material deviations.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction, described well enough to act on. C1 (authors' own C tool, repo dc22dff): compiled with gcc 12.2 and run on the shipped self-contained demo -> output BYTE-IDENTICAL to the shipped expected (SHA256 match) = clean deterministic tool reproduction. C2-C6 derived by joining the GSE168977 processed XLSX (HeLa + MEF per-transcript counts) with Ensembl release-94 annotation on «our HPC»: the central qualitative claim C5 (CDS/ORF length inversely correlates with p-RPS6 ribosome occupancy) reproduces (negative, p<1e-56); C6 (ribosomal-protein/TOP mRNAs have high p-RPS6 occupancy) reproduces (87th percentile); gene-count claims C2 (13001->14011) and C3 (10141->11078) reproduce in approach and magnitude (within ~8-9%) but not exactly; the precise low/high occupancy split C4 (1637/69) could NOT be reproduced. Root cause of the gaps: the GSE168977 supplements are depth-NORMALIZED, not raw, and omit the HeLa RNA-seq column + the transcript->gene collapse and classification rules, so exact figures need the SRA raw reads. No fabrication signal — reported values are plausible and reproducible in direction; the residual deltas are data-completeness/method-specification gaps, flagged for human review. Out of scope (not attempted): wet-lab IP specificity, polysome gradients, Western blots.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-18 ⛓ 91909995c821
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether RPS6 phosphorylation differentially affects mRNA translation depending on coding sequence (CDS/ORF) length, motivated by the open question of what molecular consequence RPS6 phosphorylation has on ribosome activity.
- ★ RPS6 becomes progressively dephosphorylated on ribosomes as they translate along an mRNA CDS finding
- ★ Average RPS6 phosphorylation is higher on mRNAs with short CDSs than on mRNAs with long CDSs finding
- ★ RPS6 phosphorylation promotes translation of short-CDS mRNAs more strongly than long-CDS mRNAs finding
- ★ mRNAs with 5' TOP motifs are not sensitive to RPS6 phosphorylation for efficient translation despite having short CDSs, implying a distinct translation mode finding
- ★ Selective ribosome footprinting was adapted to isolate and sequence footprints of ribosomes carrying phosphorylated RPS6 from endogenous mRNAs method
- ★ Dephosphorylation of RPS6 on elongating ribosomes is slow relative to ribosome movement, so it becomes appreciable only on long CDSs mechanism
- Loss of RPS6 phosphorylation causes a relative increase in ribosomal protein synthesis/levels in phospho-deficient cells finding
- The anti-p-RPS6 (Ser235/236) immunoprecipitation is specific, as Torin1 treatment (mTOR inhibition) strongly reduces recovered RNA and RPS15 resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| phospho-RPS6-selective ribosome footprinting (Ribo-seq with p-RPS6 IP) | HeLa cells | none (endogenous, Torin1 used only for IP specificity control) | position/density of p-RPS6-containing ribosome footprints on mRNAs | Illumina NextSeq 550 |
| phospho-RPS6-selective and total 80S ribosome footprinting | MEF cells (control vs Rps6 phospho-deficient, rps6 p-/-) | genetic knock-in of non-phosphorylatable Rps6 | ribosome footprint distribution and translation efficiency | Illumina NextSeq 550 |
| total 80S ribosome footprinting / RNA-seq (translation efficiency) | HeLa cells and MEF cells | none / genetic (rps6 p-/-) | normalized ribosome footprint counts and mRNA-seq counts per CDS, translation efficiency | Illumina NextSeq 550 |
| immunoblotting (Western blot) | HeLa and MEF cell lysates and IP eluates | Torin1 (100 nM, 30 min) or none | p-RPS6, RPS15 and other protein levels | Biorad ChemiDoc Imaging System |
| immunoprecipitation followed by Q-RT-PCR | HeLa cells | none | enrichment of specific mRNAs in p-RPS6 IP relative to whole-cell lysate | QuantStudio3 |
| immunofluorescence microscopy | MEF cells | none | subcellular localization/staining of p-RPS6 | standard cell culture fluorescence microscope |
| RNA-seq (total mRNA) | MEF cells (control vs rps6 p-/-) | genetic (rps6 p-/-) | total mRNA abundance per gene | Illumina NextSeq 550 |
| GO/PANTHER gene set enrichment and overrepresentation analysis | HeLa and MEF ribosome footprinting datasets (computational) | none / genetic (rps6 p-/-) | enrichment of cellular component gene sets among differentially p-RPS6-occupied or differentially translated transcripts | PANTHER suite |
- ▼ Ribosomes translating the long POLR2A mRNA show progressively decreasing p-RPS6 occupancy from start to stop codon null
- ▲ Transcripts with low p-RPS6 ribosome occupancy (n=1637) have significantly longer coding sequences than other transcripts, but not longer 5'UTRs or 3'UTRs P<0.0001
- – z-vs-z analysis identified 1637 genes with low p-RPS6 ribosome occupancy and 69 genes with high p-RPS6 occupancy n=1637 low, n=69 high
- ▼ Torin1 treatment strongly reduces RNA and RPS15 recovered in p-RPS6 immunoprecipitation compared to control, confirming IP specificity <20% of control
- ▲ Short-CDS mRNAs show greater enrichment in p-RPS6 IP (by qRT-PCR) than long-CDS mRNAs P<0.0001
- – mRNAs encoding plasma membrane proteins are enriched among low-pRPS6-occupancy transcripts, while mRNAs encoding nuclear components are enriched among high-pRPS6-occupancy transcripts null
- ▼ Grouped analysis shows progressive RPS6 dephosphorylation along the CDS is more pronounced for longer mRNAs null
- – 13001 human genes and 10141 mouse genes passed the sequencing depth threshold (>63/64 raw reads) for inclusion in analyses 13001 (human), 10141 (mouse)
- pvalue ****P<0.0001 (Mann-Whitney test comparing CDS length of low p-RPS6-occupancy transcripts vs others)
- count 1637 genes low p-RPS6 occupancy; 69 genes high p-RPS6 occupancy (z-vs-z analysis of relative p-RPS6 ribosome occupancy per transcript)
- fold_change <20% of control (RNA and RPS15 recovered in p-RPS6 IP from Torin1-treated vs control cell lysates)
- pvalue ****P<0.0001 (Mann-Whitney test on qRT-PCR enrichment of short vs long CDS mRNAs in p-RPS6 IP (two biological replicates))
- count n=4573 (Number of CDSs longer than 3000nt used in metagene plot of total and p-RPS6 80S footprints)
- pvalue *P<0.0332, **P<0.0021, ****P<0.0001 (Kruskal-Wallis tests (multiple-testing adjusted) for p-RPS6 occupancy across CDS length groups)
- count 13001 genes (human); 10141 genes (mouse) (Genes passing raw read depth threshold (64 reads/transcript) across all sequencing libraries)
- other >63 raw reads threshold (Minimum read count threshold applied for gene inclusion in correlation and enrichment analyses)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used modification-selective ribosome footprinting (Ribo-seq) coupled with RNA-seq to profile the transcriptome-wide distribution of phospho-RPS6-containing ribosomes in HeLa and MEF cells. Read alignment and quantification were performed with custom C software and standard bioinformatics tools, with translation efficiency calculated as the ratio of normalized 80S footprint counts to normalized RNA-seq counts per coding sequence. Transcript-level group comparisons were conducted with nonparametric tests (Mann-Whitney U and Kruskal-Wallis), gene-set enrichment was assessed via PANTHER, and results were visualized as box plots with IQR boxes and first-to-tenth decile whiskers; p-values were reported as threshold categories rather than exact values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Two-sided unpaired Mann-Whitney U test | Comparison of CDS length, 5′UTR length, and 3′UTR length between transcripts with low p-RPS6 occupancy and all detected transcripts (Figures 1G–I) | 1637 low-p-RPS6 transcripts versus remainder of 13001 detected genes (human dataset) | not stated |
| Two-sided unpaired Kruskal-Wallis test with multiple-testing adjustment (method unnamed) | Comparison of average p-RPS6 occupancy across CDS-length bins (Figure 2D) | 13001 detected genes distributed across CDS-length bins; exact per-bin n not stated | not stated |
| Two-sided Mann-Whitney U test | Comparison of IP enrichment (qRT-PCR) between long-CDS and short-CDS mRNAs (Figure 2E) | Two biological replicates; exact number of mRNA species tested per group not stated | not stated |
| PANTHER Gene Set Enrichment test and PANTHER overrepresentation analysis (GO cellular component) | GO enrichment among genes changing in translation between WT and rps6 p−/− cells (Figure 4A); overrepresentation among 1637 low-p-RPS6 genes (Figure 3B) | 13001 detected genes as background; foreground sizes not stated per individual test | na |
| z-vs-z analysis (empirical distribution compared to fitted normal distribution) | Identification of transcripts with significantly low or high p-RPS6 ribosome occupancy (Figure 1F) | 13001 detected genes (human dataset) | not stated |
-
Three separate Mann-Whitney tests were applied to CDS length, 5′UTR length, and 3′UTR length simultaneously (Figures 1G–I) without a stated family-wise error correction↳ Could also: Apply a Bonferroni or Benjamini-Hochberg correction across the three related comparisons — Adjusting for multiple related tests is standard practice to control the probability that any one comparison is a false positive; with n > 10 000 the tests are very high-powered so the practical impact would likely be negligible here, but stating the correction would align with reporting guidelines
-
The post-hoc procedure applied after the Kruskal-Wallis test (Figure 2D) is described only as 'adjusted for multiple testing using statistical hypothesis testing' without naming the specific method↳ Could also: Name the specific post-hoc procedure (e.g., Dunn's test with Bonferroni or Benjamini-Hochberg correction, or Steel-Dwass test) — Identifying the method by name allows readers to assess its appropriateness for the pairwise comparison structure and to reproduce the analysis exactly
-
P-values are reported only as threshold categories (* P < 0.0332, ** P < 0.0021, **** P < 0.0001) rather than as exact values↳ Could also: Report exact p-values (e.g., P = 3.7 × 10−8) — Exact p-values convey the full strength of evidence, facilitate meta-analysis, and allow readers to judge robustness to alternative correction thresholds; threshold categories lose numerical information, which is especially relevant when sample sizes are large and p-values may be very small
-
Group spread in box plots is represented with IQR boxes and first-to-tenth decile whiskers, and no effect-size measure is reported alongside any statistical test↳ Could also: Report a rank-biserial correlation (r) or common language effect size alongside each Mann-Whitney U statistic — With very large n, even negligible biological differences yield extremely low p-values; an effect-size measure communicates the magnitude of the difference independently of sample size, helping readers calibrate biological relevance
-
The z-vs-z method for identifying outlier transcripts (Figure 1F) compares the empirical distribution of relative p-RPS6 enrichment scores to a fitted normal distribution to define low- and high-occupancy genes↳ Could also: Use a count-based differential occupancy framework (e.g., DESeq2, edgeR, or limma-voom) treating p-RPS6 IP and total 80S as paired libraries — Count-based methods model overdispersion and library-size variation explicitly, provide per-transcript FDR-controlled statistics, and do not rely on the normality assumption; they are widely used for analogous immunoprecipitation-followed-by-sequencing data (e.g., RIP-seq, CLIP-seq)
-
Changes in translation efficiency between wild-type and rps6 p−/− cells were assessed by binning transcripts into CDS-length groups and comparing those bins with nonparametric tests↳ Could also: Apply a method designed for differential translation efficiency from paired Ribo-seq and RNA-seq data, such as RiboDiff, xtail, or anota2seq — These methods model the joint distribution of RNA-seq and Ribo-seq counts per transcript, account for overdispersion in both data layers, and yield per-transcript FDR-controlled statistics for differential translation, which could complement the bin-level analysis and provide transcript-level resolution
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34871442
Paper: Bohlen J, Roiuk M, Teleman AA. Phosphorylation of ribosomal protein S6 differentially affects mRNA translation based on ORF length. Nucleic Acids Res 2021;49(22):13062-13074. DOI 10.1093/nar/gkab1157. PMCID PMC8682771.
Code: https://github.com/aurelioteleman/Teleman-Lab — folder
Ribosome-Footprinting-Analysis/2021/ (authors' OWN custom C software, GPL-3.0,
default branch master). Six C programs + a self-contained demo/ with shipped
input + expected output.
Data: GEO GSE168977 (SuperSeries) → BioProject PRJNA714824, SRA SRP310845.
7 samples, 2 organisms (human HeLa, mouse MEF), Ribo-seq + RNA-seq on NextSeq 550.
Series-level processed-data supplements:
GSE168977_Processed_Data_File_HeLa.xlsx (1.7 MB),
GSE168977_Processed_Data_File_MEF.xlsx (4.2 MB).
Pipeline (as described in Methods)
Raw reads → cutadapt (--nextseq-trim=10 --discard-untrimmed -m16 -M45 -O6)
→ bowtie2 remove rRNA/tRNA → BBmap map to Ensembl transcriptome (release 94)
- genome (hg38 / mouse) with
ambiguous=all maxindel=200000→ SAM → authors' C tools (count_features_nt_range_v2etc.) count reads per transcript → normalize to library depth → derive:- TE = norm. 80S footprints in CDS / norm. RNA-seq reads in CDS
- p-RPS6 occupancy = pS6-selective footprints / total footprints per transcript → threshold at 64 raw reads/transcript in ALL libraries → bin/correlate by annotated CDS (ORF) length → Mann–Whitney / Kruskal–Wallis stats.
IN SCOPE (pipeline-derived → attempt to reproduce)
| id | result | reported | location | reproduction route |
|---|---|---|---|---|
| C1 | demo tool correctness: count_features_nt_range_v2 on shipped input.sam reproduces output_input.txt exactly |
byte-identical | repo demo README | compile C, run on shipped demo, diff vs expected — DETERMINISTIC |
| C2 | human genes passing 64-read threshold | 13001 | Methods / Fig 1E | count rows in HeLa processed table / re-map+count |
| C3 | mouse genes passing threshold | 10141 | Methods | count rows in MEF processed table / re-map+count |
| C4 | genes low vs high p-RPS6 occupancy | 1637 low / 69 high | Fig 1F legend | reclassify from processed p-RPS6 table (z-vs-z) |
| C5 | CDS length inversely correlates with avg p-RPS6 occupancy | qualitative (neg.) | Fig 2D | compute Spearman rho(CDS length, p-RPS6) from processed table |
| C6 | TOP / ribosomal-protein mRNAs: high p-RPS6 but no TE increase | qualitative | Fig 5B | subset TOP genes in processed table, check TE distribution |
Primary fast path: download the two processed-data XLSX (on «infra») and check the derived counts/relationships directly — these ARE the pipeline outputs the numbers are quoted from. Deeper path: re-run trim→filter→map→count from SRA to regenerate the per-transcript counts independently.
OUT OF SCOPE (wet-lab / manual / not a pipeline result → not attempted)
-
80% IP specificity (Torin1 RNA-reduction experiment) — wet-lab design.
- Polysome gradients, urea-PAGE size selection, Western blots, phospho-IP itself.
- rpS6P-/- vs WT ribosomal-protein level measurements by blot (Fig 5E-F) — wet-lab.
- Any claim derived from immunoprecipitation efficiency rather than read counting.
Notes
- This is the authors' own code (not third-party), so P16 third-party allowance is not needed — but the code is generic footprint-counting C, the biology lives in how counts are normalized/thresholded/binned (described in Methods, not fully scripted in the repo). Expect C1 to be exact; C2-C6 depend on faithfully re-deriving normalization + thresholding from the Methods text.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.