Ribosome A and P sites revealed by length analysis of ribosome profiling data.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH -> 1:1. Reproduced the paper's central computational result by running the authors' own GitHub pipeline (atmartens/ribosome_profiling @ 4835d77; bowtie1 + FASTX + Perl codon-composition) on the paper's own E. coli data. (1) PRIMARY (Fig 2): for E. coli LB Ribo-seq (GSE33671, SRR364368+SRR364370) the position-specific serine codon enrichment peaks EXACTLY at position 4 from the 3' footprint end = the A-site, identically across the three dominant footprint lengths -- a direct reproduction of the paper's A-site assignment. (2) CHECKPOINT: the repo's documented worked-example alignment numbers (README step 7, MOPS set GSE53767, SRR1067765-68) reproduce to ~5 sig figs (95,855,640 vs 95,854,989 alignments; 60.62% vs 60.60% unique; 158.13M vs 158.17M reads processed). Residuals <0.03%, consistent with FASTX/bowtie minor-version drift. Two 2013-era pipeline bugs had to be worked around for a modern stack (README's XM:i grep deletes all reads under bowtie 1.3.1 -> use RNAME!='*'; greedy genome-id regex needs a description-free FASTA header); both are version-drift, not fabrication, and don't change the science. NOT ATTEMPTED (80/20): yeast datasets (GSE58321/GSE50049: 5'-aligned P/A-site, proline enrichment) -- same method, more data; the 45nt-extension/GC and 2D-endpoint visualization steps; pvclust clustering. The '~28 nt canonical' length is graded partial: this E. coli LB library is genuinely length-heterogeneous (mode 36 nt), which is the premise of the paper's per-length method rather than a contradiction. No fabrication signal: all checked values are derivable from the shipped code on public GEO data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 81assessed: 2026-06-16 ⛓ 9e04660367e3
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors hypothesize that relating a ribosome footprint's sequence (the codons being translated) to its frequency in the sequencing library (pausing propensity), while accounting for footprint length variation, can reveal the locations of the ribosome A and P sites and the determinants of ribosome pauses.
- ★ Accounting for ribosome footprint length variation reveals the ribosome aminoacyl (A) and peptidyl (P) site locations within footprints. finding
- ★ In E. coli, the A site is located 4 codons from the 3' end, and footprints of different lengths should be aligned to the 3' end (extra sequence in longer footprints is at the 5' end). finding
- ★ In yeast, the P site is located 4 codons downstream of the 5' end and the A site ~5 codons downstream of the 5' end; footprints align naturally to the 5' end (extra mRNA in longer footprints extends from the 3' end). finding
- ★ GC-rich 5' end motifs occur in yeast footprints (length-dependent), not only in E. coli, calling into question the anti-Shine-Dalgarno effect's role in ribosome pausing. finding
- ★ Proline codons can induce ribosome pausing in both yeast and bacteria, with pausing in E. coli observed at extreme consecutive-proline stretches. finding
- A 2D heat-map graphical method displays footprint density across a transcript for all footprint lengths simultaneously. method
- Codon-composition matrices track codon identity, position within footprint, and footprint length for statistical analysis. method
- Source code for the analyses is provided on GitHub. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ribosome profiling (re-analysis of published data) | E. coli MG1655 grown in LB medium | LB medium (serine depletion condition) | position- and length-specific codon frequencies in footprints | — |
| ribosome profiling (re-analysis of published data) | E. coli grown in MOPS medium | none/diet (growth medium) | position- and length-specific codon and GC frequencies in footprints | — |
| ribosome profiling (re-analysis of published data) | Saccharomyces cerevisiae S288c | 3-AT (3-amino-1,2,4-triazole) treatment, cycloheximide, untreated | histidine/codon frequencies by footprint position and length | — |
| ribosome profiling (re-analysis of published data) | Saccharomyces cerevisiae | none (proline enrichment analysis) | proline codon enrichment relative to mRNA background by position/length | — |
| mRNA-seq (re-analysis of published data) | E. coli and S. cerevisiae | none | background codon frequencies / gene expression levels and nucleotide GC bias | — |
| sequence/GC content analysis | E. coli and yeast footprint sequences | none | position-specific GC content (Savitzky-Golay smoothed) | — |
| hierarchical clustering of codon occupancies | ribosome footprints (3'-aligned) | none | position-specific codon enrichment clusters | pvclust in R 3.1.1 |
| 2D heat-map visualization of footprint ends | single genes (e.g., E. coli amiB) | none | footprint 5'/3' endpoint positions and read counts along a gene | Gnuplot 4.6 |
- ▲ All six serine codons are enriched in E. coli LB footprints, peaking at position 4 upstream of the 3' end (positions 3-5 enriched), indicating the A site location.
- ▲ All four proline codons (CCN) are enriched in yeast footprints, most strongly at position 4 downstream of the 5' end, indicating the P site location. 1.5-2.5-fold
- ▲ Proline enrichment varies by codon: CCA most enriched, CCG least enriched in yeast.
- – In E. coli, individual proline codon frequencies in footprints are similar to mRNA background (CCG and to a lesser extent CCA enriched; CCU and CCC not), so no general proline pause.
- ▲ E. coli amiB gene's eight consecutive proline codons (codons 130-138) coincide with a sharp rise in ribosome footprint density.
- ▲ After 3-AT treatment, histidine codons are enriched at nearly all positions in yeast footprints, with slightly stronger enrichment 5 codons downstream of the 5' end, suggesting the A site position.
- ▲ E. coli footprint 5' ends are G/GC-rich with enrichment growing with footprint length; extended 5' regions of shorter footprints are not enriched, supporting aSD; extreme 5' position is GC-poor (unexpected).
- ▲ Yeast footprints also show GC-enriched 5' ends with length-dependent increase, unexpected under the aSD model; these biases are distinct from mRNA-seq library biases.
- fold_change 1.5-2.5-fold (overall proline codon enrichment in yeast footprints vs mRNA background)
- count 8 consecutive proline codons (codons 130-138) (longest proline stretch in E. coli amiB, coincides with footprint density rise)
- count 28 nt (canonical ribosome footprint length)
- count 64 (number of codon matrices used in the analysis (one per codon))
- other position 4 from 3' end (A site, E. coli); position 4 from 5' end (P site, yeast); position 5 from 5' end (A site, yeast) (inferred ribosome site locations within footprints)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study is an observational re-analysis of previously published yeast and Escherichia coli ribosome-profiling and mRNA-seq datasets, framed as a computational/descriptive pipeline rather than a hypothesis-testing design. Footprints were aligned (Bowtie), binned by length and position, and codon/nucleotide frequencies were tallied and normalized (e.g. footprint frequencies divided by parallel mRNA-seq background, or columns normalized to sum to 1); results are presented chiefly as fold-enrichments and 2D heat maps. The one formal inferential procedure described is hierarchical clustering of codon occupancies using pvclust with bootstrap resampling (correlation distance, average linkage, nboot = 100000).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Hierarchical clustering with bootstrap resampling (pvclust; correlation distance, average linkage) | clustering of position-specific codon occupancies to group codons by enrichment (Supplementary Figure S1, serine analysis) | nboot = 100000 bootstrap replicates; underlying sample size not stated | not stated |
| Fold-enrichment / frequency ratio (footprint codon frequency normalized to mRNA-seq background) | proline enrichment in yeast and E. coli footprints (Figures 3, 4) | — | na |
| Descriptive frequency tabulation and normalization (per-position codon/nucleotide percentages) | serine, histidine and GC-content analyses across footprint lengths (Figures 2, 6, 7, 8, 9) | — | na |
-
Codon enrichments and position assignments (e.g. A/P site locations) are reported descriptively as fold-changes and heat-map intensities without accompanying significance values or interval estimates.↳ Could also: Per-position enrichments could also be accompanied by a test or interval — for instance a binomial/Poisson or permutation test against the mRNA background, or bootstrap confidence intervals on the fold-enrichment. — Attaching an interval or p-value would also quantify how strongly the data distinguish a given position (e.g. position 4 vs neighbors) from background fluctuation, complementing the visual peak identification.
-
Experimental replicates were merged into a single pooled file (via cat) before frequency tallying.↳ Could also: One could also compute the statistic per replicate and summarize across replicates (e.g. mean with SD/SEM or a mixed-effects model), or use replicate-aware count frameworks such as DESeq2/edgeR. — Keeping replicates separate would also convey between-replicate variability and let one report dispersion, which is often preferred for gauging reproducibility of an enrichment.
-
Enrichments were examined across all 64 codons and many footprint positions and lengths, with no multiplicity adjustment described.↳ Could also: A family-wise or false-discovery-rate adjustment (e.g. Bonferroni or Benjamini-Hochberg) could also be applied if formal significance calls were made across the codon-by-position grid. — A correction would also control the rate of incidental high-enrichment cells expected when scanning a large matrix of positions and codons.
-
Hierarchical clustering used correlation distance with average linkage and pvclust bootstrap support (nboot = 100000).↳ Could also: Alternative distance metrics (e.g. Euclidean) or linkage rules (Ward, complete), or a complementary ordination such as PCA/UMAP, could also be used and compared. — Comparing linkages/metrics would also indicate how robust the serine-codon cluster is to clustering choices, beyond the bootstrap support already reported.
-
Background codon frequencies were estimated from parallel mRNA-seq, and (unlike Artieri et al.) codon-resolution read-density corrections were not performed.↳ Could also: One could also incorporate position-level read-density normalization or a model accounting for sequence-specific ligation/PCR bias when forming the background. — Such corrections would also help separate genuine ribosome occupancy from technical composition biases, which the paper itself notes are present in the library-prep steps (Figure 9).
-
GC-content trends along footprints were highlighted using Savitzky-Golay smoothing and heat-map visualization.↳ Could also: A regression of 5′ GC content on footprint length (with a slope estimate and CI) could also be reported alongside the smoothed visualization. — A fitted trend would also give an effect-size and uncertainty for the stated length-dependent GC increase, complementing the qualitative color gradient.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Eight consecutive proline codons (positions 130–138) in E. coli amiB coincide with a sharp rise in ribosome footprint density, demonstrating site-specific proline-induced pausing.other e-coli up 2015×1papers★ This paper is the founder (earliest)
-
E. coli ribosome footprint 5' ends are GC-enriched with enrichment increasing with footprint length, consistent with anti-Shine-Dalgarno base-pairing at the 5' footprint boundary.other e-coli up 2015×1papers★ This paper is the founder (earliest)
-
Proline codon frequencies in E. coli ribosome footprints are not generally elevated above mRNA background, indicating no universal proline translational pause in E. coli.other e-coli none 2015×1papers★ This paper is the founder (earliest)
-
All six serine codons are enriched at the ribosomal A site (position ~4 upstream of footprint 3' end) in E. coli ribosome profiling data under serine-limiting conditions.other e-coli up 2015×1papers★ This paper is the founder (earliest)
-
Yeast ribosome footprint 5' ends show length-dependent GC enrichment distinct from mRNA-seq library bias, which is unexpected under the anti-Shine-Dalgarno model.other saccharomyces-cerevisiae up 2015×1papers★ This paper is the founder (earliest)
-
Histidine codons are enriched at nearly all footprint positions in yeast after 3-AT treatment, with strongest enrichment at the A site (~5 codons downstream of footprint 5' end).other saccharomyces-cerevisiae up 2015×1papers★ This paper is the founder (earliest)
-
Proline codon P-site enrichment in yeast varies by codon identity: CCA is most enriched and CCG is least enriched, indicating codon-specific differences in ribosome pausing.other saccharomyces-cerevisiae mixed 2015×1papers★ This paper is the founder (earliest)
-
All four proline codons (CCN) are enriched 1.5–2.5-fold at the ribosomal P site (position ~4 downstream of footprint 5' end) in yeast ribosome profiling data.other saccharomyces-cerevisiae up 2015×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-25805170
Paper: Martens, Taylor, Hilser. Ribosome A and P sites revealed by length analysis of ribosome profiling data. Nucleic Acids Res 2015. PMID 25805170 / PMC4402525 / doi:10.1093/nar/gkv200. Code: https://github.com/atmartens/ribosome_profiling (authors' own; Perl + Gnuplot + R).
Method (in scope — pipeline-derived)
The paper's central computational result: for each ribosome-footprint length, tally position-specific codon frequencies; codons for slowly-translated amino acids (Ser, Pro, His) are enriched at the ribosomal A/P site, revealing the site position by read length. For E. coli the footprints align to the 3′ end and serine peaks at position 4 from the 3′ end = A-site (Figure 2). For yeast, footprints align to the 5′ end (P-site 4 codons from 5′; A-site ~5 from 5′).
The repo README is a faithful 15-step BASH/Perl pipeline (SRA→fastq→fastx_clipper trim →bowtie rRNA filter→bowtie genome align→simplified-SAM→PTT no-stop→elongation filter→ 3′-aligned codon-composition heatmaps). Tools: bowtie 1, FASTX-toolkit, Perl, Gnuplot.
In-scope results we attempt
- E. coli serine A-site — run the authors' full pipeline on GSE33671 (LB, SRR364368+SRR364370, the paper's E. coli set) → 3′-aligned codon-composition heatmaps; check serine % peaks at position 4 from the 3′ end. (Figure 2.)
- Canonical footprint length — footprint length distribution from the same data (claim: ~28 nt canonical; pipeline caps extension at 45 nt).
- Pipeline alignment checkpoint — reproduce the repo README §7 worked-example numbers (158,173,518 reads processed / 60.60% uniquely aligned / 95,854,989 alignments) by running the exact documented commands on GSE53767 (E. coli MOPS, SRR1067765–1067768). Deterministic; a hard 1:1 number.
Out of scope (not attempted)
- Yeast datasets (GSE58321, GSE50049) P-site/A-site (5′-aligned) — same method, more data; skipped per 80/20 (E. coli is the primary, cleanest claim).
- Gnuplot/visual figure rendering (we grade the numeric .heatmap.txt matrices, not PNGs).
- pvclust hierarchical clustering (cluster.r) — exploratory, not a pinnable number.
- Wet-lab / manual interpretation.
Notes
- Genome NC_000913.2 (the 2013 version the paper used) fetched via NCBI E-utilities;
PTT reconstructed from the NCBI feature table (PTT format deprecated). FASTA header
rewritten to
>gi|0|ref|NC_000913.2|to satisfy the scripts'\|(NC_...)\.regex. - Adapter
CTGTAGGCACCATCAAT(Weissman-lab linker, documented in README) used for both E. coli sets. - All compute on «our HPC» SLURM; data on «infra».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean 1:1 reproduction: the paper's central computational claim — E. coli serine codon enrichment peaking at position 4 from the 3' footprint end (the A-site, Fig 2) — reproduces exactly across all three dominant footprint lengths by running the authors' own GitHub pipeline on the paper's own GSE33671 data. The repo's documented alignment checkpoints reproduce to ~5 significant figures (deltas <0.03%), all consistent with bowtie/FASTX version drift, not authors' defect. The only non-match is the Introduction's general '~28 nt canonical' footprint length vs this library's broad, length-heterogeneous distribution (mode 36 nt) — but that is the very premise of the paper's per-length method, an honest divergence rather than a contradiction. No fabrication signal; every checked value is derivable from the shared code and public data.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.