Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Ribosome A and P sites revealed by length analysis of ribosome profiling data.

Nucleic Acids Res · 2015
L1 81/100 PQI 88
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
81/100
Reproducibility score
0.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 59% of all assessed papers rank 468 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH -> 1:1. Reproduced the paper's central computational result by running the authors' own GitHub pipeline (atmartens/ribosome_profiling @ 4835d77; bowtie1 + FASTX + Perl codon-composition) on the paper's own E. coli data. (1) PRIMARY (Fig 2): for E. coli LB Ribo-seq (GSE33671, SRR364368+SRR364370) the position-specific serine codon enrichment peaks EXACTLY at position 4 from the 3' footprint end = the A-site, identically across the three dominant footprint lengths -- a direct reproduction of the paper's A-site assignment. (2) CHECKPOINT: the repo's documented worked-example alignment numbers (README step 7, MOPS set GSE53767, SRR1067765-68) reproduce to ~5 sig figs (95,855,640 vs 95,854,989 alignments; 60.62% vs 60.60% unique; 158.13M vs 158.17M reads processed). Residuals <0.03%, consistent with FASTX/bowtie minor-version drift. Two 2013-era pipeline bugs had to be worked around for a modern stack (README's XM:i grep deletes all reads under bowtie 1.3.1 -> use RNAME!='*'; greedy genome-id regex needs a description-free FASTA header); both are version-drift, not fabrication, and don't change the science. NOT ATTEMPTED (80/20): yeast datasets (GSE58321/GSE50049: 5'-aligned P/A-site, proline enrichment) -- same method, more data; the 45nt-extension/GC and 2D-endpoint visualization steps; pvclust clustering. The '~28 nt canonical' length is graded partial: this E. coli LB library is genuinely length-heterogeneous (mode 36 nt), which is the premise of the paper's per-length method rather than a contradiction. No fabrication signal: all checked values are derivable from the shipped code on public GEO data.

💻 Code ↗ 🗄 Data: GSE33671

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 81
    assessed: 2026-06-16 ⛓ 9e04660367e3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The authors hypothesize that relating a ribosome footprint's sequence (the codons being translated) to its frequency in the sequencing library (pausing propensity), while accounting for footprint length variation, can reveal the locations of the ribosome A and P sites and the determinants of ribosome pauses.

Core claims
  • Accounting for ribosome footprint length variation reveals the ribosome aminoacyl (A) and peptidyl (P) site locations within footprints. finding
  • In E. coli, the A site is located 4 codons from the 3' end, and footprints of different lengths should be aligned to the 3' end (extra sequence in longer footprints is at the 5' end). finding
  • In yeast, the P site is located 4 codons downstream of the 5' end and the A site ~5 codons downstream of the 5' end; footprints align naturally to the 5' end (extra mRNA in longer footprints extends from the 3' end). finding
  • GC-rich 5' end motifs occur in yeast footprints (length-dependent), not only in E. coli, calling into question the anti-Shine-Dalgarno effect's role in ribosome pausing. finding
  • Proline codons can induce ribosome pausing in both yeast and bacteria, with pausing in E. coli observed at extreme consecutive-proline stretches. finding
  • A 2D heat-map graphical method displays footprint density across a transcript for all footprint lengths simultaneously. method
  • Codon-composition matrices track codon identity, position within footprint, and footprint length for statistical analysis. method
  • Source code for the analyses is provided on GitHub. resource
Experimental setups
Assay System Perturbation Readout Platform
ribosome profiling (re-analysis of published data) E. coli MG1655 grown in LB medium LB medium (serine depletion condition) position- and length-specific codon frequencies in footprints
ribosome profiling (re-analysis of published data) E. coli grown in MOPS medium none/diet (growth medium) position- and length-specific codon and GC frequencies in footprints
ribosome profiling (re-analysis of published data) Saccharomyces cerevisiae S288c 3-AT (3-amino-1,2,4-triazole) treatment, cycloheximide, untreated histidine/codon frequencies by footprint position and length
ribosome profiling (re-analysis of published data) Saccharomyces cerevisiae none (proline enrichment analysis) proline codon enrichment relative to mRNA background by position/length
mRNA-seq (re-analysis of published data) E. coli and S. cerevisiae none background codon frequencies / gene expression levels and nucleotide GC bias
sequence/GC content analysis E. coli and yeast footprint sequences none position-specific GC content (Savitzky-Golay smoothed)
hierarchical clustering of codon occupancies ribosome footprints (3'-aligned) none position-specific codon enrichment clusters pvclust in R 3.1.1
2D heat-map visualization of footprint ends single genes (e.g., E. coli amiB) none footprint 5'/3' endpoint positions and read counts along a gene Gnuplot 4.6
Key results
  • All six serine codons are enriched in E. coli LB footprints, peaking at position 4 upstream of the 3' end (positions 3-5 enriched), indicating the A site location.
  • All four proline codons (CCN) are enriched in yeast footprints, most strongly at position 4 downstream of the 5' end, indicating the P site location. 1.5-2.5-fold
  • Proline enrichment varies by codon: CCA most enriched, CCG least enriched in yeast.
  • In E. coli, individual proline codon frequencies in footprints are similar to mRNA background (CCG and to a lesser extent CCA enriched; CCU and CCC not), so no general proline pause.
  • E. coli amiB gene's eight consecutive proline codons (codons 130-138) coincide with a sharp rise in ribosome footprint density.
  • After 3-AT treatment, histidine codons are enriched at nearly all positions in yeast footprints, with slightly stronger enrichment 5 codons downstream of the 5' end, suggesting the A site position.
  • E. coli footprint 5' ends are G/GC-rich with enrichment growing with footprint length; extended 5' regions of shorter footprints are not enriched, supporting aSD; extreme 5' position is GC-poor (unexpected).
  • Yeast footprints also show GC-enriched 5' ends with length-dependent increase, unexpected under the aSD model; these biases are distinct from mRNA-seq library biases.
Key statistics
  • fold_change 1.5-2.5-fold (overall proline codon enrichment in yeast footprints vs mRNA background)
  • count 8 consecutive proline codons (codons 130-138) (longest proline stretch in E. coli amiB, coincides with footprint density rise)
  • count 28 nt (canonical ribosome footprint length)
  • count 64 (number of codon matrices used in the analysis (one per codon))
  • other position 4 from 3' end (A site, E. coli); position 4 from 5' end (P site, yeast); position 5 from 5' end (A site, yeast) (inferred ribosome site locations within footprints)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is an observational re-analysis of previously published yeast and Escherichia coli ribosome-profiling and mRNA-seq datasets, framed as a computational/descriptive pipeline rather than a hypothesis-testing design. Footprints were aligned (Bowtie), binned by length and position, and codon/nucleotide frequencies were tallied and normalized (e.g. footprint frequencies divided by parallel mRNA-seq background, or columns normalized to sum to 1); results are presented chiefly as fold-enrichments and 2D heat maps. The one formal inferential procedure described is hierarchical clustering of codon occupancies using pvclust with bootstrap resampling (correlation distance, average linkage, nboot = 100000).

Replicationmixed Sample sizeDescribed by listing the specific SRA/GEO run accessions reused (experimental replicates merged via cat); no formal sample-size or power calculation stated Groupstreatment vs control / footprint length classes / footprint vs mRNA background across yeast and E. coli (e.g. LB vs MOPS, 3-AT vs untreated, cycloheximide vs untreated) Pairingna Randomization/blindingna Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated (pvclust reports bootstrap approximately-unbiased/bootstrap-probability support values rather than a multiplicity correction)
Statistical tests used
Test Applied to n Assumptions
Hierarchical clustering with bootstrap resampling (pvclust; correlation distance, average linkage) clustering of position-specific codon occupancies to group codons by enrichment (Supplementary Figure S1, serine analysis) nboot = 100000 bootstrap replicates; underlying sample size not stated not stated
Fold-enrichment / frequency ratio (footprint codon frequency normalized to mRNA-seq background) proline enrichment in yeast and E. coli footprints (Figures 3, 4) na
Descriptive frequency tabulation and normalization (per-position codon/nucleotide percentages) serine, histidine and GC-content analyses across footprint lengths (Figures 2, 6, 7, 8, 9) na
Approaches that could also have been used
  • Codon enrichments and position assignments (e.g. A/P site locations) are reported descriptively as fold-changes and heat-map intensities without accompanying significance values or interval estimates.
    Could also: Per-position enrichments could also be accompanied by a test or interval — for instance a binomial/Poisson or permutation test against the mRNA background, or bootstrap confidence intervals on the fold-enrichment. — Attaching an interval or p-value would also quantify how strongly the data distinguish a given position (e.g. position 4 vs neighbors) from background fluctuation, complementing the visual peak identification.
  • Experimental replicates were merged into a single pooled file (via cat) before frequency tallying.
    Could also: One could also compute the statistic per replicate and summarize across replicates (e.g. mean with SD/SEM or a mixed-effects model), or use replicate-aware count frameworks such as DESeq2/edgeR. — Keeping replicates separate would also convey between-replicate variability and let one report dispersion, which is often preferred for gauging reproducibility of an enrichment.
  • Enrichments were examined across all 64 codons and many footprint positions and lengths, with no multiplicity adjustment described.
    Could also: A family-wise or false-discovery-rate adjustment (e.g. Bonferroni or Benjamini-Hochberg) could also be applied if formal significance calls were made across the codon-by-position grid. — A correction would also control the rate of incidental high-enrichment cells expected when scanning a large matrix of positions and codons.
  • Hierarchical clustering used correlation distance with average linkage and pvclust bootstrap support (nboot = 100000).
    Could also: Alternative distance metrics (e.g. Euclidean) or linkage rules (Ward, complete), or a complementary ordination such as PCA/UMAP, could also be used and compared. — Comparing linkages/metrics would also indicate how robust the serine-codon cluster is to clustering choices, beyond the bootstrap support already reported.
  • Background codon frequencies were estimated from parallel mRNA-seq, and (unlike Artieri et al.) codon-resolution read-density corrections were not performed.
    Could also: One could also incorporate position-level read-density normalization or a model accounting for sequence-specific ligation/PCR bias when forming the background. — Such corrections would also help separate genuine ribosome occupancy from technical composition biases, which the paper itself notes are present in the library-prep steps (Figure 9).
  • GC-content trends along footprints were highlighted using Savitzky-Golay smoothing and heat-map visualization.
    Could also: A regression of 5′ GC content on footprint length (with a slope estimate and CI) could also be reported alongside the smoothed visualization. — A fitted trend would also give an effect-size and uncertainty for the stated length-dependent GC increase, complementing the qualitative color gradient.
Software: R / pvclust R 3.1.1 · Bowtie (alignment) 1.0.0 · SRA Toolkit (fastq-dump) 2.3.2–5 · FASTX-Toolkit (fastx_clipper) 0.0.13.2 · Perl (custom processing/Savitzky-Golay smoothing) 5.18.2 · Gnuplot (heat maps) 4.6 patchlevel 3 · GNU grep / GNU cat (read filtering and replicate merging) grep 2.18; cat 8.21

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
47
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE53767 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE33671 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE50049 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE58321 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-25805170

Paper: Martens, Taylor, Hilser. Ribosome A and P sites revealed by length analysis of ribosome profiling data. Nucleic Acids Res 2015. PMID 25805170 / PMC4402525 / doi:10.1093/nar/gkv200. Code: https://github.com/atmartens/ribosome_profiling (authors' own; Perl + Gnuplot + R).

Method (in scope — pipeline-derived)

The paper's central computational result: for each ribosome-footprint length, tally position-specific codon frequencies; codons for slowly-translated amino acids (Ser, Pro, His) are enriched at the ribosomal A/P site, revealing the site position by read length. For E. coli the footprints align to the 3′ end and serine peaks at position 4 from the 3′ end = A-site (Figure 2). For yeast, footprints align to the 5′ end (P-site 4 codons from 5′; A-site ~5 from 5′).

The repo README is a faithful 15-step BASH/Perl pipeline (SRA→fastq→fastx_clipper trim →bowtie rRNA filter→bowtie genome align→simplified-SAM→PTT no-stop→elongation filter→ 3′-aligned codon-composition heatmaps). Tools: bowtie 1, FASTX-toolkit, Perl, Gnuplot.

In-scope results we attempt

  1. E. coli serine A-site — run the authors' full pipeline on GSE33671 (LB, SRR364368+SRR364370, the paper's E. coli set) → 3′-aligned codon-composition heatmaps; check serine % peaks at position 4 from the 3′ end. (Figure 2.)
  2. Canonical footprint length — footprint length distribution from the same data (claim: ~28 nt canonical; pipeline caps extension at 45 nt).
  3. Pipeline alignment checkpoint — reproduce the repo README §7 worked-example numbers (158,173,518 reads processed / 60.60% uniquely aligned / 95,854,989 alignments) by running the exact documented commands on GSE53767 (E. coli MOPS, SRR1067765–1067768). Deterministic; a hard 1:1 number.

Out of scope (not attempted)

  • Yeast datasets (GSE58321, GSE50049) P-site/A-site (5′-aligned) — same method, more data; skipped per 80/20 (E. coli is the primary, cleanest claim).
  • Gnuplot/visual figure rendering (we grade the numeric .heatmap.txt matrices, not PNGs).
  • pvclust hierarchical clustering (cluster.r) — exploratory, not a pinnable number.
  • Wet-lab / manual interpretation.

Notes

  • Genome NC_000913.2 (the 2013 version the paper used) fetched via NCBI E-utilities; PTT reconstructed from the NCBI feature table (PTT format deprecated). FASTA header rewritten to >gi|0|ref|NC_000913.2| to satisfy the scripts' \|(NC_...)\. regex.
  • Adapter CTGTAGGCACCATCAAT (Weissman-lab linker, documented in README) used for both E. coli sets.
  • All compute on «our HPC» SLURM; data on «infra».
ecoli_serine_Asite
Reported
Serine codons enriched, peaking at position 4 from the 3' footprint end = ribosomal A-site (E. coli LB, Fig 2)
Reproduced
Ser % peaks at position 4 from the 3' end for all 3 dominant footprint lengths (10/11/12 codons); 1.89x / 2.16x / 2.65x over the per-footprint mean
exact
readme_align_count
Reported
Reported 95854989 alignments (bowtie -l 23 -m 1; README step 7, GSE53767)
Reproduced
Reported 95855640 alignments (delta 0.0007%)
within tolerance
readme_align_unique_pct
Reported
60.60% uniquely aligned
Reproduced
60.62% (failed 34.45% vs 34.47%; -m suppressed 4.93% vs 4.93%)
within tolerance
readme_align_reads_processed
Reported
158173518 reads processed into genome alignment
Reproduced
158130803 reads processed (delta 0.027%)
within tolerance
canonical_fp_length
Reported
~28 nt canonical ribosome footprint length
Reproduced
GSE33671 LB length distribution broad, mode 36 nt, ~uniform 30-36 nt; no sharp 28 nt peak (consistent with paper's bacterial length-heterogeneity premise)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 81/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is a clean 1:1 reproduction: the paper's central computational claim — E. coli serine codon enrichment peaking at position 4 from the 3' footprint end (the A-site, Fig 2) — reproduces exactly across all three dominant footprint lengths by running the authors' own GitHub pipeline on the paper's own GSE33671 data. The repo's documented alignment checkpoints reproduce to ~5 significant figures (deltas <0.03%), all consistent with bowtie/FASTX version drift, not authors' defect. The only non-match is the Introduction's general '~28 nt canonical' footprint length vs this library's broad, length-heterogeneous distribution (mode 36 nt) — but that is the very premise of the paper's per-length method, an honest divergence rather than a contradiction. No fabrication signal; every checked value is derivable from the shared code and public data.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

238 k
tokens (I/O) · 20.1 M incl. cache
68 min
runtime · 4.09 CPU-h
2.5 GB
peak RAM
3
HPC jobs
hummel
machine