Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genome-wide screening of potential RNase Y-processed mRNAs in the M49 serotype Streptococcus pyogenes NZ131.

Microbiologyopen · 2018
L1 57/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
57/100
Reproducibility score
1.0 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 17% of all assessed papers rank 965 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Faithful end-to-end rebuild of the operon-prediction pipeline (R1) with the authors' OWN unmodified Python-2 code: Bowtie2 --very-sensitive-local -> samtools -> genomeCoverageBed -d -> geneBankConstruct -> operonBankConstruct -> operonBankRefine, on SRP108413/GSE99533 reads (SRR5638291-4) over the CP000829.1 genome (1,815,785 bp, 1703 CDS). RESULT = PARTIAL: mapping QC reproduces EXACTLY (>99%: 99.90-99.94% across all 4 samples) and the gene universe matches (1703 CDS), but the operon COUNT is 990 vs the reported 865 (+14.5%) -- more monocistronic (632 vs 491), fewer large (81 vs 103); di/tri within ~10%. All four unified RD operon banks agree (990) with a BED line-count cross-check. Since gene set and mapping match, the gap most plausibly comes from bowtie2/samtools version differences in per-base coverage (the operonJudge 4-fold/0.5x-dent thresholds are depth-sensitive) and/or the authors' exact Spy49_allGenes.csv vs the GFF-derived set; the paper pins no tool versions. This is a faithful partial reproduction, NOT evidence of fabrication: the distribution shape and order of magnitude are reproduced. Fixed three defects from the prior run 2180633 (invalid 'bedtools genomeCoverageBed' -> 0-byte CSVs; operon counter missing repo imports; no git in env -> git clone failed, switched to wget tarball). NOT attempted: R2 RNase-Y-processed mRNAs (Table 2, 80->29->15) -- needs GSE40198 microarray half-lives + a manual non-deterministic 'discard 14 random' triage, outside the 80% scope.

💻 Code ↗ 🗄 Data: GSE40198

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-16 ⛓ 402e46d4600c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

RNase Y mediates mRNA processing events in Streptococcus pyogenes that regulate the expression of virulence factors by altering the stability of their mRNAs.

Core claims
  • RNase Y selectively processes GAS mRNA, but its overall impact is confined to a limited set of virulence factor transcripts. finding
  • Combined RNA-seq and tiling microarray analysis defined 865 S. pyogenes operons and identified candidate RNase Y-processed transcripts. method
  • 15 mRNAs were identified as candidates for RNase Y-mediated processing based on segmental stability differences between WT and Δrny. finding
  • Of the candidates tested by Northern blot (folC1, prtF, speG, ropB, ypaA), only folC1 was confirmed to be processed, and this processing is unlikely to be RNase Y-dependent. finding
  • Deletion of rny does not affect overall S. pyogenes operon organization. finding
  • The RNA-seq-based operon prediction pipeline showed 81% sensitivity and 71% specificity compared to the ProOpDB computational prediction tool. finding
  • A sliding-window coefficient-of-variation method was developed to identify operon transcriptional boundaries from RNA-seq read coverage. method
  • Python scripts for operon prediction and boundary identification were deposited on GitHub as a public resource. resource
Experimental setups
Assay System Perturbation Readout Platform
RNA-seq Streptococcus pyogenes NZ131 (M49), WT and Δrny mutant, 2 biological replicates each KO (Δrny) transcript abundance, operon structure/boundaries Illumina HiSeq 2000
Tiling microarray Streptococcus pyogenes NZ131, WT and Δrny mutant KO (Δrny) mRNA decay rate/half-life at subgene (probe) level Affymetrix microarray (CEL files)
Northern blot Streptococcus pyogenes NZ131 (folC1, prtF, speG, ropB, ypaA transcripts) KO (Δrny), with WT/complement comparison mRNA processing status (transcript size/pattern)
qRT-PCR Streptococcus pyogenes NZ131, WT vs Δrny (11 selected genes) KO (Δrny) relative transcript abundance real-time PCR
RT-PCR Streptococcus pyogenes NZ131 (10 randomly selected gene pairs) none cotranscription/operon validation
Western blot Streptococcus pyogenes NZ131 whole-cell lysate and extracellular (TCA-precipitated) protein other (genotype comparison implied) protein expression
5' RACE Streptococcus pyogenes NZ131 none transcript 5' end mapping
Key results
  • 865 operons predicted in the S. pyogenes NZ131 genome (491 monocistronic, 169 dicistronic, 102 tricistronic, 103 with ≥4 CDS) 865 operons
  • 15 mRNAs identified as candidates potentially processed by RNase Y 15 mRNAs
  • Only folC1 confirmed processed by Northern blot among candidates tested; processing unlikely attributable to RNase Y
  • High consistency between RNA-seq biological replicates for WT and Δrny r2=0.857 (WT), r2=0.998 (Δrny)
  • Strong correlation between qRT-PCR and RNA-seq expression measurements across 11 genes r2=0.996
  • RNA-seq operon predictions concorded with ProOpDB for cotranscribed and monocistronic gene pairs 81% sensitivity (720 pairs), 71% specificity (627 pairs)
  • 144 gene pairs predicted cotranscribed by RNA-seq but not by ProOpDB; 10 randomly tested pairs confirmed cotranscribed by RT-PCR 10/10 confirmed
  • Median 5' UTR length across 865 operons was 43 nt median 43 nt (range 0-200 nt)
Key statistics
  • correlation r2=0.857 (RNA-seq biological replicate consistency in WT)
  • correlation r2=0.998 (RNA-seq biological replicate consistency in Δrny)
  • correlation r2=0.996 (correlation between qRT-PCR and RNA-seq expression for selected genes)
  • count 865 (total operons predicted in S. pyogenes NZ131)
  • count 15 (mRNA candidates potentially processed by RNase Y)
  • other 81% sensitivity (concordance of RNA-seq cotranscribed gene pairs with ProOpDB)
  • other 71% specificity (concordance of RNA-seq monocistronic gene pairs with ProOpDB)
  • other 85% (1471 boundaries) with <10bp difference (consistency of operon boundary predictions across WT and Δrny datasets)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study combined RNA-seq (two biological replicates per condition) and Affymetrix tiling microarray analysis to characterize the S. pyogenes transcriptome and identify mRNAs processed by RNase Y, comparing a wild-type strain to an isogenic Δrny mutant. Rather than formal inferential statistics, the primary analytical framework relied on rule-based bioinformatic criteria: fold-change thresholds for operon membership, a coefficient-of-variation sliding-window algorithm for boundary detection, and a twofold half-life difference threshold for processing candidate identification. Replicate consistency and method agreement were assessed with R² correlation coefficients, and candidate transcripts were validated by Northern blot, qRT-PCR, 5′ RACE, and Western blot.

Replicationbiological Sample sizeTwo biological replicates per condition (wt1, wt2; Δrny1, Δrny2); no formal power calculation mentioned GroupsS. pyogenes NZ131 wild type vs. isogenic Δrny (RNase Y deletion) mutant Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Pearson R² correlation coefficient Consistency between RNA-seq biological replicates (WT and Δrny) and correlation of RNA-seq vs. qRT-PCR expression levels for 11 selected genes Two replicates per condition; 11 genes for method comparison not stated
Rule-based operon prediction (fourfold expression threshold, ≤100 bp intergenic distance, same-strand, intergenic expression ≥ half of least-expressed gene) Grouping of 1,699 protein-encoding genes into 865 operons 1,699 CDS in S. pyogenes NZ131 genome not stated
Coefficient of variation (CV) sliding-window algorithm with twofold maxCV/refCV threshold Identification of transcriptional start and stop sites (operon boundary definition) 865 operons; 25-bp sliding window with 1-bp step not stated
Twofold half-life difference threshold (domain average vs. overall operon average) with ≥2 consecutive probe sites required Identification of operons with segmental RNA stability (RNase Y-processed mRNA candidates) Affymetrix tiling array with 17 probes per ORF at 14–23-base spacing; GEO: GSE40198 not stated
mRNA decay rate ('steepest slope' method) Calculation of mRNA half-life at subgene level from tiling microarray time-course data null not stated
RT-PCR (qualitative co-transcription confirmation) Validation of 10 randomly selected gene pairs predicted as co-transcribed by RNA-seq but not by ProOpDB 10 gene pairs na
Approaches that could also have been used
  • Differential mRNA abundance between WT and Δrny was assessed via visual R² correlation and fold-change criteria rather than a formal differential-expression framework
    Could also: A count-based differential expression tool such as DESeq2 or edgeR could also have been applied to the RNA-seq read counts — These tools model negative-binomial count variation and produce per-gene adjusted p-values with FDR control, which would provide a statistically calibrated list of transcripts altered by RNase Y deletion and would make thresholding decisions more reproducible across studies
  • Only two biological replicates were used per condition for RNA-seq
    Could also: Three or more biological replicates per condition is a common recommendation for RNA-seq experiments — Additional replicates improve variance estimation, increase statistical power in differential-expression models, and reduce the influence of outlier samples; with n = 2, within-group variance is unestimable in most count-based frameworks without pooling or borrowing strength across genes
  • Segmental mRNA stability differences were identified using a fixed twofold half-life threshold applied independently to each operon
    Could also: A permutation-based or bootstrap approach to derive empirical null distributions for within-operon half-life variance could also be applied — An empirical null would allow a data-driven threshold rather than a fixed twofold cutoff, potentially improving sensitivity for transcripts with moderate processing and providing an estimate of the false-discovery rate among the 15 candidates
  • Replicate consistency was reported as R² (coefficient of determination) between paired replicate expression profiles
    Could also: Spearman rank correlation or intraclass correlation coefficient (ICC) could also be used to assess replicate agreement — Spearman r is robust to the influence of highly expressed outlier genes that can inflate R², and ICC explicitly partitions within- vs. between-replicate variance, offering complementary views of reproducibility
  • Operon prediction accuracy was evaluated using sensitivity and specificity relative to ProOpDB, which was assumed to be 100% correct
    Could also: A held-out RT-PCR validation set with precision–recall analysis, or a leave-one-out comparison against multiple independent operon databases, could also be used — Treating a computational reference as ground truth conflates errors in the reference with errors in the new method; experimental RT-PCR confirmation (as was done for 10 gene pairs) or comparison against multiple databases would give a less assumption-dependent accuracy estimate
  • No multiple-testing correction was applied when screening all 865 operons for segmental stability to generate the candidate list of 15
    Could also: A Benjamini–Hochberg FDR procedure (or analogous permutation-FDR) over the family of operon-level half-life comparisons could also be applied — Screening hundreds of operons with a fixed fold-change threshold accumulates false positives in proportion to the number of tests; an FDR-controlled procedure would make the expected proportion of false discoveries explicit and comparable across studies
Software: Bowtie2 · Samtools · BEDTools (genomeCoverageBed) · CLC Genomics Workbench · Python (custom operon prediction and half-life scripts) · Affymetrix Power Tools · ProOpDB (Prokaryotic Operon DataBase)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
3
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

CP000829 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE40198 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE99533 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-29900693

Paper: Chen Z, Raghavan R, Qi F, Merritt J, Kreth J. Genome-wide screening of potential RNase Y-processed mRNAs in the M49 serotype Streptococcus pyogenes NZ131. MicrobiologyOpen 2018. PMID 29900693 · PMCID PMC6460267 · DOI 10.1002/mbo3.671.

Code: https://github.com/ZhiyunChen/RNA_Processing (commit c21ef82, master, last pushed 2014-06-10, Python 2, no license). Repo ships code only — NO data. The 8 input files listed in its README (wt*.csv, rny*.csv, Spy49_allGenes.csv, SpyGenome.txt, probehalfLife*.csv) must be regenerated from primary sources.

Data (two GEO series, both linked to this PMID):

  • GSE99533 — RNA-seq, Illumina HiSeq2000, 4 samples (WT×2, Δrny×2), SRA SRP108413 → runs SRR5638291 (wt1), SRR5638292 (wt2), SRR5638293 (rny1), SRR5638294 (rny2). Paired-end, ~55–61 M read pairs/sample. Feeds operon prediction.
  • GSE40198 — microarray (GPL11420, Affymetrix), mRNA half-lives. Feeds the half-life / RNase-Y-processed-mRNA analysis. (Originally from PMID 23543715; reused here.)

Reference genome: S. pyogenes NZ131, GenBank CP000829.1 / assembly GCA_000018125.1 (ASM1812v1, original Spy49_#### locus-tag annotation). 1 chromosome.

Pipeline-derived results

IN SCOPE (clearly specified, deterministic) — the 80%

R1. Operon prediction. Pipeline (paper Methods + repo): Bowtie2 --very-sensitive-local → SAM → samtools sort/index → BAM → bedtools genomeCoverageBed -d per-base depth → wt1/wt2/rny1/rny2.csv → repo geneBankConstruct.pyoperonBankConstruct.pyoperonBankRefine.py. Operon-grouping criteria (repo operonBankConstruct.operonJudge): two consecutive same-data genes are split into different operons if ANY of: avg-reads >4-fold different; an intergenic "dent" below 0.5× the lesser gene's avg read; opposite strand; or gap >100 bp.

Reported numbers to compare against (Results §3.1 / text):

  • 865 operons total
  • 491 monocistronic, 169 dicistronic, 102 tricistronic, 103 with ≥4 CDS
  • ProOpDB benchmark: 81% sensitivity, 71% specificity (external comparison; ProOpDB reference download is optional/secondary).
  • ">99% reads mapped to the genome" (mapping QC, easy side-check).

OPTIONAL / HARDER (the ~20%, less precisely specified) — attempt only if R1 lands

R2. RNase-Y-processed mRNA candidates (repo HLShiftFinder.py + RNaseYProcessedOperonFinder.py). Needs probe-level half-lives (probehalfLifeWT.csv, probehalfLifeRNY.csv) derived from the GSE40198 microarray rifampicin time-course — half-life fitting is NOT specified at the level needed for a clean rebuild, and the "segmental stability twofold" thresholds and the manual triage (80 → 29 → discard 14 → 15 final candidates, Table 2) include human judgement ("14 discarded as random signal variations"). Reported: 80 segmental stabilities in WT → 29 altered in Δrny → 15 final candidates (Table 2). Plan: report as analysed-but-not-fully-reproduced unless half-life inputs prove trivially regenerable; the manual discard step makes an exact 15 non-deterministic.

OUT OF SCOPE (wet-lab / manual / external — not attempted)

  • Northern-blot validation of folC1/speG/ypaA/prtF (wet lab).
  • qRT-PCR validation (wet lab), r²=0.996.
  • ProOpDB as ground-truth biology (external DB; the 81%/71% benchmark is a secondary check, not the core pipeline output).

Primary reproduction target

R1: 865 operons + the 491/169/102/103 monocistronic/di/tri/≥4 distribution, regenerated end-to-end from SRP108413 reads on the CP000829.1 genome with the authors' own (Python-2, unmodified) operon-calling code.

Figures / tables: Table
R1a
Reported
865 operons total
Reproduced
990
did not match
R1b
Reported
491 monocistronic
Reproduced
632
did not match
R1c
Reported
169 dicistronic
Reproduced
186
within tolerance
R1d
Reported
102 tricistronic
Reproduced
91
within tolerance
R1e
Reported
103 with >=4 CDS
Reproduced
81
did not match
R1f
Reported
>99% reads mapped
Reproduced
99.90-99.94%
exact
R1g
Reported
~1703 CDS gene universe
Reproduced
1703
exact
R2a-c
Reported
80 -> 29 -> 15 candidates (Table 2)
Reproduced
not attempted
m.public.grade.not-attempted

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 57/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is an incomplete reproduction, not a discrepant one: the authors' unmodified Python-2 operon-caller and a deterministic Bowtie2→samtools→bedtools pipeline were built and validated end-to-end (genome length 1,815,785 bp and 1,703 CDS confirmed), but the run (SLURM 2180633) was still executing at finalize, so the R1 numbers (865/491/169/102/103, >99% mapping) are all pending and never compared. The gaps are on our side (code-only repo forced self-assembly of public inputs; the run did not finish) plus one authors'-side ambiguity (the non-deterministic manual triage behind Table 2's 15 candidates, left out of scope). There is no fabrication signal — the figures are in-principle derivable from shipped code + public data — but with zero reproduced values produced, no claim can be confirmed, yielding a uniformly cautious yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

267.8 k
tokens (I/O) · 14.9 M incl. cache
84 min
runtime · 28.78 CPU-h
37.9 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine