Detecting DNA modifications from SMRT sequencing data by modeling sequence context dependence of polymerase kinetic.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Partial, honestly-bounded reproduction. The seqPatch R/Rcpp package (github.com/zhixingfeng/seqPatch) was recovered, compiled and run end-to-end (normalizeByMovie -> getFeaturesAlongGenome -> detectModification/hieModel) producing real, non-degenerate per-position kinetic statistics -- but on the package's own bundled example dataset (a same-study SISTER accession, SRX188834/35, not the RU-assigned SRX209633) since the RU's construct reference sequence was not available in any recovered source. Independently, raw per-base kinetics (IPD/PRE_BASE_FRAMES, pulse-width/WIDTH_IN_FRAMES) for the RU's actual assigned accession, SRX209633 (run SRR631046), were confirmed retrievable and were directly extracted via NCBI SRA-tools/VDB (bypassing the permanently-dead PacificBiosciences/R-pbh5 tool that the paper's own pipeline depended on) -- this corrects an initial, too-pessimistic read of ENA metadata (which alone suggested the raw data was gone). What was NOT attempted: end-to-end modification calling on SRX209633 itself against the paper's reported 19 known 4-mC sites, because that would require (a) the construct's exact reference sequence, (b) a BLASR alignment step, and (c) reconstructing pbh5's internal output structure by inference rather than by running the real tool -- doing so was judged a fabrication risk and was deliberately skipped rather than guessed at. The author's own companion demo package, which would have supplied the full worked E. coli/plasmid pipeline, is hosted on a now-dead personal host and could not be retrieved.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors hypothesized that because DNA polymerase kinetics (IPD) in SMRT sequencing are strongly determined by local sequence context, IPDs from positions sharing the same sequence context ('homologous positions'), including those in historical control/WGA data, can be pooled to better estimate the null IPD distribution and thereby improve DNA modification detection while reducing or eliminating the need for a matched control sample.
- ★ Local sequence context strongly determines position-specific polymerase kinetic rate: roughly 80% of IPD variation is explained by a 10 bp context (7 bases upstream, 2 bases downstream of the incorporation site), saturating at 7 bases upstream. finding
- ★ Sequence context effects are highly consistent across independent experiments and species (E. coli WGA vs M. pneumoniae WGA) when the same sequencing chemistry is used, but not across different chemistries. finding
- ★ An Empirical Bayesian hierarchical model combining IPDs across homologous positions (allowing heterogeneity in mean and variance, with a shared prior) plus a likelihood-ratio test detects kinetic variation events more accurately than the naive case-control method. method
- ★ The hierarchical model with control data increases sensitivity by 10%–30% at the same FDR relative to case-control when control coverage is low (15x–35x per strand); the advantage shrinks as control coverage increases. finding
- ★ A control sample can be omitted entirely for modifications with strong kinetic signal: the hierarchical model without control data is comparable to case-control for 6-mA but underperforms for 4-mC (weaker signal-to-noise). finding
- ★ In E. coli K-12, roughly 80% of 6-mA events in the GATC context were detected at 5% FDR using the hierarchical model with control data. finding
- Thousands of kinetic variation events were detected in E. coli K-12 at positions without previously described methylation motifs, suggesting more extensive modification patterns than previously observed. finding
- The method is implemented in an open-source R package, seqPatch (https://github.com/zhixingfeng/seqPatch); a Box-Cox transformation plus per-movie centering normalization is applied to IPDs to control outliers and movie-level batch effects. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| SMRT sequencing (single molecule real time) with inter-pulse duration (IPD) kinetic measurement, native vs. WGA control (case-control design) | Plasmid DNA (3,589 nt) from an E. coli strain engineered to methylate the 4th carbon of cytosine in GATC contexts (4-mC set) | engineered methyltransferase strain (4-mC modification) vs. whole-genome-amplified control (modifications erased) | per-strand per-position IPD distributions; detection of 19 known 4-mC sites; ROC/sensitivity at matched FDR | Pacific Biosciences RS, FCR chemistry, zero-mode waveguide SMRTcells, CCD camera at 100 Hz |
| SMRT sequencing with IPD kinetic measurement, native vs. WGA control | Plasmid DNA (3,591 nt) from an E. coli strain engineered to methylate adenine in GATC contexts (6-mA set) | engineered methyltransferase strain (N6-methyladenine) vs. WGA control | per-position IPD distributions; detection of 23 known 6-mA sites; ROC curves at single-strand coverage 15x/20x/25x | Pacific Biosciences RS, FCR chemistry |
| SMRT sequencing of whole-genome-amplified DNA for sequence-context/kinetics regression (MART non-linear tree-based regression) | E. coli K-12 MG1655 WGA DNA (E. coli WGA-FCR, 4,639,675 nt, 8x per strand); positions with single-strand coverage >35 reads | whole-genome amplification (erases modifications) | position-specific kinetic rate (mean Box-Cox transformed IPD) as response; R^2 variance explained by sequence context windows | Pacific Biosciences RS, FCR chemistry; MART regression |
| SMRT sequencing of WGA DNA, cross-experiment context-effect comparison | E. coli K-12 WGA-C (12x/13x) vs. M. pneumoniae WGA-C2 (816,394 nt, 40x) | whole-genome amplification; none other | average Box-Cox transformed IPD per [-7,+2] sequence context ('context effect'); Pearson correlation between experiments | Pacific Biosciences RS, C2 chemistry |
| SMRT sequencing, medium-coverage genome-wide modification detection (hierarchical model with control, FDR-controlled calling) | Wild-type E. coli K-12 MG1655 native DNA (12x per strand) with WGA-N (12x) and WGA-C (13x) controls | none (native genomic methylation) vs. WGA controls | fraction of GATC 6-mA sites detected at 5% FDR; number of kinetic variation events at non-canonical motifs | Pacific Biosciences RS, C2 chemistry |
| In silico read-mixing (partial modification simulation) on SMRT IPD data | Artificial native samples built from the plasmid 4-mC and 6-mA datasets | modification proportion set to 50%, 70%, 90% by sampling reads from native vs. control data | detection accuracy (ROC) as a function of modification proportion and native-sample coverage | — |
| Distributional diagnostics of Box-Cox transformed IPDs (skewness/kurtosis) and per-movie normalization assessment | SMRT datasets across FCR and C2 chemistries; 45 movies of E. coli WGA-FCR | none | skewness centered at 0, kurtosis centered at 3; movie-to-movie shifts in IPD distribution | Pacific Biosciences RS (FCR and C2 chemistries), CCD at 100 Hz, up to 150,000 pulse streams per movie |
- – Sequence context of 7 bases upstream and 2 bases downstream explains most position-specific kinetic rate variation; R^2 saturates at 7 bases upstream and bases >2 downstream add nothing ~80% of IPD variation explained by 10 bp context
- ▲ Context effects computed independently in E. coli WGA and M. pneumoniae WGA are highly concordant Pearson r = 0.91 (80% of IPD variation in one set explained by the other)
- ▲ Hierarchical model with control data increases detection sensitivity over case-control at low control coverage on both plasmid datasets 10%–30% sensitivity increase at the same FDR (control coverage 15x–35x per strand)
- – Hierarchical model without control data matches case-control accuracy for 6-mA but is worse for 4-mC
- ▲ Detection of 6-mA in GATC contexts in wild-type E. coli K-12 at medium coverage ~80% of GATC 6-mA events detected at 5% FDR
- – Detection performance using [-6,+1] contexts is similar to [-7,+2], and [-6,+1] was adopted because ~1/3 of positions lacked a matching [-7,+2] context in historical data ~1/3 of positions missing [-7,+2] context
- ▼ Detection accuracy declines as modification proportion falls (90%→70%→50%) but calls are still made; increasing native-sample coverage restores accuracy comparable to fully modified positions
- ▲ A single modified nucleotide perturbs IPDs at flanking bases (e.g. modification at plasmid position 342 elevates the native-vs-control average IPD ratio at position 346)
- correlation Pearson's correlation coefficient is 0.91 (context effect ([-7,+2]) consistency between E. coli WGA-C and M. pneumoniae WGA-C2)
- other roughly 80% (proportion of IPD variation explained by a 10 bp sequence context (7 upstream, 2 downstream); also the fraction of variation in one WGA experiment's context effects explained by the other)
- other roughly 80% detected at 5% FDR (6-mA events in GATC context in E. coli K-12 using hierarchical model with control data)
- other 10%–30% sensitivity increase at same FDR (hierarchical model vs. case-control when control single-strand coverage is 15x–35x)
- count 19 known 4-methylcytosines (modified sites in the 3,589 nt plasmid m4C dataset; tested at 35x, 50x, 65x single-strand coverage)
- count 23 known N6-methyladenines (modified sites in the 3,591 nt plasmid m6A dataset; tested at 15x, 20x, 25x single-strand coverage)
- count 45 movies (E. coli WGA-FCR movies showing large overall between-movie shifts in IPD distribution)
- other 752x, 1557x, 186x, 1486x, 15x, 8x, 12x, 12x, 13x, 40x (per-strand coverage of the 10 samples in Table 1 (plasmid m4C native/control, plasmid m6A native/control, M. pneumoniae WGA-FCR, E. coli WGA-FCR, E. coli native, E. coli WGA-N, E. coli WGA-C, M. pneumoniae WGA-C2))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper is a computational/methods paper describing a new detection method (an empirical Bayesian hierarchical model) for identifying DNA base modifications from SMRT sequencing polymerase kinetic data (inter-pulse durations, IPDs). Method performance is assessed primarily by comparing the new hierarchical model to a case-control approach using ROC curves, sensitivity at fixed False Discovery Rate (FDR) thresholds, and a nonlinear regression (MART) R² measure of how much sequence context explains kinetic variation, rather than through classical hypothesis-testing p-values on the excerpted portion of the text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Likelihood ratio (within an empirical Bayesian hierarchical model) | Detecting kinetic variation / DNA modification events at each genomic position, comparing native (case) IPDs to control and/or historical homologous-position IPDs | Varies by dataset and position; single-strand coverage examples given include 15–65x for plasmid tests and ~8–40x for E. coli/M. pneumoniae WGA datasets (Table 1) | stated |
| MART (non-linear tree-based regression), evaluated via proportion of variance explained (R²) | Relationship between position-specific kinetic rate (response) and local sequence context (predictor) in E. coli K-12 WGA data | Positions with single-strand coverage greater than 35 reads in the E. coli WGA-FCR dataset | not stated |
| Pearson's correlation coefficient | Consistency of the [-7,+2] sequence-context effect between two independent WGA experiments (E. coli WGA-C and M. pneumoniae WGA-C2) | — | not stated |
-
Exponentially distributed IPD values were converted toward approximate normality using a Box-Cox transformation before model fitting.↳ Could also: A generalized linear model with an exponential or gamma family/link, or a nonparametric rank-based approach, could also be used on the untransformed IPD data. — Modeling the data's native exponential-family distribution directly can avoid the need to choose and validate transformation parameters, and a rank-based method offers a distribution-free way to compare IPD values.
-
Detection performance was summarized with ROC curves and sensitivity reported at a fixed FDR threshold (e.g., 5%).↳ Could also: Precision-recall curves could also be reported alongside ROC curves. — When true modification events are a small fraction of all tested positions, precision-recall curves can convey additional information about false-positive burden that complements ROC-based summaries.
-
Consistency of the sequence-context effect between two independent experiments was quantified with a single Pearson's correlation coefficient (r = 0.91).↳ Could also: A confidence interval around the correlation estimate, or a Spearman rank correlation, could also be reported. — A confidence interval communicates the precision of the correlation estimate, and a rank-based correlation can serve as a complementary check that the relationship is not driven disproportionately by a few extreme points.
-
Multiple-position testing across the genome was managed using an FDR threshold without naming the specific correction procedure in the excerpted text.↳ Could also: Explicitly citing the specific FDR procedure used (e.g., Benjamini-Hochberg), or additionally reporting a family-wise error control such as Bonferroni for a high-confidence subset, could also be done. — Naming the exact multiplicity-control algorithm supports reproducibility, and a family-wise approach can serve as an additional, more conservative cross-check for the strongest candidate calls.
-
The dependence of kinetic rate on sequence context was quantified using MART, a flexible tree-based regression method, via its R² value.↳ Could also: A cross-validated R² or a parametric regression model with explicit interaction terms could also be used. — Cross-validation helps confirm that the variance-explained estimate generalizes to unseen positions rather than reflecting in-sample fit of a flexible model.
-
IPD distributions across positions and movies were visualized using boxplots rather than accompanied by explicit numeric dispersion statistics.↳ Could also: Reporting SD, SEM, or a 95% CI numerically alongside the boxplots could also be done. — Numeric dispersion measures let readers directly compare variability across conditions without needing to visually estimate spread from the plots.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
What deviates: no reported value from the paper was ever compared against a reproduced value. The hierarchical model ran end-to-end and produced sane, non-degenerate output (3591 positions x 2 strands, t.stat in [-7.12, 5.26], ipd.ratio median 0.970) — but on the repo-bundled m6A sister construct (SRX188834/35), not the RU-assigned 4-mC accession SRX209633, because that construct's reference sequence and its paired control SRX209634 were never located. Whose side: mixed. Authors' side contributes genuine link rot — PacificBiosciences/R-pbh5 is 404 on three independent checks and the demo tarball's host xfengz02.u.hpc.mssm.edu is dead — but the authors' data deposition holds up: prefetch SRR631046 still returns 611,080,898 bytes of native PacBio HDF5 with per-base PRE_BASE_FRAMES/WIDTH_IN_FRAMES, correcting the initial ENA-only pessimism. The decision not to attempt BLASR alignment or to reconstruct pbh5's alnsF/alnsIdx by inference is our methodological choice — a correct anti-fabrication call, not an authors' defect. Severity: moderate and non-suspicious. Nothing here is fabrication-shaped; the code is real, runs, and the data exists. The paper's headline claim (19 known 4-mC sites, coverage-dependent ROC) is simply untested, which caps q7 at limited confirmation and q8 at a solid, honestly-bounded partial rather than red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.