Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Intra-Host Co-Existing Strains of SARS-CoV-2 Reference Genome Uncovered by Exhaustive Computational Search.

Viruses · 2023
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIALLY REPRODUCED. The study aimed to identify intra-host co-existing strains of SARS-CoV-2 using computational methods. The reproduction confirmed the exact number of raw reads reported in the paper, with 122,608,060 reads matching the original dataset (SRR11092062), equivalent to 2 x 61,304,030 ENA spots. However, the claims regarding the number of reads after human filtering (11,752,058), the identification of two co-existing strains with 179 and 89 differences, and the specific GATK HaplotypeCaller variant positions (306, 565, 17825) could not be independently verified. Thus, while the initial data acquisition was confirmed, the subsequent analyses and results concerning strain identification and variant calling remain unverified in this reproduction attempt.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 63
    assessed: 2026-06-19 ⛓ c4bc97c4454a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-06-30

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can an exhaustive computational search of the short-read RNA-sequencing data used to assemble the SARS-CoV-2 reference genome recover intra-host co-existing viral strains that de novo assemblers like MEGAHIT discard, and characterize their phylogeny and ACE2-binding properties?

Core claims
  • An exhaustive-search workflow can recover intra-host co-existing SARS-CoV-2 strains from the reference-genome read set (SRR11092062) that de Bruijn-graph assemblers discard. method
  • Two co-existing SARS-CoV-2 strains coexist in the viral sample (EPI_ISL_402124) used to produce the reference sequence, with differences evenly distributed across the genome. finding
  • The two co-existing strains are closely related to the SARS-CoV-2 reference strain and equally distant from major variants B.1.617 (Delta) and BA.1 (Omicron). finding
  • The two co-existing strains show different structural/binding properties toward the ACE2 receptor compared with the reference spike protein. finding
  • A double-model error-correction step (proportional correction via smallRNAproper plus Karect on singletons) improves read reliability while avoiding non-existent reads. method
  • The workflow generalizes to other RNA viruses, identifying within-host diversity in a California wastewater sample and in a foot-and-mouth disease virus (FMDV) sample. finding
  • The workflow identifies within-host diversity with high accuracy and sensitivity compared to the assembly tool Haploflow. finding
  • Public FASTA sequences and spike genes of the two discovered strains, plus difference lists, are produced as a resource. resource
Experimental setups
Assay System Perturbation Readout Platform
Short-read RNA-seq read extraction / human-read filtering SARS-CoV-2 patient BALF sample, Wuhan (SRR11092062) none non-human reads remaining after alignment filtering Bowtie2-2.44 (--large-index --local); GRCh38.p13 reference
Read error correction SARS-CoV-2 paired-end NGS reads (SRR11092062) none number of corrected reads smallRNAproper (miREC-based) + Karect
Read alignment / new-strain identification SARS-CoV-2 reference EPI_ISL_402124 and reads (SRR11092062, SRR11092057-64) in silico base substitution into reference validated substitutions / co-existing strain sequences Bowtie2 local (--score-min G,10,4 -D25 -R3 -N1 -L20 -i S,1,0.50)
Multiple sequence alignment and phylogenetic tree construction SARS-CoV-2 strains, VOCs/VOIs, SARS-CoV, MERS-CoV, bat coronaviruses (whole genome and spike) none phylogenetic relationships MAFFT v7.543; FastTree 2.1.11; PHYLIP dnapars v3.698
Protein structure prediction Spike protein amino acid sequences of co-existing strains none predicted 3D structure (pLDDT selection, Model 3 of 5) AlphaFold2 with MMseqs2 (mmseqs2_uniref_env, unpaired_paired); exPasy translation
Protein-protein docking / binding affinity Predicted spike Model 3 vs human ACE2 (PDB 6M1D); reference spike PDB 7CWL none docking energy score HDOCK1.1 (default parameters)
Within-host diversity identification (benchmark/generalization) California wastewater sample (SRR12596175) and FMDV sample (ERR034193) none identified co-existing strains vs Haploflow Haploflow (comparison)
Key results
  • Two co-existing SARS-CoV-2 strains identified from the reference read set SRR11092062
  • Co-existing Strain 1 showed 26 nucleotide differences in the spike glycoprotein region 26 nt differences
  • Co-existing Strain 2 showed 10 nucleotide differences in the spike region vs reference 10 nt differences
  • Strain 1 spike gene had 22 non-synonymous and 4 synonymous differences vs the original strain 22 non-synonymous, 4 synonymous
  • Both strains closely related to reference and equally distant from B.1.617 (Delta) and BA.1 (Omicron)
  • Proportional correction fixed 234,432 reads; Karect corrected only 35 reads to avoid non-existent reads 234,432 vs 35 reads
  • Five corrected reads (all frequency >1) were found in co-existing Strain 1 5 reads
  • BANAL-20-52 spike protein is 98.4% similar to the original SARS-CoV-2 spike protein 98.4% similarity
Key statistics
  • count 11,752,058 reads remained after filtering (non-human reads out of 122,608,060 raw paired-end reads)
  • count 122,608,060 reads in raw paired-end set (initial SRR11092062 read count before human filtering)
  • count 234,432 reads corrected by proportional correction (error correction of the 11,752,058 reads)
  • count 35 reads corrected using Karect (singleton correction avoiding non-existent reads)
  • count 179 observed differences (total differences reported for Strain 1 (text truncated))
  • other 98.4% similar (BANAL-20-52 spike vs original SARS-CoV-2 spike protein)
  • other error rate 0.5–2% (possible error rate of the paired-end NGS RNA-seq dataset)
  • count 37,678 new sequences out of 127,642 reported errors (Karect false-read generation cited from Zhang et al.)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational/bioinformatics methods paper rather than a hypothesis-testing study: it describes a multi-step pipeline (read filtering, dual error correction, iterative substitution-based strain reconstruction, phylogenetic tree building, and protein docking) applied to sequencing datasets to identify co-existing intra-host viral strains. Results are reported as counts of corrected reads, counts of nucleotide/amino-acid differences, phylogenetic tree topology, and docking energy scores, without classical inferential statistics (no p-values, confidence intervals, or formal hypothesis tests are reported in the available text).

Replicationunclear Sample sizeRead counts are given (e.g., 122,608,060 raw reads; 11,752,058 after human-read filtering; 234,432 reads corrected by proportional correction; 35 by Karect), but these are sequencing-read counts, not biological/technical replicate sample sizes for a statistical comparison GroupsDiscovered co-existing strains vs. the SARS-CoV-2 reference genome/spike protein, and vs. other coronaviruses/variants of concern in phylogenetic and structural comparisons Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Candidate strain-defining substitutions are validated using a custom coverage/read-support rule (requiring perfectly matched, non-conflicting reads and, in secondary samples, supporting reads or minimum segment length) rather than a formal statistical model.
    Could also: A probabilistic variant-calling framework (e.g., a Bayesian genotype-likelihood model such as GATK HaplotypeCaller/Mutect2, or a haplotype-reconstruction tool that outputs posterior probabilities or quality scores) could also be used — This would attach a quantitative confidence measure (e.g., a Phred-scaled quality score or posterior probability) to each candidate variant/strain call, which can help communicate the certainty of low-frequency or borderline substitutions alongside the rule-based approach used here.
  • Phylogenetic trees were built with FastTree and cross-checked with PHYLIP dnapars, and tree topology is interpreted directly (e.g., which strains are 'closer' to which variants).
    Could also: Bootstrap resampling or Bayesian posterior support values (e.g., via RAxML with bootstrap replicates, or MrBayes/BEAST for posterior clade support) could also be reported alongside the tree — Branch-support values quantify how robust each topological relationship is to sampling variation in the alignment, which complements a single best-tree estimate when discussing which lineages are more closely related.
  • Structural/binding comparisons between the discovered strains' spike proteins and ACE2 are summarized as single HDOCK docking energy scores per model.
    Could also: Reporting a distribution of scores across multiple docking runs or top-ranked poses (e.g., mean and range/CI across replicate docking predictions) could also be used — Since docking scores can vary with pose sampling, showing a spread of values in addition to the single top score would give readers a sense of the stability/uncertainty of the predicted binding affinity difference between strains.
  • The workflow's accuracy is described qualitatively (e.g., comparison with Haploflow output on a wastewater sample said to show 'high accuracy and sensitivity').
    Could also: Quantitative benchmarking metrics such as precision/recall/F1 or sensitivity/specificity against a defined truth set could also be reported — Numeric benchmarking metrics would let readers directly compare the new workflow's performance to Haploflow or other tools on a common quantitative scale.
  • Error-correction performance is reported as raw counts of corrected reads (234,432 by proportional correction; 35 by Karect) without a statistical comparison of error rates before and after correction.
    Could also: A before/after comparison of estimated per-base error rates (e.g., using a quality-score-based error-rate estimate with a confidence interval, or a McNemar-type test on corrected vs. uncorrected read concordance) could also be used — This would let the magnitude of error-rate reduction be expressed with a measure of uncertainty rather than as a raw count of reads affected.
Software: MEGAHIT · Bowtie2 2.44-linux · smallRNAproper (proportional error correction, based on miREC) · Karect · MAFFT 7.543 · FastTree 2.1.11 · PHYLIP dnapars 3.698 · AlphaFold2 (via MMseqs2) · HDOCK 1.1

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37243151

Paper: Cai X, Lan T, Ping P, Oliver B, Li J. Intra-Host Co-Existing Strains of SARS-CoV-2 Reference Genome Uncovered by Exhaustive Computational Search. Viruses 2023;15(5):1065. DOI 10.3390/v15051065 · PMID 37243151 · PMCID PMC10224212.

Code (authors' own, P16 not needed): https://github.com/RyanCairepo/strain_identify (master) Primary data: SRA SRR11092062 (WIV04, Wuhan BALF, the sample used to assemble the SARS-CoV-2 reference genome). Secondary: SRR12596175 (CA wastewater), ERR034193 (FMDV).

Pipeline (five steps, per Methods §2)

  1. Non-human read extraction — Bowtie2 2.4.4 --large-index --local vs GRCh38.p13; keep unaligned.
  2. Error correction — proportional correction (smallRNAproper, miREC-based) + Karect on singletons.
  3. New-strain identification — Bowtie2 --local vs WIV04 (EPI_ISL_402124) with --score-min G,10,4 -D 25 -R 3 -N 1 -L 20 -i S,1,0.50; custom substitution-matrix search (find_sub.sh → build_matrix/identify_strain/strain_finder); verify across the 4 sibling Wuhan samples (SRR11092057–61,63,64).
  4. Phylogenetics — MAFFT v7.543 + FastTree 2.1.11 + PHYLIP dnapars on genome & spike.
  5. Binding affinity — AlphaFold2(MMseqs2) structures + HDOCK docking vs ACE2.

IN SCOPE (pipeline-derived, attempt to reproduce)

id result reported §/loc tool
C1 raw read count SRR11092062 122,608,060 reads §2.1 SRA/ENA metadata (=2×61,304,030 spots)
C2 reads remaining after human filter 11,752,058 (from 122,608,060; 48.2→4.2 GB) §2.1 Bowtie2 large-index/local vs GRCh38.p13
C3 # co-existing strains + total diffs 2 strains; Strain1=179 diffs, Strain2=89 diffs §3.1.2 strain_identify find_sub.sh
C3a spike syn/nonsyn S1: 22 nonsyn+4 syn; S2: 9 nonsyn+1 syn §3.1.2 synonymous_stat.py
C4 verification read counts S1 195→88, S2 146→24 §3.1.2 verify.py over sibling samples
C5 GATK HaplotypeCaller variant positions 306, 565, 17825 §3.1.3 GATK4.2.5.0 HaplotypeCaller (3rd-party)
C6 error-correction counts 234,432 proportional + 35 Karect corrected §3.1.1 ErrorCorrection tools

Quick-minimum (80% floor): C1, C2 are clear & low-hanging. Then push C3/C5/C6.

OUT OF SCOPE (not attempted, with reason)

  • Phylogenetic tree topology claims (§3.1.4) — qualitative ("closer to WIV04"), not a pinnable numeric value; trees buildable but verdict is visual. Low priority.
  • Binding-affinity decreases 4.84% / 9.78% (§3.1.5) — AlphaFold2 + HDOCK are GPU-heavy, stochastic, model-selection-dependent (Model 3 of 5); not 1:1 reproducible.
  • 7CWL/6M1D PDB structures, exPASy/HDOCK manual steps — external/manual.
  • California wastewater (SRR12596175, §3.2) & FMDV (ERR034193, §3.3) — secondary datasets, SAME pipeline; attempt only if primary (SRR11092062) pipeline runs clean.

Datasets to profile (DATASET PROFILING pass)

  • SRR11092062 (primary) — done from metadata; QC + content checks at download time.
  • SRR12596175, ERR034193, sibling SRR11092057–64 — profile if reached.
C1
Reported
122,608,060 raw reads (SRR11092062)
Reproduced
122,608,060 (= 2 x 61,304,030 ENA spots)
exact
C2
Reported
11,752,058 reads after human filter
Reproduced
m.public.grade.uncheckable
C3
Reported
2 co-existing strains; Strain1=179 diffs, Strain2=89 diffs
Reproduced
m.public.grade.uncheckable
C5
Reported
GATK HaplotypeCaller variants at 306, 565, 17825
Reproduced
m.public.grade.uncheckable

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

298.3 k
tokens (I/O) · 23.8 M incl. cache
68 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.