Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comparative Genomics of Listeria monocytogenes Isolates from Ruminant Listeriosis Cases in the Midwest United States.

Microbiol Spectr · 2022
L1 91/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
91/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 82% of all assessed papers rank 197 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. RU scoped to one isolate, TB0359 = sample SAMN24611040 (Illumina SRR17426669 + Nanopore SRR17430284). The paper assembled this isolate with Unicycler (hybrid) -> assignment correct, and we ran exactly that pipeline on «our HPC» («job», 13 min). Hybrid Unicycler v0.5.1 assembly = 2,904,175 bp, GC 38.04%, one dominant 2.6 Mb chromosomal contig (within-tol for a ~3.0 Mb L. monocytogenes genome; paper gives no per-isolate ground-truth). All deterministically-reproducible typing claims reproduce 1:1: MLST ST1 (=> CC1 = SL1, lineage I; C2+C3 exact), molecular serogroup 4b (LisSero; C4 exact), LIPI-1 present (prfA/plcA/hly/mpl/plcB full coverage; C5a exact), fosX present at 100%/100% (abricate NCBI; C5b exact). cgMLST CT8674 (C6) NOT attempted (curated Pasteur label). Dataset N reported == N observed exactly for both runs (413,782 pairs; 17,861 reads). Out of scope: 73-isolate aggregate stats, wet-lab metadata. No fabrication concern: every value is directly derivable from the shipped SRA data.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 6b2d89ed831d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study characterizes the genomic diversity of Listeria monocytogenes isolates from ruminant listeriosis cases in the Midwest/Upper Great Plains United States, testing whether phylogenetic lineage/clonal complex is associated with clinical manifestation (neurologic vs. fetal infection) and with virulence, stress, and antimicrobial resistance gene content.

Core claims
  • 73 ruminant listeriosis isolates classified by WGS/cgMLST fall into three lineages: 31.5% lineage 1, 53.4% lineage 2, 15.1% lineage 3 finding
  • Lineage 1 and 3 isolates are more strongly associated with neurologic infections, while lineage 2 isolates show a greater frequency of fetal infections finding
  • Mobile genetic elements, virulence genes, and stress/antimicrobial resistance genes underlie subgroup-specific features and may facilitate spread of hypervirulent clones such as CC1 mechanism
  • Whole-genome sequencing and cgMLST typing were used to classify and compare isolates into lineages, sublineages, and cgMLST types method
  • All isolates carry the core LIPI-1 virulence genes (prfA, plcA, plcB, actA, mpl, hly) finding
  • LIPI-3 (listeriolysin S) is present in nearly all lineage 1 isolates but only rarely in lineage 3 and absent from lineage 2 finding
  • LIPI-4 gene cluster is present in 13.7% of isolates, spanning lineages 1, 2, and 3 finding
  • Stress survival islet 1 (SSI1) and arginine deiminase genes (arcC/arcD) show lineage-specific distribution, being entirely absent in lineage 3 finding
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome sequencing (short- and long-read) L. monocytogenes isolates from cattle, sheep, and goats none genome sequence data for typing
core genome MLST (cgMLST) typing L. monocytogenes isolates (n=73) none lineage, sublineage (SL), and cgMLST type (CT) classification Institut Pasteur MLST database
Minimum spanning tree / phylogenetic clustering L. monocytogenes isolates and 3 reference genomes none clonal relatedness and clustering by CT/SL GrapeTree
Virulence gene presence/absence screening L. monocytogenes isolates none presence/absence of LIPI-1, LIPI-3, LIPI-4, and 10 internalin genes (inlAB, inlC, inlE, inlF, inlG, inlH, inlJ, inlK, inlP)
Stress and antimicrobial resistance gene screening L. monocytogenes isolates none presence/absence of cold/osmotic tolerance genes (csp, gbuABC, betL, opuCAB), general stress genes (yugI, ctc, ydaG), acid tolerance genes (glutamate decarboxylase system, arcA/B/C/D), SSI1, SSI2, sanitizer/heavy metal resistance genes (bcrA, LGI-2, LGI-3)
Statistical association analysis isolate metadata (lineage vs. clinical manifestation) none significance (P value) of association between lineage and neurologic vs. fetal clinical manifestation
Key results
  • 73 isolates distributed across lineages: 23 (31.5%) lineage 1, 39 (53.4%) lineage 2, 11 (15.1%) lineage 3
  • 91.3% of lineage 1 and 90.9% of lineage 3 isolates associated with neurologic infections vs. 51.3% of lineage 2; fetal infections significantly higher in lineage 2 (35.6%) than lineages 1/3
  • All isolates (100%, 73/73) harbored complete LIPI-1 (prfA, plcA, plcB, actA, mpl, hly) 73/73
  • LIPI-3 present in all lineage 1 isolates except TB0656, and in a single lineage 3 isolate (TB0508)
  • LIPI-4 identified in 13.7% (10/73) of isolates across lineages 1, 2, and 3 10/73
  • SSI1 found in 21.7% of lineage 1 and 51.3% of lineage 2 isolates, absent in lineage 3
  • arcC and arcD (arginine deiminase system) present only in lineage 1 and 2, absent in all lineage 3 isolates
  • Among 58 cgMLST types identified, 47 (81%) were unique to a single isolate; 11 (19%) shared allelic similarity across 2-5 isolates 47/58 (81%)
Key statistics
  • count 73 isolates total (total ruminant listeriosis isolates sequenced)
  • fold_change 23/73 (31.5%) lineage 1, 39/73 (53.4%) lineage 2, 11/73 (15.1%) lineage 3 (lineage distribution)
  • pvalue P < 0.05 (neurologic infections significantly more frequent than fetal infections overall)
  • pvalue P < 0.05 (frequency of clinical manifestations differed significantly within each lineage)
  • count 58 cgMLST types (CTs) identified; 47 (81%) unique to a single isolate (genomic diversity/CT distribution)
  • count 10/73 (13.7%) isolates carried LIPI-4 (LIPI-4 prevalence)
  • fold_change 21.7% lineage 1 vs. 51.3% lineage 2 carried SSI1 (SSI1 prevalence by lineage)
  • count 47/73 (64.4%) cattle, 17/73 (23.3%) sheep, 9/73 (12.3%) goats (host species distribution of isolates)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is a descriptive genomic/epidemiological comparison of 73 Listeria monocytogenes isolates from ruminant listeriosis cases, characterizing lineage, sublineage, cgMLST type, and virulence/stress/AMR gene content, and testing whether clinical manifestation (neurologic vs. fetal vs. other) differs in frequency by lineage. Results are reported primarily as counts and percentages, with statistical significance for group-frequency comparisons denoted by p<0.05. The specific statistical test(s), software, and correction methods used to generate these p-values are not stated in the visible text.

Replicationunclear Sample sizetotal isolate collection size (n=73) and per-lineage/sublineage counts are given as the basis for percentages; no power calculation or sample-size justification is described Groupsclinical manifestation (neurologic, fetal infection, other) frequency across three phylogenetic lineages Pairingunpaired Randomization/blindingna Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Unspecified statistical test for comparing proportions (p<0.05 reported) comparison of neurologic vs. fetal (vs. other) clinical manifestation frequencies overall and between lineages 1, 2, and 3 73 isolates total (23 lineage 1, 39 lineage 2, 11 lineage 3) not stated
Approaches that could also have been used
  • Differences in clinical manifestation frequency across lineages are reported as significant using a threshold (p<0.05) without naming the specific test.
    Could also: A chi-square test of independence or Fisher's exact test (the latter well suited to smaller cell counts, such as lineage 3's n=11) could also be used and explicitly named for comparing categorical clinical manifestation frequencies across lineages. — Naming the specific test and confirming that expected cell counts meet chi-square assumptions (or opting for Fisher's exact test when counts are small) helps readers evaluate the appropriateness of the comparison method for this categorical data.
  • Several pairwise and multi-group comparisons of clinical manifestation frequency between lineages are presented as significant.
    Could also: An omnibus test (e.g., chi-square across all three lineages) followed by a post-hoc correction for multiple comparisons (e.g., Bonferroni or Benjamini-Hochberg FDR) could also be applied when performing several pairwise contrasts. — Explicitly controlling for multiplicity across the several lineage/manifestation comparisons performed would provide additional assurance against inflated family-wise error rate from repeated testing on overlapping data.
  • Results are reported using exact counts and percentages along with significance labels (p<0.05) rather than exact p-values.
    Could also: Reporting exact p-values and/or effect size measures (e.g., odds ratios or Cramér's V with 95% confidence intervals) could also be included alongside the counts and percentages. — Exact p-values and effect sizes with confidence intervals give readers a sense of both the strength and precision of an association, beyond a binary significant/non-significant threshold.
  • The comparative genomic and gene-presence/absence data (e.g., internalin genes, LIPI-3/4, stress survival islets) are described narratively with percentages by lineage/sublineage without inferential statistics.
    Could also: Formal association testing (e.g., Fisher's exact test) between gene presence/absence and lineage, sublineage, or clinical manifestation could also be applied to these categorical genomic features. — Statistical testing of gene-presence associations would allow quantification of whether apparent lineage-specific patterns (e.g., inlG absence in lineage 1) exceed what might be expected by chance, complementing the descriptive percentages already provided.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36314928

Paper: Cardenas-Alvarez et al. (2022) Comparative Genomics of Listeria monocytogenes Isolates from Ruminant Listeriosis Cases in the Midwest United States. Microbiol Spectr. PMID 36314928 · PMCID PMC9769944 · DOI 10.1128/spectrum.01579-22.

Assigned code: https://github.com/rrwick/Unicycler (hybrid assembler). Assigned data: sra:SRR17426669.

What this RU actually is

The room is scoped to one isolate, not the whole 73-isolate study.

  • SRR17426669 → sample SAMN24611040 = isolate "LM TB0359", BioProject PRJNA794134 — Illumina MiSeq, paired 2×250, 413,782 read pairs, 207,718,564 bp (~69× for a ~3.0 Mb genome).
  • The SAME sample has a second run SRR17430284 = Oxford Nanopore MinION, 17,861 reads, 110,717,476 bp (~37×). Table 1 lists both accessions for TB0359.
  • Therefore TB0359 is one of the paper's 24 hybrid-assembled genomes, which the Methods say were assembled with Unicycler (short + long read). The Unicycler assignment is correct; reproduction = hybrid Unicycler assembly of SRR17426669 (Illumina) + SRR17430284 (Nanopore).

Note: the room's data.json only listed the Illumina run. A hybrid Unicycler run REQUIRES the Nanopore mate SRR17430284; both are added to the manifest.

Reported values for TB0359 (Table 1) — comparison targets

Field Reported
Lineage 1
Sublineage (SL) SL1
cgMLST type (CT) CT8674 (reported as a new CT)
Source Bovine
Clinical manifestation Neurologic
Year 2015
State ND
Serotype IVb (= molecular serogroup 4b)
Accessions SRR17426669 (Illumina), SRR17430284 (Nanopore)

Per-isolate assembly statistics (genome size / #contigs / N50) are NOT reported in the paper or its supplement — only an aggregate depth range (35×–70×) and the assembler choice. So assembly metrics can be produced and sanity-checked (expected L. monocytogenes chromosome ≈ 2.9–3.0 Mb, hybrid → likely 1 circular contig) but cannot be compared 1:1 to a paper number.

In scope (pipeline-derived, reproducible for TB0359)

# Result Pipeline Reported target
C1 Hybrid genome assembly Unicycler (Illumina+Nanopore) qualitative: ~3.0 Mb, ideally 1 circular chromosome; depth in 35–70× band
C2 MLST sequence type mlst / BIGSdb Pasteur scheme SL1 ⇒ expect ST1 / CC1
C3 Lineage derived from ST/CC (or in-silico) Lineage 1
C4 Serogroup/serotype in-silico (LisSero / PCR-serogroup) IVb ⇒ serogroup 4b
C5 Virulence-gene content ABRicate vs VFDB LIPI-1 present (all 73), fosX present (100%); lineage-1 marker LIPI-3 expected present

Out of scope / not attempted (and why)

  • cgMLST CT number (CT8674): assigned by the curated BIGSdb-Lm (Institut Pasteur) server. The literal CT integer is a database-curation artifact, not deterministically reproducible from raw data offline; the underlying allelic profile could in principle be computed (chewBBACA + Pasteur scheme) but the exact "CT8674" label cannot be re-derived. Recorded as partial/uncheckable.
  • All 73-isolate aggregate results (lineage distribution 31.5/53.4/15.1%, 58 CTs, prophage/plasmid prevalence, clinical associations): out of scope — this room has one isolate's data only.
  • Wet-lab / epidemiological metadata (host, year, state, clinical manifestation): not pipeline-derived → out of scope.

Fidelity caveats

  • Paper assembler version/params: "Unicycler" (no version pinned in text); short-read-only isolates used SPAdes 3.15.2. We pin a Unicycler version at run time and record it in environment.lock.
  • No per-isolate ground-truth assembly numbers → C1 graded on biological plausibility + completeness, not exact-match.
Figures / tables: Table
C1
Reported
Hybrid Unicycler assembly of TB0359 (paper gives no per-isolate size; L. monocytogenes ~2.9-3.0 Mb)
Reproduced
2,904,175 bp; GC 38.04%; 13 contigs incl. 1 dominant 2,599,351 bp chromosomal contig; N50 2,599,351
within tolerance
C2
Reported
Sublineage SL1 => MLST ST1 / CC1
Reproduced
MLST ST1 (scheme listeria_2)
exact
C3
Reported
Lineage 1 (lineage I)
Reproduced
ST1/CC1 => lineage I
exact
C4
Reported
Serotype IVb (molecular serogroup 4b)
Reproduced
LisSero serogroup 4b (4b,4d,4e; PRS/ORF2110/ORF2819 FULL, LMO0737/LMO1118 NONE)
exact
C5a
Reported
LIPI-1 present
Reproduced
prfA/plcA/hly/mpl/plcB all detected at full coverage (vfdb)
exact
C5b
Reported
fosX present (100%)
Reproduced
abricate NCBI fosX 100% coverage / 100% identity
exact
C6
Reported
cgMLST CT8674 (new CT)
Reproduced
NOT ATTEMPTED (Pasteur BIGSdb-Lm curated label, not deterministically reproducible offline)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 91/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

65.6 k
tokens (I/O) · 2.5 M incl. cache
20 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.