Complete Genome Sequencing of Lactobacillus plantarum ZLP001, a Potential Probiotic That Enhances Intestinal Epithelial Barrier Function and Defense Agai
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the core result, partially for the rest. Reproduced 1:1 the genome-level pipeline claims by de-novo assembling the paper's own PacBio reads (SRR8079316) with Flye: chromosome size 3,164,344 bp vs reported 3,164,369 (Δ25 bp), GC 44.66% vs 44.65%, chromosome rRNA 16/16 exact, tRNA 73 vs 69, subread count 73,238 exact. Chromosome size+GC independently CONFIRMED genuine against the deposited assembly GCA_003076435.1/CP021086.1 (3,164,369 bp, 44.66%) -> not fabricated. Two honest discrepancies flagged: (1) paper's 'median subread length 8,661 bp' is actually the MEAN (true median 9,289) - a mislabel, value real; (2) Flye from PacBio-only recovered 5 of 7 plasmids (missed the 67.8 kb plasmid A + one ~15 kb), so total CDS 3,131 vs 3,264 and plasmid GC differ - an assembly-completeness limit of our scoped pipeline, not a paper error. NOT ATTEMPTED (hard 20%): the named code's RAxML phylogeny on 553 single-copy orthologs - the paper specifies neither the comparison genomes nor ortholog-calling parameters, so inputs are unspecified and it is not reproducible from the methods as written; also the 16S MEGA tree and KEGG/COG counts. Did NOT run the Illumina hybrid-polish/closure step.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 68assessed: 2026-06-16 ⛓ ffaebfab3389
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhat genomic features underlie the probiotic and gut-health-promoting properties of Lactobacillus plantarum ZLP001, a strain isolated from healthy weaned piglet gut, and how does its genome compare with other L. plantarum strains?
- ★ The complete genome of L. plantarum ZLP001 comprises a single 3,164,369 bp circular chromosome (GC 44.65%) plus seven plasmids (A–G), encoding 3,264 protein-coding sequences. resource
- ★ ZLP001 carries genes (Na+:H+ antiporter, choloylglycine hydrolase, heat shock proteins, chaperones) supporting tolerance to low pH, bile salt, and stress in the gastrointestinal environment. finding
- ★ ZLP001 harbors an expanded antioxidative gene repertoire (glutathione, thioredoxin, catalase, NADH oxidase/peroxidase systems) but lacks superoxide dismutase, consistent with its high antioxidant ability. mechanism
- ★ ZLP001 contains 119 CAZyme genes across five families, more than L. plantarum KLDS1.0391, suggesting probiotic potential for pathogen defense and immune stimulation. finding
- ★ Phylogenomic analysis of 19 L. plantarum strains places ZLP001 close to BDGP2, JDM1, and LZ95 but on a relatively standalone branch, indicating distinct genomic adaptation to the gut. finding
- ★ The 19 L. plantarum genomes have an open pan-genome of 6,598 orthologous gene families and a core genome of 596 families, with 65 genes unique to ZLP001. finding
- The genome was sequenced de novo using PacBio SMRT single-molecule real-time sequencing followed by assembly and multi-database functional annotation. method
- ZLP001 possesses a transport/secretion repertoire of 306 transport genes (PTS, ABC), a Sec-SRP system, two intact prophages, and 22 CRISPR loci. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome sequencing (SMRT/PacBio) | Lactobacillus plantarum ZLP001 (from weaned piglet gastrointestinal mucosa) | none | complete genome sequence (chromosome + plasmids) | PacBio RS II, C4 chemistry, P6 polymerase; 8–12 kb library |
| De novo genome assembly | L. plantarum ZLP001 sequence reads | none | assembled contigs/genome | SOAPdenovo v2.04; SMRT Analysis v2.3.0 |
| Gene prediction and functional annotation | L. plantarum ZLP001 genome | none | CDS counts, COG and KEGG functional category assignments | Glimmer v3.02; BLAST vs COG/KEGG |
| Non-coding RNA and repeat prediction | L. plantarum ZLP001 genome | none | rRNA, tRNA, sRNA, tandem/mini/microsatellite counts | rRNAmmer v1.2, tRNAscan v1.23, Rfam v10.1, Tandem Repeat Finder v4.04 |
| Mobile genetic element analysis (prophage/CRISPR) | L. plantarum ZLP001 genome | none | prophage regions, integrases, CRISPR/Cas loci | PHAST; MinCED 3 |
| CAZyme annotation | L. plantarum ZLP001 genome | none | carbohydrate-active enzyme gene family counts | CAZy database |
| Comparative/phylogenomic and ortholog clustering analysis | 19 L. plantarum complete genomes (ZLP001 + 18 from NCBI) | none | core/pan-genome size, single-copy orthologous families, phylogenetic trees | OrthoMCL v2.0, BLASTP, MCL; MAFFT v7, RAxML, MEGA |
| Genomic DNA extraction | L. plantarum ZLP001 culture (MRS broth, 37°C, 18 h, microaerophilic) | none | purified total genomic DNA | Wizard Genomic DNA Purification Kit (Promega) |
- – Circular chromosome of 3,164,369 bp at 44.65% GC, occupying 83.28% of the genome with 3,104 chromosomal genes (avg 886 bp) 3,164,369 bp; 44.65% GC
- – Seven plasmids (A 67,802; B 48,418; C 31,389; D 27,860; E 16,139; F 15,258; G 13,837 bp) with average GC 42.05% 7 plasmids; avg GC 42.05%
- – Pan-genome of 6,598 orthologous families and core genome of 596 families (9.03%); core constitutes 19.59% of each genome pan 6,598; core 596 (9.03%)
- – 65 genes unique to ZLP001 (including 30 hypothetical proteins); strain-specific genes across 19 strains totaled 2,597 (39.36%) 65 unique; 2,597 (39.36%) strain-specific
- ▲ ZLP001 contains 119 CAZyme genes (18 CE, 13 CBM, 32 GT, 50 GH, 6 AA), more than KLDS1.0391 (14 CE, 21 CBM/CPM, 23 GT, 34 GH, 2 AA) 119 vs 94 genes
- – 1,603 CDSs assigned to 39 KEGG categories; 1,783 CDSs assigned to 20 COG categories 1,603 KEGG; 1,783 COG
- – 306 transport-related genes (54 PTS, 252 ABC), 10 complete PTS EII complexes; two intact prophages; 22 CRISPR loci 306 transport; 22 CRISPR loci
- – ZLP001 16S rRNA shows >99% similarity to other L. plantarum strains; clusters near BDGP2, JDM1, LZ95 on a standalone branch; 553 single-copy orthologous families >99% similarity; 553 single-copy families
- count 3,164,369 bp chromosome (ZLP001 chromosome size)
- count 3,264 protein-coding sequences (total CDSs in ZLP001 genome)
- other 44.65% GC (chromosome GC content)
- count 6,598 orthologous gene families (pan-genome) (pan-genome across 19 L. plantarum strains)
- count 596 (9.03%) orthologous families (core-genome) (core genome across 19 strains)
- correlation >99% 16S rRNA gene similarity (ZLP001 vs other L. plantarum strains)
- count 73,238 subreads, median length 8,661 bp (PacBio sequencing output used for assembly)
- count 119 CAZyme genes (five CAZyme families in ZLP001 genome)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genome data report describing the complete genome sequencing, assembly, annotation, and comparative/phylogenomic analysis of Lactobacillus plantarum ZLP001 against 18 other publicly available L. plantarum genomes. The work is descriptive and bioinformatic rather than experimental: results are reported as genome counts, gene-family memberships (core/pan/strain-specific), and phylogenetic relationships, with no inferential hypothesis testing of group differences. Tree reliability is the only place a statistical/resampling procedure is reported, via bootstrap support values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Neighbor-joining phylogenetic reconstruction with bootstrap resampling (1,000 replicates) | 16S rRNA gene tree of L. plantarum and other Lactobacillus (Figure 1C) | 1,000 bootstrap replications | not stated |
| Maximum-likelihood phylogenomic reconstruction (RAxML) with bootstrap resampling (1,000 replicates) | Concatenated single-copy orthologous gene-family tree of 19 strains (Figure 1D) | 1,000 bootstrap repetitions | not stated |
| BLASTP homology search with cutoffs (E-value 1e−5, percent match ≥ 50%) followed by MCL clustering (inflation 1.5) | OrthoMCL ortholog clustering; core/pan-genome and strain-specific gene determination (Figure 1E, Tables S1–S2) | 19 genomes | na |
-
Genome assembly was performed with SOAPdenovo, a short-read de Bruijn graph assembler, applied to PacBio SMRT long-read subreads.↳ Could also: Long-read-oriented assemblers such as HGAP/Canu/Flye could also have been used. — Long-read assemblers are designed around the error profile and length distribution of SMRT data and would also report assembly completeness/contiguity metrics, which can complement a short-read assembler choice.
-
Phylogenetic relationships were summarized using bootstrap support from 1,000 replicates on neighbor-joining (16S) and maximum-likelihood (concatenated orthologs) trees.↳ Could also: Bayesian inference (e.g., MrBayes/BEAST) with posterior probabilities, or additional approximate-likelihood-ratio (aLRT/SH-like) support, could also have been reported. — A second support metric or Bayesian posterior would also describe branch confidence under a different statistical framework and is commonly reported alongside bootstrap values.
-
Genome relatedness among strains was assessed via 16S rRNA similarity and concatenated single-copy ortholog phylogeny.↳ Could also: Whole-genome distance metrics such as average nucleotide identity (ANI) or digital DNA–DNA hybridization could also have been computed. — ANI/dDDH provide quantitative pairwise genome similarity values that also help resolve closely related strains where 16S has limited discriminatory power, as the paper itself notes for L. plantarum.
-
Core- and pan-genome sizes were reported as single point counts from OrthoMCL clustering of 19 genomes.↳ Could also: Pan-genome rarefaction/accumulation curves with fitted models (e.g., Heaps' law) could also have been presented. — Accumulation curves would also describe how core and pan-genome estimates change with the number of sampled genomes and characterize the open/closed nature of the pan-genome quantitatively.
-
Ortholog families were defined using fixed BLASTP cutoffs (E-value 1e−5, ≥50% match) and an MCL inflation of 1.5.↳ Could also: A sensitivity analysis varying the inflation value and identity/coverage thresholds could also have been included. — Reporting how core/pan/strain-specific counts respond to parameter choices would also convey the robustness of the clustering results to the chosen thresholds.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
ZLP001 16S rRNA shares >99% sequence similarity with other L. plantarum strains and clusters on a standalone phylogenetic branch near strains BDGP2, JDM1, and LZ95.other lactobacillus-plantarum-zlp001 2018×1papers★ This paper is the founder (earliest)
-
ZLP001 encodes 119 CAZyme genes (50 GH, 32 GT, 18 CE, 13 CBM, 6 AA), more than the 94 in L. plantarum KLDS1.0391, with enrichment in GH and CE families.other lactobacillus-plantarum-zlp001 up 2018×1papers★ This paper is the founder (earliest)
-
ZLP001 genome contains 22 CRISPR loci and 2 intact prophages alongside 306 transport-related genes (54 PTS, 252 ABC transporters) and 10 complete PTS EII complexes.other lactobacillus-plantarum-zlp001 2018×1papers★ This paper is the founder (earliest)
-
ZLP001 contains 65 strain-specific genes (including 30 hypothetical proteins) out of 2,597 total strain-specific genes (39.36%) across 19 L. plantarum strains.other lactobacillus-plantarum-zlp001 2018×1papers★ This paper is the founder (earliest)
-
Pan-genome of 19 L. plantarum strains comprises 6,598 orthologous gene families; core genome of 596 families (9.03%) constitutes 19.59% of each genome.other lactobacillus-plantarum 2018×1papers★ This paper is the founder (earliest)
-
L. plantarum ZLP001 has a complete circular chromosome of 3,164,369 bp with 44.65% GC content encoding 3,104 genes (avg 886 bp), constituting 83.28% of the genome.WGS lactobacillus-plantarum-zlp001 2018×1papers★ This paper is the founder (earliest)
-
L. plantarum ZLP001 harbors 7 plasmids ranging from 13,837 to 67,802 bp with average 42.05% GC content.WGS lactobacillus-plantarum-zlp001 2018×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — pmid-30542296
Paper: Zhang et al. 2018, Complete Genome Sequencing of Lactobacillus plantarum ZLP001, a Potential Probiotic... Front Physiol 9:1689. PMCID PMC6277807. Type: complete-genome announcement.
Data: BioProject PRJNA381357
SRR8079316— PacBio RS II SMRT subreads (73,238 reads, 634 Mb) → assembly inputSRR5407012— Illumina HiSeq 4000 PE (1.52 M reads) → used by authors for polishing- Deposited assembly
GCA_003076435.1(ASM307643v1) → independent cross-check
Code link in brief: github.com/stamatak/standard-RAxML (third-party tool the
authors used for the phylogeny; not their own pipeline). Per P16 this is valid to
reproduce, but see out-of-scope note below.
IN SCOPE — clearly-specified pipeline outputs (80%)
Pipeline: PacBio reads → de-novo assembly (paper: SOAPdenovo v2.04; we use Flye, the standard long-read assembler — SOAPdenovo is a short-read assembler and is not appropriate for PacBio data, so an identical tool match is neither possible nor sensible; genome size/GC are assembler-independent) → gene/RNA annotation (paper: Glimmer v3.02; we use Prodigal/barrnap/tRNAscan-SE).
| id | result | reported | robustness |
|---|---|---|---|
| c1 | PacBio subread count | 73,238 | exact, tool-free |
| c2 | PacBio median subread length | 8,661 bp | exact, tool-free |
| c3 | chromosome size | 3,164,369 bp | high (assembler-independent) |
| c4 | chromosome GC | 44.65% | very high (sequence-intrinsic) |
| c9 | plasmid count | 7 | high |
| c10-11 | plasmid sizes A–G | 67,802 … 13,837 bp | high |
| c12 | plasmid mean GC | 42.05% | very high |
| c5/c6 | CDS / chromosome genes | 3,264 / 3,104 | LOW — gene-caller dependent (Glimmer vs Prodigal differ) |
| c7/c8 | chromosome rRNA / tRNA | 16 / 69 | medium — tool dependent |
The strongest 1:1 evidence is c1–c4, c9–c12 (sequence-intrinsic). CDS/RNA counts are reported with an explicit tool-difference caveat — agreement within a few % is expected, not exact.
OUT OF SCOPE — the hard ~20% (not attempted, with reason)
- RAxML ML phylogeny on 553 single-copy orthologs (Figure): requires the set of other genomes compared against, which the paper does not enumerate (no species list, no accessions, no ortholog-calling pipeline parameters). Not reproducible from the described methods → skipped per 80/20 rule. The named code link (standard-RAxML) is only the tree-builder; the unspecified inputs are the blocker, not the tool.
- 16S rRNA neighbor-joining tree (MEGA) — same input-set ambiguity; wet-lab/manual.
- KEGG/COG category counts (1,603 / 1,783 CDS) — depend on the exact CDS set + database versions (2018); reproducible in principle but low-value and version-sensitive → not attempted.
- All wet-lab assays (barrier function, defense) — out of computational scope.
Verdict shape expected
A partial reproduction: strong 1:1 on genome size/GC/plasmid structure (sequence-intrinsic, anti-fabrication cross-checked against the deposited assembly), with annotation counts reported as tool-caveated approximations. The phylogeny — the result tied to the named code — is not reproducible from the paper as written.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The central, deposited genome claims reproduce essentially 1:1 — chromosome 3,164,344 bp vs reported 3,164,369, GC 44.66% vs 44.65%, rRNA 16/16, and 73,238 subreads exact — and independently match GenBank CP021086.1, confirming the values are genuine, not fabricated. The remaining deviations (5/7 plasmids, total CDS 3,131 vs 3,264, plasmid mean GC 39.44 vs 42.05) lie on our side: we deliberately skipped the authors' Illumina hybrid closure and used Prodigal instead of Glimmer. The only authors-side issue is a minor mislabel of the mean subread length (8,661 bp) as the median. Overall a solid partial reproduction with fully explainable, non-critical discrepancies.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.