Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Metagenetic and Volatilomic Approaches to Elucidate the Effect of Lactiplantibacillus plantarum Starter Cultures on Sicilian Table Olives.

Front Microbiol · 2022
L1 23/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
23/100
Reproducibility score
2.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 1% of all assessed papers rank 1160 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DECISIVE MISMATCH (data-deposit error). The in-scope 16S metabarcoding pipeline ran end-to-end on «our HPC» (QIIME2 2024.10: cutadapt -> DADA2 -> SILVA138 vsearch -> diversity) on all 15 deposited PRJNA675996 runs (1,146 ASVs, 981,226 non-chimeric reads). HEADLINE: the deposited reads are MOUSE GUT microbiota, not Sicilian table olives — dominated by Muribaculaceae 39% / Muribaculum 38% / Bacteroidota 58% (canonical murine gut), while the paper's olive taxa are essentially absent (Weissella 0.02% vs reported 85%; Enterobacter 0.01% vs reported ~57%; Lactobacillus 13%, not dominant). This CONFIRMS the SRA run metadata 'Mus musculus flora' (previously presumed a copy-paste error) and means the paper's reported 16S taxonomic results are not derivable from the cited accession. Only Good's coverage (C6, >99.98%) agrees, and that merely reflects sequencing depth. Secondary discrepancies: N=15 (C/L/H x5) vs paper 24/8; deposited amplicon is V3-V4 (341F/806R) vs the paper's stated V3-only Probio_Uni/Probio_Rev (341F/518R). FABRICATION/WRONG-DATA CONCERN flagged on C7-C9 (reported abundances absent from the data). NOT ATTEMPTED (out of scope): volatilomics (HS-SPME-GC/MS, wet-lab) and the corrplot microbiota-VOC correlation figure (needs the out-of-scope VOC table). Verdict provisional; a human reviewer should confirm and determine whether the real olive data exist under another accession.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 57cbba7689e4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether inoculation of Sicilian table olives (Nocellara Etnea cv.) with selected Lactiplantibacillus plantarum starter cultures, at reduced (5%) NaCl content, affects microbiota composition and volatile organic compound (VOC) profile compared with uninoculated controls at 5 and 8% NaCl.

Core claims
  • Inoculation with L. plantarum starter cultures (O1, O2) accelerates acidification, producing a faster and more pronounced pH drop than uninoculated controls (C5, C8). finding
  • 16S metagenetic analysis shows lactobacilli dominate the microbiota of both starter-inoculated samples (O1, O2). finding
  • Enterobacter genus occurs at high levels only in uninoculated control samples with 5% NaCl (C5). finding
  • Bacteroides, Faecalibacterium, Klebsiella, and Raoultella genera are present only in uninoculated control samples with 8% NaCl (C8). finding
  • Microbiota composition dynamics during fermentation significantly affect the volatile organic compound profile of the final products. finding
  • No off-odor-associated volatile compounds were detected in any sample investigated. finding
  • A dual approach combining 16S amplicon-based metagenetics and HS-SPME/GC-MS volatilomics can link microbial dynamics to VOC profiles in table olive fermentation. method
  • Use of low NaCl (5%) concentration together with the proposed L. plantarum starter cultures positively affects microbiota and VOCs, ensuring microbiological safety and pleasant flavor of the final product. finding
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA amplicon-based sequencing (V3 region) Table olive drupes, Nocellara Etnea cultivar L. plantarum starter culture inoculation (O1, O2) vs uninoculated control at 5% or 8% NaCl (C5, C8) Microbiota composition/taxonomy, alpha (Shannon, Chao1) and beta diversity (weighted UniFrac, PCoA) MiSeq (Illumina); Probio_Uni/Probio_Rev primers; QIIME2/DADA2; SILVA 138
Volatile organic compound analysis by HS-SPME GC-MS Table olive pulp/drupes, Nocellara Etnea cultivar L. plantarum starter culture inoculation (O1, O2) vs uninoculated control at 5% or 8% NaCl (C5, C8) VOC profile/identification and abundance Clarus 680 GC with Clarus SQ8MS single-quadrupole MS, Rtx-WAX column, PAL COMBI-xt autosampler
Classical microbiological plate counts Brine and olive drupe samples, Nocellara Etnea cultivar L. plantarum starter culture inoculation (O1, O2) vs uninoculated control at 5% or 8% NaCl (C5, C8) Log10 CFU/mL or CFU/g of total mesophilic bacteria, LAB, yeasts, Enterobacteriaceae, staphylococci, E. coli, sulfite-reducing clostridia Selective agar media (PCA, MRS, Sabouraud, VRBGA, mannitol salt agar, MacConkey, SPS agar)
pH measurement Brine samples L. plantarum starter culture inoculation (O1, O2) vs uninoculated control at 5% or 8% NaCl (C5, C8) pH value over fermentation time (0-80 days) MettlerDL25 pH meter
Key results
  • Inoculated samples (O1, O2) reached pH ~4.5 by day 24, whereas controls only reached a comparable pH by day 60.
  • At the end of fermentation, pH fell below the 4.3 safety threshold in all samples except control C5 (5% NaCl), which remained at 5.04. final pH: O1=4.23, O2=4.10, C8=4.36, C5=5.04
  • Enterobacteriaceae counts decreased significantly from day 30 in inoculated samples, especially O2, becoming undetectable (<1 log CFU/g) earlier than in controls.
  • LAB counts increased by about 1 log unit in inoculated samples by the end of fermentation, while decreasing by about 1 log unit in control samples. ~1 log CFU/g
  • Yeast population increased by approximately 6 log units by day 15 in all samples, especially C8 (reaching 8 log CFU/g), then declined by day 80 to ~3.9 (inoculated) and ~5.2 (controls) log CFU/g. ~6 log unit increase
  • E. coli and sulfite-reducing clostridia were never detected in any sample at any fermentation time point.
  • No compounds associated with off-odor metabolites were detected in any of the investigated samples.
Key statistics
  • pvalue p < 0.05 (Significance threshold for ANOVA/Tukey post hoc on pH, microbiological, and VOC data across biological replicates)
  • pvalue p < 0.05 (Welch t-test, Benjamini-Hochberg FDR corrected) (Significance threshold for differences among fermentation processes in detected genera (STAMP software))
  • mean 6.17 ± 0.08 (O1 brine pH at day 0 of fermentation)
  • mean 4.36 ± 0.06 (C8 brine pH at day 80 (end of fermentation))
  • count 7.85 ± 0.07 log10 CFU/g (LAB count in O2 olive samples at day 80)
  • count 8.08 ± 0.11 log10 CFU/g (Peak yeast count in C8 olive samples at day 15)
  • other 0 to 1 (Range of weighted UniFrac similarity values used for beta diversity/PCoA)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study employed a multi-assay design with four fermentation conditions (O1, O2, C5, C8), each conducted in biological triplicate and monitored over 80 days. Physicochemical, microbiological, and volatile organic compound (VOC) data were compared across groups using one-way ANOVA with Tukey post hoc correction, while 16S metagenetics genus-level differences were assessed with a two-sided Welch t-test under Benjamini–Hochberg FDR. Ordination methods (PCA, PCoA on UniFrac distances) and Spearman rank correlation were used to explore multivariate relationships between microbial groups and VOC profiles. Results were reported as means ± standard deviations with a significance threshold of p < 0.05.

Replicationbiological Sample sizeThree biological replicates per fermentation condition stated; no formal sample size or power calculation described GroupsO1 and O2 (inoculated, 5% NaCl) vs. C5 (uninoculated, 5% NaCl) and C8 (uninoculated, 8% NaCl) Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionTukey HSD (for ANOVA pairwise comparisons); Benjamini–Hochberg FDR (for genus-level 16S comparisons in STAMP)
Statistical tests used
Test Applied to n Assumptions
One-way ANOVA with Tukey HSD post hoc test pH values, microbiological counts (CFU/mL or CFU/g), and VOC data across fermentation groups and time points 3 biological replicates per group not stated
Two-sided Welch t-test with Benjamini–Hochberg FDR correction Genus-level differences among fermentation processes (16S rRNA metagenetics), implemented in STAMP not stated
Principal Component Analysis (PCA) VOC profiles at T15 and T80 across inoculated and control samples 3 biological replicates per group per time point na
Permutation analysis (PermutMatrix) Similarities between volatile profiles of inoculated (O1, O2) and control (C5, C8) samples at T15 and T80 na
Principal Coordinate Analysis (PCoA) on weighted UniFrac distance matrices Beta diversity among 16S rRNA amplicon samples na
Spearman rank correlation Correlation between microbial group cell densities and VOCs at T15 and T80 not stated
Shannon and Chao1 indices Alpha diversity within 16S rRNA amplicon samples (QIIME2) na
Weighted UniFrac Beta diversity between 16S rRNA amplicon samples (QIIME2) na
Approaches that could also have been used
  • One-way ANOVA was applied separately at each time point to compare groups, treating repeated longitudinal measurements (T0 through T80) as independent cross-sections
    Could also: A linear mixed-effects model or repeated-measures ANOVA could also be used, treating fermentation batch as the repeated unit and time as a within-subject factor — A mixed-effects or repeated-measures approach explicitly models the temporal autocorrelation within each fermentation replicate, which can improve statistical power and more directly address the time-by-treatment interaction that is central to the study's aims
  • Genus-level differences in the 16S metagenetics data were tested with a Welch t-test on relative-abundance or count data in STAMP
    Could also: Negative-binomial models implemented in tools such as DESeq2 or edgeR, or compositional methods such as ANCOM-BC, could also be applied to amplicon count data — 16S amplicon data are count-based and compositional; negative-binomial or compositional methods are designed for this data structure and can handle overdispersion and sequencing-depth variation without requiring transformation
  • Spearman rank correlations between microbial cell densities and VOCs were computed across many pairs without a stated multiplicity correction
    Could also: Applying a Benjamini–Hochberg FDR correction to the full set of pairwise Spearman tests would also control the expected proportion of false discoveries — With a large correlation matrix, the number of simultaneous tests inflates the family-wise false-positive rate; FDR correction—already used elsewhere in the paper for the STAMP analyses—would make the inferential standards consistent across the two analyses
  • PCA was used to visualize separation of VOC profiles among the four fermentation conditions
    Could also: A permutation-based multivariate test such as PERMANOVA (adonis in R's vegan package) could also formally test whether group membership explains a significant proportion of VOC profile variance — PCA is an exploratory visualization; PERMANOVA provides an inferential test of group separation with a p-value and an R² effect size, complementing the ordination plot with a formal statistical statement
  • Alpha diversity was summarized using Shannon entropy and Chao1 richness, with no between-group statistical test of alpha diversity explicitly described
    Could also: A Kruskal–Wallis test (or one-way ANOVA if assumptions hold) on Shannon or Chao1 values across the four conditions could also formally test whether inoculation or salt level alters within-sample diversity — Reporting index values without a significance test leaves it unclear whether observed differences in alpha diversity exceed what would be expected by chance; a formal test would distinguish meaningful diversity shifts from sampling variability
  • Results throughout were reported as means ± SD with a binary significance indicator (p < 0.05 / letter groupings), without effect sizes or confidence intervals
    Could also: Reporting 95% confidence intervals or a standardized effect size such as Cohen's d or eta-squared alongside p-values would also convey the magnitude and precision of observed differences — Effect sizes and confidence intervals communicate practical significance and estimation uncertainty, allowing readers to judge whether statistically significant differences are also biologically meaningful, which is especially useful when n = 3 per group
Software: STATISTICA 7.0 · STAMP · QIIME2 · DADA2 · R / cor.test + corrplot corrplot 0.90 · PermutMatrix

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35281313

Title: Metagenetic and Volatilomic Approaches to Elucidate the Effect of Lactiplantibacillus plantarum Starter Cultures on Sicilian Table Olives. Vaccalluzzo et al., Front Microbiol 2022. DOI: 10.3389/fmicb.2021.771636. PMCID PMC8914321. Data: PRJNA675996 (NCBI SRA). "Code": https://github.com/taiyun/corrplot (generic R correlation-plot package — third-party, used for Fig. correlation heatmaps).

What the paper reports (two methodological arms)

  1. Metagenetic (16S rRNA amplicon) — IN SCOPE (pipeline-derived).

    • Platform: Illumina MiSeq; target = V3 region of 16S rRNA, primers Probio_Uni / Probio_Rev (Milani et al. 2013).
    • Pipeline: QIIME2 + DADA2 (ASVs at 100% homology) + taxonomy vs SILVA release 138.
    • QC: keep reads length 140–400 bp, mean Q>20; drop homopolymers>7 bp & primer mismatches.
    • Reported computational outputs we will try to reproduce:
      • Total 449,654 bacterial sequences; avg 56,207 seq/sample.
      • Taxonomy: 3 phyla, 7 families, 10 genera at rel. abund. >0.5%; dominant phyla Firmicutes & Proteobacteria.
      • Alpha diversity: Good's coverage >99%; Shannon & Chao1 (Supplementary Table 2); Chao1 ↑ from d15→d80.
      • Relative abundances (Results): Weissella 85.09% (C8 d15) / 78.34% (C8 d80); Enterobacter ~57% in C5; Lactobacillus dominant & significantly higher in inoculated (O1/O2) vs controls (Welch p<0.05).
      • Beta diversity: PCoA, PC1 explains >75% of variance.
  2. Volatilomic (HS-SPME-GC/MS) — OUT OF SCOPE (wet-lab/instrumental). Volatile organic compound profiling is bench chemistry, not a bioinformatic pipeline; not reproducible from deposited data. Not attempted.

  3. Correlation analysis (corrplot) — partially in scope, depends on (1)+(2). The cited "code" repo is the generic R corrplot package used to draw microbiota↔VOC correlation heatmaps. It is a plotting tool, not an analysis pipeline, and requires the VOC table (arm 2, out of scope) as input, so the correlation figure cannot be fully reproduced. Not a primary target.

Primary reproduction target

Run the standard 16S amplicon pipeline (QIIME2/DADA2/SILVA138) on the deposited PRJNA675996 reads and compare: total/avg sequence counts, taxonomic richness (phyla/families/genera >0.5%), dominant genera & their relative abundances, and alpha diversity (Good's coverage, Shannon, Chao1).

CRITICAL DATA CAVEAT (deposit ↔ paper mismatch) — see dataset_profile.json

The paper's Data Availability cites PRJNA675996, and describes the 16S design as 4 treatments (O1,O2,C5,C8) × 2 timepoints (d15,d80) × triplicate = 24 samples (its own sequence-count arithmetic, 449,654/56,207, implies 8). The actual SRA deposit contains 15 runs (SRR13048283–297) in 3 groups of 5 with aliases 1-C1-s … 5-C5-s, 6-L1-s … 10-L5-s, 11-H1-s … 15-H5-s, and every run's description reads "16s rDNA of Mus musculus flora" (an obvious copy-paste error from another submission). The group labels C/L/H and replicate count (5) do not map cleanly onto the paper's O1/O2/C5/C8 × d15/d80 × 3 design. This is a material reproducibility issue: the deposited data cannot be unambiguously aligned to the paper's reported samples. We will still run the pipeline on the 15 deposited runs and report what is actually present, grading the N-dependent claims accordingly.

Compute plan

All heavy compute on «our HPC» (SLURM). Download 15 ENA FASTQs to «infra» on front1; build QIIME2 env on front1; submit DADA2/SILVA job; pull back small summary tables (feature table summary, taxa barplot data, alpha-diversity vectors, PCoA eigenvalues).

Status — COMPLETE (2026-06-25)

Pipeline RAN on «our HPC» (QIIME2 2024.10: cutadapt->DADA2->SILVA138 vsearch->diversity) on all 15 deposited runs. DECISIVE FINDING: the deposited data are MOUSE GUT microbiota, not table olives (Muribaculaceae 39%/Muribaculum 38%/Bacteroidota 58%; Weissella 0.02% vs reported

Figures / tables: Table
C1
Reported
449,654 total bacterial sequences
Reproduced
981,226 (15 deposited samples)
did not match
C2
Reported
56,207 avg sequences/sample
Reproduced
65,415 avg/sample
partial
C3
Reported
3 phyla >0.5% (Firmicutes/Proteobacteria)
Reproduced
6 phyla; Bacteroidota 58% dominant
did not match
C4
Reported
7 families >0.5%
Reproduced
16 families; Muribaculaceae 39% dominant (mouse-gut)
did not match
C5
Reported
10 genera >0.5%
Reproduced
21 genera; Muribaculum 38% dominant (mouse-gut)
did not match
C6
Reported
Good's coverage >99%
Reproduced
>99.98% (all 15 samples)
within tolerance
C7
Reported
Weissella 85.09% (C8 d15)
Reproduced
0.02% overall (max 0.21%)
did not match
C8
Reported
Weissella 78.34% (C8 d80)
Reproduced
0.02% overall
did not match
C9
Reported
Enterobacter ~57% (C5)
Reproduced
0.01% overall (max 0.08%)
did not match
C10
Reported
Lactobacillus dominant, higher in inoculated (p<0.05)
Reproduced
13.3% (2nd, not dominant); group mapping absent
did not match
C11
Reported
PCoA PC1 >75%
Reproduced
Bray-Curtis PC1 23.6%
did not match
C12
Reported
Shannon & Chao1 (Suppl T2); Chao1 up d15->d80
Reproduced
Shannon 5.25-6.65; Chao1 230-432; trend uncheckable
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 23/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

90.6 k
tokens (I/O) · 5.4 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.