Corpus 1,276 assessed · 1,177 scored · 644 reproduced ≥75 · 170 flagged ·∅ 74/100
← New search

Ecotype diversity and conversion in Photobacterium profundum strains.

PLoS One · 2014
77/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
77/100
Reproducibility score
at the mean
vs. all fields · 1177 studies
🎯 Scores higher than 50% of all assessed papers rank 572 of 1177 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

This is a 1:1-vs-partial reproduction. Using the paper's own third-party-style repo (flauro/3tck_comparative, 3 Perl scripts: ANI.pl, COG_scrambler.pl, scaffolding.pl) applied to the two deposited public assemblies (3TCK: GCA_000153425.1 under BioProject PRJNA13563; SS9 reference: GCA_000196255.1 under PRJNA13128), the core comparative-genomics claim reproduces closely: ANI 92.76% vs reported 92.85%, percent conserved DNA 63.25% vs reported 62.68%, both within-tolerance using the paper's own stated parameters (fragment 1020bp, min identity 30%, min alignable 714bp). Basic genome stats for 3TCK also match almost exactly: 11/11 scaffolds, 5549/5549 ORFs, length within 0.0065%, GC 40.77% vs 41.3% reported (minor discrepancy). Going beyond the floor, we additionally built an independent COG-assignment pipeline (rpsblast vs NCBI's classic COG database) and ran the repo's own COG_scrambler.pl with the paper's stated resampling parameters (subsample 4000, 10000 bootstraps); all three qualitative COG-category claims from the paper's Figure 4 reproduce in the same direction and significance (category C over-represented in 3TCK; N and L under-represented in 3TCK / over-represented in SS9), and the specific transposase-family claim (SS9 >> 3TCK) reproduces directionally (43 vs 0 for COG3436) though not at the same magnitude as the paper's reported 72 vs ~3 (different, unspecified original annotation pipeline). NOT attempted: the Ka/Ks ortholog-divergence analysis (no shipped or clearly-specified tool), the scaffolding.pl pseudomolecule-construction step (its pre-scaffold raw contig input is not public), and all wet-lab results (rRNA PFGE counts, UV survival, conjugation) which are explicitly out of scope. Dataset profiling surfaced one notable finding: the RU's 'sra:PRJNA13563' pointer is misleading -- NCBI SRA has zero run records for this BioProject; only the finished WGS assembly is public, so the underlying raw Sanger reads cannot be independently re-assembled by a third party.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-06
Rubric version
not recorded
Assessed by
Last updated
2026-08-06

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Which genomic features distinguish the deep-sea piezophilic Photobacterium profundum strain SS9 from the shallow-water non-piezophilic strain 3TCK and thereby define each strain's Hutchinsonian (depth) niche, and whether horizontal gene transfer of such features can mediate rapid bathytype conversion.

Core claims
  • No single gene restricts the environmental niche of each bathytype; instead a set of strain-specific genetic features confers depth-specific stress tolerance (temperature, pressure, nutrients). finding
  • Bathytype evolution in P. profundum is driven primarily by gene acquisition and loss (HGT/adaptive radiation) rather than by sequence substitution/positive selection. mechanism
  • Transfer of 3TCK-specific photorepair (phr) genes into SS9 demonstrates that horizontal gene transfer can provide a mechanism for rapid colonisation of new environments, i.e. bathytype conversion. finding
  • Lack of UV photorepair function is predicted to restrict colonization of shallow waters by deep bathytypes. mechanism
  • Despite 16S identity consistent with one species, SS9 and 3TCK fall below genome-level species thresholds (ANI and conserved DNA). finding
  • The draft genome sequence of P. profundum 3TCK is a new resource (NCBI BioProject PRJNA13563), with custom perl analysis scripts released at https://github.com/flauro/3tck_comparative. resource
  • Genes unique to each bathytype are concentrated on chromosome 2, the chromosome previously implicated in gene capture for environmental adaptation in Vibrionaceae. finding
  • This is the first study combining intra-specific sequence comparison with molecular genetics to address niche partitioning in piezophilic bacteria. method
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome shotgun sequencing and de novo assembly (4 kb and 40 kb insert libraries) Photobacterium profundum strain 3TCK (shallow bathytype), genomic DNA from mid-exponential culture none Draft genome sequence: scaffolds, genome length, GC content, predicted ORFs J. Craig Venter Institute sequencing; Celera assembler; NCBI PGAAP annotation pipeline
Pulsed-field gel electrophoresis of I-CeuI-digested chromosomal plugs P. profundum strains SS9, DSJ4 and 3TCK I-CeuI restriction digest (overnight) Number of ribosomal RNA (rrn) operon copies / chromosome fragment pattern 0.95% PFGE agarose in 0.5x TBE, 6 V/cm, 120° included angle, 14°C
Comparative genomics: ACT nucleotide alignment, orthology/Venn gene-content comparison, ANI, conserved DNA In silico: P. profundum 3TCK draft genome vs SS9 reference genome (PRJNA13128) none Synteny, insertions/deletions, inversions, translocations, shared vs unique gene counts, average nucleotide identity, percent conserved DNA ACT; custom perl scripts; Goris et al. ANI method (fragment 1020 bp, min identity 30%, min alignable region 714 bp)
COG functional category assignment with resampling-based statistical comparison In silico: predicted ORFs of 3TCK and SS9 none Over/under-representation of COG categories (median differences after resampling) Rodriguez-Brito method, subsample size 4000, 10,000 bootstraps
Ka/Ks (ω) selection analysis and divergence-time estimation In silico: orthologous gene pairs between 3TCK and SS9 none ω per ortholog pair, number of pairs with ω>1, Fisher exact test significance, median divergence time τ = Ks/(2λ) Reciprocal smallest distance algorithm; MUSCLE; KaKs calculator (YN00 method); λ = 8.3×10−7 SNPs/site/year
Codon usage bias analysis In silico: genes of P. profundum genomes none Per-gene codon usage vs genome-wide modal codon usage, Chi-square significance (p-value threshold 0.1) Karlin method implemented in custom perl script
Molecular cloning and tri-parental conjugation (heterologous gene transfer) P. profundum SS9 recipient; E. coli DH5α/XL1-Blue/TOP10 cloning hosts; helper E. coli with pRK2073 Introduction of 3TCK phr gene cluster (P3TCK_10673 rpoX, P3TCK_10668, P3TCK_10663 phr) in pFL122/pFL190 constructs pFL303–pFL307, including the Δ22 promoter/rpoX deletion series and arabinose-inducible phr Construct identity/insert directionality by PCR and dideoxy sequencing; conjugal transfer of UV-resistance genes Expand Long Template PCR system (Roche); NEB restriction enzymes; Applied Biosystems fluorescent-terminator sequencing
In vivo photoreactivation / UV survival assay (c.f.u. counting) P. profundum strains plated on 75% strength 2216 Marine Agar, grown at 15°C for 5 days UV irradiation (254 nm, 10 s, 220 µW/cm2) with or without 1 h blue-light recovery (350–400 nm, 20 µW/cm2); untreated control Percent survival (c.f.u. of irradiated vs untreated controls) Philips G25 T8 germicidal lamp (253.7 nm); Philips TLD 15 W/08 black light; Spectroline DM-365 XA digital radiometer
Key results
  • 3TCK draft genome: 11 scaffolds, 6,186,725 bp, 41.3% GC, 5549 ORFs; organised in two chromosomes but lacking the 80 kb dispensable plasmid of SS9 6,186,725 bp; 5549 ORFs
  • 16S rRNA gene identity between 3TCK and SS9 indicates same species, but ANI and percent conserved DNA fall below genome-level species thresholds (ANI>95%, conserved DNA>69%) ANI 92.85%; conserved DNA 62.68% vs 16S 98.73%
  • Only four orthologous gene pairs had ω>1 and none were statistically significant, indicating little detectable positive selection 4 gene pairs; P<0.01 threshold not met
  • Median divergence time between the strains was ~126,833 years, comparable to the establishment of modern thermohaline circulation 126,833 years
  • COG comparison: shallow bathytype 3TCK over-represented in energy production (C) and under-represented in motility/chemotaxis (N) and DNA replication, recombination and repair (L); category L abundance in SS9 attributed to numerous transposable elements significant at 98% significance
  • 3TCK carries at least 9 rrn operon copies, above the median for microbial genomes but fewer than SS9's 15; 3TCK operons are nearly identical whereas SS9 operons show intragenomic variation 9 copies (3TCK) vs 15 copies (SS9)
  • 3TCK has larger-than-average intergenic regions, though smaller than in the deep bathytype SS9 ~167 bp (3TCK) vs ~205 bp (SS9)
  • Genome comparison shows extreme synteny with many insertions/deletions and multiple inversions across origin/terminus of both chromosomes but few interchromosomal translocations; unique genes concentrated on chromosome 2
Key statistics
  • other 98.73% 16S rRNA gene sequence identity (3TCK vs SS9 16S identity, consistent with same species)
  • other ANI 92.85%; percent conserved DNA 62.68% (Genome-level comparison of 3TCK vs SS9, below species thresholds (ANI>95%, conserved DNA>69%))
  • count 6,186,725 bp total in 11 scaffolds; 41.3% GC; 5549 ORFs (3TCK draft genome general features)
  • count 4 orthologous gene pairs with ω>1, none significant by Fisher exact test (P<0.01) (Global Ka/Ks analysis of positive selection)
  • other median divergence time 126,833 years; λ = 8.3×10−7 SNPs/site/year (τ = Ks/(2λ) medianed across all ortholog pairs)
  • count at least 9 rrn operon copies in 3TCK vs 15 in SS9 (I-CeuI PFGE estimate of ribosomal RNA operon copy number)
  • mean ~167 bp (3TCK) vs ~205 bp (SS9) average intergenic region length (Intergenic region size comparison between bathytypes)
  • other UV dose 220 µW/cm2 for 10 s at 253.7 nm; photoreactivation 20 µW/cm2 for 1 h at 350–400 nm (In vivo photoreactivation assay conditions, triplicate dilution series per strain)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper reports a comparative genomics study of two Photobacterium profundum bathytypes (deep-water SS9 and shallow-water 3TCK), combined with a small in vivo UV-survival/photoreactivation experiment. Statistical treatment consists of a handful of targeted significance tests applied to specific comparisons (COG category representation, dN/dS ratios of orthologous gene pairs, and per-gene codon usage bias) rather than a unified statistical framework, and results are reported mainly as significance thresholds and percentages rather than as full test statistics.

Replicationunclear Sample sizeNo formal sample-size or power justification is described; the genome comparison involves single draft genomes per strain, and the UV survival assay used a triplicate dilution series per strain/condition. GroupsTwo P. profundum bathytypes (deep SS9 vs. shallow 3TCK) for genomic/COG/Ka-Ks/codon-usage comparisons; multiple strains across untreated, UV-irradiated, and UV+photoreactivation conditions for the survival assay Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Resampling/bootstrap-based comparison (method of Rodriguez-Brito) Statistical comparison of COG category representation between the SS9 and 3TCK genomes (Figure 4) subsample size of 4000, 10,000 bootstraps, evaluated at 98% significance not stated
Fisher's exact test Significance of orthologous gene pairs with a Ka/Ks ratio (ω) > 1 four gene pairs identified with ω>1; total ortholog pairs tested not stated; threshold P<0.01 not stated
Chi-square test Per-gene codon usage bias compared to the genome-wide mode of codon usage not stated (applied per gene across the genome); p-value threshold set to 0.1 not stated
Approaches that could also have been used
  • COG category over/under-representation between the two bathytypes was assessed with a bootstrap resampling method at a 98% significance threshold, applied across multiple COG categories.
    Could also: A false-discovery-rate procedure such as Benjamini-Hochberg could also be applied across the full set of COG category tests. — This would explicitly control the expected proportion of false positives when many categories are tested simultaneously, complementing the per-category bootstrap threshold already used.
  • Significance of elevated Ka/Ks ratios (ω>1) for orthologous gene pairs was assessed with a Fisher exact test.
    Could also: A likelihood ratio test comparing models with and without positive selection (e.g., as implemented in PAML/codeml) could also be used. — Likelihood-based selection tests can incorporate site-to-site or branch-to-branch variation in ω and provide a model-based framework for testing positive selection, complementing the pairwise Fisher exact test approach.
  • Codon usage bias per gene was tested against the genome-wide mode using a chi-square test with a p-value threshold of 0.1.
    Could also: Applying a multiple-testing correction (e.g., FDR) across the many genes tested, or using an exact test for smaller counts, could also be used. — With many genes evaluated in parallel, an FDR-adjusted threshold would help contextualize how many of the nominally significant genes might be expected by chance, and an exact test can be more appropriate than chi-square when expected counts are small.
  • UV survival differences between strains were summarized as percent survival from triplicate platings, without a stated formal statistical comparison between conditions or strains.
    Could also: Reporting a dispersion measure (e.g., SD or 95% CI) across replicate platings and applying a t-test or ANOVA (with post-hoc correction for multiple strain comparisons) could also be used. — This would allow readers to gauge variability across the triplicate measurements and provide a formal statistical comparison of survival between strains and treatment conditions, in addition to the reported percentages.
  • The time of divergence between strains was estimated as a single median value from a fixed substitution rate, without an accompanying interval estimate.
    Could also: A bootstrap-based or Bayesian confidence/credible interval around the divergence time estimate could also be reported. — This would convey the uncertainty around the point estimate of divergence time, complementing the single median value derived from the fixed molecular clock rate.
Software: Custom Perl scripts (resampling/bootstrap method of Rodriguez-Brito; codon usage analysis) · KaKs_Calculator (YN00 method)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

3TCK_scaffold_count
Reported
11 scaffolds (Results text)
Reproduced
11 scaffolds (GCA_000153425.1 / WGS AAPH00000000, downloaded from NCBI)
exact
3TCK_genome_length
Reported
6,186,725 bp total length
Reproduced
6,186,325 bp (deposited assembly; ANI.pl reports query length 6,186,325 after N-stripping)
within tolerance
3TCK_gc_content
Reported
41.3% average GC
Reproduced
40.77% GC (own parse of deposited genomic FASTA)
within tolerance
3TCK_orf_count
Reported
5549 ORFs
Reproduced
5549 protein records in deposited annotation (GCA_000153425.1)
exact
ani_3TCK_vs_SS9
Reported
ANI = 92.85%
Reproduced
ANI = 92.7554% (repo's own ANI.pl, default params matching paper: fragment size 1020 bp, min identity 30%, min alignable region 714 bp, run on deposited 3TCK/SS9 assemblies)
within tolerance
pcd_3TCK_vs_SS9
Reported
Percent conserved DNA = 62.68%
Reproduced
63.2546%
within tolerance
cog_category_C_over_in_3TCK
Reported
COG category C (energy production/conversion) over-represented in 3TCK vs SS9 (Fig 4)
Reproduced
Independent COG assignment (rpsblast vs NCBI classic COG db) + repo's COG_scrambler.pl (subsample 4000, 10000 bootstraps, script default 97% CI vs paper's stated 98%): median diff for C = -1.283 (negative = over-rep in query=3TCK), null 97%CI=[-0.961,0.957] -> outside interval, significant, same direction as paper
within tolerance
cog_category_N_under_in_3TCK
Reported
COG category N (motility/chemotaxis) significantly decreased in 3TCK vs SS9
Reproduced
median diff for N = +0.900 (positive = over-rep in reference=SS9, i.e. under-rep in 3TCK), null 97%CI=[-0.676,0.661] -> outside interval, same direction as paper
within tolerance
cog_category_L_under_in_3TCK
Reported
COG category L (DNA replication/recombination/repair) significantly decreased in 3TCK vs SS9
Reproduced
median diff for L = +4.887, null 97%CI=[-0.917,0.931] -> far outside interval, same direction as paper (also the category most consistent with the reported TE excess in SS9)
within tolerance
transposase_cog3436_excess_in_SS9
Reported
SS9's largest COG-categorized transposable-element family, COG3436, has 72 members; 3TCK has only 3 COG-categorized TEs in total across all families
Reproduced
Independent rpsblast/COG assignment: COG3436 count = 43 in SS9, 0 in 3TCK. Same direction (SS9 >> 3TCK) but different magnitude, likely due to a different COG-assignment tool/e-value cutoff/COG database vintage than the original (unspecified) annotation method
partial
rrna_operon_copy_number
Reported
3TCK >=9, SS9 = 15 rRNA operon copies (PFGE with I-CeuI digestion)
Reproduced
not attempted
partial
kaks_orthologs_divergence_time
Reported
Ka/Ks (omega) via reciprocal smallest distance + MUSCLE + KaKs_Calculator YN00; only 4 gene pairs with omega>1 (none significant, Fisher exact P<0.01); estimated divergence time 126,833 years
Reproduced
not attempted
partial
scaffolding_pl_pseudomolecule
Reported
scaffolding.pl (repo script) used to join JCVI shotgun contigs into the final 11-scaffold pseudomolecule via a 6-frame stop-codon spacer
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.