Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

An archaeal histone-like protein regulates gene expression in response to salt stress.

Nucleic Acids Res · 2021
L1 94/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (partial->strong, 1:1 where the deposit allows). Authors' own repo (amyschmid/HpyA_codes, HEAD 7babe18) ships commented R-markdown alongside its processed inputs and final supplementary tables. All 10 in-scope pipeline claims attempted and graded on «our HPC» with paper-pinned environments (DESeq2 1.30.1 / R 4.0.5 «job»; grofit 1.1-1 from shipped tarball; IRanges/GenomicRanges/rtracklayer): 6 EXACT, 4 WITHIN-TOL, 0 mismatch. RNA-seq DEG counts 168/143/46/21 are exactly consistent with shipped TableS4 (Condition tally), and an independent DESeq2 1.30.1 re-run recovers 140/143 (98%) of the final low-salt DEGs (the final = intersection of the shipped 21-sample arm with a 5-outlier arm whose input combineddata_5reps.csv is NOT deposited, so the single shipped matrix yields larger pre-intersection counts 221/52/248). ChIP 59 peaks is exact in shipped TableS3 (Basic_S3=59 rows); independent IRanges annotation of the shipped BEDs gives 64 peaks / 96 genes vs reported 59 / 86, the ~8-12% excess being the documented manual artifact-curation. Growth: grofit logistic on shipped curves gives WT 88-93% (paper 89%, exact when the one flagged outlier is excluded), KO 66.4% median (paper 67%), and t-test P=0.0011-0.0013 satisfying the reported P<0.008. No fabrication signal: every reported value is auditable against the deposited final tables. NOT attempted (out of scope): raw-read upstream alignment/trimming/peak-calling from GSE182514 (TrimGalore/Bowtie2/HTSeq/MACS2), wet-lab assays, microscopy/circularity, qPCR.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 94
    assessed: 2026-06-21 ⛓ e4b1a320c48e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Given the unusual negatively charged, highly ionic haloarchaeal cytoplasm, HpyA's non-canonical gene-regulatory function (rather than DNA-packaging) is hypothesized to be linked to the unique hypersaline cytoplasmic environment of Halobacterium salinarum.

Core claims
  • HpyA is important for maintaining wild-type growth rate under reduced salinity finding
  • HpyA is important for maintaining rod-shaped cell morphology under reduced salinity finding
  • HpyA preferentially binds DNA at ~60 discrete genomic sites under low salt in a reproducible, salt-dependent manner finding
  • HpyA binding is too sparse to coat or compact the genome, unlike canonical archaeal histones mechanism
  • High prevalence of HpyA binding within gene bodies suggests a regulatory mechanism distinct from canonical transcription factors mechanism
  • HpyA directly regulates iron/ion uptake genes via DNA binding mechanism
  • HpyA globally and indirectly activates other ion uptake, purine biosynthesis, and DNA replication/repair pathways in a salt-dependent manner finding
  • HpyA functions as a specific transcriptional regulator of metal ion balance rather than a genome-packaging histone mechanism
Experimental setups
Assay System Perturbation Readout Platform
growth curve phenotyping Hbt. salinarum WT (MDK407, Δura3) and ΔhpyA (KAD100) hpyA gene deletion; optimal (4.2M NaCl) vs reduced (3.4M NaCl) salt maximum instantaneous growth rate (μmax) R package grofit, logistic regression
phase contrast microscopy / cell circularity Hbt. salinarum WT, ΔhpyA, and complemented ΔhpyA/pKAD17 (KAD128) hpyA deletion and complementation; optimal vs reduced salt cell circularity Zeiss Axio Scope A1 microscope, MicrobeJ/ImageJ
ChIP-seq Hbt. salinarum strains AKS134 (empty vector control) and KAD128 (HpyA-HA) HA-tagged HpyA overexpression vs empty vector; exponential vs stationary phase genome-wide HpyA-DNA binding sites/peaks Illumina HiSeq4000; MACS2, bedtools, IRanges
RNA-seq Hbt. salinarum WT (MDK407) and ΔhpyA (KAD100) hpyA deletion; optimal (4.2M NaCl) vs low salt (3.4M NaCl) differential gene expression Illumina Novaseq6000; DESeq2, HTSeq
Key results
  • WT growth rate in reduced salt drops to 89% of optimal-salt rate, while ΔhpyA drops to only 67% 89% (WT) vs 67% (ΔhpyA)
  • Growth impairment of ΔhpyA under reduced salt is statistically significant relative to WT P < 0.008
  • ΔhpyA cells are significantly rounder than WT under optimal salt (non-overlapping 95% CIs)
  • ΔhpyA morphology under reduced salt is the most circular of all strain-condition combinations
  • ChIP-seq identifies reproducible, salt-dependent HpyA binding at discrete genomic sites, reproducible across ≥2 biological replicates ~60 sites
Key statistics
  • other 89% of optimal growth rate (WT μmax in reduced salt relative to optimal salt)
  • other 67% of optimal growth rate (ΔhpyA μmax in reduced salt relative to optimal salt)
  • pvalue P < 0.008 (unpaired two-sample t-test comparing ΔhpyA vs WT growth impairment in reduced salt)
  • count ~60 discrete binding sites (reproducible HpyA ChIP-seq peaks across biological replicates)
  • pvalue BH-adjusted Wald test P < 0.05 (DESeq2 criterion for significant differential expression across RNA-seq contrasts)
  • other RIN > 9.0 (RNA quality threshold for RNA-seq samples)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study employs a multi-assay quantitative design to characterize HpyA function across optimal and reduced salt conditions in Halobacterium salinarum. Growth rate differences were assessed by unpaired t-test on μmax values derived from logistic regression of nine biological replicates per strain. Cell circularity was compared across four strain-by-condition groups using non-parametric bootstrap confidence intervals of the median. ChIP-seq peaks were called with MACS2 (FDR q ≤ 0.05) and retained only if reproducible in at least two biological replicates; binding-location enrichments were tested with hypergeometric or Fisher's exact tests. RNA-seq differential expression was conducted with DESeq2 Wald tests across three pairwise contrasts, with Benjamini-Hochberg FDR correction applied throughout.

Replicationbiological Sample size9 biological replicates for growth assays; 6 biological replicates per strain-condition for RNA-seq (exact post-QC n not reported); 4 biological replicates for ChIP-seq HpyA-HA, 1 for empty-vector control; no formal power calculation described GroupsWT (Δura3) vs ΔhpyA under optimal (4.2 M NaCl) and reduced salt (3.4 M NaCl); complemented strain (ΔhpyA + hpyA-HA) assessed for morphology only Pairingunpaired Randomization/blindingnot stated DispersionCI Exact p-valuesno Effect sizesno Confidence intervalsyes Multiplicity correctionBenjamini-Hochberg FDR for DESeq2 differential expression and for functional enrichment hypergeometric tests; MACS2 internal FDR (q ≤ 0.05) for ChIP-seq peaks; no correction stated for the single growth-rate t-test or the bootstrap CI comparisons
Statistical tests used
Test Applied to n Assumptions
Unpaired two-sample t-test Comparison of maximum instantaneous growth rate (μmax) between WT and ΔhpyA in reduced salt (Figure 1B) 9 biological replicates per strain not stated
Non-parametric bootstrap resampling of median (1000 resamples, bias-corrected adjusted percentile 95% CI; inference by CI non-overlap) Comparison of cell circularity distributions across four strain × salt-condition combinations (Figure 2) na
MACS2 peak calling (internal FDR, q-value ≤ 0.05 cutoff, nomodel mode) ChIP-seq peak identification across four HpyA-HA biological replicates versus one empty-vector control replicate 4 biological replicates (HpyA-HA); 1 biological replicate (empty vector control) not stated
Hypergeometric test Enrichment of ChIP-seq peak center locations in genic versus intergenic regions; enrichment of overlap between HpyA peaks and TrmB binding locations not stated
Fisher's exact test (BEDtools 'fisher' function) Enrichment of overlap between HpyA ChIP-seq peaks and RosR binding locations not stated
DESeq2 Wald test with Benjamini-Hochberg FDR adjustment (adjusted P < 0.05) RNA-seq differential expression: three pairwise contrasts (ΔhpyA vs WT in optimal salt; ΔhpyA vs WT in reduced salt; reduced vs optimal salt in WT background) 6 biological replicates per strain-condition combination before Strong-PCA outlier removal; post-removal n not stated not stated
Hypergeometric test with Benjamini-Hochberg FDR adjustment Functional enrichment (arCOG categories) of differentially expressed genes from RNA-seq not stated
Approaches that could also have been used
  • Cell circularity across four strain-by-condition groups was compared by visual inspection of bootstrap 95% CI overlap around medians, without a formal omnibus or post-hoc test
    Could also: A Kruskal-Wallis test followed by Dunn's post-hoc test (or pairwise Wilcoxon tests) with Bonferroni or BH correction across all pairwise comparisons could also have been used — A formal omnibus test with post-hoc correction yields explicit p-values per comparison and controls the family-wise or false-discovery error rate across all six pairwise cells; it complements CI-overlap visualization and makes the inference criterion transparent to readers
  • Growth rate (μmax) differences were assessed with an unpaired two-sample t-test on nine replicates per strain, without a stated normality check
    Could also: A Mann-Whitney U (Wilcoxon rank-sum) test could also have been applied to the same μmax values — With n = 9 per group, a non-parametric alternative requires no distributional assumption about μmax; it is equally straightforward and guards against sensitivity to skew that logistic-regression-derived rates can sometimes exhibit
  • A separate t-test was reported for the reduced-salt comparison while the optimal-salt comparison was described descriptively as 'indistinguishable', treating the two salt conditions independently
    Could also: A two-way ANOVA (factors: genotype × salt condition) followed by Tukey HSD post-hoc tests could also have been used — A two-way factorial design directly tests the genotype × condition interaction — the quantity of central biological interest — within a single model and adjusts for the family of four group means simultaneously, avoiding the need for separate tests per condition
  • RNA-seq differential expression was modeled with DESeq2's negative binomial framework
    Could also: edgeR (quasi-likelihood F-test or exact test) or limma-voom could also have been applied to the same count matrix — Both are widely validated alternatives for small-n RNA-seq experiments; comparing the overlap of significant genes across two tools is sometimes used to identify a high-confidence consensus set and to assess method sensitivity
  • The growth-rate t-test result is reported as P < 0.008 and no standardized effect size is provided for any phenotypic comparison
    Could also: Cohen's d (or Hedges' g for small n) with a 95% CI could also have been reported alongside the p-value for the growth-rate comparison — Standardized effect sizes convey the magnitude of the difference independently of sample size, aiding interpretation of biological relevance and enabling future meta-analyses or power calculations
  • Outlier RNA-seq samples were identified and excluded using Strong PCA prior to DESeq2 analysis; the number and identity of removed samples and the exclusion threshold are not stated in the methods
    Could also: DESeq2's built-in Cook's distance flagging, or hierarchical clustering of sample-to-sample Euclidean or Poisson distances, could also have served as a primary or complementary quality-control approach — DESeq2's internal Cook's distance approach flags per-gene outlier observations rather than removing entire samples, preserving more data; explicitly reporting how many samples were removed and by what criterion is standard practice that aids reproducibility
Software: R/grofit · R/boot · ImageJ/MicrobeJ · FastQC 2015 · Trim Galore! 2015 · Bowtie2 · SAMtools · MACS2 2.1.1 · BEDtools · R/IRanges · Operon-Mapper · HTSeq · Strong PCA · R/DESeq2 · R/factoextra · R/ggplot2 · R/pheatmap

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34883507 (Sakrikar & Schmid 2021, NAR, gkab1175)

Paper: An archaeal histone-like protein regulates gene expression in response to salt stress. Organism: Halobacterium salinarum NRC-1 (RefSeq GCF_000006805.1 / ASM680v1). Repo: https://github.com/amyschmid/HpyA_codes (default branch main, HEAD 7babe18750a9a555d556f625977fe527711aeec0, 2021-11-19) — AUTHORS' OWN code. Data: GEO GSE182514 (RNA-seq + ChIP-seq raw reads).

Key structural fact

The repo ships its processed/intermediate input files alongside commented .Rmd analysis scripts. Each analysis directory (RNA-seq/, ChIP-seq/, growth/) contains the inputs its Rmd reads. This means the downstream statistical results are reproducible from the shipped data with R only — no raw-read realignment required for those endpoints.

In scope (pipeline-derived, attempted)

A. RNA-seq differential expression (HIGH tractability) — PRIMARY TARGET

  • Pipeline (downstream): DESeq2 1.30.1 (R 4.1.1) on shipped count matrix combineddata_21samples.csv (2621 genes x 21 samples) + combinedmeta_21samples.csv.
  • Design: ~Batch+Salt+Genotype+Salt:Genotype; DEG = padj < 0.05 (BH Wald), no log2FC cut.
  • Script: RNA-seq/2021Combined_hpyAsalt_newcounts_all-AKS.Rmd.
  • Reproduces claims: RNA_DEG_total(168), lowsalt(143), optimal(46), unique_low(121), all_conditions(21).
  • Compute: trivial (seconds). Runs on «our HPC» per hard rules; data/clone on «infra».

B. ChIP-seq peak annotation (MEDIUM tractability)

  • Downstream annotation from shipped FINAL manually-curated peak BEDs (Finalmanual_*.bed) + genome gff via ChIP-seq/2021-05-21-peak_annotate_IRanges.Rmd (IRanges/GenomicRanges).
  • Reproduces: CHIP_peaks_total(59), CHIP_genes_near(86, within 500bp).
  • NOTE: upstream peak CALLING (MACS2 v2.1.1 nomodel qval=0.05, >=2 reps + manual curation) needs raw GEO/SRA reads — heavier, depends on GSE182514 download. Attempt only if shipped beds insufficient.

C. Growth-rate analysis (MEDIUM tractability)

  • grofit 1.1-1 (shipped as growth/grofit_1.1.1-1.tar.gz, archived/removed from CRAN) on shipped growth/allgrowthcurves.xlsx / combined_grofit_*.xlsx via growth/2021-09-15-hpyA-analysis-figs-phenotypes-rev.Rmd.
  • Reproduces: GROWTH_WT_pct(89%), GROWTH_KO_pct(67%), GROWTH_ttest_p(P<0.008).

Out of scope (not attempted / lower priority)

  • Raw-read upstream steps (TrimGalore!, Bowtie2 alignment, HTSeq counting, MACS2/mosaics peak calling) from GSE182514 — only attempted if a downstream endpoint cannot be reached from shipped data.
  • Wet-lab assays, microscopy/circularity imaging, qPCR, protein work — manual/experimental, not pipeline.
  • Figure cosmetics (heatmaps, volcano styling) — not quantitative claims.

Discrepancy to flag

README/dependencies list mosaics (R) for ChIP, while Methods text says MACS2 v2.1.1; both appear in the dependency files. Will note which the shipped beds actually came from.

Execution plan

  1. («our HPC» front1) clone repo on «infra» work dir; build conda R 4.1.x env (DESeq2, IRanges, grofit-from-tarball).
  2. Run A (DESeq2) first — submit small SLURM job (or front1 R, compute is seconds). Pull DEG counts.
  3. Run B + C. Compare to claims.tsv; fill agreement.json + AUDIT.md.
  4. Profile GSE182514 in dataset_profile.json (N reported vs observed, completeness, QC).

Blocker as of first pass: «our HPC» «host» ssh timing out (tunnel down). Not touching VPN; retrying. All prep (claims, scope, driver scripts) done offline so the run is one submit away when the tunnel returns.

Figures / tables: Fig 1B
RNA_DEG_total
Reported
168
Reproduced
168 (shipped TableS4); 248 from-scratch 21-sample union
exact
RNA_DEG_lowsalt_KOvsWT
Reported
143
Reproduced
143 (TableS4); 221 from-scratch; 140/143 final recovered in my run
exact
RNA_DEG_optsalt_KOvsWT
Reported
46
Reproduced
46 (TableS4); 52 from-scratch
exact
RNA_DEG_all_conditions
Reported
21
Reproduced
21 (TableS4 Optimal_low)
exact
RNA_DEG_unique_low
Reported
121
Reproduced
122 (TableS4 Low-only)
within tolerance
CHIP_peaks_total
Reported
59
Reproduced
59 (shipped TableS3 Basic_S3); 64 from-scratch IRanges
exact
CHIP_genes_near
Reported
86
Reproduced
88-89 (TableS3 Full_S3); 96 from-scratch
within tolerance
GROWTH_WT_pct
Reported
89%
Reproduced
93.2% mean / 88.4% excl-outlier / 93.5% median
within tolerance
GROWTH_KO_pct
Reported
67%
Reproduced
66.4% median / 63.8% mean
within tolerance
GROWTH_ttest_p
Reported
P<0.008
Reproduced
P=0.0011 (equal-var) / 0.0013 (Welch)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

295.6 k
tokens (I/O) · 22.9 M incl. cache
84 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.