High-throughput sequencing SELEX for the determination of DNA-binding protein specificities in vitro.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH + 1:1. STAR Protocols methods paper for the eme_selex HT-SELEX k-mer analysis (SALL4 ZFC4). The repo (commit df08bb3) ships the 21 trimmed input FASTA AND the expected output tables, so the pipeline half is fully reproducible. Re-ran eme_selex v0.4 kmer_fraction_from_file(k=5) on the 21 shipped FASTA on a «our HPC» compute node («job») and compared to the shipped tables: R1 mean 5-mer fractions (3584 values) and R2 library-paired fold-change (3584 values) both reproduced BIT-EXACT (max abs diff ~1e-16, Pearson=1.0); value_std matches at pandas ddof=1 (deduced). R3 AT-rich progressive enrichment through ZFC4 cycles 1->3->6 reproduced as a monotonic trend (1.62->3.82->4.80) with AT-rich top motifs, no-protein control flat - qualitatively confirms Fig 7 (graded partial; no single printed number). R4 (Fig 9 coverage Spearman R^2) NOT attempted - needs the high-coverage E-MTAB-11484 deposit; honestly recorded, not fabricated. Dataset E-MTAB-9236 profiled: open, 42 assays in deposit vs 21 analysed (subset shipped in repo), parses cleanly, ~25k reads/sample (~20k stated). Verdict: pipeline-derived results reproduce 1:1.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-26
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnet- ★ HT-SELEX enables unbiased, in vitro determination of preferred DNA target motifs for DNA-binding proteins by iterative selection and PCR amplification of bound oligonucleotides method
- ★ The eme_selex bioinformatic pipeline detects promiscuous DNA binding by quantifying enrichment of all possible k-mers resource
- ★ The protocol is optimized for Illumina sequencing compatibility using barcoded primers for multiplexed HT-SELEX samples method
- ★ SALL4 C2H2 zinc-finger cluster 4 (ZFC4) promiscuously binds multiple AT-rich sequences, requiring more SELEX cycles than typical transcription factors finding
- A hexahistidine tag is preferred for bait proteins because of its small size and cost-efficient purification by IMAC, unlike larger tags (GST, MBP) that may impair DNA binding method
- A 20 bp random insert covers binding motifs for the vast majority of sequence-specific DNA-binding proteins method
- Two to three SELEX cycles are usually sufficient for transcription factors, while SALL4 ZFC4 required up to six cycles finding
- Oligonucleotides for the random library must use standard desalting only, without PAGE/HPLC purification, to avoid biasing sequence randomness method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| PCR amplification of random oligonucleotide library (cycle 0 generation) | synthetic single-stranded DNA templates (Random library 1/2/3) | none | double-stranded DNA library formation | Phusion DNA Polymerase, Alpha Cycler 4 PCR machine |
| Polyacrylamide gel electrophoresis (10%) | PCR-generated DNA libraries | none | band size and absence of heteroduplexes | Mini-PROTEAN electrophoresis system, ethidium bromide staining |
| Spectrophotometry | purified SELEX DNA libraries | none | DNA concentration and integrity | NanoDrop ND-1000 |
| Recombinant protein purification (IMAC) | histidine-tagged SALL4 ZFC4 expressed in bacteria | none | purified affinity-tagged DNA-binding protein | Ni Sepharose 6 Fast Flow resin |
| SELEX (protein-DNA complex selection) | purified His-SALL4 ZFC4 with random/enriched DNA library | protein addition vs. no-protein negative control | protein-bound DNA sequences captured on resin | Ni Sepharose 6 Fast Flow resin, rotating wheel incubation |
| PCR amplification of enriched (bead-bound) SELEX DNA | protein-DNA complexes on resin | variable PCR cycle number (8x/12x etc.) | amplified enriched DNA library | Phusion DNA Polymerase |
| High-throughput sequencing (HT-SELEX) | barcoded, pooled SELEX libraries | SELEX cycle number (0-6) | sequence reads for k-mer/motif enrichment analysis | Illumina sequencing |
| Microfluidic capillary electrophoresis (library QC) | HT-SELEX sequencing libraries | none | library size distribution and quality | Agilent 2100 Bioanalyzer, High Sensitivity DNA Kit |
- – Cycle 0 PCR libraries showed a single band at 83 bp with no detectable heteroduplexes on 10% polyacrylamide gel 83 bp
- – Each PCR reaction of the random library yields approximately 500 ng of DNA ~500 ng/reaction
- – 24x PCR reactions generate enough material (≈12 μg DNA) for 8x SELEX samples ≈12 μg total
- – SALL4 ZFC4 required up to 6 SELEX cycles due to promiscuous AT-rich sequence binding, more than the 2-3 cycles typical for transcription factors up to 6 cycles vs 2-3
- – eme_selex quantifies enrichment of all possible k-mers to identify promiscuous DNA binding
- other 83 bp (expected size of cycle 0 double-stranded DNA library band on gel)
- mean ~500 ng DNA per PCR reaction (yield of cycle 0 library amplification)
- count ≈12 μg DNA from 24x PCR reactions (total library DNA generated for 8x SELEX samples)
- other 1.5 μg library/sample (amount of cycle 0 random library DNA needed to initiate first SELEX cycle)
- other 200 ng (amount of SELEX library used from cycle N-1 for subsequent SELEX cycles)
- count 2-3 cycles (typical TFs) vs up to 6 cycles (SALL4 ZFC4) (number of SELEX cycles required for successful HT-SELEX)
- other 20 bp (length of random insert in SELEX library oligonucleotides)
- other at least two mismatches (required difference between 8 bp barcodes in Seqlib RV primers)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a STAR Protocols methods paper describing the step-by-step wet-lab and bioinformatic procedure for HT-SELEX (high-throughput sequencing SELEX) to characterize DNA-binding protein specificities, rather than a hypothesis-testing study. Three independent random oligonucleotide libraries are used as technical replicates, iterative selection/PCR cycles enrich bound sequences, and a custom pipeline (eme_selex) quantifies k-mer enrichment. No formal inferential statistical tests, p-values, or measures of dispersion are described in this excerpt; analysis is framed as bioinformatic enrichment/ranking of k-mers rather than classical significance testing.
-
K-mer enrichment is assessed via the custom eme_selex pipeline without a stated formal statistical test for calling enrichment significant.↳ Could also: A formal enrichment statistic such as a hypergeometric test, binomial test, or z-score comparing observed vs. expected k-mer frequencies from the random (cycle 0) library — This would provide a quantitative significance threshold for labeling a k-mer as enriched, beyond ranking or raw frequency counts.
-
Three independent random-oligonucleotide libraries are described as technical replicates without a stated method for quantifying agreement between them.↳ Could also: Reporting a reproducibility metric such as Pearson/Spearman correlation of k-mer frequencies or a coefficient of variation across the three replicate libraries — This would let readers directly gauge how consistent motif enrichment is across the independent libraries.
-
A no-protein negative control library is included, but the text does not state how it is statistically compared to protein-bound samples.↳ Could also: A differential enrichment test between protein and no-protein libraries, such as Fisher's exact test on k-mer/read counts or a DESeq2-style Wald test — This would formally distinguish protein-specific sequence enrichment from background PCR or amplification bias captured by the negative control.
-
No measure of uncertainty (e.g., confidence interval) is reported alongside k-mer enrichment values in this excerpt.↳ Could also: Bootstrapped confidence intervals on k-mer frequency estimates — This would communicate the precision of enrichment estimates given finite sequencing depth per library.
-
The number of SELEX cycles performed (2-6) is chosen based on prior experience/expected convergence rather than a stated quantitative stopping rule.↳ Could also: A quantitative convergence metric, such as a plateau in Shannon entropy or enrichment score tracked across successive cycles — This would offer an objective, reproducible criterion for determining when the selection process has converged, applicable across different proteins.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 35776646
Title: High-throughput sequencing SELEX for the determination of DNA-binding
protein specificities in vitro.
Venue: STAR Protocols (2022) · DOI 10.1016/j.xpro.2022.101490 · PMCID PMC9243297
Code: https://github.com/kashyapchhatbar/eme_selex (default branch main, pushed 2023-07-12)
Data: ArrayExpress E-MTAB-9236 (raw HT-SELEX fastq, 21 samples, SALL4 ZFC4, human)
Supplementary high-coverage: E-MTAB-11484 (~3M reads/sample, coverage analysis)
What kind of paper
This is a STAR Protocols methods paper, not a discovery paper. It documents a
wet-lab HT-SELEX protocol PLUS a bioinformatic analysis workflow (eme_selex).
The bioinformatic half is fully shipped and reproducible: the repo carries the
input FASTA, the analysis code, and the expected output tables.
Pipeline-derived results (IN SCOPE)
The repo docs/ ships a complete worked example for SALL4 ZFC4 HT-SELEX:
- Input: 21 quality-trimmed FASTA files
docs/data/fasta/RV{01,02,03,06,07,10,11,14,15,18,19,22,23,26,27,30,31,34,35,38,39}.fa.gz(one per metadataSampleName). - Tool:
eme_selex.eme.kmer_fraction_from_file(file, k, top=50)→ per-samplecounts(reads containing each canonical k-mer, ≤1 per read),fractions(count / n_reads),models(top-50 PFMs). - Expected outputs shipped in repo to compare against:
| Result | Expected file | Pipeline | Figure/role |
|---|---|---|---|
| R1: mean 5-mer fractions per (cycle, protein) | docs/data/melt_fractions_mean.tsv |
eme_selex kf(k=5) → mean over 3 libs | Fig 7 enrichment |
| R2: mean 5-mer fold-change vs cycle 0 | docs/data/melt_fold_change_mean.tsv |
R1 / cycle-0 mean | Fig 7 enrichment |
| R3: AT-rich k-mer progressive enrichment through cycles 1→3→6 (ZFC4) | derived from R2 (AT column) | trend test | Fig 7 (qualitative) |
| R4: Spearman R² vs sub-sampling depth & k (5–10) | docs/data/spearman_r2_k_5_10.tsv |
sub-sample reads, rank-corr | Fig 9 coverage |
R1–R3 reproduce from the 21 shipped FASTA alone (clean 1:1 against shipped tables).
R4 needs the high-coverage E-MTAB-11484 deposit (the full/500000/50000/miseq
depths) — attempted as a stretch; if that deposit is impractical to obtain, R4 is
recorded as partially/not attempted with the reason (NOT fabricated).
Stretch — from-raw validation
Download raw fastq from E-MTAB-9236, run the documented flexbar trim
(--post-trim-length 20 --min-read-length 20 --qtrim-threshold 30 --fasta-output),
and verify the trimmed FASTA reproduce the repo's shipped docs/data/fasta/*.fa.gz
(validates the preprocessing step end-to-end). This closes the loop from the public
accession to the shipped intermediate.
OUT OF SCOPE (wet-lab / manual / not pipeline)
- All wet-lab steps (PCR cycling, bead binding, library prep, Bioanalyzer profiles Fig 6, gel images) — physical protocol, not computational.
- Motif-logo rendering (logomaker) is cosmetic visualization of R1's PFMs; the PFM models themselves are in scope, the rendered logo image is not graded.
- Read-count / sequencing-depth statements ("20,000 reads/sample average", "10,000–50,000 sufficient") are dataset properties — checked under DATASET PROFILING, not as pipeline outputs.
Datasets to profile
- E-MTAB-9236 — raw HT-SELEX fastq, 21 samples (3 cycle0 libs + cycle 1/3/6 × {ZFC4, no-protein} × 3 libs). Profiled for N reported (21) vs observed.
- E-MTAB-11484 — high-coverage resequencing (coverage analysis). Profiled if reached.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.