Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

High-throughput sequencing SELEX for the determination of DNA-binding protein specificities in vitro.

STAR Protoc · 2022
65/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
65/100
Reproducibility score
0.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 27% of all assessed papers rank 843 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH + 1:1. STAR Protocols methods paper for the eme_selex HT-SELEX k-mer analysis (SALL4 ZFC4). The repo (commit df08bb3) ships the 21 trimmed input FASTA AND the expected output tables, so the pipeline half is fully reproducible. Re-ran eme_selex v0.4 kmer_fraction_from_file(k=5) on the 21 shipped FASTA on a «our HPC» compute node («job») and compared to the shipped tables: R1 mean 5-mer fractions (3584 values) and R2 library-paired fold-change (3584 values) both reproduced BIT-EXACT (max abs diff ~1e-16, Pearson=1.0); value_std matches at pandas ddof=1 (deduced). R3 AT-rich progressive enrichment through ZFC4 cycles 1->3->6 reproduced as a monotonic trend (1.62->3.82->4.80) with AT-rich top motifs, no-protein control flat - qualitatively confirms Fig 7 (graded partial; no single printed number). R4 (Fig 9 coverage Spearman R^2) NOT attempted - needs the high-coverage E-MTAB-11484 deposit; honestly recorded, not fabricated. Dataset E-MTAB-9236 profiled: open, 42 assays in deposit vs 21 analysed (subset shipped in repo), parses cleanly, ~25k reads/sample (~20k stated). Verdict: pipeline-derived results reproduce 1:1.

💻 Code ↗ 🗄 Data: E-MTAB-9236

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-26
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Core claims
  • HT-SELEX enables unbiased, in vitro determination of preferred DNA target motifs for DNA-binding proteins by iterative selection and PCR amplification of bound oligonucleotides method
  • The eme_selex bioinformatic pipeline detects promiscuous DNA binding by quantifying enrichment of all possible k-mers resource
  • The protocol is optimized for Illumina sequencing compatibility using barcoded primers for multiplexed HT-SELEX samples method
  • SALL4 C2H2 zinc-finger cluster 4 (ZFC4) promiscuously binds multiple AT-rich sequences, requiring more SELEX cycles than typical transcription factors finding
  • A hexahistidine tag is preferred for bait proteins because of its small size and cost-efficient purification by IMAC, unlike larger tags (GST, MBP) that may impair DNA binding method
  • A 20 bp random insert covers binding motifs for the vast majority of sequence-specific DNA-binding proteins method
  • Two to three SELEX cycles are usually sufficient for transcription factors, while SALL4 ZFC4 required up to six cycles finding
  • Oligonucleotides for the random library must use standard desalting only, without PAGE/HPLC purification, to avoid biasing sequence randomness method
Experimental setups
Assay System Perturbation Readout Platform
PCR amplification of random oligonucleotide library (cycle 0 generation) synthetic single-stranded DNA templates (Random library 1/2/3) none double-stranded DNA library formation Phusion DNA Polymerase, Alpha Cycler 4 PCR machine
Polyacrylamide gel electrophoresis (10%) PCR-generated DNA libraries none band size and absence of heteroduplexes Mini-PROTEAN electrophoresis system, ethidium bromide staining
Spectrophotometry purified SELEX DNA libraries none DNA concentration and integrity NanoDrop ND-1000
Recombinant protein purification (IMAC) histidine-tagged SALL4 ZFC4 expressed in bacteria none purified affinity-tagged DNA-binding protein Ni Sepharose 6 Fast Flow resin
SELEX (protein-DNA complex selection) purified His-SALL4 ZFC4 with random/enriched DNA library protein addition vs. no-protein negative control protein-bound DNA sequences captured on resin Ni Sepharose 6 Fast Flow resin, rotating wheel incubation
PCR amplification of enriched (bead-bound) SELEX DNA protein-DNA complexes on resin variable PCR cycle number (8x/12x etc.) amplified enriched DNA library Phusion DNA Polymerase
High-throughput sequencing (HT-SELEX) barcoded, pooled SELEX libraries SELEX cycle number (0-6) sequence reads for k-mer/motif enrichment analysis Illumina sequencing
Microfluidic capillary electrophoresis (library QC) HT-SELEX sequencing libraries none library size distribution and quality Agilent 2100 Bioanalyzer, High Sensitivity DNA Kit
Key results
  • Cycle 0 PCR libraries showed a single band at 83 bp with no detectable heteroduplexes on 10% polyacrylamide gel 83 bp
  • Each PCR reaction of the random library yields approximately 500 ng of DNA ~500 ng/reaction
  • 24x PCR reactions generate enough material (≈12 μg DNA) for 8x SELEX samples ≈12 μg total
  • SALL4 ZFC4 required up to 6 SELEX cycles due to promiscuous AT-rich sequence binding, more than the 2-3 cycles typical for transcription factors up to 6 cycles vs 2-3
  • eme_selex quantifies enrichment of all possible k-mers to identify promiscuous DNA binding
Key statistics
  • other 83 bp (expected size of cycle 0 double-stranded DNA library band on gel)
  • mean ~500 ng DNA per PCR reaction (yield of cycle 0 library amplification)
  • count ≈12 μg DNA from 24x PCR reactions (total library DNA generated for 8x SELEX samples)
  • other 1.5 μg library/sample (amount of cycle 0 random library DNA needed to initiate first SELEX cycle)
  • other 200 ng (amount of SELEX library used from cycle N-1 for subsequent SELEX cycles)
  • count 2-3 cycles (typical TFs) vs up to 6 cycles (SALL4 ZFC4) (number of SELEX cycles required for successful HT-SELEX)
  • other 20 bp (length of random insert in SELEX library oligonucleotides)
  • other at least two mismatches (required difference between 8 bp barcodes in Seqlib RV primers)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a STAR Protocols methods paper describing the step-by-step wet-lab and bioinformatic procedure for HT-SELEX (high-throughput sequencing SELEX) to characterize DNA-binding protein specificities, rather than a hypothesis-testing study. Three independent random oligonucleotide libraries are used as technical replicates, iterative selection/PCR cycles enrich bound sequences, and a custom pipeline (eme_selex) quantifies k-mer enrichment. No formal inferential statistical tests, p-values, or measures of dispersion are described in this excerpt; analysis is framed as bioinformatic enrichment/ranking of k-mers rather than classical significance testing.

Replicationtechnical Sample sizeThree independent random-oligonucleotide libraries (library 1/2/3) are used as technical replicates per condition, plus a no-protein negative control for each library; no formal sample-size or power calculation is described. Groupsprotein-bound SELEX-selected libraries vs. a no-protein negative control, across successive selection cycles (2-6 cycles) Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • K-mer enrichment is assessed via the custom eme_selex pipeline without a stated formal statistical test for calling enrichment significant.
    Could also: A formal enrichment statistic such as a hypergeometric test, binomial test, or z-score comparing observed vs. expected k-mer frequencies from the random (cycle 0) library — This would provide a quantitative significance threshold for labeling a k-mer as enriched, beyond ranking or raw frequency counts.
  • Three independent random-oligonucleotide libraries are described as technical replicates without a stated method for quantifying agreement between them.
    Could also: Reporting a reproducibility metric such as Pearson/Spearman correlation of k-mer frequencies or a coefficient of variation across the three replicate libraries — This would let readers directly gauge how consistent motif enrichment is across the independent libraries.
  • A no-protein negative control library is included, but the text does not state how it is statistically compared to protein-bound samples.
    Could also: A differential enrichment test between protein and no-protein libraries, such as Fisher's exact test on k-mer/read counts or a DESeq2-style Wald test — This would formally distinguish protein-specific sequence enrichment from background PCR or amplification bias captured by the negative control.
  • No measure of uncertainty (e.g., confidence interval) is reported alongside k-mer enrichment values in this excerpt.
    Could also: Bootstrapped confidence intervals on k-mer frequency estimates — This would communicate the precision of enrichment estimates given finite sequencing depth per library.
  • The number of SELEX cycles performed (2-6) is chosen based on prior experience/expected convergence rather than a stated quantitative stopping rule.
    Could also: A quantitative convergence metric, such as a plateau in Shannon entropy or enrichment score tracked across successive cycles — This would offer an objective, reproducible criterion for determining when the selection process has converged, applicable across different proteins.
Software: eme_selex (custom pipeline) · Flexbar 3.5.0 · Snakemake · Jupyterlab · Pandas · Seaborn

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 35776646

Title: High-throughput sequencing SELEX for the determination of DNA-binding protein specificities in vitro. Venue: STAR Protocols (2022) · DOI 10.1016/j.xpro.2022.101490 · PMCID PMC9243297 Code: https://github.com/kashyapchhatbar/eme_selex (default branch main, pushed 2023-07-12) Data: ArrayExpress E-MTAB-9236 (raw HT-SELEX fastq, 21 samples, SALL4 ZFC4, human) Supplementary high-coverage: E-MTAB-11484 (~3M reads/sample, coverage analysis)

What kind of paper

This is a STAR Protocols methods paper, not a discovery paper. It documents a wet-lab HT-SELEX protocol PLUS a bioinformatic analysis workflow (eme_selex). The bioinformatic half is fully shipped and reproducible: the repo carries the input FASTA, the analysis code, and the expected output tables.

Pipeline-derived results (IN SCOPE)

The repo docs/ ships a complete worked example for SALL4 ZFC4 HT-SELEX:

  • Input: 21 quality-trimmed FASTA files docs/data/fasta/RV{01,02,03,06,07,10,11,14,15,18,19,22,23,26,27,30,31,34,35,38,39}.fa.gz (one per metadata SampleName).
  • Tool: eme_selex.eme.kmer_fraction_from_file(file, k, top=50) → per-sample counts (reads containing each canonical k-mer, ≤1 per read), fractions (count / n_reads), models (top-50 PFMs).
  • Expected outputs shipped in repo to compare against:
Result Expected file Pipeline Figure/role
R1: mean 5-mer fractions per (cycle, protein) docs/data/melt_fractions_mean.tsv eme_selex kf(k=5) → mean over 3 libs Fig 7 enrichment
R2: mean 5-mer fold-change vs cycle 0 docs/data/melt_fold_change_mean.tsv R1 / cycle-0 mean Fig 7 enrichment
R3: AT-rich k-mer progressive enrichment through cycles 1→3→6 (ZFC4) derived from R2 (AT column) trend test Fig 7 (qualitative)
R4: Spearman R² vs sub-sampling depth & k (5–10) docs/data/spearman_r2_k_5_10.tsv sub-sample reads, rank-corr Fig 9 coverage

R1–R3 reproduce from the 21 shipped FASTA alone (clean 1:1 against shipped tables). R4 needs the high-coverage E-MTAB-11484 deposit (the full/500000/50000/miseq depths) — attempted as a stretch; if that deposit is impractical to obtain, R4 is recorded as partially/not attempted with the reason (NOT fabricated).

Stretch — from-raw validation

Download raw fastq from E-MTAB-9236, run the documented flexbar trim (--post-trim-length 20 --min-read-length 20 --qtrim-threshold 30 --fasta-output), and verify the trimmed FASTA reproduce the repo's shipped docs/data/fasta/*.fa.gz (validates the preprocessing step end-to-end). This closes the loop from the public accession to the shipped intermediate.

OUT OF SCOPE (wet-lab / manual / not pipeline)

  • All wet-lab steps (PCR cycling, bead binding, library prep, Bioanalyzer profiles Fig 6, gel images) — physical protocol, not computational.
  • Motif-logo rendering (logomaker) is cosmetic visualization of R1's PFMs; the PFM models themselves are in scope, the rendered logo image is not graded.
  • Read-count / sequencing-depth statements ("20,000 reads/sample average", "10,000–50,000 sufficient") are dataset properties — checked under DATASET PROFILING, not as pipeline outputs.

Datasets to profile

  • E-MTAB-9236 — raw HT-SELEX fastq, 21 samples (3 cycle0 libs + cycle 1/3/6 × {ZFC4, no-protein} × 3 libs). Profiled for N reported (21) vs observed.
  • E-MTAB-11484 — high-coverage resequencing (coverage analysis). Profiled if reached.
Figures / tables: Fig 7Fig 9
R1
Reported
melt_fractions_mean.tsv (3584 mean 5-mer fraction values, 7 cycle/protein groups x 512 canonical 5-mers)
Reproduced
3584/3584 reproduced, max_abs_diff=9.97e-17, Pearson=1.0; value_std matches at pandas ddof=1
exact
R2
Reported
melt_fold_change_mean.tsv (3584 library-paired fold-change values vs cycle-0 no-protein)
Reproduced
3584/3584 reproduced, max_abs_diff=4.44e-16, Pearson=1.0; library-paired definition confirmed
exact
R3
Reported
SALL4 ZFC4 progressively enriches AT-rich 5-mers across cycles 1->3->6 (Fig 7)
Reproduced
ZFC4 fully-AT 5-mer mean fold-change 1.62->3.82->4.80 (monotonic); no-protein 1.06->1.22->1.43; top cycle-6 motifs AATAT/ATATT/TAATA (AT-rich)
partial
R4
Reported
spearman_r2_k_5_10.tsv (Spearman R^2 across coverage depths, k=5..10, Fig 9)
Reproduced
not attempted (needs high-coverage deposit E-MTAB-11484 download + subsampling)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.