oPOSSUM-3: advanced analysis of regulatory motif over-representation across genes or ChIP-Seq datasets.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
FINAL verdict: reproduced the self-contained SSA (Sequence Set Analysis) mode of oPOSSUM-3 against the paper's own Sox2 and c-Myc GSE11431 mouse ES-cell ChIP-seq case studies, and graded all claims resolvable within this session's scope. Sox2 SSA correctly recovers the Sox2/Pou5f1::Sox2 motif as the top over-represented hit (within-tol). The paper's specific quantitative claims are a mixed bag under close scrutiny: Z-score/Fisher-score top-10 concordance for Sox2 falls well short of the paper's 82% (we get 50%, mismatch); background-size robustness (Spearman rho) mostly clears the paper's >=0.90 floor except at the undersized 0.5x ratio (partial); same-size replicate stability splits by score type, with Fisher passing (0.9156 vs >=0.86) but Z-score falling short (0.9025 vs >=0.97, partial). Gene-based analysis (a separate oPOSSUM-3 feature) is out of scope: it depends on an external Wasserman-lab MySQL database unreachable from this environment, with no self-contained data mirror available (error/documented external blocker). The c-Myc case, including the paper's strongest claimed result (100% Z/Fisher top-10 concordance), could NOT be graded: at finalization time, a direct sacct/squeue/log check confirmed «job» (the R-PATH-fixed rerun) was still mid-scan (~50 min elapsed of a run whose predecessor took 8h11m to reach the same point) -- not 'done for hours' as assumed when the finalize instruction was issued. Rather than keep perfecting or wait out the remaining hours, this result is finalized now as instructed: an honest partial reproduction with one claim explicitly left ungraded and disclosed, not fabricated.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper asks whether a new generation of motif over-representation software — incorporating multi-species phastCons conservation, an updated JASPAR profile set, TFBS profile clustering, and sequence-based (ChIP-Seq) input — can identify the transcription factors mediating regulation of co-expressed gene sets or co-recovered ChIP-Seq sequences.
- ★ oPOSSUM-3 is a web-accessible system that identifies over-represented TFBS and TFBS families in DNA sequences of co-expressed genes or in sequences from high-throughput methods such as ChIP-Seq. resource
- ★ Validation with known sets of co-regulated genes and published ChIP-Seq data demonstrates the capacity of oPOSSUM-3 to identify mediating transcription factors for co-regulated genes or co-recovered sequences. finding
- ★ TFBS Cluster Analysis (TCA) and anchored Combination TFBS Cluster Analysis (aCTCA) address the challenge of homologous TFs with highly similar or identical binding specificity by reporting results as TFBS sequence patterns rather than individual profile names. method
- ★ Over-representation is measured by two complementary statistics — a Z-score based on the normal approximation to the binomial distribution (counting predicted binding site nucleotides) and a Fisher score from a one-tailed Fisher exact probability (counting sequences/genes with the site). method
- ★ For sequence-based analysis, a third measure, a Kolmogorov-Smirnoff centrality score, compares the empirical distributions of TFBS distances to the maximum confidence position (DistMCP) between target and background sets; functional TFBSs are expected to cluster around zero in target sequences and be randomly distributed in background sequences. method
- ★ SSA in version 3 replaces pairwise sequence alignment-based phylogenetic footprinting with phastCons scores; conserved regions are genomic regions of at least 20 nucleotides with average phastCons above threshold, and TFBS searches are restricted to motifs within or overlapping (by at least 1 bp) these regions. method
- anchored Combination Site Analysis (aCSA) restricts over-representation assessment to sequence regions proximal to predicted TFBS of a user-specified anchor TF, counting pairings of the anchor TFBS with all secondary TFBSs within an inter-binding site distance; self-interactions can be observed when anchor TFBSs occur in clusters. method
- Clustering TF binding profiles within annotated JASPAR structural families (rather than by structural class alone) is necessary because the extent of binding sequence similarity varies among classes — e.g. zinc fingers have low profile similarity among members. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| TFBS motif over-representation analysis, gene-based Single Site Analysis (SSA) with phastCons conservation filtering | Human (hg19), mouse (mm9), fruit fly (dm3), nematode C. elegans (ce6), and yeast S. cerevisiae genome/gene annotation databases; reference sets of co-regulated genes | none | Z-scores (predicted binding site nucleotide counts) and Fisher scores (number of genes containing the TFBS) for each JASPAR profile, with FDR-adjusted p-values | oPOSSUM-3 web system; Ensembl v64 (v54 for C. elegans) annotation and sequence; Wormbase WS200 operon annotation; UCSC phastCons46wayPlacental / phastCons30wayPlacental / phastCons15way / phastCons6way; JASPAR 2010 CORE and PBM collections plus custom PENDING collection; R statistics package |
| Sequence-based TFBS motif over-representation analysis (SSA) on ChIP-Seq-derived sequence sets | Published large-scale ChIP-Seq sequence collections, including the Nfe2L2 ChIP-Seq dataset of Malhotra et al. (2010) | none (user-supplied foreground TF-bound vs. background control sequences; no conservation filtering applied) | Z-scores, Fisher scores, and Kolmogorov-Smirnoff centrality scores based on TFBS distance to the maximum confidence position (DistMCP) | oPOSSUM-3 sequence-based system; JASPAR 2010 profiles; R statistics package |
| anchored Combination Site Analysis (aCSA) / anchored Combination TFBS Cluster Analysis (aCTCA) — over-representation of TFBS pairs | Gene-based (annotated genomes: human, mouse, fruit fly, nematode, yeast) or user-supplied sequence sets | user-specified anchoring TF and inter-binding site distance parameter | Counts and over-representation scores for pairings of the anchor TFBS with secondary TFBSs (or cluster pairs) within the specified inter-binding site distance | oPOSSUM-3; JASPAR 2010 profiles |
| TFBS profile pairwise similarity scoring and clustering (TFBS Cluster Analysis, TCA) | JASPAR 2010 TFBS profile collections annotated for TF structural class and family | clustering parameters: cluster score threshold T and radius margin R | Pairwise similarity score table and resulting TFBS clusters; overlapping TFBS hits of the same cluster merged into a single cluster hit for over-representation counting | MatrixAligner similarity scoring program (Sandelin et al. 2003) |
| Pre-computation of putative TFBS locations across species-specific promoter search regions | Human and mouse (10,000 bp upstream and 10,000 bp downstream of Ensembl-annotated TSS); fruit fly (3000 bp each direction); nematode (1500 bp each direction); yeast (1000 bp upstream of TSS to the 3' end of each gene) | none; exons excluded for invertebrates in gene-based pre-calculations | Pre-computed TFBS hit locations and conserved regions used for gene-based analyses | Ensembl v64 (C. elegans v54); UCSC Genome Browser phastCons score sets |
| Web service usage tracking | oPOSSUM-2 public web server users | none | Number of unique monthly users, excluding automated internet search software | — |
- – Assessment against reference sets of co-regulated genes and large-scale ChIP-Seq sequence collections exemplifies the utility of oPOSSUM-3 for identifying mediating TFs.
- – oPOSSUM-2 was found to perform well in an independent assessment of motif over-representation analysis tools (Meng et al. 2010).
- ▲ In the Nfe2L2 ChIP-Seq dataset (Malhotra et al. 2010), target-set binding site distances to the maximum confidence position are clustered around zero relative to background, illustrating centrality enrichment.
- – Clustering of JASPAR profiles by family similarity condenses redundant, near-identical profiles so that independent enriched profiles can be identified rather than results being dominated by subsets of nearly identical profiles.
- ▲ Replacement of pairwise-alignment phylogenetic footprinting with multi-species phastCons scores improves alignment quality and greatly increases the number of genes available for analysis.
- ▲ The JASPAR 2010 update provides a significant increase in non-vertebrate profiles, permitting extension of regulatory analysis software to non-vertebrate species such as insects.
- – FDR-adjusted p-values were computed for Z-scores and Fisher scores in Tables 2–5 for completeness but do not affect the relative rankings.
- count 340 unique users per month (on average, excluding automated internet search software) (oPOSSUM-2 web service usage)
- count 250 profiles from the JASPAR 2010 CORE collection, 184 from the PBM collection, and 4 from the custom PENDING collection (Profiles analyzed in the TFBS clustering procedure)
- other cluster score threshold T = 1.8; radius margin R = 0.1 (MatrixAligner-based TFBS profile clustering parameters, selected empirically from the distribution of pairwise similarity scores)
- other phastCons thresholds of 0.4 and 0.6; conserved regions of at least 20 nucleotides; TFBS overlap of at least 1 bp (Definition of conserved regions for gene-based SSA)
- other continuity correction of 0.5; Z = (x−µ−0.5)/σ, µ = BC, C = n/N, P = B/N, σ = sqrt(nP(1−P)) (Z-score calculation using normal approximation to the binomial distribution)
- count 10,000 bp upstream and 10,000 bp downstream of TSS (human, mouse); 3000 bp each direction (fruit fly); 1500 bp each direction (nematode); 1000 bp upstream to 3' gene end (yeast) (Species-specific search regions for gene-based TFBS pre-computation)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper describes a bioinformatics software system (oPOSSUM-3) that ranks transcription factor binding site (TFBS) motifs by comparing their occurrence in a foreground (target) gene or sequence set against a background (control) set. Over-representation is quantified with two independent scoring methods (a Z-score and a Fisher score) and, for sequence-based ChIP-Seq analysis, a third Kolmogorov-Smirnov (KS)-based centrality score. Benjamini-Hochberg FDR-adjusted p-values are reported alongside the raw scores in the validation tables, though the authors state these are provided for completeness and do not affect the relative rankings used to interpret results.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Z-score based on normal approximation to the binomial distribution (with 0.5 continuity correction) | TFBS over-representation in foreground vs. background nucleotide counts (SSA, aCSA, TCA, aCTCA; Tables 2-5) | total nucleotides in foreground (n) and background (N) sequence sets, as defined in the Z-score formula | stated |
| Fisher exact probability test (hypergeometric distribution) | presence/absence of a TFBS across foreground vs. background sequences, same analyses as above (Tables 2-5) | number of foreground and background sequences containing/lacking the TFBS | stated |
| Kolmogorov-Smirnov (KS) test | comparison of empirical distributions of TFBS distance-to-maximum-confidence-position (DistMCP) between target and background sequences, sequence-based/ChIP-Seq analysis (e.g., Nfe2L2 dataset, Figure S2) | not explicitly stated as an n; based on the set of TFBS-to-MCP distances in target vs. background sequences | not stated |
-
TFBS enrichment in nucleotide counts was assessed with a Z-score derived from a normal approximation to the binomial distribution with a continuity correction.↳ Could also: An exact binomial test or exact Poisson test for count-based enrichment — An exact test avoids relying on the normal approximation and could also be used directly, particularly useful when foreground counts are small, where the approximation is least accurate.
-
Presence/absence enrichment of TFBS across sequences was assessed with a one-tailed Fisher exact test based on the hypergeometric distribution.↳ Could also: A chi-square test of independence, or a hypergeometric-based gene-set enrichment approach such as GSEA-style rank statistics — A chi-square test is a widely used alternative for the same 2x2 contingency structure when expected counts are reasonably large, and rank-based enrichment methods can additionally incorporate the strength/order of TFBS signal rather than a binary presence call.
-
Multiple-testing correction was performed with the Benjamini-Hochberg FDR procedure applied to Z-score- and Fisher-score-derived p-values, and the authors note the underlying tests may not fully capture the non-random structure of genomic sequences.↳ Could also: A permutation-based (empirical) null distribution generated by repeatedly resampling or shuffling background/foreground sequence labels — A permutation-derived null can also account for genome-specific sequence composition and non-random structure directly from the data, which could complement the parametric FDR approach the authors describe as imperfect for this purpose.
-
Positional enrichment of TFBS relative to the ChIP-Seq peak maximum confidence position was assessed with a Kolmogorov-Smirnov test comparing target and background DistMCP distributions.↳ Could also: A permutation test on the DistMCP distributions, or a Wilcoxon rank-sum test on distances to the peak center — A permutation test can also be tailored to the exact experimental design without distributional assumptions, and a rank-sum test offers an alternative for comparing central tendency of the distance distributions when full distributional shape comparison is not the primary interest.
-
Results were reported as ranked Z-scores, Fisher scores, and KS scores without accompanying confidence intervals or standardized effect sizes.↳ Could also: Reporting a standardized effect size (e.g., fold-enrichment or odds ratio) with a bootstrap or asymptotic confidence interval alongside the ranking scores — Effect sizes with confidence intervals could also convey the magnitude and precision of enrichment for a given TFBS, complementing the rank-based presentation used in the paper.
-
Two independent scoring methods (Z-score and Fisher score) were calculated and presented separately for the same TFBS over-representation question.↳ Could also: A formal rank-aggregation or meta-analytic combination of the two scores into a single unified statistic — Combining the two complementary measures statistically could also provide a single consolidated ranking, which may be useful when the two scores disagree on relative ordering of candidate TFBS.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The self-contained SSA mode reproduces qualitatively well: on the paper's own GSE11431 Sox2 peak set, Pou5f1::Sox2 ranks #1 and the entire top-10 is Sox/Pou-family — the tool does what it claims. The quantitative validation claims fall short: the reported 82% Z-score/Fisher top-10 concordance came out 50% (5/10; 75% at N=20), and the reported >=0.97 same-size replicate Z-score correlation came out 0.9025 (Fisher 0.9156 passes its >=0.86 bar); background-size robustness clears >=0.90 for all ratios >=1x but not at 0.5x (0.8332/0.8890). These gaps sit mostly on our side / underspecification: a necessary JASPAR2026 substitution (2,633 matrices, 872–879 scored) for the ~2010 matrix set, plus self-constructed background sets the paper never describes — not evidence of authors' error, and nothing here is fabrication-suspect. Coverage is genuinely incomplete and honestly disclosed: the paper's strongest claim (c-Myc 100% concordance) was left ungraded because the rerun was still mid-scan at forced finalization, and the gene-based/aCSA mode depends on an unreachable, undeposited Wasserman-lab MySQL database — an external-infrastructure reproducibility defect on the authors' side that nevertheless is not a defect in the reported values.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.