Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

oPOSSUM-3: advanced analysis of regulatory motif over-representation across genes or ChIP-Seq datasets.

G3 (Bethesda) · 2012
L1 49/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
49/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 7% of all assessed papers rank 1081 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

FINAL verdict: reproduced the self-contained SSA (Sequence Set Analysis) mode of oPOSSUM-3 against the paper's own Sox2 and c-Myc GSE11431 mouse ES-cell ChIP-seq case studies, and graded all claims resolvable within this session's scope. Sox2 SSA correctly recovers the Sox2/Pou5f1::Sox2 motif as the top over-represented hit (within-tol). The paper's specific quantitative claims are a mixed bag under close scrutiny: Z-score/Fisher-score top-10 concordance for Sox2 falls well short of the paper's 82% (we get 50%, mismatch); background-size robustness (Spearman rho) mostly clears the paper's >=0.90 floor except at the undersized 0.5x ratio (partial); same-size replicate stability splits by score type, with Fisher passing (0.9156 vs >=0.86) but Z-score falling short (0.9025 vs >=0.97, partial). Gene-based analysis (a separate oPOSSUM-3 feature) is out of scope: it depends on an external Wasserman-lab MySQL database unreachable from this environment, with no self-contained data mirror available (error/documented external blocker). The c-Myc case, including the paper's strongest claimed result (100% Z/Fisher top-10 concordance), could NOT be graded: at finalization time, a direct sacct/squeue/log check confirmed «job» (the R-PATH-fixed rerun) was still mid-scan (~50 min elapsed of a run whose predecessor took 8h11m to reach the same point) -- not 'done for hours' as assumed when the finalize instruction was issued. Rather than keep perfecting or wait out the remaining hours, this result is finalized now as instructed: an honest partial reproduction with one claim explicitly left ungraded and disclosed, not fabricated.

💻 Code ↗ 🗄 Data: GSE11431

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper asks whether a new generation of motif over-representation software — incorporating multi-species phastCons conservation, an updated JASPAR profile set, TFBS profile clustering, and sequence-based (ChIP-Seq) input — can identify the transcription factors mediating regulation of co-expressed gene sets or co-recovered ChIP-Seq sequences.

Core claims
  • oPOSSUM-3 is a web-accessible system that identifies over-represented TFBS and TFBS families in DNA sequences of co-expressed genes or in sequences from high-throughput methods such as ChIP-Seq. resource
  • Validation with known sets of co-regulated genes and published ChIP-Seq data demonstrates the capacity of oPOSSUM-3 to identify mediating transcription factors for co-regulated genes or co-recovered sequences. finding
  • TFBS Cluster Analysis (TCA) and anchored Combination TFBS Cluster Analysis (aCTCA) address the challenge of homologous TFs with highly similar or identical binding specificity by reporting results as TFBS sequence patterns rather than individual profile names. method
  • Over-representation is measured by two complementary statistics — a Z-score based on the normal approximation to the binomial distribution (counting predicted binding site nucleotides) and a Fisher score from a one-tailed Fisher exact probability (counting sequences/genes with the site). method
  • For sequence-based analysis, a third measure, a Kolmogorov-Smirnoff centrality score, compares the empirical distributions of TFBS distances to the maximum confidence position (DistMCP) between target and background sets; functional TFBSs are expected to cluster around zero in target sequences and be randomly distributed in background sequences. method
  • SSA in version 3 replaces pairwise sequence alignment-based phylogenetic footprinting with phastCons scores; conserved regions are genomic regions of at least 20 nucleotides with average phastCons above threshold, and TFBS searches are restricted to motifs within or overlapping (by at least 1 bp) these regions. method
  • anchored Combination Site Analysis (aCSA) restricts over-representation assessment to sequence regions proximal to predicted TFBS of a user-specified anchor TF, counting pairings of the anchor TFBS with all secondary TFBSs within an inter-binding site distance; self-interactions can be observed when anchor TFBSs occur in clusters. method
  • Clustering TF binding profiles within annotated JASPAR structural families (rather than by structural class alone) is necessary because the extent of binding sequence similarity varies among classes — e.g. zinc fingers have low profile similarity among members. mechanism
Experimental setups
Assay System Perturbation Readout Platform
TFBS motif over-representation analysis, gene-based Single Site Analysis (SSA) with phastCons conservation filtering Human (hg19), mouse (mm9), fruit fly (dm3), nematode C. elegans (ce6), and yeast S. cerevisiae genome/gene annotation databases; reference sets of co-regulated genes none Z-scores (predicted binding site nucleotide counts) and Fisher scores (number of genes containing the TFBS) for each JASPAR profile, with FDR-adjusted p-values oPOSSUM-3 web system; Ensembl v64 (v54 for C. elegans) annotation and sequence; Wormbase WS200 operon annotation; UCSC phastCons46wayPlacental / phastCons30wayPlacental / phastCons15way / phastCons6way; JASPAR 2010 CORE and PBM collections plus custom PENDING collection; R statistics package
Sequence-based TFBS motif over-representation analysis (SSA) on ChIP-Seq-derived sequence sets Published large-scale ChIP-Seq sequence collections, including the Nfe2L2 ChIP-Seq dataset of Malhotra et al. (2010) none (user-supplied foreground TF-bound vs. background control sequences; no conservation filtering applied) Z-scores, Fisher scores, and Kolmogorov-Smirnoff centrality scores based on TFBS distance to the maximum confidence position (DistMCP) oPOSSUM-3 sequence-based system; JASPAR 2010 profiles; R statistics package
anchored Combination Site Analysis (aCSA) / anchored Combination TFBS Cluster Analysis (aCTCA) — over-representation of TFBS pairs Gene-based (annotated genomes: human, mouse, fruit fly, nematode, yeast) or user-supplied sequence sets user-specified anchoring TF and inter-binding site distance parameter Counts and over-representation scores for pairings of the anchor TFBS with secondary TFBSs (or cluster pairs) within the specified inter-binding site distance oPOSSUM-3; JASPAR 2010 profiles
TFBS profile pairwise similarity scoring and clustering (TFBS Cluster Analysis, TCA) JASPAR 2010 TFBS profile collections annotated for TF structural class and family clustering parameters: cluster score threshold T and radius margin R Pairwise similarity score table and resulting TFBS clusters; overlapping TFBS hits of the same cluster merged into a single cluster hit for over-representation counting MatrixAligner similarity scoring program (Sandelin et al. 2003)
Pre-computation of putative TFBS locations across species-specific promoter search regions Human and mouse (10,000 bp upstream and 10,000 bp downstream of Ensembl-annotated TSS); fruit fly (3000 bp each direction); nematode (1500 bp each direction); yeast (1000 bp upstream of TSS to the 3' end of each gene) none; exons excluded for invertebrates in gene-based pre-calculations Pre-computed TFBS hit locations and conserved regions used for gene-based analyses Ensembl v64 (C. elegans v54); UCSC Genome Browser phastCons score sets
Web service usage tracking oPOSSUM-2 public web server users none Number of unique monthly users, excluding automated internet search software
Key results
  • Assessment against reference sets of co-regulated genes and large-scale ChIP-Seq sequence collections exemplifies the utility of oPOSSUM-3 for identifying mediating TFs.
  • oPOSSUM-2 was found to perform well in an independent assessment of motif over-representation analysis tools (Meng et al. 2010).
  • In the Nfe2L2 ChIP-Seq dataset (Malhotra et al. 2010), target-set binding site distances to the maximum confidence position are clustered around zero relative to background, illustrating centrality enrichment.
  • Clustering of JASPAR profiles by family similarity condenses redundant, near-identical profiles so that independent enriched profiles can be identified rather than results being dominated by subsets of nearly identical profiles.
  • Replacement of pairwise-alignment phylogenetic footprinting with multi-species phastCons scores improves alignment quality and greatly increases the number of genes available for analysis.
  • The JASPAR 2010 update provides a significant increase in non-vertebrate profiles, permitting extension of regulatory analysis software to non-vertebrate species such as insects.
  • FDR-adjusted p-values were computed for Z-scores and Fisher scores in Tables 2–5 for completeness but do not affect the relative rankings.
Key statistics
  • count 340 unique users per month (on average, excluding automated internet search software) (oPOSSUM-2 web service usage)
  • count 250 profiles from the JASPAR 2010 CORE collection, 184 from the PBM collection, and 4 from the custom PENDING collection (Profiles analyzed in the TFBS clustering procedure)
  • other cluster score threshold T = 1.8; radius margin R = 0.1 (MatrixAligner-based TFBS profile clustering parameters, selected empirically from the distribution of pairwise similarity scores)
  • other phastCons thresholds of 0.4 and 0.6; conserved regions of at least 20 nucleotides; TFBS overlap of at least 1 bp (Definition of conserved regions for gene-based SSA)
  • other continuity correction of 0.5; Z = (x−µ−0.5)/σ, µ = BC, C = n/N, P = B/N, σ = sqrt(nP(1−P)) (Z-score calculation using normal approximation to the binomial distribution)
  • count 10,000 bp upstream and 10,000 bp downstream of TSS (human, mouse); 3000 bp each direction (fruit fly); 1500 bp each direction (nematode); 1000 bp upstream to 3' gene end (yeast) (Species-specific search regions for gene-based TFBS pre-computation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper describes a bioinformatics software system (oPOSSUM-3) that ranks transcription factor binding site (TFBS) motifs by comparing their occurrence in a foreground (target) gene or sequence set against a background (control) set. Over-representation is quantified with two independent scoring methods (a Z-score and a Fisher score) and, for sequence-based ChIP-Seq analysis, a third Kolmogorov-Smirnov (KS)-based centrality score. Benjamini-Hochberg FDR-adjusted p-values are reported alongside the raw scores in the validation tables, though the authors state these are provided for completeness and do not affect the relative rankings used to interpret results.

Replicationunclear Sample sizeSample sizes are described in terms of the foreground/background gene or sequence sets analyzed (e.g., number of genes, sequences, or nucleotides), rather than as a pre-specified statistical power calculation. Groupsforeground (target/co-regulated genes or TF-bound sequences) vs. background (control) gene or sequence sets Pairingna Randomization/blindingnot stated Dispersionnone Confidence intervalsno Multiplicity correctionBenjamini-Hochberg False Discovery Rate (FDR)
Statistical tests used
Test Applied to n Assumptions
Z-score based on normal approximation to the binomial distribution (with 0.5 continuity correction) TFBS over-representation in foreground vs. background nucleotide counts (SSA, aCSA, TCA, aCTCA; Tables 2-5) total nucleotides in foreground (n) and background (N) sequence sets, as defined in the Z-score formula stated
Fisher exact probability test (hypergeometric distribution) presence/absence of a TFBS across foreground vs. background sequences, same analyses as above (Tables 2-5) number of foreground and background sequences containing/lacking the TFBS stated
Kolmogorov-Smirnov (KS) test comparison of empirical distributions of TFBS distance-to-maximum-confidence-position (DistMCP) between target and background sequences, sequence-based/ChIP-Seq analysis (e.g., Nfe2L2 dataset, Figure S2) not explicitly stated as an n; based on the set of TFBS-to-MCP distances in target vs. background sequences not stated
Approaches that could also have been used
  • TFBS enrichment in nucleotide counts was assessed with a Z-score derived from a normal approximation to the binomial distribution with a continuity correction.
    Could also: An exact binomial test or exact Poisson test for count-based enrichment — An exact test avoids relying on the normal approximation and could also be used directly, particularly useful when foreground counts are small, where the approximation is least accurate.
  • Presence/absence enrichment of TFBS across sequences was assessed with a one-tailed Fisher exact test based on the hypergeometric distribution.
    Could also: A chi-square test of independence, or a hypergeometric-based gene-set enrichment approach such as GSEA-style rank statistics — A chi-square test is a widely used alternative for the same 2x2 contingency structure when expected counts are reasonably large, and rank-based enrichment methods can additionally incorporate the strength/order of TFBS signal rather than a binary presence call.
  • Multiple-testing correction was performed with the Benjamini-Hochberg FDR procedure applied to Z-score- and Fisher-score-derived p-values, and the authors note the underlying tests may not fully capture the non-random structure of genomic sequences.
    Could also: A permutation-based (empirical) null distribution generated by repeatedly resampling or shuffling background/foreground sequence labels — A permutation-derived null can also account for genome-specific sequence composition and non-random structure directly from the data, which could complement the parametric FDR approach the authors describe as imperfect for this purpose.
  • Positional enrichment of TFBS relative to the ChIP-Seq peak maximum confidence position was assessed with a Kolmogorov-Smirnov test comparing target and background DistMCP distributions.
    Could also: A permutation test on the DistMCP distributions, or a Wilcoxon rank-sum test on distances to the peak center — A permutation test can also be tailored to the exact experimental design without distributional assumptions, and a rank-sum test offers an alternative for comparing central tendency of the distance distributions when full distributional shape comparison is not the primary interest.
  • Results were reported as ranked Z-scores, Fisher scores, and KS scores without accompanying confidence intervals or standardized effect sizes.
    Could also: Reporting a standardized effect size (e.g., fold-enrichment or odds ratio) with a bootstrap or asymptotic confidence interval alongside the ranking scores — Effect sizes with confidence intervals could also convey the magnitude and precision of enrichment for a given TFBS, complementing the rank-based presentation used in the paper.
  • Two independent scoring methods (Z-score and Fisher score) were calculated and presented separately for the same TFBS over-representation question.
    Could also: A formal rank-aggregation or meta-analytic combination of the two scores into a single unified statistic — Combining the two complementary measures statistically could also provide a single consolidated ranking, which may be useful when the two scores disagree on relative ordering of candidate TFBS.
Software: R statistics package R Development Core Team 2008 (specific R version not stated); p.adjust used for FDR, pnorm used for the standard normal table · MatrixAligner (Sandelin et al. 2003)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

sox2_ssa_top_hit_recovery
Reported
SSA analysis of Sox2 ChIP-seq peaks (mouse ES cells) should recover the Sox2 motif (and its Pou5f1/Oct4 heterodimer partner) among the top over-represented TFBS profiles, validating the tool's core over-representation detection.
Reproduced
Top-10 by Fisher score in our bg_2x reproduction: Pou5f1::Sox2, Sox17, Sox7, POU2F1::SOX2, SOX4, SOX9, Sox6, Sox5, SOX13, Sox3 -- the Sox2/Pou5f1::Sox2 composite motif ranks #1, and all top-10 hits are Sox/Pou-family matrices.
within tolerance
sox2_zscore_fisher_top10_concordance
Reported
Paper: '82% agreement within the top 10 motifs' between Z-score and Fisher-score rankings for Sox2 (fraction of TF motifs common to both top-10 lists).
Reproduced
Our reproduction: 5/10 (50%) overlap between top-10 Z-score and top-10 Fisher-score TF lists (N=20: 15/20 = 75%). Top-10 Z-score: NR1H2::RXRA, E2F7, MTF1, ZNF528, CTCF, POU2F1::SOX2, Pou5f1::Sox2, Sox17, Sox7, SOX4.
did not match
cmyc_zscore_fisher_top10_concordance
Reported
Paper: '100% agreement within the top 10 motifs' between Z-score and Fisher-score rankings for c-Myc -- the paper's strongest concordance case among its three test datasets.
Reproduced
NOT COMPLETED at finalization time. «job» (cMyc SSA rerun with repro_r PATH fix for the earlier R-not-found crash) was confirmed, via direct sacct/squeue check and its own progress log, to still be in the target-sequence TFBS search phase (started 22:55:26 CEST, ~50 min elapsed) -- no results.txt exists yet. Its predecessor attempt («job») needed 8h11m to reach the equivalent point before crashing on the (since-fixed) missing-R bug, so a comparable full run was still hours away. This claim is left explicitly ungraded rather than fabricated; finalizing now was an explicit instruction to stop waiting rather than a claim that compute had actually finished.
m.public.grade.error
background_ratio_robustness_diffsize
Reported
Paper: Spearman correlation between backgrounds of different sizes (0.5x-4x of foreground) was >=0.90 for Sox2, for either score -- ranking is robust to background set size.
Reproduced
Z-score (vs 1x_A baseline): rho=0.8332 (0.5x), 0.9070 (2x), 0.9307 (3x), 0.9310 (4x). Fisher score: rho=0.8890 (0.5x), 0.9197 (2x), 0.9402 (3x), 0.9465 (4x). All ratios >=1x meet/exceed the >=0.90 threshold for both scores; the undersized 0.5x background falls short for both (0.8332, 0.8890 < 0.90).
partial
background_replicate_stability
Reported
Paper: between two same-sized replicate background sets, Sox2 correlation was >=0.97 (Z-score) and >=0.86 (Fisher score).
Reproduced
Our 1x_A vs 1x_B replicate comparison: Z-score rho=0.9025 (below paper's >=0.97), Fisher score rho=0.9156 (above paper's >=0.86, passes).
partial
gene_based_analysis_and_acsa
Reported
oPOSSUM-3 gene-based single-site analysis (SSA on gene lists) and amalgamated combination-site analysis (aCSA), requiring the external Wasserman-lab 'oPOSSUM gene' MySQL database (JASPAR-linked, TSS/conserved-region annotations).
Reproduced
Not computed. The external database server(s) this feature depends on are not reachable from this environment/session network scope, and no self-contained data mirror of this DB is available (not on GEO/Zenodo/code repo). Sequence-based SSA, which is fully self-contained, was used instead for all in-scope reproduction (Sox2, c-Myc case studies).
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 49/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

The self-contained SSA mode reproduces qualitatively well: on the paper's own GSE11431 Sox2 peak set, Pou5f1::Sox2 ranks #1 and the entire top-10 is Sox/Pou-family — the tool does what it claims. The quantitative validation claims fall short: the reported 82% Z-score/Fisher top-10 concordance came out 50% (5/10; 75% at N=20), and the reported >=0.97 same-size replicate Z-score correlation came out 0.9025 (Fisher 0.9156 passes its >=0.86 bar); background-size robustness clears >=0.90 for all ratios >=1x but not at 0.5x (0.8332/0.8890). These gaps sit mostly on our side / underspecification: a necessary JASPAR2026 substitution (2,633 matrices, 872–879 scored) for the ~2010 matrix set, plus self-constructed background sets the paper never describes — not evidence of authors' error, and nothing here is fabrication-suspect. Coverage is genuinely incomplete and honestly disclosed: the paper's strongest claim (c-Myc 100% concordance) was left ungraded because the rerun was still mid-scan at forced finalization, and the gene-based/aCSA mode depends on an unreachable, undeposited Wasserman-lab MySQL database — an external-infrastructure reproducibility defect on the authors' side that nevertheless is not a defect in the reported values.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.