Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Ensembl 2013.

Nucleic Acids Res · 2012
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced both in-scope pipeline claims of Flicek et al. 2013 (Ensembl 2013): VEP functional variant annotation (HGVS/regulatory/protein-domain/cross-reference annotation on 2045 ClinVar GRCh37 variants, 19065 consequence calls across expected categories) and BodyMap2 (E-MTAB-513) RNA-seq reprocessing into transcript/gene models (HISAT2+StringTie on ERR030890, 95.75% mapping rate, 58657 transcripts, 89597 gene-level quant entries). E-MTAB-513 dataset profiled: 48 ENA runs / 19 sources (16 tissues + 1 mixture in triplicate) / 3,736,859,003 total reported reads / ~209.5GB fastq, matching the paper's described BodyMap2 scope. Out-of-scope items (full internal Ensembl genebuild/eHive pipeline, Regulatory Build across 532 ENCODE datasets, DGVa structural variant merging, COSMIC somatic import) were not reproduced as they require proprietary internal pipelines, licensed/registration-gated data (COSMIC), or computational scale far beyond a single-room reproduction; third-party substitutes (HISAT2/StringTie for genebuild, VEP for functional annotation) were used where a faithful open substitute existed.

💻 Code ↗ 🗄 Data: E-MTAB-513

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Core claims
  • Ensembl (http://www.ensembl.org) provides genome information for sequenced chordate genomes, currently supporting 70 species with a focus on human, mouse, zebrafish and rat. resource
  • Ensembl provides evidence-based gene sets for all supported species, whole-genome multiple species alignments across vertebrates plus clade-specific alignments (eutherian mammals, primates, birds, fish), variation data for 17 species, and regulation annotations based on ENCODE and other data sets. resource
  • RNA-seq data are now routinely incorporated as supporting evidence in Ensembl gene annotation, and a new RNA-seq update pipeline allows existing gene sets to be updated by merging standard annotation models with RNA-seq-based models. method
  • The RNA-seq update pipeline improves gene sets by lengthening truncated genes, merging adjacent gene fragments and splitting artificially merged genes; it is particularly effective for species distantly related to well-annotated mammals and those with little species-specific sequence data. finding
  • Ensembl includes GRC 'fix' and 'novel' human assembly patches and uses its comparative genomics infrastructure (LASTZ self-alignment) to compare patches against the reference human genome, showing how patches alter annotation. method
  • The human gene set is updated each release by merging Ensembl automatic annotation with Havana manual annotation to produce the GENCODE gene set, including all current human CCDS models. method
  • Variant consequence annotation uses defined Sequence Ontology terms for all descriptions, a standard also adopted by the UCSC genome browser and ICGC to enable comparison of variation annotation. method
  • Combined Segway and ChromHMM segmentation (developed for ENCODE) classifies the human genome into functional segment types from 12 specific assays, giving a single-track summary of functional architecture. method
Experimental setups
Assay System Perturbation Readout Platform
Evidence-based automated gene annotation (genebuild) merged with manual annotation Human genome (GRCh37.p8); mouse, zebrafish and selected pig regions none Gene/transcript models (GENCODE gene set, CCDS models) Ensembl genebuild pipeline plus Havana manual annotation
RNA-seq (read alignment to genome, RNA-seq-based gene model building, intron-spanning read features) 13 species: zebrafish, chimpanzee, Nile tilapia, dog, Chinese softshell turtle, pig, ferret, platyfish, coelacanth, Tasmanian devil, orang-utan, opossum, platypus; plus human Illumina BodyMap 2.0 tissues none RNA-seq gene models, BAM files, intron features/supported splice sites, tissue-specific expression Illumina Human BodyMap 2.0 (ArrayExpress E-MTAB-513); Ensembl RNA-seq and RNA-seq update pipelines
ChIP-seq and DNase-seq 13 human and 5 mouse cell lines (segmentation in GM12878, K562, H1-hESC, HepG2, HeLa-S3, HUVEC) none Genomic locations of histone modifications and TF binding regions; regulatory features; genome segmentation states ENCODE data sets; Segway and ChromHMM segmentation; JASPAR binding matrices; raw reads in the European Nucleotide Archive
TF binding motif scanning within ChIP-seq binding regions Human genome regulatory build none Positions of high-probability TF-binding sites at 5% False Discovery Rate JASPAR database matrices
Variation data import, merging and QC (SNPs, in-dels, structural variants, genotypes) 17 species; human, rat, chimpanzee, orang-utan, zebrafish, pig, dog, macaque updated this year; mouse remapped to GRCm38 none rsIDs, locations, allele frequencies, genotypes, structural variants, somatic mutations, phenotype associations, clinical significance dbSNP, DGVa, 1000 Genomes Project, NHLBI Exome Sequencing Project, HGMD, COSMIC, OMIM, EGA, NHGRI GWAS catalog; eHive-based QC pipeline
Variant effect / consequence prediction on transcripts and regulatory features All supported species (human focus) none Sequence Ontology consequence terms; amino acid change impact predictions; overlap with regulatory features and TF binding motifs Variant Effect Predictor (VEP), SIFT, PolyPhen
Genotyping array probe mapping / variant chip annotation Human variants none Flagging of variants present on genotyping chips Affymetrix GeneChip 100K, GeneChip 500K, GenomeWideSNP_6.0; nine Illumina chips (CytoSNP12v1, Human660W-quad, Human1M-duoV3, CardioMetaboChip, HumanOmni1-Quad, HumanHap650, HumanHap550, HumanOmni2.5, Human610_Quad)
Comparative genomics: whole-genome multiple and pairwise alignments, self-alignments, gene tree (protein and ncRNA) inference, gene family expansion/contraction and split-gene analysis Vertebrates and clade-specific sets (eutherian mammals, primates, birds, fish); human self-alignment and assembly patch comparison; new species coelacanth and lamprey none Alignment blocks, orthology/paralogy gene trees, super-trees, gene family expansions/contractions, gene split annotations, ancestral alleles LASTZ; Ensembl Compara gene tree pipeline; CAFE; eHive workflow management system
Key results
  • Ensembl release 69 (October 2012) supports 70 species, 61 fully supported on the main site, with full gene annotations for 58 chordates (43 high-coverage, 15 low-coverage) plus imported annotation for 3 non-chordate model organisms. 70 species; 58 chordate gene sets
  • Five new species gained full support in the past year (Atlantic cod, coelacanth, ferret, Nile tilapia, Chinese softshell turtle) and six new species were added to the Ensembl Pre! site. 5 new fully supported; 6 new Pre! species (9 total on Pre!)
  • The RNA-seq update pipeline was used to improve the existing opossum, platypus and orang-utan gene sets for Ensembl release 69. 3 species gene sets updated
  • Regulation database contains 532 ChIP-seq and DNase-seq data sets covering 49 histone modification types and binding regions of 113 TFs, 40 of which have JASPAR binding matrices. 532 data sets; 49 modifications; 113 TFs; 40 with matrices
  • Regulatory Build coverage increased by 15% in the past year and now annotates 270 Mb of the human genome in 518 020 regulatory features across experiments in 13 cell lines. +15%; 270 Mb; 518 020 features
  • Human structural variation data are more comprehensive than all other species combined, with >6 million variants of which 5624 are somatic; structural variation data are available for human, mouse, horse, zebrafish, cow and macaque. >6 million variants; 5624 somatic
  • Human variation resources now include ~79 000 HGMD mutation locations, >135 000 COSMIC somatic mutation positions and phenotype data for >287 000 variants. ~79 000; >135 000; >287 000
  • Inclusion of the novel patch HSCHR9_1_CTG35 adds sequence missing from the original GRCh37 assembly, corrects an inversion and relocates the RNA gene RP11-548B3.3 from 5′ of APBA1 into its second intron, without altering downstream annotation.
Key statistics
  • count 70 species supported; 61 fully supported on main site (Ensembl release 69 species support)
  • count 58 chordates with full gene annotation (43 high-coverage, 15 low-coverage) plus 3 non-chordate model organisms (Gene annotation coverage)
  • count 13 species incorporate RNA-seq data (RNA-seq evidence in gene annotation)
  • count 532 ChIP-seq and DNase-seq data sets from 13 human and 5 mouse cell lines (Ensembl regulation database content)
  • count 49 histone modification types; 113 TFs; 40 TFs with JASPAR matrices (Regulatory data types represented)
  • count 518 020 regulatory features covering 270 Mb (Human Regulatory Build size)
  • other 15% increase in Regulatory Build coverage in the past year; motif sites called at 5% False Discovery Rate (Regulatory Build growth and motif calling threshold)
  • count >6 million human structural variants, 5624 somatic; ~79 000 HGMD locations; >135 000 COSMIC positions; >287 000 variants with phenotype data; variation for 17 species (Variation resource scale)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a database/resource-description paper (the annual Ensembl project update) rather than a hypothesis-driven experimental study; it describes genome annotation pipelines, data integration, and software infrastructure, and reports counts of species, genes, variants, and regulatory features without inferential statistical testing or group comparisons. The one quantitative threshold mentioned is a 5% False Discovery Rate used to call high-probability transcription-factor binding sites from JASPAR motif data within ChIP-seq-defined binding regions.

Replicationunclear Groupsna — no experimental groups are statistically compared; the text describes species, datasets, and annotation pipelines Pairingna Randomization/blindingna Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correction5% False Discovery Rate threshold
Approaches that could also have been used
  • High-probability transcription-factor binding sites within ChIP-seq regions were called using a 5% False Discovery Rate threshold on JASPAR motif matches.
    Could also: A different FDR cutoff (e.g., 1%) or a Bonferroni-type family-wise error correction could also be applied to the same motif-scanning procedure. — Different multiple-testing correction choices and stringency levels trade off sensitivity for specificity in genome-wide motif calling, and reporting results at more than one threshold is a common way to convey how call sets depend on this choice.
  • Gene family expansion and contraction across the phylogeny were assessed using the CAFE tool.
    Could also: Other likelihood-based birth-death models of gene family size evolution (implemented in tools such as BadiRate or CAFE's alternative model variants) could also be used. — Alternative implementations can differ in their assumptions about rate heterogeneity across lineages or gene families, which can be informative to compare when characterizing family-size evolution.
  • Large-scale pipeline outputs (e.g., numbers of structural variants, regulatory features, or gene models) are reported as point counts without accompanying uncertainty measures.
    Could also: Reporting an estimated false-discovery or error rate, or a confidence interval, for automatically generated call sets (e.g., structural variant calls, regulatory feature predictions) could also be included alongside the counts. — Quantifying uncertainty around large, pipeline-derived counts can help users calibrate how much confidence to place in specific automatically generated annotations.
  • The RNA-seq-based gene annotation procedure combines evidence from intron-spanning reads, cDNA/EST alignments, and protein-to-genome alignments through a described filtering and merging workflow.
    Could also: A formal probabilistic or Bayesian evidence-integration framework could also be used to combine these heterogeneous evidence types into gene models. — Explicit probabilistic models can yield calibrated per-model confidence scores, which some annotation pipelines use to help distinguish highly supported models from more tentative ones.
  • Quality control (QC) of variation data is described as leveraging the eHive workflow system without specifying particular QC statistics.
    Could also: Standard population-genetics QC metrics, such as Hardy-Weinberg equilibrium testing or genotype call-rate thresholds, could also be reported as part of the QC procedure. — These are widely used, standardized statistics for flagging potentially unreliable variant calls or genotypes in large variation databases, and reporting them can make QC criteria more transparent to users.
Software: CAFE (gene family expansion/contraction analysis) · Segway (genome segmentation) · ChromHMM (genome segmentation) · SIFT (amino acid change prediction) · PolyPhen (amino acid change prediction) · eHive (workflow/QC pipeline management)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

vep-functional-annotation
Reported
Ensembl 2013 describes VEP (Variant Effect Predictor) providing transcript, protein and regulatory consequence annotation for variants, including HGVS nomenclature, canonical transcript flags, protein domains, and cross-references to known variants (dbSNP/ClinVar/COSMIC/1000 Genomes).
Reproduced
Ran VEP v109.3 in --offline --cache mode (GRCh37 cache release 109) against 2045 ClinVar GRCh37 variants with --hgvs --regulatory --symbol --biotype --canonical --protein --domains --check_existing. Produced 19,065 consequence annotations across categories: missense_variant(2869), synonymous_variant(1088), regulatory_region_variant(1288), TF_binding_site_variant(295), TFBS_ablation(69), splice-related(566), frameshift_variant(98), plus intronic/UTR/non-coding categories, all with HGVSc/HGVSp, SYMBOL, BIOTYPE, CANONICAL, ENSP, DOMAINS, CLIN_SIG fields populated.
within tolerance
bodymap-rnaseq-reprocessing
Reported
Ensembl 2013 describes reprocessing of Illumina BodyMap2 (E-MTAB-513) RNA-seq data across 16 human tissues (plus a mixture) into aligned reads and transcript/gene models as part of the Ensembl annotation pipeline.
Reproduced
Downloaded ERR030890 (brain tissue, single-end, 64,313,204 reads) from ENA, built a GRCh37 (Ensembl release 69) HISAT2 index from the Ensembl FTP archive, aligned with HISAT2 (95.75% overall/primary alignment rate: 61,582,790/64,313,204 primary mapped, 56,881,675 unique + 4,701,115 multi-mapped), sorted/indexed with samtools, then assembled transcripts and quantified gene expression with StringTie guided by the Ensembl GRCh37.69 GTF: 58,657 transcripts assembled, 89,597 gene-level quantification entries produced.
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

Ensembl 2013 is a resource/database update paper, so both in-scope claims are capability statements with no published numbers. The reproduction ran VEP v109.3 offline on a self-chosen ClinVar GRCh37 set (2045 variants -> 19,065 consequence annotations with HGVSc/HGVSp, SYMBOL, BIOTYPE, CANONICAL, ENSP, DOMAINS, CLIN_SIG all populated) and HISAT2+StringTie on one BodyMap2 run (ERR030890, 95.75% alignment, 58,657 transcripts, 89,597 gene-level entries) — both confirm the described functionality, but neither can be put against a reported value, hence q2=red. The deviations that exist are all on our side: newer reference versions (release 109 cache vs the paper's release 69), a self-defined variant input, 1 of 48 runs processed, and HISAT2/StringTie standing in for Ensembl's unpublished eHive genebuild. No sign of an authors'-side problem; the untested parts (Regulatory Build, DGVa, COSMIC) are scope/access limits, not evidence against the paper — overall solid but demonstrative rather than a 1:1 numeric reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.