The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced both in-scope pipeline claims of Flicek et al. 2013 (Ensembl 2013): VEP functional variant annotation (HGVS/regulatory/protein-domain/cross-reference annotation on 2045 ClinVar GRCh37 variants, 19065 consequence calls across expected categories) and BodyMap2 (E-MTAB-513) RNA-seq reprocessing into transcript/gene models (HISAT2+StringTie on ERR030890, 95.75% mapping rate, 58657 transcripts, 89597 gene-level quant entries). E-MTAB-513 dataset profiled: 48 ENA runs / 19 sources (16 tissues + 1 mixture in triplicate) / 3,736,859,003 total reported reads / ~209.5GB fastq, matching the paper's described BodyMap2 scope. Out-of-scope items (full internal Ensembl genebuild/eHive pipeline, Regulatory Build across 532 ENCODE datasets, DGVa structural variant merging, COSMIC somatic import) were not reproduced as they require proprietary internal pipelines, licensed/registration-gated data (COSMIC), or computational scale far beyond a single-room reproduction; third-party substitutes (HISAT2/StringTie for genebuild, VEP for functional annotation) were used where a faithful open substitute existed.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opus- ★ Ensembl (http://www.ensembl.org) provides genome information for sequenced chordate genomes, currently supporting 70 species with a focus on human, mouse, zebrafish and rat. resource
- ★ Ensembl provides evidence-based gene sets for all supported species, whole-genome multiple species alignments across vertebrates plus clade-specific alignments (eutherian mammals, primates, birds, fish), variation data for 17 species, and regulation annotations based on ENCODE and other data sets. resource
- ★ RNA-seq data are now routinely incorporated as supporting evidence in Ensembl gene annotation, and a new RNA-seq update pipeline allows existing gene sets to be updated by merging standard annotation models with RNA-seq-based models. method
- ★ The RNA-seq update pipeline improves gene sets by lengthening truncated genes, merging adjacent gene fragments and splitting artificially merged genes; it is particularly effective for species distantly related to well-annotated mammals and those with little species-specific sequence data. finding
- ★ Ensembl includes GRC 'fix' and 'novel' human assembly patches and uses its comparative genomics infrastructure (LASTZ self-alignment) to compare patches against the reference human genome, showing how patches alter annotation. method
- The human gene set is updated each release by merging Ensembl automatic annotation with Havana manual annotation to produce the GENCODE gene set, including all current human CCDS models. method
- Variant consequence annotation uses defined Sequence Ontology terms for all descriptions, a standard also adopted by the UCSC genome browser and ICGC to enable comparison of variation annotation. method
- Combined Segway and ChromHMM segmentation (developed for ENCODE) classifies the human genome into functional segment types from 12 specific assays, giving a single-track summary of functional architecture. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Evidence-based automated gene annotation (genebuild) merged with manual annotation | Human genome (GRCh37.p8); mouse, zebrafish and selected pig regions | none | Gene/transcript models (GENCODE gene set, CCDS models) | Ensembl genebuild pipeline plus Havana manual annotation |
| RNA-seq (read alignment to genome, RNA-seq-based gene model building, intron-spanning read features) | 13 species: zebrafish, chimpanzee, Nile tilapia, dog, Chinese softshell turtle, pig, ferret, platyfish, coelacanth, Tasmanian devil, orang-utan, opossum, platypus; plus human Illumina BodyMap 2.0 tissues | none | RNA-seq gene models, BAM files, intron features/supported splice sites, tissue-specific expression | Illumina Human BodyMap 2.0 (ArrayExpress E-MTAB-513); Ensembl RNA-seq and RNA-seq update pipelines |
| ChIP-seq and DNase-seq | 13 human and 5 mouse cell lines (segmentation in GM12878, K562, H1-hESC, HepG2, HeLa-S3, HUVEC) | none | Genomic locations of histone modifications and TF binding regions; regulatory features; genome segmentation states | ENCODE data sets; Segway and ChromHMM segmentation; JASPAR binding matrices; raw reads in the European Nucleotide Archive |
| TF binding motif scanning within ChIP-seq binding regions | Human genome regulatory build | none | Positions of high-probability TF-binding sites at 5% False Discovery Rate | JASPAR database matrices |
| Variation data import, merging and QC (SNPs, in-dels, structural variants, genotypes) | 17 species; human, rat, chimpanzee, orang-utan, zebrafish, pig, dog, macaque updated this year; mouse remapped to GRCm38 | none | rsIDs, locations, allele frequencies, genotypes, structural variants, somatic mutations, phenotype associations, clinical significance | dbSNP, DGVa, 1000 Genomes Project, NHLBI Exome Sequencing Project, HGMD, COSMIC, OMIM, EGA, NHGRI GWAS catalog; eHive-based QC pipeline |
| Variant effect / consequence prediction on transcripts and regulatory features | All supported species (human focus) | none | Sequence Ontology consequence terms; amino acid change impact predictions; overlap with regulatory features and TF binding motifs | Variant Effect Predictor (VEP), SIFT, PolyPhen |
| Genotyping array probe mapping / variant chip annotation | Human variants | none | Flagging of variants present on genotyping chips | Affymetrix GeneChip 100K, GeneChip 500K, GenomeWideSNP_6.0; nine Illumina chips (CytoSNP12v1, Human660W-quad, Human1M-duoV3, CardioMetaboChip, HumanOmni1-Quad, HumanHap650, HumanHap550, HumanOmni2.5, Human610_Quad) |
| Comparative genomics: whole-genome multiple and pairwise alignments, self-alignments, gene tree (protein and ncRNA) inference, gene family expansion/contraction and split-gene analysis | Vertebrates and clade-specific sets (eutherian mammals, primates, birds, fish); human self-alignment and assembly patch comparison; new species coelacanth and lamprey | none | Alignment blocks, orthology/paralogy gene trees, super-trees, gene family expansions/contractions, gene split annotations, ancestral alleles | LASTZ; Ensembl Compara gene tree pipeline; CAFE; eHive workflow management system |
- ▲ Ensembl release 69 (October 2012) supports 70 species, 61 fully supported on the main site, with full gene annotations for 58 chordates (43 high-coverage, 15 low-coverage) plus imported annotation for 3 non-chordate model organisms. 70 species; 58 chordate gene sets
- ▲ Five new species gained full support in the past year (Atlantic cod, coelacanth, ferret, Nile tilapia, Chinese softshell turtle) and six new species were added to the Ensembl Pre! site. 5 new fully supported; 6 new Pre! species (9 total on Pre!)
- ▲ The RNA-seq update pipeline was used to improve the existing opossum, platypus and orang-utan gene sets for Ensembl release 69. 3 species gene sets updated
- ▲ Regulation database contains 532 ChIP-seq and DNase-seq data sets covering 49 histone modification types and binding regions of 113 TFs, 40 of which have JASPAR binding matrices. 532 data sets; 49 modifications; 113 TFs; 40 with matrices
- ▲ Regulatory Build coverage increased by 15% in the past year and now annotates 270 Mb of the human genome in 518 020 regulatory features across experiments in 13 cell lines. +15%; 270 Mb; 518 020 features
- ▲ Human structural variation data are more comprehensive than all other species combined, with >6 million variants of which 5624 are somatic; structural variation data are available for human, mouse, horse, zebrafish, cow and macaque. >6 million variants; 5624 somatic
- ▲ Human variation resources now include ~79 000 HGMD mutation locations, >135 000 COSMIC somatic mutation positions and phenotype data for >287 000 variants. ~79 000; >135 000; >287 000
- – Inclusion of the novel patch HSCHR9_1_CTG35 adds sequence missing from the original GRCh37 assembly, corrects an inversion and relocates the RNA gene RP11-548B3.3 from 5′ of APBA1 into its second intron, without altering downstream annotation.
- count 70 species supported; 61 fully supported on main site (Ensembl release 69 species support)
- count 58 chordates with full gene annotation (43 high-coverage, 15 low-coverage) plus 3 non-chordate model organisms (Gene annotation coverage)
- count 13 species incorporate RNA-seq data (RNA-seq evidence in gene annotation)
- count 532 ChIP-seq and DNase-seq data sets from 13 human and 5 mouse cell lines (Ensembl regulation database content)
- count 49 histone modification types; 113 TFs; 40 TFs with JASPAR matrices (Regulatory data types represented)
- count 518 020 regulatory features covering 270 Mb (Human Regulatory Build size)
- other 15% increase in Regulatory Build coverage in the past year; motif sites called at 5% False Discovery Rate (Regulatory Build growth and motif calling threshold)
- count >6 million human structural variants, 5624 somatic; ~79 000 HGMD locations; >135 000 COSMIC positions; >287 000 variants with phenotype data; variation for 17 species (Variation resource scale)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a database/resource-description paper (the annual Ensembl project update) rather than a hypothesis-driven experimental study; it describes genome annotation pipelines, data integration, and software infrastructure, and reports counts of species, genes, variants, and regulatory features without inferential statistical testing or group comparisons. The one quantitative threshold mentioned is a 5% False Discovery Rate used to call high-probability transcription-factor binding sites from JASPAR motif data within ChIP-seq-defined binding regions.
-
High-probability transcription-factor binding sites within ChIP-seq regions were called using a 5% False Discovery Rate threshold on JASPAR motif matches.↳ Could also: A different FDR cutoff (e.g., 1%) or a Bonferroni-type family-wise error correction could also be applied to the same motif-scanning procedure. — Different multiple-testing correction choices and stringency levels trade off sensitivity for specificity in genome-wide motif calling, and reporting results at more than one threshold is a common way to convey how call sets depend on this choice.
-
Gene family expansion and contraction across the phylogeny were assessed using the CAFE tool.↳ Could also: Other likelihood-based birth-death models of gene family size evolution (implemented in tools such as BadiRate or CAFE's alternative model variants) could also be used. — Alternative implementations can differ in their assumptions about rate heterogeneity across lineages or gene families, which can be informative to compare when characterizing family-size evolution.
-
Large-scale pipeline outputs (e.g., numbers of structural variants, regulatory features, or gene models) are reported as point counts without accompanying uncertainty measures.↳ Could also: Reporting an estimated false-discovery or error rate, or a confidence interval, for automatically generated call sets (e.g., structural variant calls, regulatory feature predictions) could also be included alongside the counts. — Quantifying uncertainty around large, pipeline-derived counts can help users calibrate how much confidence to place in specific automatically generated annotations.
-
The RNA-seq-based gene annotation procedure combines evidence from intron-spanning reads, cDNA/EST alignments, and protein-to-genome alignments through a described filtering and merging workflow.↳ Could also: A formal probabilistic or Bayesian evidence-integration framework could also be used to combine these heterogeneous evidence types into gene models. — Explicit probabilistic models can yield calibrated per-model confidence scores, which some annotation pipelines use to help distinguish highly supported models from more tentative ones.
-
Quality control (QC) of variation data is described as leveraging the eHive workflow system without specifying particular QC statistics.↳ Could also: Standard population-genetics QC metrics, such as Hardy-Weinberg equilibrium testing or genotype call-rate thresholds, could also be reported as part of the QC procedure. — These are widely used, standardized statistics for flagging potentially unreliable variant calls or genotypes in large variation databases, and reporting them can make QC criteria more transparent to users.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Ensembl 2013 is a resource/database update paper, so both in-scope claims are capability statements with no published numbers. The reproduction ran VEP v109.3 offline on a self-chosen ClinVar GRCh37 set (2045 variants -> 19,065 consequence annotations with HGVSc/HGVSp, SYMBOL, BIOTYPE, CANONICAL, ENSP, DOMAINS, CLIN_SIG all populated) and HISAT2+StringTie on one BodyMap2 run (ERR030890, 95.75% alignment, 58,657 transcripts, 89,597 gene-level entries) — both confirm the described functionality, but neither can be put against a reported value, hence q2=red. The deviations that exist are all on our side: newer reference versions (release 109 cache vs the paper's release 69), a self-defined variant input, 1 of 48 runs processed, and HISAT2/StringTie standing in for Ensembl's unpublished eHive genebuild. No sign of an authors'-side problem; the untested parts (Regulatory Build, DGVa, COSMIC) are scope/access limits, not evidence against the paper — overall solid but demonstrative rather than a 1:1 numeric reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.