Improving the annotation of the cattle genome by annotating transcription start sites in a diverse set of tissues and populations using Cap Analysis Gene Expres
The main results reproduced: recomputed values matched the published ones within tolerance.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Salavati et al 2023, G3 - cattle CAGE TSS atlas via nf-cage (tagDust2->bowtie2) + CAGEfightR. DESCRIBED WELL ENOUGH and REPRODUCES (within tolerance). Data accession corrected: PRJEB43235 (109 Bos taurus CAGE runs, confirmed via ENA) - the brief's PRJEB34864 is the SHEEP project. (1) Figshare deposit 21769649 downloaded+md5-verified+parsed on «our HPC»: coexpression links (15,600), superenhancer stretches (3,379), longest stretch (54,732bp/18 enh) reproduce EXACTLY; TSS atlas 52,701 vs 51,295, enhancers 2,355 vs 2,328, genes 15,448 vs 15,364, transcripts 27,911 vs 27,588, novel/annotated split, full txType breakdown, per-tissue means and mean Kendall (0.35) all within ~5%. Deposit (v3) runs ~2-3% LARGER than manuscript text = minor final-filter drift, opposite direction to fabrication. (2) nf-cage pipeline RERUN on «our HPC»: conda env on «infra» + tagDust2 2.33 compiled from source + bowtie2 index of ARS-UCD1.2, then bowtie2 --very-sensitive on 8-run/6-animal subset -> mean mapping rate 93.23% (range 91.9-94.3%) vs reported 94% (within-tol). NOT ATTEMPTED: C4/C5 pre-filter CAGEfightR intermediates (not deposited); C21 cross-species (out of scope). 17/20 in-scope claims reproduced (3 exact, 14 within-tol). No fabrication indicators.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-19 ⛓ a6ee38a4a1c6
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCAGE sequencing across a diverse set of tissues and populations can precisely map transcription start sites (TSS) and their associated short-range enhancers in the cattle genome, revealing tissue-specific, population-specific, and cattle-specific regulatory features that improve annotation of the ARS-UCD1.2 reference assembly.
- ★ CAGE sequencing of 24 tissues from 3 cattle populations (dairy, beef-dairy cross, Kinsella composite) defines TSS and coexpressed short-range enhancers in the ARS-UCD1.2 reference genome finding
- ★ 51,295 TSS and 2,328 TSS-Enhancer regions were identified as shared across the 3 cattle populations finding
- ★ Cross-species comparison of CAGE data from 7 other species (including sheep) revealed a set of TSS and TSS-Enhancers specific to cattle finding
- ★ Superenhancer stretches identified in the CAGE data set overlap with previously reported CNV regions (CNV6, CNV28, CNV33) associated with milk production traits finding
- ★ Population-specific TSS and TSS-Enhancer signatures (e.g. a Holstein-specific signature) were established by comparing HOL, Charolais x Holstein, and Kinsella composite populations finding
- TSS and TSS-Enhancer clusters were predicted using the uni- and bidirectional clustering algorithms of the CAGEfightR v1.16.0 package within a NextFlow-based nf-cage analysis pipeline method
- The CAGE data set and annotation tracks are provided as a resource to be combined with other transcriptomic data for the same tissues to build a high-resolution transcript diversity map for the BovReg project resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| CAGE sequencing | 24 tissues from 6 cattle (3 populations: Belgian Holstein dairy, German Charolais x Holstein F2 beef-dairy cross, Canadian Kinsella composite beef) | none | TSS and TSS-Enhancer cluster locations and normalized expression (CTPM) | Illumina NextSeq 550 (50-nt single end) |
| RNA extraction and quality control | tissue samples from all 3 cattle populations | none | RNA quantity (Nanodrop) and RNA integrity number (RIN) | Agilent Bioanalyzer system |
| Coexpression linkage analysis (Kendall correlation) | predicted TSS and TSS-Enhancer regions genome-wide | none | correlation coefficient and genomic gap between linked TSS-enhancer pairs, classified as cis/trans/novel | CAGEfightR v1.16.0 findLinks function |
| Superenhancer stretch identification (hierarchical clustering / window scan) | TSS-Enhancer regions genome-wide | none | clusters of >=3 enhancers within a 10-kb window | CAGEfightR v1.16.0 findStretches |
| CNV overlap analysis | cattle genome (UMD3.1 lifted to ARS-UCD1.2) | none | overlap of superenhancer stretches with milk-trait-associated CNV regions (CNV6, CNV28, CNV33) | IGVtools |
| Tissue-specificity index analysis | CTPM expression matrix across all tissue types | none | tissue specificity index (TSI, 0-1) per TSS, visualized as heatmap | tspex v0.6.1; pheatmap v1.0.12 |
| Population-specific TSS/TSS-Enhancer clustering | tissue samples grouped by population (HOL, Charolais x Holstein F2, Kinsella composite) | none | population-restricted TSS/TSS-Enhancer sets (100% within-population support) | CAGEfightR v1.16.0 |
| Cross-species comparative CAGE analysis | cattle plus 7 other species including sheep | none | species-specific vs shared TSS and TSS-Enhancers | — |
- – 51,295 TSS identified across the CAGE data set shared across 3 cattle populations 51,295 TSS
- – 2,328 TSS-Enhancer regions identified as shared across the 3 cattle populations 2,328 TSS-Enhancer regions
- – A subset of TSS and TSS-Enhancers were found to be cattle-specific relative to 7 other species
- – Superenhancer stretches overlapped with the milk-trait-associated CNV6, CNV28, and CNV33 regions
- – A Holstein-specific TSS/TSS-Enhancer signature was established relative to Charolais x Holstein and Kinsella composite signatures
- – 204 base-pair-resolution bigWig files (positive/negative strand) generated for 102 samples were used for downstream TSS/enhancer analysis n=204 files, 102 samples
- – Suitable RNA samples were obtained for 43 dairy, 33 beef-dairy cross, and 33 composite beef samples for CAGE library preparation 43/33/33 samples
- – 7 CAGE libraries were excluded from downstream analysis due to poor clustering, low RIN/mapping rate, or castration confound 7 libraries excluded
- count 51,295 TSS (TSS identified and shared across 3 cattle populations)
- count 2,328 TSS-Enhancer regions (TSS-Enhancer regions shared across 3 populations)
- pvalue P < 0.05 (significance threshold for Kendall correlation test of TSS-Enhancer coexpression)
- other FDR < 0.01 (Benjamini-Hochberg adjusted significance threshold for coexpression links)
- count n = 204 bigWig files for 102 samples (base-pair resolution mapped output files used for CAGEfightR analysis)
- count minimum 10 reads per CTSS; 2/3 sample support (66/102 tissues) (filtration criteria for defining TSS and TSS-Enhancer regions)
- count 43 (dairy), 33 (beef-dairy cross), 33 (composite beef KC) (number of RNA samples per population passing QC for CAGE library preparation)
- other 400-1,000 bp (distance window used to define bidirectional TSS-Enhancer clusters on opposing strands)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper describes a descriptive genomic annotation study (CAGE-Seq) rather than a classic hypothesis-testing experiment: transcription start sites (TSS) and TSS-enhancer pairs were identified computationally using the CAGEfightR pipeline, with count-based filtration thresholds (minimum reads per cluster, minimum sample support) used to define reproducible features across 24 tissues and 3 cattle populations (n=2 individuals, 1 male/1 female, per population). The main inferential statistical procedure was a Kendall correlation test used to identify significant coexpression links between TSS and putative enhancers, with results filtered at P<0.05 and then corrected for multiple testing using Benjamini-Hochberg FDR (<0.01). Tissue specificity was summarized using a continuous tissue-specificity index (TSI) rather than a formal statistical test, and superenhancer regions were identified via window-based clustering followed by the same Kendall correlation approach.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Kendall correlation test | Testing coexpression links between predicted TSS and TSS-Enhancer regions (findLinks, CAGEfightR) | 102 samples (204 bigWig files, positive/negative strand) | not stated |
| Kendall correlation test | Coexpression analysis of superenhancer stretches (expression matrix, CTPM values) | — | not stated |
-
Coexpression between TSS and enhancer regions was tested using Kendall's rank correlation.↳ Could also: Spearman's rank correlation or Pearson's correlation could also be used — Spearman is a widely used nonparametric alternative that is often more familiar to readers and computationally lighter for large datasets, while Pearson's correlation would capture linear-magnitude relationships directly if the expression data met linearity/normality assumptions; each offers a slightly different lens on the same coexpression question.
-
Multiple testing across correlation tests was controlled using Benjamini-Hochberg FDR adjustment.↳ Could also: A Bonferroni or Holm correction could also be applied — These methods control the family-wise error rate more conservatively than FDR, which can be preferred when minimizing any false positives is prioritized over maximizing statistical power, though this comes at some cost to sensitivity for detecting true links.
-
TSS/Enhancer sets and population signatures were defined using presence/support thresholds (e.g., minimum reads, 2/3 sample support, 100% support within a population) rather than a formal statistical test of differential presence.↳ Could also: A count-based differential expression/detection framework (e.g., DESeq2 or edgeR modeling tissue or population as a factor) could also be used — Such frameworks would allow formal significance testing and effect-size/confidence-interval estimation for tissue- or population-specific differences, complementing the descriptive presence-based filtering approach used here.
-
Tissue specificity was summarized with a continuous tissue-specificity index (TSI) per TSS.↳ Could also: A formal statistical test of tissue enrichment (e.g., analysis of variance across tissues, or a generalized linear model) could also be used alongside the index — This would provide a p-value or interval-based measure of confidence for tissue specificity claims in addition to the descriptive TSI score.
-
Biological replication in this design consists of 2 individuals (1 male, 1 female) per population.↳ Could also: Where feasible, additional biological replicates per sex/population combination could also be incorporated — Additional replicates would support variance estimation and interval-based inference (e.g., via mixed-effects models), complementing the current single-pair-per-population design, which is a common and practical constraint in large-animal FAANG-style studies.
-
Correlation coefficients are reported as point estimates without confidence intervals.↳ Could also: Bootstrapped or analytic confidence intervals around the Kendall correlation coefficients could also be reported — Intervals would convey the precision of the coexpression correlation estimates in addition to the point value and significance threshold already provided.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37216666 (Salavati et al. 2023, G3; cattle CAGE TSS annotation)
Paper in one line
CAGE-seq of 102 cattle libraries (24 tissues, 6 animals, 3 populations) mapped to ARS-UCD1.2_Btau5.0.1Y; TSS + TSS-Enhancer clusters called with CAGEfightR; new TSS atlas for the cattle genome.
Code & data provenance
- Code: https://github.com/mazdax/nf-cage (DSL2 Nextflow; demux FASTX → tagDust2 trim →
bowtie2 map → bedGraph/bigWig). Container
mazdax/nf-cage:latest. Release v1.0.0. Downstream CAGEfightR analysis is described in Methods (R scripts not in nf-cage repo; also referenced bitbucket cagewrap_public). - Data (CORRECTED): PRJEB43235 (ERP127181) "CAGE sequencing of BOVREG cattle tissues", 109 runs on ENA. The brief listed PRJEB34864 — WRONG: that is the sheep CAGE project (Ovis aries, Rambouillet "Benz 2616", Salavati 2020). Registry harvesting error; documented.
- Processed outputs: figshare 10.6084/m9.figshare.21769649 (tissue-level GFF3, coexpression links/stretches/superenhancers, population-level GFF3/BED12) — the authors' final results.
In scope (pipeline-derived)
- nf-cage upstream: reads/sample (C1), bowtie2 mapping rate (C2), library count (C3).
- CAGEfightR downstream: TSS/enhancer counts and annotation (C4–C20).
Out of scope
- C21 cross-species comparison (external multi-species genome/annotation join) — not a single-pipeline output; not attempted.
- Wet-lab: tissue collection, RNA/CAGE library prep, sequencing.
- Biological interpretation (TSI tissue-specificity narrative, GO, figures of expression).
Reproduction tiers (80% floor = T1+T2, then keep going T3+T4)
- T1 reads/sample from ENA metadata. [done — partial]
- T2 verify figshare deposited outputs vs reported counts (delivers-promised audit). [«our HPC»/«infra»]
- T3 run nf-cage bowtie2 on a FASTQ subset → mapping rate ~94%. [«our HPC» SLURM]
- T4 (hard) CAGEfightR recompute of 51,295 TSS / 2,328 enhancers from bigWigs + Ensembl v106.
Blocker log
- «our HPC» ssh tunnel down at start (Connection closed); per brief, waiting for central VPN fix.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.