Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Manual curation for improved genome annotation of the functionally extinct northern white rhinoceros (Ceratotherium simum cottoni).

PLoS One · 2026
L1 69/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
69/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 35% of all assessed papers rank 745 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to partially reproduce via the paper's OWN shipped annotation (P16). The repo (eruggeri/Northern-White-Rhinoceros-Annotation @2807cc3) ships only the final annotation as Git-LFS GTF/GFF3, no scripts (Galaxy workflow is a PDF). 1:1 result: 3 of 4 numeric Table-2 claims reproduce within a few percent when the curated subset is isolated from the file - transcripts 32975 vs 34385 (-4.1%), avg transcript length 37615 vs 36981 bp (+1.7%, and this confirms 'length' means genomic span incl. introns), genome coverage 49.7% vs 48% (genome 2,495,808,350 bp). Gene count partial: the file's ids mix SwissProt entry-names and HGNC symbols (23793 raw tokens) so 15738 is not cleanly recoverable without an external mapping DB. KEY finding for auditors: the shipped GTF/GFF3 is the FULL un-curated StringTie merged assembly (475,980 transcripts / 430k MSTRG genes) - a SUPERSET; Table 2 describes the curated subset embedded within it, so a naive feature count of the repo file gives ~14x the reported transcripts. No fabrication evidence: reported numbers are consistent with the curated subset of the shipped data. BUSCO (C5) NOT reproduced - BUSCO 5.8.2 transcriptome mode crashed with a runtime BatchFatalError (no aligner initialised); env built and 32267 transcripts extracted fine, so it is a tool bug not a data issue, and the paper never stated BUSCO version/lineage/mode anyway. NOT attempted (out of scope/hard-20%): the manual curation itself (subjective BLAST-based; the headline +81% genes depends on it) and re-running HISAT2+StringTie2 from raw GEO FASTQ (heavy, and would not reproduce post-curation Table-2 numbers).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 69
    assessed: 2026-06-14 ⛓ 44e0e61fbaba
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can RNA sequencing of granulosa cells combined with de novo transcript assembly and extensive manual curation substantially improve the limited and error-prone genome annotation of the functionally extinct northern white rhinoceros (Ceratotherium simum cottoni)?

Core claims
  • Manual curation of RNA-seq-derived de novo transcripts increased the number of functional genes in the NWR annotation by 81% (from 8,701 to 15,738). finding
  • The number of annotated functional transcripts increased by 141% (from 14,274 to 34,385) in the new annotation. finding
  • Using in vivo collected granulosa cells across developmental stages for RNA-seq and StringTie de novo assembly yields a more diverse, complete annotation. method
  • The improved annotation corrects gene nomenclature to HGNC/VGNC conventions and fixes erroneous, protein-named, and bacterial gene assignments, making it operational for transcriptional studies. resource
  • BUSCO analysis showed 93.4% completeness, supporting the quality and completeness of the transcriptome/annotation. finding
  • The original BRAKER3-based annotation was flawed: only 51% of transcripts were called to the correct genetic sequence, with many protein-named, misassigned, or bacterial genes. finding
  • Total transcript sequence length covers 48% of the genome assembly, consistent with total-RNA sequencing capturing unspliced transcripts, repeats, and possible pseudogenes. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (total RNA, stranded, Ribo-Zero) southern white rhinoceros (Ceratotherium simum simum) in vivo collected mural granulosa cells; reads mapped to NWR genome CerSimCot1.0 none transcript reads for de novo transcript assembly and annotation Illumina TruSeq Stranded Total RNA with Ribo-Zero Plus; 100 bp paired-end on Illumina NovaSeq 6000 (GSE261038, 6 samples)
bulk RNA-seq (total RNA, stranded, Ribo-Zero) southern white rhinoceros granulosa cells; reads mapped to NWR genome CerSimCot1.0 none transcript reads for de novo transcript assembly and annotation Illumina Stranded Total RNA with Ribo-Zero Plus (NovaSeq X plus prep); 150 bp paired-end on Illumina NovaSeq 6000 (GSE300824, 8 libraries)
RNA isolation and quality control granulosa cells in RNAlater none RNA quantity and RNA integrity number (RIN > 6.0 threshold) Arcturus PicoPure RNA Isolation Kit; Qubit 4 Fluorometer; Agilent 4150 TapeStation
read alignment and de novo transcript assembly NWR genome CerSimCot1.0 (GCA_021442165.1) none aligned reads and merged transcript annotation file Galaxy (usegalaxy.org); HISAT2; StringTie2
manual curation via sequence homology search de novo transcript sequences vs Ceratotherium simum simum, Diceros bicornis, Equus caballus none gene assignment at ≥80% percent identity; corrected nomenclature NCBI BLAST+
transcriptome completeness assessment assembled NWR transcriptome none BUSCO completeness (single-copy, duplicated) BUSCO / OrthoDB
transrectal ovum pickup (OPU) sample collection four southern white rhinoceros females; ten follicles (2 growing, 6 dominant, 2 pre-ovulatory) none mural granulosa cells collected ultrasound-guided probe with double-lumen needles
Key results
  • Functional genes increased from 8,701 to 15,738 in the new annotation 81% increase
  • Functional transcripts increased from 14,274 to 34,385 141% increase
  • BUSCO completeness of the transcriptome 93.4% (single copy 89.4%, duplicated 3.9%)
  • Percent of genome covered by exons rose from 1% to 48% 1% to 48%
  • Average transcript length remained essentially unchanged between annotations 36,981 bp vs 39,878 bp
  • Only 51% of original annotated transcripts were correctly called to genetic sequence 7,299/14,274 (51%)
  • Average sequencing depth per sample 37,364,052 reads
  • New gene count approaches other mammalian genomes (cow Ensembl release 113: 20,848 gene models) 15,738 vs 20,848
Key statistics
  • count 15,738 functional genes (new) vs 8,701 (original) (annotated genes, 81% increase)
  • count 34,385 transcripts (new) vs 14,274 (original) (annotated transcripts, 141% increase)
  • other 93.4% (single copy 89.4%, duplicated 3.9%) (BUSCO completeness of transcriptome)
  • mean 37,364,052 (average sequencing depth per sample across 14 samples)
  • count 7,299/14,274 (51%) (original transcripts called to correct genetic sequence; 6,763 to protein names, 212 misassigned, 455 to bacterial genes)
  • other 48% (percent of genome assembly covered by total transcript sequence/exons (vs 1% original))
  • other ≥80% percent identity (BLAST threshold for assigning a sequence to a known gene)
  • count 20,848 gene models (cow gene models, Ensembl release 113, for comparison)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a bioinformatics methods paper describing the manual curation and improvement of the northern white rhinoceros genome annotation using RNA-seq from 14 granulosa cell samples collected from four southern white rhinoceros females. Reads were aligned with HISAT2 and assembled into transcripts with StringTie2 on the Galaxy platform, then manually curated via NCBI BLAST homology searches (≥80% nucleotide identity) against related species. All results are reported descriptively as raw counts and percentages comparing the original and new annotations; transcriptome completeness was assessed with BUSCO (93.4%). No inferential statistical tests were performed.

Replicationmixed Sample size4 southern white rhinoceros females; granulosa cells from 10 follicles yielding 14 RNA samples total; 4 animals run in technical replicates producing 8 samples (GSE261038); 6 additional samples without technical replicates (GSE300824) GroupsNew curated annotation vs. original BRAKER3 annotation (descriptive count comparison, no inferential group contrast) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
BUSCO completeness assessment Quality evaluation of the final curated transcriptome annotation na
Approaches that could also have been used
  • The average sequencing depth across 14 samples (37,364,052 reads) is reported without a measure of spread; individual values in Table 1 range from ~23 M to ~64 M reads
    Could also: Report standard deviation or range alongside the mean sequencing depth — A dispersion measure would convey whether coverage was consistent across samples, which is relevant context for assessing whether lower-depth samples could have limited transcript discovery for the merged annotation
  • Transcript assembly was performed with StringTie2 on each sample individually and then merged into a union annotation
    Could also: Use a dedicated genome annotation pipeline such as MAKER2, Liftoff, or EVidenceModeler (EVM) to integrate RNA-seq evidence with protein homology evidence in a more automated framework — Structured annotation pipelines weight multiple evidence types systematically and produce standardized quality scores (e.g., annotation edit distance), which could reduce the manual curation burden and make the annotation process more reproducible
  • Homology-based manual curation applied a single fixed 80% nucleotide identity threshold for gene assignment via BLASTn
    Could also: Complement nucleotide BLAST with protein-level searches (BLASTx or BLASTp against UniProt/OrthoFinder) — Protein-level homology is more sensitive at typical evolutionary distances between rhinoceros and horse, and can identify conserved genes where nucleotide identity falls below the 80% threshold while the encoded protein remains recognisable
  • Read alignment was performed with HISAT2
    Could also: Use STAR (Spliced Transcripts Alignment to a Reference) as an alternative splice-aware aligner — STAR is widely benchmarked alongside HISAT2 and can differ in sensitivity for novel splice junctions, which matters when annotating a genome with limited prior transcript evidence; comparing aligners is a common quality-control step in annotation projects
  • Annotation completeness was evaluated solely with BUSCO
    Could also: Complement BUSCO with protein-level completeness tools such as OMArk, or with annotation consistency metrics such as annotation edit distance (AED) scores — BUSCO captures single-copy orthologue representation but does not assess isoform accuracy or gene-model structural quality; additional metrics provide complementary views of annotation completeness and correctness
  • Granulosa cells across follicular developmental stages were the sole tissue source for transcript discovery
    Could also: Incorporate RNA-seq from additional tissue types (e.g., blood, fibroblasts, liver) to broaden transcriptome coverage — Many genes are expressed in a tissue-restricted manner; multi-tissue sampling is a standard strategy in genome annotation projects to approach the full complement of expressed genes, which would likely increase the BUSCO completeness score beyond the 93.4% achieved here
Software: Galaxy/usegalaxy.org · HISAT2 · StringTie2 · NCBI BLAST+ · BUSCO

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GCA_021442165.1 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE261038 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE300824 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41490125

Paper: Ruggeri et al. (PLoS One, 2026). Manual curation for improved genome annotation of the functionally extinct northern white rhinoceros. DOI 10.1371/journal.pone.0340594 · PMCID PMC12768360

Shipped artifacts: GitHub eruggeri/Northern-White-Rhinoceros-Annotation @2807cc3 ships ONLY the final annotation as Git-LFS files (GCA_021442165_annotation.gtf 225 MB, .gff3 154 MB) + README. No scripts. The Galaxy workflow is a PDF supplement (S1 File). Raw RNA-seq: GEO GSE261038 (6 samples) + GSE300824 (8 samples).

In scope (pipeline-derived, attempted)

Result Pipeline Reproducible?
Table 2: # transcripts (34,385) StringTie2 merged assembly → count from shipped GTF YES — count features in shipped GTF
Table 2: avg transcript length (36,981 bp) derived from GTF coords YES — compute from shipped GTF
Table 2: % genome covered (48%) transcript span / genome size YES — GTF + genome FASTA
Table 2: # genes (15,738) StringTie2 + curation gene grouping PARTIAL — file naming mixes SwissProt names + symbols
BUSCO completeness (93.4%) BUSCO on transcriptome YES (soft) — re-run BUSCO; lineage/version/mode unspecified in paper
Shipped-file integrity git-lfs SHA256 YES

Out of scope (wet-lab / manual / unspecified — NOT attempted)

  • Manual curation itself (BLAST-based, ≥80% identity, subjective gene-by-gene validation; the headline "+81% genes" claim depends on this). Not parameterized, not reproducible.
  • HISAT2 read alignment of the 14 RNA-seq samples from raw FASTQ. Heavy (the hard 20%); and it would NOT reproduce the curated Table-2 numbers anyway (those come post-curation). Deliberately skipped per 80/20 — see AUDIT.md.
  • Original-annotation correction counts (7,299 aligned / 6,763 / 212 / 455 bacterial genes removed): products of manual curation, not pipeline-derivable.

Approach

Reproduce by re-deriving the Table-2 pipeline statistics directly from the paper's own shipped annotation (P16: applying tools to the paper's data is valid), and by running BUSCO (third-party) on the annotation transcriptome. All compute on «our HPC»; data on «infra».

Figures / tables: Table
C1
Reported
34385 transcripts (Table 2)
Reproduced
32975 curated-subset transcripts
within tolerance
C2
Reported
36981 bp avg transcript length (Table 2)
Reproduced
37615 bp (genomic span)
within tolerance
C3
Reported
15738 annotated genes (Table 2)
Reproduced
23793 raw symbol tokens (not cleanly collapsible)
partial
C4
Reported
48% of genome covered (Table 2)
Reproduced
49.7% (1.240 Gb / 2.496 Gb genome)
within tolerance
C5
Reported
BUSCO 93.4% complete (89.4% single, 3.9% dup)
Reproduced
ERROR - BUSCO 5.8.2 transcriptome mode crashed (no aligner)
did not match
C6
Reported
shipped GTF/GFF3 annotation
Reproduced
sha256 99e1dfb5...203518 matches Git-LFS pointer; file is full uncurated StringTie superset (475980 tx)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Reproducing directly from the paper's own sha256-verified shipped annotation, 3 of 4 numeric Table-2 claims match within a few percent (transcripts 32,975 vs 34,385, length 37,615 vs 36,981 bp, coverage 49.7% vs 48%) once the curated symbol-bearing subset is isolated — no fabrication evidence, the reported values are consistent with the shared data. The main weaknesses are on the authors'/packaging side: the repo ships the full uncurated StringTie superset (475,980 tx) rather than the curated set, no scripts are provided (Galaxy workflow is a PDF), and the 15,738 gene count is not cleanly derivable due to mixed SwissProt/HGNC naming. BUSCO (93.4%) could not be tested due to a tool crash, and the headline +81%-genes manual-curation claim was out of scope, so the central conclusion is only partially confirmed. Overall a solid partial reproduction with explainable, mostly input/definition-level deviations — yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

138.7 k
tokens (I/O) · 9.6 M incl. cache
26 min
runtime · 1.32 CPU-h
8.3 GB
peak RAM
1
HPC jobs
hummel
machine