Structure of the intergenic spacers in chicken ribosomal DNA.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce 1:1 on the structural results, via a third-party-tools-on-the-paper's-data approach (no authors' code exists; P16). The deposited rDNA repeat unit MG967540 (27,000 bp) independently reproduces the paper's repeat-unit decomposition to the base pair (gene cluster 11,830 + IGS 15,170) and the per-repeat-family GC contents to within ~2-4 points (SV 78->80.5%, AL 15->17.6%, EL/VAL ~71-76%), including the striking GC-rich-vs-AT-rich SV/AL contrast. TRF independently recovers the dominant central period-93 tandem array (~9.6 kb, the paper's main IGS-length-variation block). The genuine assembly step: Canu reassembly of the raw PacBio reads (SRR10270572) recovers the rDNA-unit SEQUENCE at >99.4% BLAST identity to MG967540, but yields 6 contigs/174 kb rather than the paper's single 100,614 bp contig - an expected difference between modern Canu and the authors' 2014 HGAP3/Celera/Quiver pipeline on a tandem array. No fabrication detected: all reported structural values are derivable from the shipped public data. NOT attempted (hard 20%): MEGA7 phylogenetic trees, the 33 polymorphic positions across 3 clusters, transcriptome coverage tracks, byte-exact HGAP3/Quiver replication (software unobtainable). Overall: partial but strongly faithful.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 80assessed: 2026-06-14 ⛓ 72e3ba584926
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe complex, highly repetitive structure of the chicken ribosomal DNA intergenic spacer (IGS) has prevented complete sequencing with short-read methods, so the authors used long-read PacBio sequencing to determine whether the full IGS structure could be resolved and characterized, including its regulatory and repeat elements.
- ★ Long-read PacBio RSII sequencing of a BAC clone plus HGAP assembly can resolve the complete, highly repetitive chicken IGS structure that short-read (Illumina) sequencing previously failed to assemble. method
- ★ Three complete IGS sequences (from the BAC clone) and one IGS sequence (from a red junglefowl contig) were identified, each containing three internal tandem repeat blocks (5′, central, 3′) plus small and large unique regions (SUR, LUR). finding
- ★ The central repeat block (Elena/EL repeats), ranging from 9297 to 14,414 bp, is the main source of length variation among the different chicken IGS. finding
- ★ Transcription initiation and termination sites map to the small and large unique regions (SUR and LUR), respectively, while no functionally significant sites were found within the tandem repeats themselves. finding
- ★ Elena repeats are organized into highly structured tetrads (EL2–EL1–EL3–EL3) and follow the overall pattern (EL2F)5+(EL2–EL1–EL3–EL3)n+(EL2)2+(EL2F)3. finding
- ★ The chicken IGS structure differs from that of human, apes, Xenopus, and fish, but shares molecular organization features with turtles, another Sauropsida lineage. mechanism
- ★ Combined with the previously determined rRNA gene cluster sequence, this work provides a complete, confidently assembled chicken rDNA repeat unit sequence. resource
- Three distinct tandem repeat families were named and characterized: Svetlana (SV) and Alsu (AL) repeats at the 5′ block, Elena (EL) repeats in the central block, and Valerie (VAL) repeats at the 3′ block. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Long-read PacBio RSII single-molecule sequencing | BAC clone WAG137G04 (White Leghorn chicken, Wageningen BAC library) | none | raw sequence reads for rDNA-containing BAC clone | PacBio RSII with P6-C4 polymerase |
| Genome assembly (HGAP3/Celera/Quiver workflow) | PacBio reads from BAC clone WAG137G04 | none | assembled contig (WAG137G4_utg0) representing rDNA repeat unit and IGS | SMRT Analysis v2.2.0 / HGAP / Celera assembler / Quiver |
| BLAST sequence homology search | Gallus_gallus-5.0 whole-genome assembly (red junglefowl, contig AADN04001305.1) | none | identification of WGS contigs homologous to chicken rDNA repeat unit | NCBI BLAST |
| Sequence alignment, repeat annotation, CpG island analysis | Assembled chicken IGS sequences (WAG137G4_utg0 and AADN04001305.1) | none | identification/annotation of tandem repeat blocks, unique regions, CpG distribution | Geneious 9.0.5 |
| Phylogenetic reconstruction (maximum likelihood, Kimura 2-parameter model) | IGS-specific tandem repeat sequences (SV, EL, VAL groups) | none | clade structure and evolutionary relationships among repeat variants | MEGA 7.0 |
| Transcriptome (RNA-seq) data analysis | Red junglefowl tissues: testis, ovary, kidney, liver, heart (Chickspress project) | none | presence/absence of IGS transcription | Chickspress raw RNA-seq data (miRNeasy Mini Kit extraction) |
- – Three complete IGS sequences from the BAC clone and one from the red junglefowl contig were identified, each with three tandem repeat blocks and conserved unique regions
- – Central repeat block length varied substantially and was the primary driver of overall IGS length differences 9297-14,414 bp
- – 5′ repeat block length varied among the four IGS analyzed 950-2290 bp
- – 3′ repeat block length varied among the four IGS analyzed 2400-3766 bp
- – Pairwise alignment of the three rRNA gene clusters found variable nucleotide positions, concentrated in the 5′ETS with none in 5.8S rRNA 33 variable positions (13 SNPs, 20 InDel)
- – Elena repeats form highly organized tetrads combining EL2, EL1, and two EL3 repeats 372-373 bp total tetrad length
- – Assembled contig from the BAC clone was obtained with high average sequencing coverage, though coverage was irregular across repeat regions 500X average coverage, 100,614 bp contig
- count 33 variable nucleotide positions (13 SNPs, 20 InDel) (pairwise alignment of the three rRNA gene clusters in WAG137G4_utg0)
- count 100,614 bp (length of assembled WAG137G4_utg0 contig)
- other 500X average coverage (coverage of WAG137G4_utg0 contig assembly, though irregular across repeat regions)
- count 7168 reads, mean read length 9 kb (PacBio reads obtained for WAG137G04 BAC clone)
- fold_change central repeat block 9297-14,414 bp vs. 5′ block 950-2290 bp vs. 3′ block 2400-3766 bp (relative size and variability of the three IGS tandem repeat blocks)
- other (C+G) content: SV repeats 78%, AL repeats 15%, EL repeats 65-76%, VAL repeats 67-80% (GC content of the different IGS tandem repeat families)
- other Bootstrap test with 500 replications (reliability assessment of maximum likelihood phylogenetic tree topology)
- count 776 preassembled reads, 5,031,668 bp (long, highly accurate reads generated during HGAP preassembly step)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a descriptive molecular genomics study characterizing the structure of chicken intergenic spacer (IGS) sequences from a single BAC clone sequenced by PacBio RSII long-read technology. Statistical methods were limited to phylogenetic reconstruction of repeat relationships (maximum likelihood with Kimura 2-parameter model, bootstrap-validated) and tabulation of nucleotide diversity and GC content of repeat families. No inferential group-comparison tests were applied; the work is primarily sequence annotation and evolutionary description.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Maximum likelihood phylogenetic reconstruction with Kimura 2-parameter model + Gamma distribution (5 categories, parameter = 2.8698) | Figures 4 and 6: phylogenetic relationships among Elena (EL) and Valerie (VAL) IGS repeat variants | Number of repeat sequences analyzed not explicitly stated as a single count; derived from repeat copies across three IGS in the WAG137G4_utg0 contig | not stated |
| Bootstrap test (500 replications) for phylogenetic tree topology reliability | Figures 4 and 6: Elena and Valerie repeat phylogenetic trees | 500 bootstrap replicates | not stated |
| Pairwise nucleotide sequence alignment (variable site counting) | Identification of 33 variable nucleotide positions (13 SNPs, 20 InDels) across three rRNA gene clusters in WAG137G4_utg0 | 3 complete rRNA gene cluster sequences | na |
| BLAST homology search | Identification of WGS contigs homologous to the chicken rDNA unit in Gallus_gallus-5.0 assembly; E. coli read filtering | null | na |
| Dot-plot self-similarity analysis | Identification of tandem repeat block structure and unique regions within each IGS (Additional file 6: Figure S2) | 4 IGS sequences (3 from BAC clone, 1 from AADN04001305.1) | na |
-
Phylogenetic trees for Elena and Valerie repeats were constructed using the maximum likelihood method with the Kimura 2-parameter (K2P) substitution model↳ Could also: Bayesian inference (e.g., MrBayes or BEAST) with the same or a more general substitution model (e.g., GTR+G+I) could also be used for phylogenetic reconstruction — Bayesian inference provides posterior probability support values (an alternative reliability metric to bootstrap), and richer substitution models like GTR+G+I allow more parameters to be estimated from the data, which can matter when repeat sequences have highly unequal base frequencies (as the paper's GC-rich repeats do)
-
Bootstrap reliability of tree topology was assessed with 500 replications↳ Could also: 1000 bootstrap replications is a widely used standard, and ultrafast bootstrap approximation (UFBoot, implemented in IQ-TREE) can achieve more accurate support values with comparable or lower computational cost — Larger bootstrap replication numbers or ultrafast bootstrap can improve the precision and calibration of support values, particularly for shallow nodes that may be unstable
-
Nucleotide diversity within each repeat family is reported as a single percentage in Table 1 with no measure of dispersion or confidence↳ Could also: A 95% confidence interval or standard deviation across pairwise diversity estimates could also accompany the point estimate — Reporting spread alongside the mean diversity would convey how variable the diversity estimate is within a repeat family, which is informative given the small per-family copy numbers
-
IGS structural variants were compared by visual dot-plot analysis and manual annotation rather than by a quantitative similarity metric↳ Could also: Percent identity matrices or pairwise distance matrices computed from multiple-sequence alignments (e.g., via MUSCLE or MAFFT followed by a distance summary) could also quantify structural similarity among the four IGS sequences — Quantitative pairwise distances would complement the dot-plot visualization by making the degree of divergence between IGS variants explicit and reproducible
-
Transcriptome read mapping to the IGS reference was used to infer IGS transcription across five tissues, using reanalysis of public Chickspress data↳ Could also: A formal read-count or coverage-based analysis (e.g., featureCounts plus a normalized coverage depth statistic) could also be applied to quantify relative transcriptional activity across annotated IGS regions and tissues — Quantitative coverage summaries would allow the extent of IGS transcription to be compared across tissues and regions in a reproducible, numerically precise way rather than relying solely on presence/absence of mapped reads
-
A single substitution model (K2P) was selected for all repeat groups without a formal model-selection step↳ Could also: A model selection procedure (e.g., jModelTest, ModelTest-NG, or the built-in MEGA model selection) could also be run to identify the best-fitting substitution model for each repeat alignment prior to tree construction — Formal model selection ensures that the chosen model fits the observed substitution patterns, which can affect branch-length estimates and topology support, particularly for GC-biased sequences like the Elena repeats
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-31655542
Paper: Dyomin et al. 2019, Genet Sel Evol — "Structure of the intergenic spacers in chicken ribosomal DNA." DOI 10.1186/s12711-019-0501-7.
What the paper did (pipeline chain)
- PacBio RS II long-read sequencing of BAC clone WAG137G04 (chicken rDNA).
- Assembly: SMRT Analysis v2.2.0 / HGAP3 → BLASR filter (remove E. coli
K-12, quality<0.75, len<500) → preassembly → Celera assembler draft →
Quiver polish. Result: contig
WAG137G4_utg0, 100,614 bp, ~500× cov, containing 3 complete + 1 partial rDNA repeat units. - Annotation (Geneious 9.0.5), BLAST vs Gallus_gallus-5.0 + Repbase.
- Structural dissection of the intergenic spacer (IGS): three tandem-repeat blocks (named SV / AL / EL / VAL repeat families) + a Large Unique Region (LUR) + Small Unique Region (SUR).
- Deposited one complete rDNA repeat unit: GenBank MG967540 = 27,000 bp (gene cluster ~11,830 bp + IGS 15,170 bp — the shortest of three variants).
- Phylogenetics (MEGA7) of repeat families; transcriptome coverage (Chickspress).
Data & code availability
- Raw reads: SRA
PRJNA577229→ run SRR10270572 (PacBio RS II, 8,654 reads, 79.4 Mbp, 23 MB fastq.gz). PUBLIC, obtainable from ENA. ✅ - Assembled reference: GenBank MG967540 (27,000 bp). PUBLIC. ✅
- Code: NO code-availability statement in the paper. The BRIEF's repo link
(
PacificBiosciences/Bioinformatics-Training) is a generic PacBio training repo, not the authors' analysis code. Per P16, we reproduce by applying standard third-party tools to the paper's own public data.
IN SCOPE (pipeline-derived, reproducible)
| id | reported result | reproduction approach |
|---|---|---|
| C1 | rDNA repeat unit length = 27,000 bp (= gene cluster 11,830 + IGS 15,170) | measure MG967540 length (seqkit) |
| C2 | IGS / repeat families are GC-rich & CpG-island-rich (SV 78% GC, EL 65–76%, VAL 68–80%; AL low 15%) | compute GC% of IGS region + per-repeat-family (seqkit fx2tab / TRF) |
| C3 | IGS built from tandem-repeat blocks: SV 137–158 bp, AL 209–303 bp, EL 89–94 bp, VAL 72–95 bp period | run Tandem Repeats Finder (TRF) on MG967540 IGS → recover period sizes |
| C4 | Assembly of the BAC reads → ~100 kb contig with multiple rDNA repeat units | reassemble SRR10270572 with Canu (direct successor of the Celera assembler used) on «our HPC»; report contig length & rDNA content (BLAST vs MG967540) |
OUT OF SCOPE (not attempted, why)
- Phylogenetic trees (MEGA7) — manual/GUI, K2P+Γ params given but tree topology is a figure, not a pinnable scalar. (non_pipeline-ish / no_expected_scalar)
- Transcriptome coverage tracks (Chickspress) — qualitative figure.
- 33 polymorphic positions across the 3 clusters — requires the full 100 kb multi-unit assembly + manual alignment; deep into the hard 20%.
- Exact HGAP3/SMRT v2.2.0/Quiver replication — that software (2014) is effectively unobtainable; Canu is the faithful modern stand-in.
Heavy compute
All assembly + analysis on «our HPC» SLURM, data on «infra». «host» holds only small result values + pointers.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All reported structural results are independently re-derivable from the authors' own public data (MG967540 + SRR10270572): the 27,000 bp repeat unit decomposes to the base pair (11,830 + 15,170), per-family GC contents reproduce within 2-4 points, and TRF recovers the dominant central period-93 ~9.6 kb array — no fabrication. The only substantive deviation is the BAC assembly contiguity (6 contigs/174 kb vs a single 100,614 bp contig), which is an expected assembler/version artifact (modern Canu vs the 2014 HGAP3/Celera+Quiver pipeline) on a tandem array, with the rDNA sequence still recovered at >99.4% identity — this is on the methods/tooling side, not the authors'. The central conclusion (GC-rich tandem-block IGS structure with the central block as the main length-variation source) holds fully, so this is a solid reproduction with explainable, technical deviations.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.