Structure of the intergenic spacers in chicken ribosomal DNA.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce 1:1 on the structural results, via a third-party-tools-on-the-paper's-data approach (no authors' code exists; P16). The deposited rDNA repeat unit MG967540 (27,000 bp) independently reproduces the paper's repeat-unit decomposition to the base pair (gene cluster 11,830 + IGS 15,170) and the per-repeat-family GC contents to within ~2-4 points (SV 78->80.5%, AL 15->17.6%, EL/VAL ~71-76%), including the striking GC-rich-vs-AT-rich SV/AL contrast. TRF independently recovers the dominant central period-93 tandem array (~9.6 kb, the paper's main IGS-length-variation block). The genuine assembly step: Canu reassembly of the raw PacBio reads (SRR10270572) recovers the rDNA-unit SEQUENCE at >99.4% BLAST identity to MG967540, but yields 6 contigs/174 kb rather than the paper's single 100,614 bp contig - an expected difference between modern Canu and the authors' 2014 HGAP3/Celera/Quiver pipeline on a tandem array. No fabrication detected: all reported structural values are derivable from the shipped public data. NOT attempted (hard 20%): MEGA7 phylogenetic trees, the 33 polymorphic positions across 3 clusters, transcriptome coverage tracks, byte-exact HGAP3/Quiver replication (software unobtainable). Overall: partial but strongly faithful.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 80assessed: 2026-06-14 ⛓ 72e3ba584926
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhat is the detailed structure and organization of the intergenic spacers (IGS) in chicken (Gallus gallus) ribosomal DNA, which are normally absent from genome assemblies due to their high repeat content?
- ★ Long-read PacBio sequencing of a chicken NOR-containing BAC clone resolved three complete IGS sequences plus rRNA gene clusters that are otherwise missing from genome assemblies. resource
- ★ Chicken IGS contains three blocks of tandem repeats (5′ block, central block, 3′ block) forming highly organized GC-rich arrays, with variation in IGS length mainly driven by the central repeat block. finding
- ★ Transcription initiation and termination sites of rDNA are located within small (SUR) and large (LUR) unique regions respectively, while no functionally significant sites occur within the tandem repeats. finding
- ★ The central repeat block is composed of phylogenetically related Elena (EL) repeats organized as (EL2F)5 + (EL2–EL1–EL3–EL3)n + (EL2)2 + (EL2F)3 tetrads. mechanism
- ★ The structure of chicken IGS differs from human, ape, Xenopus and fish IGS but shares organizational features with turtles, another Sauropsida representative. finding
- ★ Three named tandem repeat families (Svetlana/SV, Alsu/AL at 5′ end; Elena/EL central; Valerie/VAL at 3′ end) were identified and characterized. finding
- Combined with the previously reported rRNA gene cluster, the data allow the complete chicken rDNA sequence to be considered assembled with confidence in molecular DNA organization. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| long-read DNA sequencing (single-molecule) | WAG137G04 BAC clone (Wageningen library) from female White Leghorn chicken; contains chicken NOR fragment | none | BAC nucleotide sequence / rDNA repeat structure | PacBio RSII (SMRTBell Prep Kit 1.0, P6-C4 polymerase, BluePippin size selection) |
| de novo genome assembly | PacBio RSII reads from WAG137G04 BAC pool | none | assembled contig WAG137G4_utg0 (100,614 bp) | HGAP3 workflow, SMRT Analysis v2.2.0, Celera assembler, Quiver polishing |
| comparative genomics / BLAST homology search | Gallus_gallus-5.0 red junglefowl assembly (contig AADN04001305.1) | none | WGS contigs homologous to chicken rDNA repeat unit | NCBI BLAST, Repbase Update library |
| sequence alignment, repeat annotation, CpG island analysis | chicken IGS sequences | none | repeat identification, GC content, nucleotide diversity, IGS transcription | Geneious 9.0.5 |
| phylogenetic analysis (maximum likelihood) | IGS specific repeats (SV, AL, EL, VAL) | none | phylogenetic tree / evolutionary relationships between repeats | MEGA v7.0, Kimura 2-parameter + Gamma model, 500 bootstrap replicates |
| transcriptome (RNA-seq) data analysis | red junglefowl tissues: testis, ovary, kidney, liver, heart | none | whether IGS is transcribed | Chickspress project data; RNA extracted with miRNeasy Mini Kit (Qiagen) |
- – Three complete IGS in White Leghorn (WAG137G4_utg0 contig) plus one IGS in red junglefowl contig AADN04001305.1 were detected. 3 + 1 IGS
- – WAG137G4_utg0 contig contained three full rRNA gene clusters (11,871; 11,830; 11,855 bp) and one incomplete cluster, plus three intercalary IGS. 3 full clusters ~11.8 kb
- – Central repeat block length varied most among IGS, ranging from 9297 to 14,414 bp, representing the main source of IGS length differences. 9297–14,414 bp
- – 5′ repeat block ranged 950–2290 bp and 3′ repeat block ranged 2400–3766 bp across the three IGS. 5′:950–2290 bp; 3′:2400–3766 bp
- – Elena repeats fell into three clades (EL1, EL2, EL3); EL1 and EL3 monophyletic, EL2 polyphyletic; EL2–EL1–EL3–EL3 tetrads of 372–373 bp. tetrad 372–373 bp
- – Pairwise alignment of the three rRNA gene clusters detected 33 variable nucleotide positions (13 SNPs and 20 InDels), most in the 5′ETS. 33 variable positions
- – Chicken IGS structure was very similar to the 14,002-bp IGS in red junglefowl contig AADN04001305.1. 14,002 bp
- count average number of rDNA repeat copies in diploid sets ranged from 279 to 368 (prior work by Delany and Krupkin on rDNA copy number)
- other rDNA repeat unit length variability ranged from 11 to 50 kb (intra- and inter-NOR variability mainly due to IGS size heterogeneity (prior work))
- count 33 variable nucleotide positions (13 SNPs, 20 InDels) (pairwise alignment of three rRNA gene clusters)
- count 7168 reads; mean read length 9 kb; quality value 48 (PacBio RSII sequencing of WAG137G04 BAC)
- other longest contig 100,614 bp at average coverage 500X (HGAP3 assembly WAG137G4_utg0)
- other central block Elena repeats 92–94 bp, (C+G) content 65–76% (central repeat block composition)
- other Svetlana repeats 137–158 bp, 78% GC; Alsu repeats 209–303 bp, 15% GC (5′ repeat block tandem repeats)
- count 8390 filtered reads (78,347,167 bp); 776 preassembled reads (5,031,668 bp) (read filtering and preassembly steps)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a descriptive molecular genomics study characterizing the structure of chicken intergenic spacer (IGS) sequences from a single BAC clone sequenced by PacBio RSII long-read technology. Statistical methods were limited to phylogenetic reconstruction of repeat relationships (maximum likelihood with Kimura 2-parameter model, bootstrap-validated) and tabulation of nucleotide diversity and GC content of repeat families. No inferential group-comparison tests were applied; the work is primarily sequence annotation and evolutionary description.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Maximum likelihood phylogenetic reconstruction with Kimura 2-parameter model + Gamma distribution (5 categories, parameter = 2.8698) | Figures 4 and 6: phylogenetic relationships among Elena (EL) and Valerie (VAL) IGS repeat variants | Number of repeat sequences analyzed not explicitly stated as a single count; derived from repeat copies across three IGS in the WAG137G4_utg0 contig | not stated |
| Bootstrap test (500 replications) for phylogenetic tree topology reliability | Figures 4 and 6: Elena and Valerie repeat phylogenetic trees | 500 bootstrap replicates | not stated |
| Pairwise nucleotide sequence alignment (variable site counting) | Identification of 33 variable nucleotide positions (13 SNPs, 20 InDels) across three rRNA gene clusters in WAG137G4_utg0 | 3 complete rRNA gene cluster sequences | na |
| BLAST homology search | Identification of WGS contigs homologous to the chicken rDNA unit in Gallus_gallus-5.0 assembly; E. coli read filtering | null | na |
| Dot-plot self-similarity analysis | Identification of tandem repeat block structure and unique regions within each IGS (Additional file 6: Figure S2) | 4 IGS sequences (3 from BAC clone, 1 from AADN04001305.1) | na |
-
Phylogenetic trees for Elena and Valerie repeats were constructed using the maximum likelihood method with the Kimura 2-parameter (K2P) substitution model↳ Could also: Bayesian inference (e.g., MrBayes or BEAST) with the same or a more general substitution model (e.g., GTR+G+I) could also be used for phylogenetic reconstruction — Bayesian inference provides posterior probability support values (an alternative reliability metric to bootstrap), and richer substitution models like GTR+G+I allow more parameters to be estimated from the data, which can matter when repeat sequences have highly unequal base frequencies (as the paper's GC-rich repeats do)
-
Bootstrap reliability of tree topology was assessed with 500 replications↳ Could also: 1000 bootstrap replications is a widely used standard, and ultrafast bootstrap approximation (UFBoot, implemented in IQ-TREE) can achieve more accurate support values with comparable or lower computational cost — Larger bootstrap replication numbers or ultrafast bootstrap can improve the precision and calibration of support values, particularly for shallow nodes that may be unstable
-
Nucleotide diversity within each repeat family is reported as a single percentage in Table 1 with no measure of dispersion or confidence↳ Could also: A 95% confidence interval or standard deviation across pairwise diversity estimates could also accompany the point estimate — Reporting spread alongside the mean diversity would convey how variable the diversity estimate is within a repeat family, which is informative given the small per-family copy numbers
-
IGS structural variants were compared by visual dot-plot analysis and manual annotation rather than by a quantitative similarity metric↳ Could also: Percent identity matrices or pairwise distance matrices computed from multiple-sequence alignments (e.g., via MUSCLE or MAFFT followed by a distance summary) could also quantify structural similarity among the four IGS sequences — Quantitative pairwise distances would complement the dot-plot visualization by making the degree of divergence between IGS variants explicit and reproducible
-
Transcriptome read mapping to the IGS reference was used to infer IGS transcription across five tissues, using reanalysis of public Chickspress data↳ Could also: A formal read-count or coverage-based analysis (e.g., featureCounts plus a normalized coverage depth statistic) could also be applied to quantify relative transcriptional activity across annotated IGS regions and tissues — Quantitative coverage summaries would allow the extent of IGS transcription to be compared across tissues and regions in a reproducible, numerically precise way rather than relying solely on presence/absence of mapped reads
-
A single substitution model (K2P) was selected for all repeat groups without a formal model-selection step↳ Could also: A model selection procedure (e.g., jModelTest, ModelTest-NG, or the built-in MEGA model selection) could also be run to identify the best-fitting substitution model for each repeat alignment prior to tree construction — Formal model selection ensures that the chosen model fits the observed substitution patterns, which can affect branch-length estimates and topology support, particularly for GC-biased sequences like the Elena repeats
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-31655542
Paper: Dyomin et al. 2019, Genet Sel Evol — "Structure of the intergenic spacers in chicken ribosomal DNA." DOI 10.1186/s12711-019-0501-7.
What the paper did (pipeline chain)
- PacBio RS II long-read sequencing of BAC clone WAG137G04 (chicken rDNA).
- Assembly: SMRT Analysis v2.2.0 / HGAP3 → BLASR filter (remove E. coli
K-12, quality<0.75, len<500) → preassembly → Celera assembler draft →
Quiver polish. Result: contig
WAG137G4_utg0, 100,614 bp, ~500× cov, containing 3 complete + 1 partial rDNA repeat units. - Annotation (Geneious 9.0.5), BLAST vs Gallus_gallus-5.0 + Repbase.
- Structural dissection of the intergenic spacer (IGS): three tandem-repeat blocks (named SV / AL / EL / VAL repeat families) + a Large Unique Region (LUR) + Small Unique Region (SUR).
- Deposited one complete rDNA repeat unit: GenBank MG967540 = 27,000 bp (gene cluster ~11,830 bp + IGS 15,170 bp — the shortest of three variants).
- Phylogenetics (MEGA7) of repeat families; transcriptome coverage (Chickspress).
Data & code availability
- Raw reads: SRA
PRJNA577229→ run SRR10270572 (PacBio RS II, 8,654 reads, 79.4 Mbp, 23 MB fastq.gz). PUBLIC, obtainable from ENA. ✅ - Assembled reference: GenBank MG967540 (27,000 bp). PUBLIC. ✅
- Code: NO code-availability statement in the paper. The BRIEF's repo link
(
PacificBiosciences/Bioinformatics-Training) is a generic PacBio training repo, not the authors' analysis code. Per P16, we reproduce by applying standard third-party tools to the paper's own public data.
IN SCOPE (pipeline-derived, reproducible)
| id | reported result | reproduction approach |
|---|---|---|
| C1 | rDNA repeat unit length = 27,000 bp (= gene cluster 11,830 + IGS 15,170) | measure MG967540 length (seqkit) |
| C2 | IGS / repeat families are GC-rich & CpG-island-rich (SV 78% GC, EL 65–76%, VAL 68–80%; AL low 15%) | compute GC% of IGS region + per-repeat-family (seqkit fx2tab / TRF) |
| C3 | IGS built from tandem-repeat blocks: SV 137–158 bp, AL 209–303 bp, EL 89–94 bp, VAL 72–95 bp period | run Tandem Repeats Finder (TRF) on MG967540 IGS → recover period sizes |
| C4 | Assembly of the BAC reads → ~100 kb contig with multiple rDNA repeat units | reassemble SRR10270572 with Canu (direct successor of the Celera assembler used) on «our HPC»; report contig length & rDNA content (BLAST vs MG967540) |
OUT OF SCOPE (not attempted, why)
- Phylogenetic trees (MEGA7) — manual/GUI, K2P+Γ params given but tree topology is a figure, not a pinnable scalar. (non_pipeline-ish / no_expected_scalar)
- Transcriptome coverage tracks (Chickspress) — qualitative figure.
- 33 polymorphic positions across the 3 clusters — requires the full 100 kb multi-unit assembly + manual alignment; deep into the hard 20%.
- Exact HGAP3/SMRT v2.2.0/Quiver replication — that software (2014) is effectively unobtainable; Canu is the faithful modern stand-in.
Heavy compute
All assembly + analysis on «our HPC» SLURM, data on «infra». «host» holds only small result values + pointers.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All reported structural results are independently re-derivable from the authors' own public data (MG967540 + SRR10270572): the 27,000 bp repeat unit decomposes to the base pair (11,830 + 15,170), per-family GC contents reproduce within 2-4 points, and TRF recovers the dominant central period-93 ~9.6 kb array — no fabrication. The only substantive deviation is the BAC assembly contiguity (6 contigs/174 kb vs a single 100,614 bp contig), which is an expected assembler/version artifact (modern Canu vs the 2014 HGAP3/Celera+Quiver pipeline) on a tandem array, with the rDNA sequence still recovered at >99.4% identity — this is on the methods/tooling side, not the authors'. The central conclusion (GC-rich tandem-block IGS structure with the central block as the main length-variation source) holds fully, so this is a solid reproduction with explainable, technical deviations.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.