Evolution of Highly Repetitive Silk Genes in the Luna Moth, Actias luna.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> 1:1 reproduction of the deterministic, shipped-data-derived computational claims. Via P16 (recompute from the authors' deposited figshare data: SericinAlignmentMAFFT.fasta protein alignment + Aluna.gff3 annotation), all 8 sericin protein lengths, all 8 molecular weights, and all 8 exon counts in Table 1 reproduce EXACTLY using Biopython ProtParam + GFF parsing on «our HPC»/«infra»; the Results-text amino-acid-composition ranges equal the min/max of the computed per-gene values (e.g. glycine 13.2-20.5% = ser1..serG endpoints exactly). Isoelectric points reproduce only partially (serG within-tol, but serF and the group-4 acidic proteins are offset 0.2-0.5 because pI is pKa-scale dependent; relative ranking preserved). No fabrication indicators: Table 1 is fully regenerable from the deposited sequences/annotation. NOT attempted (out of scope or the hard 20%): manual Geneious repeat-number extraction (R1/R2), the SplitsTree phylogenetic network and 4-group assignment, ChromSyn/telociraptor synteny maps, the edgeR instar differential-expression patterns (qualitative, would need full SRA->Subread->featureCounts->edgeR rerun), and a full genome BUSCO completeness run (the reported 5286 is just the lepidoptera_odb10 lineage size).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 89assessed: 2026-06-16 ⛓ 0ac292cd8e98
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBecause sericins—highly repetitive, serine-rich proteins forming the outer silk coating—may diversify silk properties within and across lepidopteran species, the authors test how sericin gene duplications have evolved in the Luna moth (Actias luna) and other Saturniidae, and whether duplication has led to subfunctionalization that enables shifts in silk composition.
- ★ Eight sericin genes were identified in the Actias luna genome, including two clusters of closely related paralogs. finding
- ★ Sericin gene duplications in A. luna and other Saturniidae have led to convergent subfunctionalization, enabling changes to silk composition. mechanism
- ★ A. luna sericins exhibit considerable variation in repeat number and amino acid composition and display distinct gene expression patterns across life stages. finding
- ★ Sericin gene duplications enable dynamic shifts in silk composition both within and between species, potentially reflecting adaptive responses to ecological/functional demands. mechanism
- ★ Based on protein-level similarity and synteny across nine Bombycoidea, sericins fall into four distinct groups within Saturniidae (Sericin 1 Group, Group 2, Group 3, and others). finding
- A. luna ser1 is an ortholog of B. mori sericin 1, sharing a conserved N-terminal CXCX motif characteristic of Sericin 1 across Lepidoptera. finding
- A multi-method strategy (sequence similarity via BLAST, shared sequence motifs via nhmmer, and genomic proximity) was used to identify sericin genes, with manual annotation using short-read RNA-seq and ISO-seq data. method
- Group 3 sericins share a conserved 38-amino-acid repeat motif and are well represented across saturniids but absent in B. mori. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Genome BLAST sequence-similarity search | Actias luna genome assembly (queried with 22 known Bombycoidea sericins) | none | location/identification of candidate sericin genes | BLAST (Altschul et al. 1990; Camacho et al. 2009) |
| HMM profile / sequence-motif search | Actias luna genome | none | sericin genes matching conserved motifs from aligned known sericins | nhmmer (Wheeler and Eddy 2013) |
| Manual gene annotation | Actias luna genome, silk gland transcripts | none | gene structure/exon-intron annotation of eight sericins | BRAKER3 automated annotation; Geneious |
| Long-read (ISO-seq) RNA sequencing transcriptome | Actias luna silk glands across different life stages | none (developmental life-stage comparison) | silk-gland-specific expression and stage-specific expression of sericins | ISO-seq (PacBio) |
| Short-read RNA sequencing | Actias luna silk glands | none | transcript evidence for manual annotation | — |
| Genome-wide synteny mapping | Genomes of A. luna, A. yamamai, S. ricini, and B. mori | none | conserved single-copy gene synteny blocks and sericin gene chromosomal locations | BUSCO genes; ChromSyn |
| Protein sequence phylogenetic network analysis | Translated sericin proteins from nine Bombycoidea species | none | pairwise protein identity / sericin groupings with bootstrap support | SplitsTree |
| In silico protein property prediction | A. luna encoded sericin proteins (signal peptide removed) | none | protein length, molecular weight, isoelectric point, repeat number/length | Geneious |
- – Eight putative sericin genes (serA, ser1, serB, serC, serD, serE, serF, serG) identified in A. luna; BLAST and motif methods both yielded the same six, with two more found by genomic proximity. 8 genes
- – Silk-gland-specific expression confirmed for five of eight sericins (ser1, serA-D) via the silk gland ISO-seq transcriptome. 5 of 8
- – Six sericins (serB-G) lie in two distinct gene clusters within a ~1.5 Mb region on contig ptg000081l. ~1.5 Mb region
- – Sericins form four distinct groups within Saturniidae based on protein similarity and synteny; Sericin 1 Group and Group 3 had high bootstrap support. 4 groups
- – A. luna SerE-G (Group 3) are highly similar at the protein level despite considerable variation in repeat number (serE:37, serF:65, serG:47). repeats 37–65
- – Manual annotation improved the automated BRAKER3 annotation for seven of eight sericins; only serD was previously correctly annotated. 7 of 8
- ▲ A. yamamai genome shows an expansion of Group 3 with at least four distinct members (Src2-5). ≥4 members
- – serB-D vary strongly in exon number (4 to 19), repeat number (21 to 89), and length (241 to 1439 aa). exons 4–19; repeats 21–89
- count eight sericin genes (sericin genes identified in A. luna genome)
- count 22 previously characterized sericins from eight species (reference sericin dataset used for queries)
- count six sericin genes described in Bombyx mori (known B. mori sericins; three in cocoon silk, three in early larval silk)
- other 20% to 50% (proportion of silk fiber made up by the outer sericin-rich coating)
- other ~70 million years (estimated divergence time of Saturniidae from Bombycidae)
- count Protein length serA 2031 aa; ser1 3039 aa; serB 966 aa; serC 260 aa; serD 2458 aa; serE 1457 aa; serF 2765 aa; serG 2099 aa (encoded protein lengths of A. luna sericins (Table 1))
- other first cluster ~105 kb (serB-D); second cluster <55 kb (serE-G) (genomic span of the two sericin gene clusters on ptg000081l)
- other bootstrap value over 10 (threshold for splits shown in the SplitsTree phylogenetic network)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This comparative genomics and molecular evolution study identified eight sericin silk genes in Actias luna using sequence similarity (BLAST) and profile HMM (nhmmer) searches against 22 reference sericins from eight Bombycoidea species. Genomic relationships were assessed via BUSCO-based synteny mapping across four species, and inter-species sericin groupings were inferred through a pairwise-distance phylogenetic network (SplitsTree, bootstrap threshold >10). Life-stage gene expression was assessed qualitatively from IsoSeq long-read silk-gland transcriptomes, with no formal statistical modeling of expression differences reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| BLAST sequence similarity search | Identification of A. luna sericin genes from the genome using 22 reference sericins | 22 reference sericins from 8 Bombycoidea species | na |
| nhmmer profile HMM search | Identification of A. luna sericin genes using conserved sequence motifs derived from the 22-sericin alignment | 22 reference sericins used to build the profile | na |
| SplitsTree neighbor-net phylogenetic network with bootstrap support (threshold >10) | Grouping of sericins from 9 Bombycoidea species by pairwise protein identity (Fig. 2) | Translated sericin protein sequences from 9 species; exact sequence count not stated in excerpt | not stated |
| BUSCO-based genome-wide synteny mapping (ChromSyn) | Comparison of sericin genomic locations across A. luna, A. yamamai, S. ricini, and B. mori (Figs. 1, S1) | 4 genome assemblies | na |
| Pairwise protein identity comparison (tabulated) | Assessment of similarity among A. luna and comparator sericins (Table S2) | Sericin protein sequences from 9 Bombycoidea species | not stated |
| IsoSeq transcript detection (qualitative presence/absence) | Confirmation of silk-gland expression for 5 of 8 A. luna sericins across life stages (Table 1) | Number of biological replicates not stated | not stated |
-
Sericin evolutionary relationships were inferred using a phylogenetic network (SplitsTree) based on pairwise protein distances, with splits displayed only above a bootstrap threshold of >10↳ Could also: A maximum-likelihood or Bayesian phylogenetic tree (e.g., IQ-TREE, RAxML, MrBayes) with an explicit amino-acid substitution model could also have been used, or the network could be reported with a more conservative bootstrap display threshold (e.g., >50) — Tree-based methods provide rooted, bifurcating topologies with branch-length estimates and posterior probabilities; they are more interpretable for directional ancestral inference, while networks are well-suited when reticulate evolution is expected — presenting both alongside higher bootstrap thresholds would allow readers to assess which groupings are robust
-
Life-stage gene expression was assessed qualitatively as transcript presence or absence in the IsoSeq silk-gland transcriptome↳ Could also: Quantitative differential expression analysis using replicated short-read RNA-seq with tools such as DESeq2 or edgeR could also have been applied — Quantitative approaches with biological replicates yield expression-level estimates, fold-changes, and statistical significance, enabling formal comparisons of expression magnitude across life stages and sericin genes rather than presence/absence calls alone
-
Genomic synteny across four species was assessed using BUSCO single-copy ortholog anchors (ChromSyn)↳ Could also: Whole-genome alignment tools such as MUMmer/nucmer, minimap2, or LAST with synteny visualization (e.g., SyRI, D-GENIES) could also have been used — Whole-genome alignment captures fine-scale synteny and structural variation independently of conserved gene subsets, which may be sparsely distributed in repetitive or rapidly evolving loci such as sericin clusters
-
Sericins were grouped using overall pairwise protein identity without a formal multiple-sequence alignment-based phylogenetic model↳ Could also: Domain-aware or repeat-aware multiple sequence alignment (e.g., MAFFT with iterative refinement, or MUSCLE) followed by phylogenetic inference on conserved non-repeat regions could also have been applied — Highly repetitive proteins are challenging for global alignment; isolating conserved flanking or signal-peptide regions for tree inference can reduce alignment artefacts introduced by variable repeat number, yielding topologies that better reflect shared ancestry rather than compositional similarity
-
Gene identification relied on two complementary search strategies (BLAST and nhmmer) plus manual annotation; a third tier used genomic proximity↳ Could also: An automated genome annotation pipeline specifically trained on insect repeat-rich genes (e.g., MAKER2 or Helixer with insect models) integrated with RNA evidence could also have been used as an independent validation layer — Automated pipelines with explicit training data improve reproducibility and reduce the subjectivity of manual curation; benchmarking automated predictions against the manual annotations would quantify sensitivity and specificity for this gene family
-
Protein characteristics (length, isoelectric point, repeat number) were reported as single point values per gene without cross-gene statistical summaries↳ Could also: Descriptive statistics (mean, SD or range) across sericin groups or species, combined with a non-parametric comparison (e.g., Kruskal-Wallis) of repeat number across groups, could also have been reported — Summary statistics across sericin groups would allow readers to quantitatively compare compositional variation between groups and assess whether observed differences (e.g., in repeat number) are systematic rather than idiosyncratic to individual genes
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41738778
Paper: Evolution of Highly Repetitive Silk Genes in the Luna Moth, Actias luna. Genome Biol Evol 2026. DOI 10.1093/gbe/evag036. PMCID PMC12962854.
Data shipped (public):
- BioProject PRJNA1072661 (SRA): short-read RNA-seq (5 instars) + IsoSeq silk-gland long reads.
- Genome assembly GCA_039707435.1 (Actias luna, ~532 Mb, separate 2024 paper PMC11258402).
- figshare 10.6084/m9.figshare.29847707:
SericinAlignmentMAFFT.fasta(176 KB) — MAFFT protein alignment, 34 sericins (incl. the 8 A. luna).Aluna.gff3(44 MB) — genome feature file (annotation).SupplTableS2.xlsx(21 KB) — supplementary table.
Code link in brief: github.com/harvardinformatics/TranscriptomeAssemblyTools — a third-party read-cleaning wrapper set (Trimmomatic/Rcorrector helpers), NOT the authors' analysis code. Per P16 we reproduce by running standard tools on the paper's own shipped data.
In scope (pipeline-derived, attempted)
| id | reported result | location | pipeline | feasibility |
|---|---|---|---|---|
| C1 | Sericin protein lengths (aa) for SerA,Ser1,SerB–G | Table 1 | extract ungapped seq from shipped MAFFT alignment | easy, deterministic |
| C2 | Sericin protein MW (kDa) | Table 1 | Biopython ProtParam molecular_weight on shipped seq | easy, deterministic |
| C3 | Isoelectric points (Group4 3.01–3.87; serF 7.15; serG 5.24) | Results text | ProtParam isoelectric_point | easy, deterministic |
| C4 | Amino-acid composition (Ser/Thr/Gly %, acidic %) per gene | Results text | aa-composition count from shipped seq | easy, deterministic |
| C5 | Exon counts per sericin (Table 1 "Exons") | Table 1 | parse shipped Aluna.gff3 | easy, deterministic |
| C6 | BUSCO single-copy count = 5,286 | Methods | BUSCO v5.3.0 lepidoptera_odb10 on GCA_039707435.1 | medium (genome BUSCO job) |
C1–C5 are directly derivable from the shipped figshare files → ideal auditable 1:1 checks (and a fabrication probe: do Table 1's reported MW/length match the actual deposited sequences?). C6 is a heavier but clean single-number genome pipeline output.
Out of scope / not attempted (with reason)
- Repeat numbers R1/R2 (Table 1) — manual extraction in Geneious v11.1.5; subjective motif boundary calls, not a deterministic pipeline. (manual)
- Phylogenetic network (94 splits / 83 shown, 4 groups) — MAFFT→IQ-TREE→SplitsTree; the alignment is shipped but exact SplitsTree split set + bootstrap is the hard last 20% and topology grouping is interpretive. (partial-feasible, deprioritized)
- Chromosome synteny / telomere maps (ChromSyn, telociraptor) — visual, 4-genome, no single reported number to grade. (figure-only)
- edgeR differential expression / instar expression patterns — qualitative ("drop to near zero by L5"); needs full SRA→Subread→featureCounts→edgeR rerun; no pinned numeric claim to grade 1:1. (qualitative, heavy)
- Signal-peptide / BRAKER3 / sericin discovery — wet-lab-guided manual annotation in Geneious; not a reproducible automated entrypoint. (manual)
- Caterpillar rearing, sequencing — wet lab. (out of scope)
Plan
Single «our HPC» job on «infra»: download the 3 figshare files, build a tiny env (python+biopython), compute C1–C5, write small JSON back to «host». Then (if green) a second BUSCO job for C6. Drops are valid; do not chase the phylo/synteny 20%.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All deterministic, shipped-data-derived claims reproduce 1:1 from the authors' deposited figshare files: 8/8 protein lengths, 8/8 molecular weights (SerB off by 0.01 kDa rounding), 8/8 exon counts, and aa-composition ranges that equal the exact min/max of the shipped sequences — strong positive evidence Table 1 is genuinely data-derived, no fabrication signal. The only deviation is in the pI values (serF 7.15→6.64; group-4 3.01–3.87→4.05–4.08), an expected pKa-scale/tool-choice difference on our side that preserves the paper's relative ordering. Severity is negligible and the central reproduced conclusion holds; overall yellow only because of the explainable pI offset and the large interpretive portions (phylogenetic network, edgeR expression, synteny, repeat counts) that were out of scope.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.