Revised annotations, sex-biased expression, and lineage-specific genes in the Drosophila melanogaster group.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced this paper's pipeline-derived results via two complementary routes: (1) an internal-consistency audit recounting tested/significant differentially-expressed-gene counts directly from the authors' own deposited Cuffdiff output for D. ananassae, D. yakuba, and D. simulans (Table 6) plus gene counts (Table 2) and ortholog counts (Table 4); and (2) a genuinely independent, from-raw-SRA-reads re-execution of the D. melanogaster HISAT2+Cuffdiff differential expression pipeline (substituting HISAT2 for the paper's now-unobtainable 2012-2013 Tophat2/Bowtie2, and a modern Ensembl BDGP6.22.98 annotation for the paper's ~2013 FlyBase build). Table 6 audit: 10/12 species-comparisons match exactly, 2/12 match exactly on significant-gene counts but diverge on the tested/background count (flagged for human review). Table 2: D. ananassae and D. yakuba gene totals reconstruct to within 1% of the reported value using the paper's stated 'Augustus genes + added FlyBase orphans' model; D. simulans does not fit this model (25.5% off), a genuine unresolved discrepancy. Table 4: naive rank-1-BLAST-hit ortholog counts systematically overcount vs. the paper's true fuzzy-reciprocal-best-hit method by 35-95% — an expected, documented methodological mismatch, not a bug (full fuzzy-RBH reimplementation was out of scope for this pass). The independent D. melanogaster re-run reproduces the paper's background/tested gene universe within ~4-9% across all 4 tissue comparisons, but reproduces the significant-DE-gene counts well for only 1/4 comparisons (ovary vs. carcass); the other 3 diverge substantially (66-651 observed vs. 220-370 reported), most plausibly due to aligner and reference-annotation-version differences given the single-replicate-per-tissue design. GO-enrichment (DAVID) results were scoped out of reproduction entirely: DAVID is a third-party hosted service with no archived 2014 database snapshot, making a present-day rerun methodologically uninterpretable. No completeness claim is made; several items (Table 4 fuzzy-RBH reimplementation, de novo Trinity/Augustus annotation for D. ananassae/yakuba/simulans) were deliberately not attempted this pass and remain open for a future, deeper pass.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-08-03
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-03no human curator yet
- Last updated
- 2026-08-03
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan RNA-seq-based reannotation of Drosophila outgroup species (D. simulans, D. yakuba, D. ananassae) produce improved gene models—with UTRs and empirically verified intron-exon boundaries—that reveal lineage-specific genes and sex-/tissue-biased expression missed by homology-based comparative annotation? The paper tests this by generating whole-transcriptome annotations and performing differential expression testing across reproductive tissues and carcass.
- ★ Revised RNA-seq-based gene models for D. ananassae, D. yakuba, and D. simulans include UTRs, empirically verified intron-exon boundaries, and previously unannotated novel exons, improving on r1.3 comparative-genomics annotations that lack UTRs. resource
- ★ Thousands of genes are differentially expressed across tissues in D. yakuba and D. simulans, with roughly 60% agreement in tissue-biased expression patterns of reciprocal-best-hit orthologs between the two species. finding
- ★ Hundreds of lineage-specific genes on major chromosomal arms exist in each species with no BLASTp hit among transcripts of other Drosophila species, representing candidates for neofunctionalized proteins and a source of genetic novelty. finding
- ★ Putative polycistronic transcripts are widespread, reflecting a combination of transcriptional read-through and putative gene fusion/fission events across taxa; whether such transcription is functional or a stochastic byproduct of transcriptional errors is unclear. finding
- ★ Ortholog groups across D. melanogaster, D. simulans, D. yakuba, and D. ananassae were identified using fuzzy reciprocal-best-hit BLASTp, where genes with E-values within a single log-unit are assigned the same rank (cutoff E ≤ 1e-10). method
- ★ Differential expression of lineage-specific genes in at least one of four germline/somatic tissue comparisons indicates they are not solely artifacts of gene annotation software. finding
- ★ Annotation pipeline combines Trinity de novo transcriptome assembly (after digital normalization) with Augustus v2.5.5 gene prediction using Trinity assemblies as hints, then merges in orphaned FlyBase/w501 gene models with FPKM ≥ 2 lacking an Augustus match. method
- The larger number of genes in D. ananassae reflects FlyBase models failing to match RNA-seq-supported models, highlighting the difficulty of annotation via comparative genomics across large phylogenetic distances. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-transcriptome RNA-seq (76-bp reads, paired-end with dual indexing; one female-tissue replicate single-end; 4-plex barcoded; 32 samples across the equivalent of one eight-lane flow cell) | Dissected adult tissues (female ovaries, female headless gonadectomized carcass from 2 flies; testes plus accessory glands, male headless gonadectomized carcass from 4 flies) of D. yakuba (14021-0261.01), D. simulans w501, D. ananassae (14024-0371.13), D. melanogaster (14021-0231.36); virgin flies collected within 2 hr of eclosion, aged 2-5 d, reared at 25° and ~40% humidity | none (tissue and sex comparison; 3 biological replicates for D. yakuba and D. simulans, 1 replicate each for D. ananassae and D. melanogaster) | Transcript abundance / read coverage used for de novo transcript assembly and gene-model annotation | Illumina HiSeq 2500; Trizol (Invitrogen) + Direct-Zol RNA MiniPrep (Zymo), Turbo DNase (Invitrogen), RNA Clean and Concentrator-5 (Zymo), Ovation RNA-Seq System V2 (Nugen, SPIA), Nextera DNA Sample Preparation Kit (Illumina), Kapa Library Quant Kits, Bioanalyzer 2100 (Agilent), Qubit RNA/DNA HS Assay Kit (Invitrogen) |
| De novo transcriptome assembly and ab initio gene prediction (Trinity with digital normalization --max_cov 30; Augustus v2.5.5 --species=fly with hints file and extrinsic.ME.cfg) | Concatenated ovary, female carcass, testes, and male carcass reads of D. yakuba, D. simulans, and D. ananassae | none (computational) | Assembled transcripts and predicted gene models (transcripts, genes, UTRs, intron-exon boundaries, polycistronic transcripts) | Trinity (trinityrnaseq); Augustus v2.5.5 |
| Read mapping and expression quantification / differential expression testing (Tophat v2.0.6 + Bowtie2 v2.0.2 with -G reference-guided mapping; Cufflinks 2.0.2 / Cuffdiff with upper-quantile normalization -N) | D. yakuba, D. simulans, D. ananassae revised GFF annotations plus unmodified D. melanogaster r5.45 models; four tissue comparisons (female ovary vs female carcass, female ovary vs male testes, male carcass vs male testes, male carcass vs female carcass) | none (tissue/sex contrast) | FPKM expression levels and counts of significantly differentially expressed genes at FDR ≤ 0.1 | Tophat v.2.0.6, Bowtie2 v.2.0.2, Cufflinks 2.0.2 suite (gffread) |
| All-by-all BLASTp of translated sequences (fuzzy reciprocal best hit; E ≤ 1e-10, low-complexity filter off -F F; CDS overlap requirement of ≥85% CDS match with ≥90% amino acid similarity for matching Augustus models to prior annotations) | Reference proteome translations of D. melanogaster, D. simulans, D. yakuba, D. ananassae gene model predictions; FlyBase r1.3 (D. yakuba, D. ananassae), D. melanogaster r5.45, D. simulans w501 models of Hu et al. (2012) | none (computational) | First-order ortholog counts and counts of lineage-specific genes with no cross-species hit | BLASTp |
| Gene ontology enrichment analysis (clustering threshold set to Low) | D. melanogaster orthologs (FlyBase functional classes) of genes with tissue-specific expression from cufflinks comparisons (female carcass vs female ovaries; male carcass vs male testes) at genomewide FDR ≤ 0.10 | none | Overrepresented functional categories among sex-specific/tissue-specific genes | DAVID gene ontology analysis software |
| Coverage visualization of quantile-normalized RNA-seq data at individual loci | D. yakuba reference strain, Adh locus; carcass and ovary tissues (three replicates) | none | Per-base normalized coverage distinguishing introns, exons, and 5′/3′ UTRs; comparison of FlyBase vs revised gene models | — |
- ▲ Revised annotations total 22,989 transcripts for 16,278 genes in D. simulans, 20,315 transcripts for 17,579 genes in D. yakuba, and 22,420 transcripts for 20,580 genes in D. ananassae (vs FlyBase r1.3: 15,413, 16,077, and 15,069 genes respectively) 16,278 / 17,579 / 20,580 genes (revised) vs 15,413 / 16,077 / 15,069 (FlyBase)
- – Putative polycistronic transcripts identified: 2529 in D. yakuba, 2379 in D. ananassae, 561 in D. simulans 2529 / 2379 / 561
- – Roughly 60% of genes with tissue-biased expression in D. yakuba that have a reciprocal best-hit ortholog in D. simulans show the same tissue-specific bias in D. simulans, with marginally greater agreement for carcass-biased than reproductive-tissue-biased genes in both sexes ~60% agreement
- – Thousands of genes are differentially expressed by tissue in D. yakuba and D. simulans (e.g., 5420 of 8689 tested for female ovary vs female carcass in D. yakuba; 5053 of 8967 in D. simulans), but only hundreds in D. melanogaster and D. ananassae, attributed to increased power from biological replicates D. yakuba 5420/8689; D. simulans 5053/8967; D. melanogaster 112/11,326; D. ananassae 203/7537
- – Lineage-specific genes: 1340 total in D. yakuba (230 on major chromosomes), 1314 in D. simulans (369 on major chromosomes), 2977 in D. ananassae 230 / 369 on major chromosomal arms
- – Lineage-specific genes showing differential expression in at least one of four tissue comparisons: 334 in D. yakuba, 222 in D. simulans, 118 in D. ananassae 334 / 222 / 118
- ▲ First-order orthologs with D. melanogaster identified for 12,127 genes in D. simulans, 11,425 in D. yakuba, and 11,348 in D. ananassae; the increase for D. simulans is attributed to improved annotation and the improved w501 assembly 12,127 / 11,425 / 11,348 genes
- – GO categories overrepresented between ovary and female carcass involve reproduction, chromosome segregation, and DNA synthesis/repair, whereas testes vs carcass genes involve sperm development, cell division, and energy production
- count 22,989 transcripts / 16,278 genes (D. simulans); 20,315 transcripts / 17,579 genes (D. yakuba); 22,420 transcripts / 20,580 genes (D. ananassae) (Total revised gene models (Table 2))
- other 72.3% (D. ananassae), 79.4% (D. yakuba), 80.05% (D. simulans) (Percent of revised gene models with ≥60% of features (exons, UTRs) supported by RNA-seq data (Table 3))
- count 10,369 (D. yakuba genes with reciprocal best-hit orthologs in D. simulans retained for cross-species differential expression comparison)
- other ~60% (Agreement in tissue-biased expression of orthologs between D. yakuba and D. simulans)
- count 5420 significant of 8689 tested (D. yakuba female ovary vs female carcass); 5868/10,202 (ovary vs testes); 3065/10,412 (male carcass vs testes); 724/9430 (male vs female carcass) (Differentially expressed genes at FDR ≤ 0.1, D. yakuba (Table 6))
- count 5053/8967, 5741/10,222, 4628/10,679, 611/9566 (Differentially expressed genes at FDR ≤ 0.1, D. simulans, four tissue comparisons (Table 6))
- other 48% (D. yakuba), 46% (D. ananassae) (1:1 concordance of revised models with FlyBase gene models)
- other E ≤ 10^-10; ≥85% CDS match with ≥90% amino acid similarity; FDR ≤ 0.1; FPKM ≥ 2 (Thresholds used for BLASTp ortholog/lineage-specific calls, model matching, differential expression, and inclusion of orphaned FlyBase models)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper uses RNA-seq based gene annotation and differential expression analysis in four Drosophila species. Differential expression between tissue pairs (e.g., ovary vs. carcass, testes vs. carcass) within each species was tested using the Cufflinks/Cuffdiff suite, with significance reported as gene counts/percentages passing a genome-wide false discovery rate (FDR) threshold of ≤0.1, rather than as individual p-values or effect sizes. Orthology and lineage-specific gene calls were made via reciprocal best-hit BLASTp with a fixed E-value cutoff, and functional enrichment was assessed with DAVID gene ontology analysis. Results are presented primarily as tabulated counts of genes/transcripts and of differentially expressed or lineage-specific genes across species and tissue comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Cufflinks/Cuffdiff differential expression testing (FPKM-based statistical test with FDR thresholding) | Pairwise tissue comparisons (female carcass vs. female ovary; male carcass vs. male testes; female carcass vs. male carcass; female ovary vs. male testes) within each species (Table 6) | Three biological replicates for D. yakuba and D. simulans references; one replicate each for D. ananassae and D. melanogaster references, with replicates grouped for reference genomes | not stated |
| Fuzzy reciprocal-best-hit BLASTp (E ≤ 10^-10 cutoff) for ortholog assignment and lineage-specific gene classification | Ortholog identification across D. melanogaster, D. simulans, D. yakuba, D. ananassae (Table 4); lineage-specific gene calls (Table 5) | — | not stated |
| DAVID gene ontology overrepresentation analysis (clustering threshold set to Low) | Functional category enrichment among genes with sex-specific or tissue-specific expression (File S1) | — | not stated |
-
Differential expression between tissues was tested using the Cufflinks/Cuffdiff suite with a genome-wide FDR ≤0.1 threshold, reported as counts of significant genes rather than per-gene p-values or fold changes.↳ Could also: A count-based negative-binomial framework such as DESeq2 or edgeR — These tools model gene-wise dispersion directly from raw read counts and are widely used alternatives to the FPKM/Cuffdiff approach; they would also allow explicit reporting of log2 fold changes and per-gene p-values/FDR-adjusted q-values alongside significance calls.
-
Some species-level comparisons were based on a single biological replicate per sex/tissue (D. ananassae and D. melanogaster), while D. yakuba and D. simulans had three replicates.↳ Could also: Increasing biological replication, or explicitly modeling replicate number in the dispersion estimate — Additional replicates would also allow calculation of within-group variability (e.g., SD/SEM or confidence intervals on expression estimates) and give more precise estimates of differential expression significance than can be obtained from a single replicate.
-
Multiple testing across thousands of genes per comparison was addressed with a genome-wide FDR ≤0.1 cutoff from Cuffdiff's default settings, without naming the specific correction algorithm.↳ Could also: An explicitly named and justified correction procedure such as Benjamini-Hochberg (with the target FDR reported) or a more conservative family-wise error control like Bonferroni — Naming the exact multiple-testing procedure and reporting the resulting adjusted q-values for each gene would also let readers directly evaluate the strength of evidence for individual candidate genes.
-
Orthology and lineage-specificity were inferred from BLASTp sequence similarity alone, using a fixed E-value cutoff (E ≤ 10^-10).↳ Could also: Synteny-aware orthology tools (e.g., OrthoFinder, or reciprocal-best-hit combined with genomic collinearity information) — Incorporating synteny/gene order alongside sequence similarity would also help distinguish true lineage-specific gene gain from cases of rapid sequence divergence or assembly fragmentation that a similarity-only cutoff cannot separate.
-
Differentially expressed gene counts were summarized as raw counts/percentages per comparison (Table 6) without confidence intervals or effect-size estimates in the reported text.↳ Could also: Reporting log2 fold-change distributions with confidence intervals, as produced natively by DESeq2/edgeR/Cuffdiff output tables — Effect sizes and confidence intervals would also convey the magnitude and precision of expression differences, complementing the binary significant/not-significant FDR calls.
-
Gene ontology enrichment was performed with DAVID using a clustering threshold set to Low, without specifying the underlying statistical test or multiple-testing correction for enrichment.↳ Could also: Enrichment methods with an explicit statistical model and correction, such as permutation-based GSEA or topGO with Fisher's exact test plus Benjamini-Hochberg correction — These approaches would also make the significance criteria for 'overrepresented' categories explicit and adjustable, and allow direct comparison of enrichment strength across functional categories.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Two different reproduction modes were run. The internal-consistency audit of the authors' deposited *.expdiff files reproduces 11 of 13 Table 6 rows exactly (e.g. 203/7537, 200/8282, 3065/10412), which is strong evidence the reported values are derivable; the two exceptions are tested/background counts only (Dana male tissues 8349 vs 9349, Dyak female tissues 8689 vs 8969) while their significant counts (1013, 5420) match perfectly — an authors-side inconsistency between the printed table and the deposited snapshot, not a fabrication signal. The independent from-raw-reads arm for D. melanogaster deviates more (112->106, and 370->651 significant genes; backgrounds within ~4-5%), but this is squarely attributable to our own forced substitution of HISAT2 and Ensembl BDGP6.22.98 for the paper's 2013 Tophat2/Bowtie2 + FlyBase stack. Severity is moderate: magnitudes and directions hold and the central claim of pervasive, gonad-dominated sex-biased expression is confirmed, while exact counts are not — hence yellow across the board rather than red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.