Foster thy young: enhanced prediction of orphan genes in assembled genomes.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the DOWNSTREAM analytical results. The full de-novo BIND/MIND/BRAKER/MAKER annotation (up to 595 GB RNA-Seq, weeks of compute) was OUT OF SCOPE; instead I verified that the paper's central reported numbers - the per-phylostratum sensitivity at recovering Araport11 genes (Figure 3 / SuppTable S11), the headline orphan-recovery percentages, and the Ribo-Seq translation-evidence percentages - are derivable from and consistent with the SHIPPED supplementary data. Result: MOSTLY 1:1 for the main figure/table (MAKER-Typical 21%, MAKER-Pool 53%, MAKER-Orphan 68%, BRAKER-Pool 41%, DirInf-Orphan 63%, MIND-Orphan 76%, Ribo 98%/56% all reproduce exactly/within-tol). HOWEVER I found genuine internal inconsistencies in the shipped FigureS2 upset-plot source files used for the audit: (1) the BIND column is byte-identical (965/1318 = 73.2%) in BOTH pool_orphan.txt and orphan_orphan.txt, matching neither the reported BIND-Pool 66% nor BIND-Orphan 76%; (2) the 'Maker' column in pool_orphan.txt corresponds to a different MAKER scenario (Make-Ara11+prots 77%) than its label. These anomalies are confined to the upset-plot inputs, not the main Figure 3/SuppTable S11. The '>85% combined BIND+MIND' claim reproduced at only 80-82% (likely depressed by the BIND anomaly). NOT attempted: independent re-run of the gene-prediction pipeline (gff generation) and mikado-compare gff->membership regeneration (feasible on «our HPC»; the VPN tunnel was down during this session). All verdicts are PROVISIONAL and must be independently checked by a human.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 84assessed: 2026-06-19 ⛓ 05c9e8f0ca86
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-19
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan ab initio and evidence-based gene-prediction pipelines accurately predict newly-emerged 'orphan' genes and other young lineage-specific genes, and can combining ab initio prediction with direct RNA-Seq inference improve detection of orphan genes across phylostrata?
- ★ Each of the five gene-prediction pipelines under-predicts orphan genes, with detection as low as 11% under one prediction scenario. finding
- ★ Combined pipelines BIND (BRAKER+Direct Inference) and MIND (MAKER+Direct Inference) yield the best overall predictions; BIND identifies 68% of annotated orphan genes and 99% of ancient genes in Arabidopsis. finding
- ★ Increasing RNA-Seq diversity (e.g., orphan-rich datasets) greatly improves gene prediction efficacy, especially for younger genes. finding
- ★ Direct Inference is an evidence-based pipeline that predicts gene structures from genome-guided alignment of RNA-Seq data, useful for young/orphan genes that homology- and ab initio-based methods miss. method
- ★ BIND and MIND pipelines use Mikado to integrate Direct Inference predictions with BRAKER or MAKER ab initio predictions. method
- Homology-based methods inherently cannot predict orphan genes because orphans are species-specific and lack orthologs. mechanism
- ★ A FAIR, automated, reproducible, well-documented Direct Inference/BIND/MIND solution implemented with pyrpipe and snakemake is provided. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Direct Inference gene prediction (genome-guided RNA-Seq alignment/assembly) | Arabidopsis thaliana Col0 (Araport11 genome) | none | predicted gene structures / orphan and ancient gene detection across phylostrata | pyrpipe; multiple assemblers (e.g. Trinity v2.6.6); STAR/HiSat2 |
| BRAKER ab initio gene prediction | Arabidopsis thaliana (Araport11) | none | predicted genes / sensitivity by phylostratum | BRAKER v2.1.2 (GeneMark-ET v4.33, AUGUSTUS v3.3.1); HiSat2 v2.1.0; SAMTools v1.9 |
| MAKER ab initio gene prediction | Arabidopsis thaliana (Araport11) | none | predicted genes / orphan and ancient gene detection | MAKER v2.31.10; Trinity v2.6.6; orfipy/TransDecoder v3.0.1 |
| BIND combined prediction (BRAKER + Direct Inference) | Arabidopsis thaliana (Araport11) | none | orphan/ancient gene detection, sensitivity | Mikado |
| MIND combined prediction (MAKER + Direct Inference) | Arabidopsis thaliana (Araport11) | none | orphan/ancient gene detection, sensitivity | Mikado |
| Gene prediction benchmarking (gold-standard) | Saccharomyces cerevisiae (yeast, genome R64-1-1) | none | predicted vs annotated genes across phylostrata | — |
| Gene prediction cross-validation | Oryza sativa (rice, GCA_009797565.1) | none | predicted vs NCBI-annotated genes | — |
| Phylostratigraphic classification of genes | Arabidopsis, yeast, rice proteomes | none | gene age / phylostratal origin (orphan vs ancient) | phylostratr (R); BLAST |
- ▼ Pipelines under-predict orphan genes, as few as 11% detected under one scenario 11%
- ▲ BIND identifies 68% of annotated orphan genes in Arabidopsis 68%
- ▲ BIND identifies 99% of ancient genes in Arabidopsis 99%
- ▲ BIND gives the highest sensitivity score regardless of dataset in Arabidopsis
- ▲ Increasing RNA-Seq diversity greatly improves prediction efficacy, particularly for younger genes
- ▲ Over 50% of novel transcriptional active regions in rice were identified by transcriptome profiling >50%
- – BRAKER predictions using RNA-Seq plus protein evidence were virtually identical to RNA-Seq evidence alone
- count 11% (lowest fraction of orphan genes predicted under one scenario)
- count 68% (annotated orphan genes identified by BIND in Arabidopsis)
- count 99% (ancient genes identified by BIND in Arabidopsis)
- count 38 RNA-Seq samples, each >60% of all annotated orphan transcripts (Orphan-rich dataset composition for A. thaliana and yeast)
- count >50% (novel transcriptional active regions in rice identified by transcriptome profiling (2010))
- count ~two thirds (proportion of annotated Arabidopsis genes that are very ancient)
- count 8.7 million eukaryotic species × 1000 orphans per eukaryote (estimate of extant protein-coding orphan genes across eukaryotes)
- count 12.8 / 241.4 / 595.1 GB (data sizes of Typical (12 SRR), Pooled (77 SRR), Orphan-rich (38 SRR) Arabidopsis datasets)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational benchmarking study comparing five gene prediction pipelines (BRAKER, MAKER, Direct Inference, BIND, MIND) on their ability to detect protein-coding genes stratified by phylostratigraphic age, with special focus on orphan (species-specific) genes. Gold-standard annotations for Arabidopsis and yeast served as reference benchmarks, with cross-validation in rice. Performance was evaluated using sensitivity scores and percentage of annotated genes detected across three RNA-Seq dataset sizes and compositions per species. Results were reported as descriptive percentages without inferential statistical tests or uncertainty estimates in the text provided.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Sensitivity (recall) and related gene-prediction performance metrics — descriptive benchmarking, not inferential hypothesis tests | Comparison of all five pipelines across phylostrata (orphan vs. ancient genes) in Arabidopsis, yeast, and rice | — | not stated |
-
Sensitivity and detection percentages were reported as single point estimates per pipeline-dataset combination, with no uncertainty quantification↳ Could also: Bootstrap resampling over RNA-Seq sample subsets, or Clopper-Pearson exact confidence intervals on proportions, could accompany each point estimate — Uncertainty bounds around performance metrics let readers assess whether observed differences between pipelines are consistent or sensitive to the particular RNA-Seq samples included in each dataset
-
Pipelines were compared by inspecting percentage differences in detected genes without formal inferential tests↳ Could also: McNemar's test or a paired proportion test could be applied when comparing two pipelines on the same annotated gene set, since each gene is a paired binary outcome (detected vs. not detected) for both pipelines — A formal test of paired proportions complements descriptive differences by quantifying whether the gap in detection rates exceeds what would be expected by chance given the number of genes evaluated
-
Fifteen Arabidopsis pipeline-dataset scenarios were evaluated, with differences described numerically without multiple-comparison adjustment↳ Could also: If inferential tests were applied across pipeline-dataset combinations, a Benjamini-Hochberg FDR or Bonferroni correction could control the family-wise error rate — Acknowledging the multiplicity of comparisons in a benchmarking study, and correcting for it when tests are applied, increases transparency about which performance differences are robust
-
RNA-Seq dataset size and diversity were treated as three qualitative categories (Typical, Pool, Orphan-rich)↳ Could also: A regression or dose-response model could treat dataset size or diversity as a continuous predictor of sensitivity — Modeling detection rate as a continuous function of dataset size would allow estimation of how much additional RNA-Seq data is needed to achieve a target improvement, potentially generalizing beyond the three tested levels
-
Cross-species validation in rice was assessed qualitatively, comparing pipeline rankings observed in Arabidopsis↳ Could also: A rank-correlation (e.g., Spearman's rho) across species, or a mixed-effects model with species as a random effect, could quantify cross-species consistency — Formally measuring how well pipeline rankings in the gold-standard organisms predict rankings in rice provides a more systematic estimate of generalizability across genomes
-
Performance was summarized with a single sensitivity metric; precision (positive predictive value) was mentioned but the full precision-recall trade-off was not summarized with a composite measure↳ Could also: The F1 score (harmonic mean of precision and recall) or the area under a precision-recall curve could provide a single balanced summary across the sensitivity-precision trade-off — In gene prediction benchmarking, sensitivity and precision are in tension; a composite metric like F1 facilitates direct comparison across pipelines without requiring readers to jointly weigh two separate statistics
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 34928390 "Foster thy young: enhanced prediction of orphan genes in assembled genomes"
Singh & Wurtele, Nucleic Acids Research 2022, e37. DOI 10.1093/nar/gkab1238.
Repo: github.com/eswlab/orphan-prediction @ 7fa99822f8cc540e02ae9f5648acfc506a83455d.
What the paper does
Introduces two annotation workflows — BIND (BRAKER + Direct-Inference, merged with Mikado) and MIND (MAKER + Direct-Inference) — and shows they recover more orphan genes (species-specific, youngest phylostratum) than BRAKER or MAKER alone. Benchmarked against Araport11 (Arabidopsis), SGD R64-1-1 (yeast), and rice (GCA_009797565.1). Key idea: selecting RNA-Seq samples enriched for orphan transcripts ("Orphan-rich" dataset) improves orphan recovery.
In scope (pipeline-derived, attempted)
The downstream evaluation results — i.e. whether the paper's reported numbers are derivable from / consistent with the shipped derived data:
| result | pipeline | how reproduced |
|---|---|---|
| Fig 3 / SuppTable S11-A — per-phylostratum % of Araport11 genes recovered by each method × dataset | mikado compare → phylostrata binning | recompute % from shipped per-gene membership matrices (FigureS2/*_orphan.txt, FigureS5/ara11_all.txt); cross-check vs figure3_ps.csv and SuppTable S11 |
| Headline orphan-recovery % (MAKER 21/53/68, BRAKER 41, DI 63, BIND 76, MIND 76) | same | same |
| ">85% combined BIND+MIND orphans" (Suppl Fig S2) | membership union | compute BIND∪MIND over Araport11 orphans |
| Ribo-Seq translation evidence (98% match, 56% novel, 97% BIND) | ribotricer/ribo pipeline | read shipped SuppTable S6-A |
These need only small shipped text/xlsx files; no heavy compute. Done on «host» in-memory (nothing large persisted); the deeper gff→membership step is for «our HPC».
Out of scope (not attempted — too heavy / honest blocker)
- Full de-novo annotation (BIND/MIND/BRAKER/MAKER from raw RNA-Seq). Inputs are up to 595 GB RNA-Seq (Orphan-rich 38 samples) + 241 GB (Pooled 77) and the pipeline chains STAR/HiSat2 → Trinity → TransDecoder → BRAKER/MAKER → Class2/ StringTie/Cufflinks → PortCullis → Mikado. Weeks of compute, TB storage.
- phylostratr from scratch (diamond all-vs-all vs ~100 proteomes) — the strata assignment is shipped in the membership matrices, so we reuse it.
- Yeast and rice benchmarks (same pipeline class; Arabidopsis is the headline).
Partially attempted / deferred
mikado compareof the shippedprediction_gff3/arabidopsis_gff3/*_orph.gff3vs the Araport11 reference, to regenerate the membership matrices from the GFFs and adjudicate the BIND/MAKER membership-file anomalies and the novel-gene counts (14,739 / 18,114). Feasible on «our HPC» (5 mikado-compare jobs on ~80 MB GFFs, minutes) — deferred because the «our HPC» VPN tunnel was down this session.
Key references in repo
prediction_gff3/arabidopsis_gff3/{BIND,BRAKER,DI,MAKER,MIND}_orph.gff3(final models, orphan-rich dataset)SuppTables/Supplementary_Table_11-percent_match.xlsx(S11 = source of Fig 3)SuppTables/Supplementary_Table_6-RiboSeq.xlsx(S6 = Ribo-Seq)plots_publication/Figure3/{fig3.R,figure3_ps.csv},FigureS2/*_orphan.txt,FigureS5/ara11_all.txt
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.