Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Foster thy young: enhanced prediction of orphan genes in assembled genomes.

Nucleic Acids Res · 2022
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the DOWNSTREAM analytical results. The full de-novo BIND/MIND/BRAKER/MAKER annotation (up to 595 GB RNA-Seq, weeks of compute) was OUT OF SCOPE; instead I verified that the paper's central reported numbers - the per-phylostratum sensitivity at recovering Araport11 genes (Figure 3 / SuppTable S11), the headline orphan-recovery percentages, and the Ribo-Seq translation-evidence percentages - are derivable from and consistent with the SHIPPED supplementary data. Result: MOSTLY 1:1 for the main figure/table (MAKER-Typical 21%, MAKER-Pool 53%, MAKER-Orphan 68%, BRAKER-Pool 41%, DirInf-Orphan 63%, MIND-Orphan 76%, Ribo 98%/56% all reproduce exactly/within-tol). HOWEVER I found genuine internal inconsistencies in the shipped FigureS2 upset-plot source files used for the audit: (1) the BIND column is byte-identical (965/1318 = 73.2%) in BOTH pool_orphan.txt and orphan_orphan.txt, matching neither the reported BIND-Pool 66% nor BIND-Orphan 76%; (2) the 'Maker' column in pool_orphan.txt corresponds to a different MAKER scenario (Make-Ara11+prots 77%) than its label. These anomalies are confined to the upset-plot inputs, not the main Figure 3/SuppTable S11. The '>85% combined BIND+MIND' claim reproduced at only 80-82% (likely depressed by the BIND anomaly). NOT attempted: independent re-run of the gene-prediction pipeline (gff generation) and mikado-compare gff->membership regeneration (feasible on «our HPC»; the VPN tunnel was down during this session). All verdicts are PROVISIONAL and must be independently checked by a human.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.4054262

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-19 ⛓ 05c9e8f0ca86
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can ab initio and evidence-based gene-prediction pipelines accurately predict newly-emerged 'orphan' genes and other young lineage-specific genes, and can combining ab initio prediction with direct RNA-Seq inference improve detection of orphan genes across phylostrata?

Core claims
  • Each of the five gene-prediction pipelines under-predicts orphan genes, with detection as low as 11% under one prediction scenario. finding
  • Combined pipelines BIND (BRAKER+Direct Inference) and MIND (MAKER+Direct Inference) yield the best overall predictions; BIND identifies 68% of annotated orphan genes and 99% of ancient genes in Arabidopsis. finding
  • Increasing RNA-Seq diversity (e.g., orphan-rich datasets) greatly improves gene prediction efficacy, especially for younger genes. finding
  • Direct Inference is an evidence-based pipeline that predicts gene structures from genome-guided alignment of RNA-Seq data, useful for young/orphan genes that homology- and ab initio-based methods miss. method
  • BIND and MIND pipelines use Mikado to integrate Direct Inference predictions with BRAKER or MAKER ab initio predictions. method
  • Homology-based methods inherently cannot predict orphan genes because orphans are species-specific and lack orthologs. mechanism
  • A FAIR, automated, reproducible, well-documented Direct Inference/BIND/MIND solution implemented with pyrpipe and snakemake is provided. resource
Experimental setups
Assay System Perturbation Readout Platform
Direct Inference gene prediction (genome-guided RNA-Seq alignment/assembly) Arabidopsis thaliana Col0 (Araport11 genome) none predicted gene structures / orphan and ancient gene detection across phylostrata pyrpipe; multiple assemblers (e.g. Trinity v2.6.6); STAR/HiSat2
BRAKER ab initio gene prediction Arabidopsis thaliana (Araport11) none predicted genes / sensitivity by phylostratum BRAKER v2.1.2 (GeneMark-ET v4.33, AUGUSTUS v3.3.1); HiSat2 v2.1.0; SAMTools v1.9
MAKER ab initio gene prediction Arabidopsis thaliana (Araport11) none predicted genes / orphan and ancient gene detection MAKER v2.31.10; Trinity v2.6.6; orfipy/TransDecoder v3.0.1
BIND combined prediction (BRAKER + Direct Inference) Arabidopsis thaliana (Araport11) none orphan/ancient gene detection, sensitivity Mikado
MIND combined prediction (MAKER + Direct Inference) Arabidopsis thaliana (Araport11) none orphan/ancient gene detection, sensitivity Mikado
Gene prediction benchmarking (gold-standard) Saccharomyces cerevisiae (yeast, genome R64-1-1) none predicted vs annotated genes across phylostrata
Gene prediction cross-validation Oryza sativa (rice, GCA_009797565.1) none predicted vs NCBI-annotated genes
Phylostratigraphic classification of genes Arabidopsis, yeast, rice proteomes none gene age / phylostratal origin (orphan vs ancient) phylostratr (R); BLAST
Key results
  • Pipelines under-predict orphan genes, as few as 11% detected under one scenario 11%
  • BIND identifies 68% of annotated orphan genes in Arabidopsis 68%
  • BIND identifies 99% of ancient genes in Arabidopsis 99%
  • BIND gives the highest sensitivity score regardless of dataset in Arabidopsis
  • Increasing RNA-Seq diversity greatly improves prediction efficacy, particularly for younger genes
  • Over 50% of novel transcriptional active regions in rice were identified by transcriptome profiling >50%
  • BRAKER predictions using RNA-Seq plus protein evidence were virtually identical to RNA-Seq evidence alone
Key statistics
  • count 11% (lowest fraction of orphan genes predicted under one scenario)
  • count 68% (annotated orphan genes identified by BIND in Arabidopsis)
  • count 99% (ancient genes identified by BIND in Arabidopsis)
  • count 38 RNA-Seq samples, each >60% of all annotated orphan transcripts (Orphan-rich dataset composition for A. thaliana and yeast)
  • count >50% (novel transcriptional active regions in rice identified by transcriptome profiling (2010))
  • count ~two thirds (proportion of annotated Arabidopsis genes that are very ancient)
  • count 8.7 million eukaryotic species × 1000 orphans per eukaryote (estimate of extant protein-coding orphan genes across eukaryotes)
  • count 12.8 / 241.4 / 595.1 GB (data sizes of Typical (12 SRR), Pooled (77 SRR), Orphan-rich (38 SRR) Arabidopsis datasets)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational benchmarking study comparing five gene prediction pipelines (BRAKER, MAKER, Direct Inference, BIND, MIND) on their ability to detect protein-coding genes stratified by phylostratigraphic age, with special focus on orphan (species-specific) genes. Gold-standard annotations for Arabidopsis and yeast served as reference benchmarks, with cross-validation in rice. Performance was evaluated using sensitivity scores and percentage of annotated genes detected across three RNA-Seq dataset sizes and compositions per species. Results were reported as descriptive percentages without inferential statistical tests or uncertainty estimates in the text provided.

Replicationunclear Sample sizeThree RNA-Seq dataset sizes for Arabidopsis: Typical (12 SRR samples, 12.8 GB), Pool (77 samples, 241.4 GB), Orphan-rich (38 samples, 595.1 GB); analogous sets constructed for yeast and rice GroupsFive prediction pipelines × three RNA-Seq dataset types (Typical, Pool, Orphan-rich) × three species, stratified by phylostratum (orphan vs. ancient genes) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Sensitivity (recall) and related gene-prediction performance metrics — descriptive benchmarking, not inferential hypothesis tests Comparison of all five pipelines across phylostrata (orphan vs. ancient genes) in Arabidopsis, yeast, and rice not stated
Approaches that could also have been used
  • Sensitivity and detection percentages were reported as single point estimates per pipeline-dataset combination, with no uncertainty quantification
    Could also: Bootstrap resampling over RNA-Seq sample subsets, or Clopper-Pearson exact confidence intervals on proportions, could accompany each point estimate — Uncertainty bounds around performance metrics let readers assess whether observed differences between pipelines are consistent or sensitive to the particular RNA-Seq samples included in each dataset
  • Pipelines were compared by inspecting percentage differences in detected genes without formal inferential tests
    Could also: McNemar's test or a paired proportion test could be applied when comparing two pipelines on the same annotated gene set, since each gene is a paired binary outcome (detected vs. not detected) for both pipelines — A formal test of paired proportions complements descriptive differences by quantifying whether the gap in detection rates exceeds what would be expected by chance given the number of genes evaluated
  • Fifteen Arabidopsis pipeline-dataset scenarios were evaluated, with differences described numerically without multiple-comparison adjustment
    Could also: If inferential tests were applied across pipeline-dataset combinations, a Benjamini-Hochberg FDR or Bonferroni correction could control the family-wise error rate — Acknowledging the multiplicity of comparisons in a benchmarking study, and correcting for it when tests are applied, increases transparency about which performance differences are robust
  • RNA-Seq dataset size and diversity were treated as three qualitative categories (Typical, Pool, Orphan-rich)
    Could also: A regression or dose-response model could treat dataset size or diversity as a continuous predictor of sensitivity — Modeling detection rate as a continuous function of dataset size would allow estimation of how much additional RNA-Seq data is needed to achieve a target improvement, potentially generalizing beyond the three tested levels
  • Cross-species validation in rice was assessed qualitatively, comparing pipeline rankings observed in Arabidopsis
    Could also: A rank-correlation (e.g., Spearman's rho) across species, or a mixed-effects model with species as a random effect, could quantify cross-species consistency — Formally measuring how well pipeline rankings in the gold-standard organisms predict rankings in rice provides a more systematic estimate of generalizability across genomes
  • Performance was summarized with a single sensitivity metric; precision (positive predictive value) was mentioned but the full precision-recall trade-off was not summarized with a composite measure
    Could also: The F1 score (harmonic mean of precision and recall) or the area under a precision-recall curve could provide a single balanced summary across the sensitivity-precision trade-off — In gene prediction benchmarking, sensitivity and precision are in tension; a composite metric like F1 facilitates direct comparison across pipelines without requiring readers to jointly weigh two separate statistics
Software: BRAKER v2.1.2 · MAKER v2.31.10 · Trinity v2.6.6 · HiSat2 v2.1.0 · SAMTools v1.9 · GeneMark-ET v4.33 · AUGUSTUS v3.3.1 · TransDecoder v3.0.1 · orfipy · Mikado · pyrpipe · snakemake · R/phylostratr · SRA-toolkit v2.8.0

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 34928390 "Foster thy young: enhanced prediction of orphan genes in assembled genomes"

Singh & Wurtele, Nucleic Acids Research 2022, e37. DOI 10.1093/nar/gkab1238. Repo: github.com/eswlab/orphan-prediction @ 7fa99822f8cc540e02ae9f5648acfc506a83455d.

What the paper does

Introduces two annotation workflows — BIND (BRAKER + Direct-Inference, merged with Mikado) and MIND (MAKER + Direct-Inference) — and shows they recover more orphan genes (species-specific, youngest phylostratum) than BRAKER or MAKER alone. Benchmarked against Araport11 (Arabidopsis), SGD R64-1-1 (yeast), and rice (GCA_009797565.1). Key idea: selecting RNA-Seq samples enriched for orphan transcripts ("Orphan-rich" dataset) improves orphan recovery.

In scope (pipeline-derived, attempted)

The downstream evaluation results — i.e. whether the paper's reported numbers are derivable from / consistent with the shipped derived data:

result pipeline how reproduced
Fig 3 / SuppTable S11-A — per-phylostratum % of Araport11 genes recovered by each method × dataset mikado compare → phylostrata binning recompute % from shipped per-gene membership matrices (FigureS2/*_orphan.txt, FigureS5/ara11_all.txt); cross-check vs figure3_ps.csv and SuppTable S11
Headline orphan-recovery % (MAKER 21/53/68, BRAKER 41, DI 63, BIND 76, MIND 76) same same
">85% combined BIND+MIND orphans" (Suppl Fig S2) membership union compute BIND∪MIND over Araport11 orphans
Ribo-Seq translation evidence (98% match, 56% novel, 97% BIND) ribotricer/ribo pipeline read shipped SuppTable S6-A

These need only small shipped text/xlsx files; no heavy compute. Done on «host» in-memory (nothing large persisted); the deeper gff→membership step is for «our HPC».

Out of scope (not attempted — too heavy / honest blocker)

  • Full de-novo annotation (BIND/MIND/BRAKER/MAKER from raw RNA-Seq). Inputs are up to 595 GB RNA-Seq (Orphan-rich 38 samples) + 241 GB (Pooled 77) and the pipeline chains STAR/HiSat2 → Trinity → TransDecoder → BRAKER/MAKER → Class2/ StringTie/Cufflinks → PortCullis → Mikado. Weeks of compute, TB storage.
  • phylostratr from scratch (diamond all-vs-all vs ~100 proteomes) — the strata assignment is shipped in the membership matrices, so we reuse it.
  • Yeast and rice benchmarks (same pipeline class; Arabidopsis is the headline).

Partially attempted / deferred

  • mikado compare of the shipped prediction_gff3/arabidopsis_gff3/*_orph.gff3 vs the Araport11 reference, to regenerate the membership matrices from the GFFs and adjudicate the BIND/MAKER membership-file anomalies and the novel-gene counts (14,739 / 18,114). Feasible on «our HPC» (5 mikado-compare jobs on ~80 MB GFFs, minutes) — deferred because the «our HPC» VPN tunnel was down this session.

Key references in repo

  • prediction_gff3/arabidopsis_gff3/{BIND,BRAKER,DI,MAKER,MIND}_orph.gff3 (final models, orphan-rich dataset)
  • SuppTables/Supplementary_Table_11-percent_match.xlsx (S11 = source of Fig 3)
  • SuppTables/Supplementary_Table_6-RiboSeq.xlsx (S6 = Ribo-Seq)
  • plots_publication/Figure3/{fig3.R,figure3_ps.csv}, FigureS2/*_orphan.txt, FigureS5/ara11_all.txt
Figures / tables: Fig 3TableFig S2Fig 4Fig 6
maker_typical_orphan
Reported
MAKER predicted 21% of annotated Arabidopsis orphan genes (Typical dataset)
Reproduced
20.7% from shipped membership matrix; SuppTable S11-A Make-Typical = 0.207
exact
maker_pool_orphan
Reported
MAKER predicted 53% of annotated orphans (Pooled dataset)
Reproduced
SuppTable S11-A Make-Pool = 0.529; figure3_ps.csv MAKER_Pool = 53 (membership file pool_orphan.txt 'Maker' col is a DIFFERENT scenario, 76.9% = Make-Ara11+prots)
exact
maker_orphan_orphan
Reported
MAKER predicted 68% of annotated orphans (Orphan-rich dataset)
Reproduced
68.4% from membership matrix; SuppTable S11-A Make-Orphan = 0.684
exact
braker_pool_orphan
Reported
BRAKER predicted only 41% of orphan genes (Pooled dataset)
Reproduced
41.2% from membership matrix; SuppTable S11-A Brake-Pool = 0.412
exact
dirinf_orphan_orphan
Reported
Direct Inference predicted 63% of annotated orphans (Orphan-rich dataset)
Reproduced
62.6% from membership matrix; SuppTable S11-A DirInf-Orphan = 0.626
exact
bind_orphan_orphan
Reported
BIND predicted 76% of annotated orphans (Orphan-rich dataset; Fig 3 / S11)
Reproduced
SuppTable S11-A BIND-Orphan = 0.764 (=figure3_ps.csv 76); BUT independent membership matrix gives 73.2% (BIND column identical 965/1318 in BOTH pool and orphan files - shipped data inconsistency, flagged)
partial
mind_orphan_orphan
Reported
MIND predicted ~76% of annotated orphans (Orphan-rich dataset)
Reproduced
75.2% from membership matrix; SuppTable S11-A MIND-Orphan = 0.762
within tolerance
bind_mind_combined_orphan
Reported
Over 85% of annotated orphan genes predicted by combining BIND and MIND (Orphan-rich or Pooled)
Reproduced
BIND_or_MIND union from membership = 80.2% (Orphan-rich) / 81.7% (Pooled) - below 85%; likely affected by the anomalous duplicated BIND column
partial
novel_genes_bind_mind
Reported
MIND predicted 18,114 genes / BIND 14,739 genes not matching any TAIR-annotated gene
Reproduced
NOT cleanly reproducible from shipped membership matrices (ara11_all.txt novel rows: BIND 7,793 / MIND 10,681 at transcript level); requires re-running 'mikado compare' on prediction_gff3 vs Araport11 (deferred to «our HPC»)
partial
riboseq_match_translation
Reported
98% of predictions matching annotated genes had translation evidence; ~56% of novel genes
Reproduced
SuppTable S6-A 'All prediction' columns: Match-to-Araport11 0.987 (=98%); Novel-in-BIND 0.58 (=~56%)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

208.4 k
tokens (I/O) · 9.2 M incl. cache
19 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.