Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

auts2 Features and Expression Are Highly Conserved during Evolution Despite Different Evolutionary Fates Following Whole Genome Duplication.

Cells · 2022
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

FINAL. The paper's only computational pipeline (TrimGalore trim -> Salmon quant -> per-gene TPM, feeding Fig 2B) was reproduced on the full medaka tissue set. Pipeline: TrimGalore 0.6.6 (--paired -q20 --clip_R1/R2 13 --three_prime_clip_R1/R2 2, per Methods 2.4) + Salmon 1.10.3 (paper said 1.8.0; its pufferfish segfaulted on the node) vs Ensembl Oryzias latipes ASM223467v1 cdna. ALL 11 medaka runs SRR1524271-281 quantified («job», exit 0): auts2a TPM ranges 0.52-56.3 (>100-fold), with 4 high-expressing and 7 low/near-absent runs -- a strongly tissue-restricted pattern that qualitatively matches the paper's claim of auts2a predominant expression in a subset of tissues (embryo/brain/gonads). auts2b is absent from the current Ensembl annotation (0 in every run), consistent with the WGD 'different evolutionary fates' theme. Honest limits: the paper prints no per-tissue numeric TPM (heatmap) and SRA/ENA carry no tissue labels (need Table S3), so the match is on DISTRIBUTION SHAPE not per-tissue identity, and all grades are qualitative 'partial' rather than exact. Out of scope (wet-lab/manual, not attempted): synteny/phylogeny, qRT-PCR (Fig5), RNAscope (Fig6).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ a3359275d300
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study aims to comprehensively characterize the evolutionary history of auts2 genes in teleost fish—determining whether the two zebrafish auts2 genes (auts2a, auts2b) originate from the teleost-specific whole genome duplication (TGD)—and to assess whether auts2 expression patterns are conserved across teleosts and with mammals.

Core claims
  • auts2a and auts2b originate from the teleost-specific whole genome duplication (TGD) finding
  • auts2a, highly similar to human AUTS2, was almost systematically retained following TGD finding
  • auts2b, encoding a shorter protein similar to a short human AUTS2 isoform, was lost more frequently and independently across teleost lineages finding
  • RNA-seq across 10 species reveals a highly conserved auts2a/auts2b expression profile, predominant in embryo, brain, and gonads finding
  • The long human AUTS2 isoform's functions were likely retained by auts2a, while short isoform functions were retained by auts2a and/or auts2b depending on lineage mechanism
  • auts2a shows a burst in expression during medaka brain formation, localized to brain regions associated with neurodevelopmental disorders finding
  • Additional lineage-specific WGD events in Salmonidae and Cyprininae further reshaped auts2 gene repertoires finding
  • Non-functional auts2b gene remnants are detectable in Acanthomorphata and other genomes, confirming ancient gene loss rather than absence of duplication finding
Experimental setups
Assay System Perturbation Readout Platform
comparative genomics / synteny analysis 78 fish species (incl. spotted gar, bowfin, teleosts) none conserved gene order/orthology around auts2a/auts2b loci Genomicus v03.01, Ensembl v106, BLAST
TBLASTN gene remnant search genomic sequences of selected teleost species (e.g., Mexican tetra, electric eel, channel catfish, iridescent shark catfish) none detection of non-functional auts2a/auts2b remnants TBLASTN
phylogenetic analysis Auts2 full-length protein sequences from 33 species none tree topology separating auts2a and auts2b clades MEGA v11.0.11 (CLUSTALW alignment; maximum likelihood, neighbor-joining, minimum likelihood, 500 bootstrap replicates)
bulk RNA-seq 11 tissues from 10 species (medaka, Atlantic cod, Northern pike, zebrafish, Mexican tetra Pachon/surface morphs, iridescent shark catfish, allis shad, European eel, spotted gar, bowfin) none gene expression levels in TPM TrimGalore + Salmon v1.8
protein sequence alignment/domain analysis multiple species Auts2 amino acid sequences none conserved protein domains and isoform length BioEdit v7.0.5.3
RT-qPCR Japanese medaka embryos (11 developmental stages, 4 biological replicates) developmental stage (none/timepoint) relative auts2a mRNA abundance (full-length vs. full-length+short isoform primers), normalized to luciferase spike-in Light Cycler 480, PowerUp SYBR Green
RNAscope in situ hybridization (Multiplex Fluorescent V2) Japanese medaka stage 29 embryos (n=5) none spatial localization of auts2a transcript TCS SP8 confocal microscope, ACDBio RNAscope kit
Key results
  • Conserved synteny between Holostei (spotted gar, bowfin) and teleosts shows two distinct auts2 genomic regions, indicating TGD origin
  • auts2a retained in 52 Acanthomorphata species while auts2b was lost in this lineage ~150 Mya, with detectable non-functional remnants in 8 tested species n=52 species retained auts2a only; n=8 species with confirmed remnants
  • auts2b absent (no detectable remnants) in 3 Osteoglossomorpha species, consistent with loss ~200-250 Mya n=3 species
  • Highly variable auts2a/auts2b retention across Otomorpha clade (both genes, only auts2a, or only auts2b depending on species)
  • RNA-seq shows predominant auts2a/auts2b expression in embryo, brain, and gonads across all 10 species examined
  • auts2a expression increases sharply during medaka brain formation stages
  • Cyprininae species retain both auts2a and auts2b in duplicated form (auts2a1/a2, auts2b1/b2) due to additional lineage-specific WGD
Key statistics
  • count 78 species (species included in synteny/accession analysis)
  • count 76 teleost species (comprehensive gene repertoire analysis)
  • count 52 species (Acanthomorphata species retaining only auts2a)
  • count 33 species (species used in phylogenetic analysis of Auts2 protein sequences)
  • count 10 species (species with RNA-seq expression data across 11 tissues)
  • other ~150 Mya (estimated divergence time of Acanthomorphata lineage, associated with auts2b loss)
  • other 200–250 Mya (estimated timeframe of auts2b loss in Osteoglossomorpha)
  • pvalue p < 0.05 (significance threshold for Kruskal test and pairwise Wilcoxon signed rank test comparing auts2a expression across medaka developmental stages)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is primarily a comparative-genomics/phylogenetics analysis (synteny mapping, TBLASTN gene-remnant searches, maximum-likelihood/neighbor-joining/minimum-likelihood phylogenies) described descriptively without inferential statistics, combined with a quantitative developmental expression analysis in medaka. auts2a mRNA levels were measured by RT-qPCR across 11 embryonic stages (4 biological replicates per stage, each analyzed in technical quadruplicate) and compared using a Kruskal-Wallis test across all stages, followed by pairwise Wilcoxon signed-rank tests between individual stage pairs, both evaluated against a p<0.05 threshold. RNA-seq expression (TPM) across 10 species/11 tissues and RNAscope in situ hybridization images were reported descriptively, without a stated inferential test for those comparisons.

Replicationmixed Sample size4 biological replicates (pools of embryos) per developmental stage (pools of 60–95 embryos at early stages, 28–44 at later stages); each biological replicate additionally run in technical quadruplicate by qPCR Groupsauts2a expression levels across 11 medaka embryonic developmental stages, per primer set (exon 2–3 junction vs. exon 18–19 junction) Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Kruskal-Wallis test (referred to as 'Kruskal test') comparison of auts2a qPCR expression levels across 11 medaka embryonic developmental stages, performed separately for each primer set 4 biological replicates (pooled embryos) per stage not stated
Pairwise Wilcoxon signed rank test two-by-two comparisons of auts2a expression between individual developmental stages, following the omnibus Kruskal-Wallis test 4 biological replicates per stage not stated
Approaches that could also have been used
  • Expression differences across 11 developmental stages were assessed with a Kruskal-Wallis test followed by pairwise Wilcoxon tests at p<0.05, without a stated multiple-comparison adjustment.
    Could also: A formal multiplicity correction such as Dunn's test with Benjamini-Hochberg FDR or Bonferroni adjustment applied to the pairwise p-values. — With 11 stages generating many pairwise comparisons, an explicit correction step is a standard way to control the family-wise error rate or false discovery rate across the full set of contrasts.
  • Pairwise comparisons between distinct, independently sampled embryo pools (different biological replicates per stage) were labeled a 'Wilcoxon signed rank test', a term conventionally used for paired/matched samples.
    Could also: The Mann-Whitney U test (Wilcoxon rank-sum test), the standard non-parametric test for two independent (unpaired) groups. — When the compared samples are separate biological replicate pools rather than matched pairs, the rank-sum/Mann-Whitney form is the test conventionally used for independent-sample comparisons.
  • Non-parametric tests (Kruskal-Wallis and Wilcoxon) were used to compare qPCR-derived expression levels across stages.
    Could also: A parametric approach such as one-way ANOVA with a Tukey HSD or Dunnett post-hoc test, potentially after log-transforming the normalized expression values. — If normality and variance-homogeneity assumptions are reasonably met, parametric tests can offer greater statistical power with small per-group sample sizes such as n=4.
  • Sample size was 4 biological replicates per stage, with no stated power analysis or sample-size justification.
    Could also: Reporting an a priori or post-hoc power calculation, or presenting effect sizes with confidence intervals alongside the hypothesis tests. — This helps convey how precisely differences between stages could be detected given the chosen replicate number, particularly useful for interpreting non-significant comparisons.
  • RNA-seq expression levels (TPM) across species and tissues were presented descriptively without an accompanying inferential statistical test.
    Could also: A dedicated differential-expression framework such as DESeq2 or edgeR with appropriate multiple-testing correction, if statistical comparison of expression levels across species/tissues were intended. — Such tools model count-based variance-mean relationships and provide adjusted p-values, which is the standard approach when formal statistical comparison of RNA-seq expression is the goal rather than descriptive profiling.
  • Dispersion around expression estimates (e.g., in qPCR or RNA-seq figures) is not described in the methods text provided.
    Could also: Reporting standard deviation, SEM, or a 95% confidence interval alongside means in figures. — Explicitly stating the dispersion measure used helps readers gauge the variability and precision of the expression estimates at each stage.
Software: MEGA 11.0.11 · Salmon 1.8 · TrimGalore 0.6.6 · BioEdit 7.0.5.3 · Light Cycler 480 software 1.5.1.62 · ImageJ (FIJI) 1.53f51

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36078102

Paper: Merdrignac et al. 2022, Cells 11(17):2694. "auts2 Features and Expression Are Highly Conserved during Evolution Despite Different Evolutionary Fates Following Whole Genome Duplication." DOI 10.3390/cells11172694 · PMCID PMC9454499.

Pipeline-derived results (IN SCOPE)

The only computational/bioinformatic pipeline in the paper is Section 2.4 "Read Processing and Estimation of Expression Levels", which feeds Figure 2B (and the qualitative claims in Results §3.3 about tissue distribution of auts2a/auts2b).

Pipeline as described:

  1. Download RNA-seq from SRA. Japanese medaka = SRR1524271–81 (11 runs / 11 tissues, libs L_Ol_2L_Ol_13); plus 10 other species (cod, pike, zebrafish, 2× Mexican tetra, catfish, allis shad, European eel, spotted gar, bowfin). Data origin: PhyloFish (Pasquier et al. 2016, ref [26], BioProject PRJNA255889).
  2. TrimGalore v0.6.6 — trim known 3′ adaptor + low-quality bases (Phred < 20). Stated parameters: -r_clip 13 -three_prime_clip 2 → maps to TrimGalore --clip_R1 13 --three_prime_clip_R1 2 (and the R2 equivalents for paired-end; the paper's shorthand is ambiguous on R2).
  3. Salmon v1.8 (defaults) — map trimmed reads to each species' reference genome + annotated transcripts (Table S4).
  4. Expression in TPM (transcripts per kilobase million).

Reproducible target: the auts2a (and auts2b) TPM expression profile across medaka tissues = the medaka portion of Figure 2B, and the qualitative claim "predominant expression in embryo, adult brain, and gonads."

Reproduction unit assigned by BRIEF

  • Code: https://github.com/FelixKrueger/TrimGalore (third-party tool — valid per P16).
  • Data: SRR1524271 (one medaka run, lib L_Ol_2, paired-end, HiSeq 2000, 33.5 M spots, 2×100 bp). Minimum unit = run TrimGalore+Salmon on this run and obtain its auts2a/auts2b TPM. Stretch: process all 11 medaka runs to reconstruct the medaka tissue profile of Fig 2B.

OUT OF SCOPE (wet-lab / manual / external — not attempted)

  • §2.1 Synteny analysis, §2.2 gene-remnant identification — manual/BLAST/genome browser.
  • §2.3 Phylogenetic analysis (MEGA 11, CLUSTALW) — Fig 1, Fig 3. Manual, GUI tool.
  • §2.5 Sequence-feature alignment (BioEdit GUI) — Fig 4. Manual.
  • §2.6–2.10 Embryo sampling, RNA extraction, qRT-PCR (Fig 5), RNAscope ISH (Fig 6) — pure wet-lab.

Honesty caveats

  • No numeric ground-truth in text. Fig 2B is a heatmap/bar figure; the paper reports NO per-tissue TPM numbers in text or in the main tables we can read. Comparison is therefore qualitative (pattern: high in brain/gonad/embryo, low elsewhere) unless Supplementary Table S3/S4 ship numeric TPM (to be checked).
  • Tissue identity of SRR1524271 is not pinnable from SRA metadata. Libs are named L_Ol_2L_Ol_13 with no tissue field; the tissue↔run map lives in the paper's Table S3 / PhyloFish metadata. Recorded as a limitation.
  • Reference transcriptome version for medaka (Ensembl Oryzias latipes) at run time may differ from the authors' 2022 build → affects exact TPM (within-tol at best).
Figures / tables: Fig 2B
auts2a-expressed-medaka
Reported
auts2a is expressed in medaka tissues (Oryzias latipes), part of conserved auts2 expression (Fig 2B / Results 3.3); no numeric TPM printed in text
Reproduced
auts2a detected and expressed in all 11 medaka runs; e.g. SRR1524271 (lib L_Ol_2) gene-level TPM=56.29 (2937 reads); mapping rates 63.7-84.1%
partial
auts2b-annotation-fate
Reported
auts2 ohnologs have different evolutionary fates following WGD (title/theme)
Reproduced
auts2b NOT annotated in current Ensembl medaka build (ASM223467v1); 0 TPM/0 reads in all 11 runs; only auts2a present/quantifiable
partial
auts2-tissue-pattern
Reported
auts2 (auts2a) predominant expression in embryo, adult brain and gonads; conserved tissue distribution (Fig 2B)
Reproduced
auts2a TPM across all 11 medaka runs spans 0.52-56.3 (>100-fold): 4 high runs (56.3, 21.4, 16.1, 15.3), 7 low (<4, several <1) -> strongly tissue-restricted, few-high/many-low pattern matching predominant expression in a subset of tissues
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

This is an explicitly unfinished reproduction: the SRA data (SRR1524271, n_observed 11) is in hand and the §2.4 pipeline (TrimGalore v0.6.6 → Salmon v1.8 → TPM) is specified, but nothing has been computed (reproduced='PENDING', agreement.json='not-run-yet', grade=null). The endpoint itself is problematic for verification: Fig 2B is a heatmap with no printed numbers, so even when the pipeline runs the comparison can only be qualitative (q2 red). The shortfall sits on our/methodological side (single library chosen, job not yet run) plus a paper-presentation limit (no numeric ground-truth), not on a detected discrepancy or fabrication — hence yellow on derivability/core-claim/overall rather than red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

46.5 k
tokens (I/O) · 1.6 M incl. cache
20 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.