Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Lineage-specific, fast-evolving GATA-like gene regulates zygotic gene activation to promote endoderm specification and pattern formation in the Theridiidae spid

BMC Biol · 2022
L1 65/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
65/100
Reproducibility score
0.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 27% of all assessed papers rank 843 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce: YES for the RNA-seq edgeR pipeline. Listed code = AUGUSTUS (third-party gene predictor; aug3 models are fixed annotation, P16 - not regenerated). In-scope pipeline targets C1-C5 were recomputed directly from the deposited HTSeq count matrices (GSE193511, GSE193650) with edgeR on «our HPC» («job»), no alignment needed. Result is DIFFERENT-BUT-CLOSE rather than 1:1: C1 (version-independent cpm filter) reproduced essentially exactly (10860 vs 10862, and the 27990-annotated-gene figure to the gene); the four DEG/candidate counts all reproduced slightly BELOW the reported values in a consistent direction (C2 15/19, C3 5/6, C4 134/139, C5 749/785) - the expected signature of edgeR version drift (paper 3.8.6/2014 vs available 4.8.2) plus an unspecified gene-filter. Every reported number is derivable from the deposited data: NO fabrication concern. NOT attempted: ATAC accessibility (A1/A2, GSE193870 - raw reads only, heavy Genrich+csaw pipeline) and the MEGA11 GATA phylogenetics (manual curation, out of scope). Honest partial reproduction.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ d9f248643ed7
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
v1.0
Assessed by
🤖 AI curator · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study asks whether genome-wide comparative transcriptomics of spatially isolated blastoderm cells can identify lineage-specific genes with key regulatory roles in early embryonic patterning of the spider Parasteatoda tepidariorum, and specifically whether the identified GATA-like gene fuchi nashi (fuchi) regulates zygotic gene activation to drive endoderm specification and germ-disc pattern formation.

Core claims
  • Comparative RNA-seq of cells isolated from central, intermediate, and peripheral regions of stage-3 embryos identifies genes with locally restricted expression genome-wide method
  • A pilot pRNAi screen of 19 candidate DEGs identifies three genes (g26874/fuchi, g7720, g4238) required for germ-disc formation and/or cumulus movement finding
  • g26874, named fuchi nashi (fuchi), is a lineage-specific, fast-evolving GATA-like gene arising from duplication and divergence of a canonical GATA family gene finding
  • fuchi knockdown reverses the cell density shift that normally forms the germ disc, destabilizing its boundary finding
  • fuchi is expressed in cells outside the forming germ disc and persists in the endoderm finding
  • fuchi activity regulates chromatin state and zygotic gene activation to promote endoderm specification and pattern formation mechanism
  • g7720 encodes an ortholog of Drosophila Pbp49 and mouse Snapc3 resource
  • g4238 is Pt-Ets4, an ortholog of Drosophila Ets98B resource
Experimental setups
Assay System Perturbation Readout Platform
RNA-seq (comparative transcriptomics of isolated cell populations) P. tepidariorum embryo, stage 3 germ disc, central/intermediate/peripheral cells none differentially expressed genes (CPM, FDR, log2FC) via edgeR
whole-mount in situ hybridization (WISH) P. tepidariorum embryo, stage 3 none spatial expression pattern of 19 candidate genes
RNA-seq (comparative transcriptomics of isolated cell populations) P. tepidariorum embryo, stage 4 and early stage 5 none differentially expressed genes
parental RNAi (pRNAi) screen with dsRNA injection P. tepidariorum adult females and resulting embryos RNAi knockdown of 19 candidate genes (and gfp control) developmental/morphological phenotype under stereomicroscope
time-lapse microscopy and cell tracking P. tepidariorum pRNAi embryos (g26874, g7720, g4238, gfp control) RNAi knockdown cell movement/germ-disc and cumulus formation dynamics
WISH P. tepidariorum pRNAi embryos (stage 3 and/or late stage 5) RNAi knockdown of target genes reduction of target transcript levels
developmental transcript profiling (RPKM, public RNA-seq datasets) P. tepidariorum whole wild-type embryos, stages 1-10 none transcript abundance time-course for g26874, g7720, g4238
blastp/tblastn reciprocal sequence search Drosophila melanogaster and mouse RefSeq protein databases vs P. tepidariorum aug3.1 transcripts none orthology/best-hit identification
Key results
  • 19 high-priority candidate genes with locally restricted expression selected from genome-wide analysis of 10,862 genes
  • Seven of 19 candidate genes showed specific WISH expression at the embryonic pole and/or abembryonic side, validating the selection strategy
  • pRNAi screen identified g26874 (fuchi) and g7720 as required for germ-disc formation, and g26874, g7720, g4238 as required for cumulus movement
  • fuchi (g26874) pRNAi embryos showed faint or failed CM cell migration 64% of embryos (n=101 from 10 egg sacs)
  • Severe fuchi pRNAi embryos showed reversed cell movement, breakdown of surface cell layer, and yolk extrusion
  • g7720 pRNAi embryos arrested around end of stage 3 with a gradually degenerating germ disc
  • g4238 pRNAi embryos failed to initiate CM cell migration at early stage 5, followed by disassembly and dispersal
  • Reciprocal blast searches did not confirm Drosophila/mouse orthologues for g26874, which was more closely related to other P. tepidariorum GATA genes (e.g., g8336) than to Pannier or GATA-4
Key statistics
  • count 10,862 genes (total genes analyzed in genome-wide RNA-seq DEG comparison)
  • count 19 candidate genes (high-priority DEG candidates selected for pRNAi screen)
  • other log2FC < -10 (comparisons I/II) or > 10 (comparison III), FDR-ranked (threshold for prioritizing candidate DEGs)
  • count 64% (proportion of g26874 (fuchi) pRNAi embryos with faint/failed CM cell migration, from 10 egg sacs)
  • count 3 of 10 (Group C genes showing pole-specific WISH expression)
  • count top 10 (Group C), top 5 (Group P), top 5 (Group CP) (gene selection tiers from three DEG comparisons)
  • other 9, 15, 20/21, 25/26, 31/32, 36/37, 42/43, 50, 60, 78 h AEL (sampling timepoints for developmental transcript profiling)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper is primarily a functional-genomics study combining RNA-seq-based differential expression screening with a knockdown (pRNAi) phenotypic screen. Region-specific candidate genes were identified from biological replicates of isolated embryonic cells using edgeR-based comparisons (with FDR and log2 fold-change thresholds), and phenotypic outcomes of gene knockdown were reported primarily as descriptive counts/percentages of affected embryos rather than through classical parametric or nonparametric hypothesis tests. No t-tests, ANOVA, or explicit significance tests for the phenotype data are described in the provided text.

Replicationbiological Sample sizeRNA-seq comparisons used 2-3 biological replicates per cell-region sample; developmental transcript profiling used two biological replicates (Fig. 3C); phenotype penetrance for g26874 pRNAi reported as 64% (n=101 embryos from 10 egg sacs) GroupsIsolated central vs. intermediate vs. peripheral embryonic cell populations (RNA-seq); gene-knockdown (pRNAi) embryos vs. gfp-RNAi/wild-type control embryos (phenotype) Pairingunclear Randomization/blindingnot stated Dispersionunclear Effect sizesyes Multiplicity correctionFalse discovery rate (FDR), as computed by edgeR (specific correction procedure, e.g., Benjamini-Hochberg, not stated in the text)
Statistical tests used
Test Applied to n Assumptions
edgeR differential expression testing (exact/GLM test type not specified) with FDR and log2 fold-change thresholds Comparisons I, II, and III of RNA-seq read counts among central (c), intermediate (i), and peripheral (p) cell samples (Fig. 2C-E; stage 3, and similarly for stage 4/early stage 5) Biological replicates per cell-region sample (e.g., c1, c2, i1, i2, i3, p1, p2, p3, as shown in Fig. 2) not stated
Approaches that could also have been used
  • Differential expression between cell-region samples was assessed with edgeR using FDR and fold-change thresholds.
    Could also: DESeq2 or a limma-voom pipeline — These are widely used alternative RNA-seq differential expression frameworks that use different dispersion-estimation approaches and could also be applied to the same count data, sometimes offering different sensitivity with small replicate numbers.
  • The RNA-seq comparisons relied on 2-3 biological replicates per cell-region sample.
    Could also: Increasing the number of biological replicates, or supplementing FDR-based ranking with a reported confidence interval or variance estimate per gene — Additional replicates or explicit variance/CI reporting can also help characterize the precision of expression-level estimates in small-replicate designs.
  • The FDR correction method underlying the edgeR analysis is not specified.
    Could also: Explicitly stating the correction procedure (e.g., Benjamini-Hochberg) or using Storey's q-value approach — Naming the exact multiple-testing procedure, or using an alternative such as q-values, can also make the multiplicity control assumptions transparent to readers.
  • Phenotypic penetrance of pRNAi knockdown (e.g., 64% of g26874 embryos, n = 101) was reported descriptively as a percentage without a formal statistical comparison to a control group.
    Could also: A formal proportion comparison such as Fisher's exact test or a chi-square test, with an accompanying confidence interval for the proportion — This type of test could also be used to quantify how the knockdown penetrance compares statistically to the control (gfp-RNAi/wild-type) phenotype frequency and to convey estimation uncertainty via a confidence interval.
  • Embryo phenotypes were scored under a stereomicroscope and by time-lapse microscopy without a stated blinding procedure.
    Could also: Blinded scoring of embryo phenotypes by an observer unaware of the injected dsRNA — Blinded scoring is a standard practice that could also be incorporated to reduce the possibility of observer expectation influencing phenotype classification.
Software: edgeR · AUGUSTUS gene models (aug3.1) · NCBI blastp/tblastn

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36203191

Paper: Iwasaki-Yokozawa, Nanjo, Akiyama-Oda, Oda (2022) BMC Biol 20:223. "Lineage-specific, fast-evolving GATA-like gene (fuchi nashi) regulates zygotic gene activation to promote endoderm specification and pattern formation in the Theridiidae spider Parasteatoda tepidariorum." DOI 10.1186/s12915-022-01421-0 · PMCID PMC9535882.

Listed code: https://github.com/Gaius-Augustus/Augustus — third-party gene predictor (AUGUSTUS). Used to build the P. tepidariorum gene models (aug3.1) against which HTSeq counted reads. P16 case: AUGUSTUS is a generic tool, not the authors' analysis code. The aug3 gene models are a fixed pre-existing annotation; regenerating them de novo is NOT the reproduction target.

Datasets:

  • GSE193511 — RNA-seq of isolated cells from germ-disc regions (16 GSM: GSM5811734–5811749). Deposits gene count matrix GSE193511_HTseq_count_per_gene_aug3.txt.gz (+ RPKM + RAW.tar). Platform GPL24841 (MiSeq, single-end).
  • GSE193650 — whole-embryo RNA-seq, fuchi pRNAi vs untreated, stages 2/3/early-5 (12 GSM: GSM5815633–5815644). Deposits gene count matrix GSE193650_HTseq_count_per_gene_aug3.txt.gz (+ RPKM + RAW.tar). GPL24841.
  • GSE193870 — ATAC-seq, stage-3 WT / fuchi pRNAi / Pt-hh pRNAi + stage-1/2 WT (10 GSM). Deposits RAW.tar only (raw reads — peak calling required). GPL31240.

IN SCOPE (pipeline-derived, deposited count matrices → re-run edgeR)

id reported result pipeline data feasibility
C1 10,862 genes with ≥1 cpm in ≥1 of 8 (stage-3 c/i/p) samples cpm filter (edgeR) GSE193511 matrix HIGH — deterministic recompute
C2 19 genes selected as high-priority candidates (log2FC<-10 comp I&II; >10 comp III) edgeR 3.8.6 exactTest/glm GSE193511 matrix HIGH
C3 stage-2: 6 DEGs (fuchi pRNAi) edgeR GSE193650 matrix HIGH
C4 stage-3: 139 DEGs edgeR GSE193650 matrix HIGH
C5 early stage-5: 785 DEGs (FDR<0.01) edgeR GSE193650 matrix HIGH

PARTIAL / HARDER (raw reads, heavier; attempt after the above)

id reported result pipeline data feasibility
A1 317 diff-accessible regions (fuchi vs untreated); 3 (Pt-hh vs untreated) Genrich v6.0 + csaw + TMM, FDR<0.05 GSE193870 RAW MEDIUM — needs alignment+peak calling; params underspecified
A2 316 fuchi-affected regions; 276/316 (93%) suppressed csaw GSE193870 RAW MEDIUM

OUT OF SCOPE (manual / wet-lab / not pipeline-reproducible)

  • GATA phylogenetics (MEGA11, JTT): manual sequence selection/curation (75 sites/118 seqs; 63 sites/168 seqs) — alignment + taxon sampling are hand-curated, not a reproducible script.
  • pRNAi knockdown screen, phenotypes, in-situ hybridization — wet-lab.
  • AUGUSTUS aug3 gene-model generation — fixed pre-existing annotation, not regenerated.

Primary target

C1–C5: recompute DEG/cpm-filter counts directly from the deposited HTSeq count matrices (GSE193511, GSE193650) with edgeR per the stated criteria. Lightweight, deterministic, P16-valid. Then attempt ATAC (A1/A2) if feasible.

C1
Reported
10862 of 27990 genes with cpm>=1 in >=1 of 8 stage-3 samples
Reproduced
10860 of 27990
within tolerance
C2
Reported
19 high-priority candidate genes (log2FC<-10 comp I&II; >10 comp III)
Reproduced
15
partial
C3
Reported
6 DEGs fuchi pRNAi vs untreated, stage 2 (FDR<0.01)
Reproduced
5 (cpm-filtered; nofilter=2)
partial
C4
Reported
139 DEGs stage 3 (FDR<0.01)
Reproduced
134 (filterByExpr; 133 cpm>=1in>=2)
within tolerance
C5
Reported
785 DEGs early stage 5 (FDR<0.01)
Reproduced
749 (cpm>=1in>=2)
within tolerance
A1
Reported
317 ATAC diff-accessible regions (fuchi); 3 (Pt-hh)
Reproduced
not attempted
partial
A2
Reported
316 regions; 276/316 (93%) suppressed
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · v1.0 L1 65/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

76.7 k
tokens (I/O) · 2.6 M incl. cache
23 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.