Eye in a Disk: eyeIntegration Human Pan-Eye and Body Transcriptome Database Version 1.0.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough and reproduced 1:1. Re-ran the authors' published manuscript R code (eyeIntegration_v1_app_manuscript @917dedca) on the shipped intermediate data/*.Rdata objects on «our HPC». 19/23 pinned numeric claims reproduce EXACTLY (sample counts 835/1314, QC removals 81/61, per-tissue counts, 41/45 gene/transcript types, 882 DE tests, DEG range 1-33380, new-sample counts 448/207/655, 24 new studies). 2 statistical-power values are within ~1 percentage point (shipped ssizeRNA simulation, power_data.Rdata: 84.1% vs 83%, 89.5% vs ~90%). NOT attempted: R2=0.89 TPM correlation and 67.3e9 total reads - these require the full raw Salmon re-processing of 2291 SRA samples / the 23.8GB SQLite DB, far beyond the quick minimum and not in the shipped repo data (reproducible-in-principle from the Zenodo DB). Key methodological note: counts are n_distinct(sample_accession), not table rows (data is run-level). No fabrication flags - every value is directly derivable from shipped data/code.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 94assessed: 2026-06-16 ⛓ 0b04dbf5777a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a reproducible, versioned RNA-seq transcriptome database of healthy human eye tissues (alongside GTEx body tissues) be built and served via an interactive web app to enable querying of gene and transcript expression across eye and body tissues?
- ★ EiaD is a reproducible, versioned pan-eye and body RNA-seq transcriptome dataset built from 916 eye and 1375 GTEx samples via a Snakemake pipeline output as a single SQLite database. resource
- ★ The rewritten eyeIntegration v1.0 web app interactively serves the EiaD dataset across 19 eye tissues and 54 body tissues. resource
- ★ A Snakemake-based reproducible pipeline automates transcript/gene quantification, QC, normalization, 882 differential expression tests, and GO term enrichment. method
- ★ Fetal retina and organoid retina are highly similar at a pan-transcriptome level but differ in specific pathways and gene families such as protocadherin and HOXB. finding
- Temporal gene expression patterns in fetal retina are recapitulated in retina organoids. finding
- Differentially expressed protocadherin and HOXB family genes between organoid and embryonic retina suggest targetable pathways to improve organoid differentiation. mechanism
- The app supports versioned datasets, custom URL shortcuts, new visualizations, data download, and local install with three commands. resource
- A prototype tool displays single-cell RNA-seq data for cell type-specific gene expression across murine retinal development. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (transcript/gene quantification) | human eye-related tissues (cornea, lens, retina, RPE/choroid; adult, fetal, iPSC-derived, organoid, immortalized cell lines) | none | transcript and gene expression (length-scaled TPM) | Salmon (gencode v28 index), tximport |
| bulk RNA-seq (reference body tissues) | 54 human body tissues from GTEx (1375 samples) | none | gene and transcript expression | Salmon, tximport |
| differential expression testing | EiaD eye and body tissue samples | none | 882 differential expression tests | limma/edgeR (R) |
| GO term enrichment analysis | EiaD differentially expressed gene sets | none | GO term enrichment | — |
| pan-transcriptome similarity / outlier and clustering analysis (t-SNE) | fetal retina vs retina organoid samples | none | transcriptome similarity, differentially expressed processes/gene families | qsmooth, edgeR, limma (R) |
| single-cell RNA-seq (prototype visualization) | murine retinal development | none | cell type-specific gene expression | — |
- – Organoid retina is highly similar to early fetal retina tissue at a pan-transcriptome level
- – Protocadherin and HOXB family genes are differentially expressed between organoid and embryonic retina
- – Temporal gene expression patterns of fetal retina are recapitulated in organoids
- – eyeIntegration portal visualizes transcriptomes across 19 eye tissues and 54 body tissues 19 eye tissues; 54 body tissues
- count 1375 RNA-seq samples across 54 tissues (GTEx) (GTEx noneye reference set)
- count 916 eye samples (curated human eye-related RNA-seq samples)
- count 882 differential expression tests (performed by the pipeline)
- count 19 eye tissues and 54 body tissues (tissues served by eyeIntegration v1.0)
- count 3000 genes with highest variance (selected for outlier identification)
- other median count >200 across all samples (gene filtering threshold)
- other mapping rate less than 40% (QC threshold for sample removal)
- other transcripts accounting for 5% or less of parent gene expression removed (transcript filtering per Soneson et al.)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a bioinformatics resource paper describing a reproducible Snakemake pipeline that assembles a pan-eye and body RNA-seq transcriptome database (EiaD) from 916 eye and 1375 GTEx samples. The statistical workflow centers on transcript/gene quantification (Salmon, tximport), library-size and cross-study normalization (edgeR calcNormFactors, qsmooth), batch/mapping-rate correction (limma), a correlation-based outlier-detection procedure adapted from Wright et al., and 882 differential expression tests plus GO term enrichment, with results served interactively rather than reported as a fixed set of inferential comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| differential expression testing (882 tests; specific test framework not stated in the available text) | across eye and body tissue/subtissue comparisons, including organoid vs. fetal/embryonic retina | — | not stated |
| GO term enrichment analysis (method not stated) | differentially expressed gene sets, e.g. protocadherin and HOXB family pathways | — | na |
| correlation-based outlier detection (per-sample average correlation within subtissue, adapted from Wright et al.) | quality control across subtissue types using the 3000 highest-variance genes | — | not stated |
-
Differential expression was assessed with 882 tests within a limma-based normalization/correction workflow.↳ Could also: Count-based negative-binomial frameworks such as DESeq2 or edgeR's glm/quasi-likelihood tests, or limma-voom with precision weights, could also have been used. — These approaches model RNA-seq count distributions explicitly and are widely used; reporting which was chosen helps readers match the test to the data type and reproduce results.
-
The text describes performing many differential expression tests across numerous tissue comparisons.↳ Could also: An explicit family-wise or false-discovery-rate correction (e.g., Benjamini-Hochberg FDR) applied across the family of tests could also be reported alongside the results. — Stating the multiplicity scope and method makes the rate of expected false positives transparent when many genes and comparisons are evaluated.
-
Cross-study technical variation was handled with library-size normalization, quantile smoothing, and limma batch correction.↳ Could also: Latent-variable methods such as RUVseq, svaseq/ComBat, or including study as a covariate in the DE model could also account for unmodeled batch structure. — These alternatives estimate hidden sources of variation directly and can be compared to confirm robustness of tissue-level signals across heterogeneous public datasets.
-
Outliers were identified using an average within-subtissue correlation approach adapted from Wright et al. with a fixed 3000 high-variance gene set.↳ Could also: Distribution-based or model-based outlier metrics (e.g., PCA/Mahalanobis distance, arrayQualityMetrics-style scores, or median-absolute-deviation thresholds) could also flag atypical samples. — Combining complementary outlier criteria can provide convergent evidence and reduce sensitivity to any single threshold or gene-set choice.
-
A mapping-rate cutoff of 40% and a median count >200 gene-expression filter were selected from the observed distributions.↳ Could also: Data-driven thresholds (e.g., filterByExpr in edgeR) or sensitivity analyses across several cutoffs could also be reported. — Showing how conclusions hold across alternative thresholds conveys how robust the included-sample and expressed-gene sets are to the chosen values.
-
High-dimensional tissue relationships were visualized with t-SNE on corrected expression values.↳ Could also: Complementary projections such as PCA or UMAP could also be presented. — PCA preserves global variance structure and UMAP often better preserves global topology, so multiple embeddings can corroborate the reported similarity between organoid and fetal retina.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Temporal gene expression patterns of fetal retinal development are recapitulated in retina organoidsRNA-seq human fetal retina none 2019×1papers★ This paper is the founder (earliest)
-
PCDH (protocadherin) and HOXB family genes are differentially expressed between retina organoids and embryonic retinaRNA-seq human retina mixed 2019×1papers★ This paper is the founder (earliest)
-
Retina organoids are highly similar to early fetal retina at the pan-transcriptome levelRNA-seq human retina none 2019×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-31343654 (Eye in a Disk / eyeIntegration v1.0)
Swamy & McGaughey, IOVS 2019. A database/resource paper. EiaD = a pan-eye +
GTEx body transcriptome database built by a Snakemake pipeline (EiaD_build)
from public SRA RNA-seq, served via a Shiny app (eyeIntegration_app). The
manuscript's reported numbers are regenerated by an R-Markdown repo
(eyeIntegration_v1_app_manuscript) that ships the intermediate data/*.Rdata
objects + src/*.R.
In scope (pipeline-derived, reproducible)
We re-run the authors' published analysis code on the SHIPPED intermediate data
(the manuscript repo's data/*.Rdata). This is the "third-party tool / author
code on the paper's own data" path (P16) and regenerates the manuscript's
reported counts/statistics directly:
- Post-QC sample counts: 835 eye, 1314 GTEx across 54 tissues
- QC removals: 81 eye, 61 GTEx
- Per-tissue eye counts: 6 iPSC/ESC, 56 cornea, 4 lens, 648 retina, 121 RPE
- Gene/transcript type diversity: 41 gene types, 45 transcript types
- 882 differential-expression tests
- DEG range per comparison (min..max)
- Statistical power: 83% @ 20 samples, ~90% @ 30 samples (power_data.Rdata)
- GTEx vs EiaD TPM R² = 0.89 (if reference TPM is shipped)
Out of scope (NOT attempted, with reason)
- Full raw re-processing of 2291 SRA samples (67,315,523,736 reads) through
Salmon (gencode v28) -> tximport -> qsmooth -> limma. This is the
EiaD_buildSnakemake pipeline: hundreds of TB-scale SRA downloads + weeks of compute. Out of scope as a "quick minimum"; we instead validate the published derived objects. The 0.89 R² and read total trace to this stage. - Wet-lab / external: none (pure bioinformatics paper).
- The live web app + visitor analytics (visitor_stats) — not a scientific claim.
Pipelines named per result
- Sample QC counts, tissue tables, gene/tx types, DE-test count, DEG range:
limma/edgeR/tximport pipeline, summarized in
data_pull.Rdata. - Power: ssizeRNA on edgeR dispersions ->
power_data.Rdata. - R²=0.89: Salmon-TPM vs GTEx v7 reference correlation.
Repos / data pointers
- Code: github.com/davemcg/eyeIntegration_v1_app_manuscript (manuscript numbers), /EiaD_build (Snakemake), /eyeIntegration_app (Shiny).
- Data: zenodo 10.5281/zenodo.3238677 (eyeIntegration_v104_04.tar.gz 23.8GB = full DB; + repo archives). Manuscript intermediate data ships in the repo.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Reproduction re-ran the authors' own manuscript R code on their shipped Zenodo intermediate objects, so input data is identical 1:1. 19/23 pinned values (sample counts 835/1314, QC removals 81/61, per-tissue counts, 882 DE tests, DEG range 1–33,380, new-sample counts 448/207/655) reproduced exactly; the only deviations are two power values within ~1pp (ssizeRNA simulation rounding). The two unattempted claims (R²=0.89 TPM correlation, 67.3e9 reads) are out-of-scope for compute but reproducible-in-principle from the shipped DB — not failures and not on the authors' side. No fabrication or substantive discrepancy: a clean, green reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
<synthetic>Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.