A curated collection of transcriptome datasets to investigate the molecular mechanisms of immunoglobulin E-mediated atopic diseases.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Curation/web-tool paper (GXB Gene Expression Browser over 33 GEO datasets). The linked repo BenaroyaResearch/gxbrowser @5a63c361 is the viewer WEB APP (Tomcat6/Grails2.1/Java7/MySQL/Mongo/R) and produces no number in the paper, so it was NOT stood up (out of scope, obsolete stack). The only pipeline-derived result is GXB's mean two-group fold-change (FC = mean(exp)/mean(ctrl), linear scale, genes ranked by FC), which the authors validate against two literature FCs. Per P16 I reimplemented that exact described method on the paper's own GEO data («our HPC» «job»: R4.3.3 + GEOquery2.70 on «infra», getGEO GSE8507/GSE88796; local python FC arithmetic). C1 (CD151/GSE8507, Job's PBMC vs healthy): single probe 204306_s_at, linear Affy signal, disease label from GSM description (Job's 56 / healthy 85). Baseline PBMC contrast -> FC 1.81 vs reported 1.7 (within-tol, delta 0.11/6.5%; reported value is itself rounded; FC>1 up-in-Job's in every grouping, range 1.45-1.81). C2 (CEACAM1/GSE88796, egg-allergic vs tolerant): 3 Illumina probes, log2->linearized; reported 1.69 is bracketed only by single probe ILMN_1716815 in egg-stimulated allergic-vs-control contrasts (1.68-1.74), while the mean of all 3 probes gives 1.06-1.47 -> graded partial because the paper pins neither the exact sample subset (which stimulation/control group) nor the probe-selection, so 1.69 is plausible/derivable but not uniquely reproducible. NOT attempted (hard 20% / out of scope): deploying the GXB web app, re-curating all 33 datasets, and independently recomputing the 33-dataset/1860-profile curation tallies (manual curation, not a pipeline). No fabrication signs: both reported GXB FCs are derivable from the public GEO data and agree in direction and order of magnitude; C2's residual gap is text ambiguity, not a non-derivable number. Described well enough for a 1:1 on the single-probe Affy case (C1); under-specified on the multi-probe stimulation case (C2).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 62assessed: 2026-06-14 ⛓ 0b035a2c6677
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan publicly available GEO transcriptomic datasets relevant to IgE-mediated atopic diseases be systematically curated and made interpretable through an online visualization platform to facilitate cross-study mechanistic discovery and validation of complex disease signatures?
- ★ A curated collection of 33 GEO transcriptome datasets relevant to IgE-mediated atopic diseases (allergies to primary immunodeficiencies) was assembled, encompassing 1860 transcriptome profiles. resource
- ★ The datasets were made available on the Gene Expression Browser (GXB), an open-source web application for query, visualization and annotation of metadata, with ranked gene lists and sample grouping. resource
- ★ A seven-strategy GEO search workflow with merging, deduplication, platform/species filtering and manual relevance curation was used to select the datasets. method
- ★ Dataset validation against associated publications showed good concordance between GXB gene expression trend and fold-change. finding
- IL-4, IL-13 and STAT6 are key mediators of Th2 responses and IgM class switch to IgE. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Expression profiling by microarray and NGS (curated GEO datasets) | Human samples (PBMC, T cells, B cells, epithelium, skin, blood, etc.) | various (disease vs healthy, stimulated vs unstimulated, drug/immunotherapy) | gene expression (ranked gene lists, signal, fold-change) | Affymetrix, Illumina, Agilent arrays and Illumina HiSeq (GPL accessions listed) |
| RNA-seq (Illumina HiSeq 2500) | Human PBMC/T cells | Allergen-specific immunotherapy (low vs high allergic) | gene expression | Illumina HiSeq 2500 (GPL16791) |
| Microarray (Affymetrix U133A) | Human bronchial biopsy | Allergic asthma vs healthy | gene expression | Affymetrix Human Genome U133A Array (GPL96) |
| Microarray (Affymetrix Gene 1.0 ST) | Human nasal epithelium cells | IL4 stimulation vs control / rhinitis vs healthy | gene expression | Affymetrix Human Gene 1.0 ST Array (GPL6244) |
| Microarray (Illumina HumanHT-12 V4.0) | Human T cells / B cells / PBMC | immunotherapy / disease vs control | gene expression | Illumina HumanHT-12 V4.0 (GPL10558) |
| Microarray (Affymetrix U133 Plus 2.0) | Human PBMC, neutrophils, skin, sputum, mast cells | disease vs control / drug stimulation (e.g. dexamethasone, FK506) | gene expression | Affymetrix Human Genome U133 Plus 2.0 Array (GPL570) |
| Microarray (Agilent SurePrint G3 8x60K) | Human skin/whole blood, T cells | chronic spontaneous urticaria / seasonal allergic rhinitis vs control | gene expression | Agilent SurePrint G3 Human GE 8x60K (GPL16699/GPL14550) |
| RNA-seq (Illumina HiSeq 2000) | Human B lymphocytes | HDM allergy vs control | gene expression (IL4R increase) | Illumina HiSeq 2000 (GPL11154) |
- – Seven independent search strategies yielded 435 merged results, reduced to 196 and 117 datasets across two queries. 435; 196; 117
- ▼ Manual filtering to human microarray/NGS retained 115 datasets, then 53 relevant to IgE-related atopic disease, then 33 in the final collection. 115 → 53 → 33
- – Final collection comprises 33 datasets and 1860 transcriptome profiles. 33 datasets; 1860 profiles
- – Multiple datasets showed strong validation of gene expression trend/fold-change against associated publications (e.g. SERPINB2/CX3CR1/C7; PI3/S100A8/S100A7).
- – Heritable components of allergic diseases and atopy estimated at 33%–76%. 33%–76%
- – Largest dataset (asthma exacerbation PBMC, GSE19301) contained 685 samples. 685 samples
- count 33 (datasets in final curated collection)
- count 1860 (transcriptome profiles in collection)
- count 435 (merged results from seven search strategies)
- count 196 and 117 (datasets from two queries after deduplication)
- count 115; 53 (datasets after platform/species filter; after IgE relevance filter)
- other 33%–76% (estimated heritability of allergic diseases and atopy)
- other ~20% (allergic disease prevalence in developed nations)
- count 685 (samples in asthma exacerbation PBMC dataset (GSE19301))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a database/resource paper describing the curation, upload, and web-based visualization of 33 publicly available GEO transcriptome datasets (1860 profiles) relevant to IgE-mediated atopic diseases. The paper's own analytical content is limited to a systematic search-and-filter workflow for dataset selection and a qualitative concordance check comparing gene expression trends and fold-changes in GXB against values reported in the original publications. No new inferential statistical analyses were performed by the authors; statistical methods used to generate the underlying datasets belonged to the original depositing studies.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Qualitative trend and fold-change concordance check (informal visual/manual comparison, not a formal statistical test) | Validation of GXB-displayed gene expression against values reported in associated publications (Table 1, 'Trend validation' and 'FC validation' columns) | — | not stated |
-
Validation of GXB data against published results was performed qualitatively, categorizing concordance as 'Strong', 'Good', or similar ordinal labels for a small set of manually selected genes per dataset.↳ Could also: A quantitative validation metric such as Spearman rank correlation or Pearson correlation of log2 fold-changes between GXB-derived values and published values could also have been computed across all reported genes. — A numeric correlation coefficient would provide a continuous, reproducible measure of concordance and would allow readers to compare validation quality across datasets on a common scale rather than relying on subjective category assignment.
-
The dataset selection workflow was described narratively with step-by-step counts of datasets retained or excluded at each filter stage.↳ Could also: A PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) or similar structured reporting framework could also have been used to document the search and selection process. — PRISMA-style reporting provides a standardized, externally reproducible audit trail for systematic searches and is increasingly expected for literature- or database-curation studies to facilitate replication and updating of the collection.
-
Seven independent search strategies were merged and duplicates removed, but inter-rater agreement for the manual relevance-filtering step is not described.↳ Could also: A Cohen's kappa or percent-agreement statistic computed between two independent reviewers could also have been reported for the manual inclusion/exclusion step. — Inter-rater reliability metrics make explicit how reproducible the subjective curation decisions are, which is a recognized quality indicator for systematic curation studies.
-
The 33 curated datasets are described individually with sample-size information (Table 1) but no aggregate summary statistics across the collection are provided (e.g., distribution of sample sizes, platform types).↳ Could also: Descriptive statistics (median and IQR of sample sizes, counts and proportions by platform, disease category, and experimental design type) could also have been presented in a summary table or figure. — Aggregate distributional summaries help readers quickly assess the coverage and potential biases of the collection (e.g., platform over-representation) without inspecting all 33 rows of Table 1.
-
Fold-change values used for validation are taken directly from GXB output without a stated normalization pipeline specific to this paper.↳ Could also: A standardized re-normalization pipeline (e.g., RMA for Affymetrix, limma-voom or DESeq2 with a common normalization for RNA-seq) applied uniformly across all 33 datasets could also have been used prior to cross-dataset comparison. — Applying a consistent normalization method across heterogeneous datasets reduces platform- and study-specific technical variation and is a common approach in multi-study meta-analyses, potentially improving the comparability of fold-change estimates across datasets.
-
The paper describes the collection thematically by disease category and platform but does not assess cross-dataset consistency for any shared gene signatures.↳ Could also: A cross-study meta-analysis (e.g., using the metaMA or MetaDE R package, or a random-effects model on log fold-changes) for genes that appear differentially expressed in multiple datasets could also have been performed. — Formal cross-study synthesis would allow the collection to support pooled effect-size estimates and heterogeneity assessment, going beyond visualization to quantitative evidence aggregation — a natural next step given the stated goal of biomarker discovery.
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- CoINcIDE: A framework for discovery of patient... L1 87/100
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Unveiling prognostics biomarkers of tyrosine m...⚑ L1 51/100 ⚑
- Meta-analysis of gene expression profiles of l... L1 78/100
- Colorectal Cancer Prediction Based on Weighted...⚑ L1 80/100 ⚑
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Construction and Validation of an Immune Infil...⚑ L1 51/100 ⚑
- Identification of a novel 10 immune-related ge...
- Exploration of the shared diagnostic genes and... L1 76/100
- IRSN-23 gene diagnosis enhances breast cancer... L1 71/100
- Molecular Classification Models for Triple Neg... L1 86/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Autoencoder Networks Decipher the Association... L1 74/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
- Discovery and validation of molecular patterns... L1 83/100
- Comparative profiling of skeletal muscle model... L1 64/100
- PulmonDB: a curated lung disease gene expressi...⚑ L1 53/100 ⚑
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Meta-analysis of gene expression profiles of l... L1 78/100
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
- Comparative profiling of skeletal muscle model... L1 64/100
- CoINcIDE: A framework for discovery of patient... L1 87/100
- PulmonDB: a curated lung disease gene expressi...⚑ L1 53/100 ⚑
- Meta-analysis of gene expression profiles of l... L1 78/100
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Screening of Diagnostic Biomarkers and Immune...⚑ L1 48/100 ⚑
- IRSN-23 gene diagnosis enhances breast cancer... L1 71/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- PulmonDB: a curated lung disease gene expressi...⚑ L1 53/100 ⚑
- Curation of over 10 000 transcriptomic studies... L1 80/100
- VIGET: A web portal for study of vaccine-induc... L1 64/100
- Exploration of the shared diagnostic genes and... L1 76/100
- Exploring the key genetic association between... L1 91/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-31290545
Paper: Huang et al. 2019, Database (Oxford) baz066 — "A curated collection of transcriptome datasets to investigate the molecular mechanisms of IgE-mediated atopic diseases." A data-curation / web-tool paper.
Artifacts:
- Code: https://github.com/BenaroyaResearch/gxbrowser — the GXB web application (Grails/Tomcat/MySQL/Mongo/R viewer stack). It is infrastructure to host & display curated expression data, NOT an analysis pipeline that produces the paper's reported numbers. Standing it up reproduces nothing about the claims → out of scope.
- Data: 33 GEO datasets (GSE...), incl. GSE87399 (brief's nominal accession), GSE8507, GSE88796 (the two used in the paper's own FC validation).
What is pipeline-derived (in scope)
The paper's only computational result is the fold-change (FC) ranking GXB performs: for a user-defined two-group comparison (experimental vs control), each gene's mean fold-change is computed in linear scale, and genes are ranked by FC. The authors validate this against published literature FCs with two concrete data points (Figure 2 / text):
| # | dataset | comparison | gene | GXB FC (reported) | literature FC | lit source |
|---|---|---|---|---|---|---|
| C1 | GSE8507 (PBMC) | Job's syndrome vs healthy controls | CD151 | 1.7 | 2.0 | Holland et al. 2007 |
| C2 | GSE88796 | egg allergic vs tolerant controls | CEACAM1 | 1.69 | 1.6 | Kosoy et al. 2016 |
In-scope reproduction (P16 = reimplement the described method on the paper's own data, equally valid): download the GEO series matrices, reconstruct the two-group comparison, compute the mean linear fold-change for CD151 / CEACAM1, and compare to the GXB-reported 1.7 / 1.69. This directly tests whether the paper's headline validation numbers are derivable from the public data.
Secondary (descriptive curation counts)
The collection-level numbers (33 datasets, 1,860 transcriptome profiles, 3 RNA-seq
- 12 microarray platforms, 20 in-vitro + 13 ex-vivo, 7 disease categories) are manual-curation tallies, not pipeline outputs. Verifiable in principle by summing Table 1, but not a bioinformatic pipeline result → noted, lightly checked if Table 1 is machine-readable, not the focus.
Out of scope (not attempted)
- Standing up the GXB Grails web app (no claim depends on it; obsolete stack: Tomcat 6 / Java 7 / Grails 2.1 / Mongo 2 / MySQL 5.1).
- Re-curating all 33 datasets.
- The interactive browser features (URL sharing, KEGG filtering, export).
Known reproduction risk (the hard 20%)
Both validation datasets are multi-group / multi-condition (GSE8507: 141 samples, PMN+PBMC × stim × timepoint; GSE88796: 132 samples, 3 phenotype groups × ±egg-stim). The paper does not pin the exact sample subset (which timepoint / stimulation state / which "tolerant control" group) behind each FC. Group definition is therefore the main uncertainty; computed FC is reported per the most natural grouping with the assumption flagged.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
A data-curation/web-tool paper whose only computational outputs are two literature-validation fold-changes; the linked repo is just the viewer app and produces no claim value. C1 (CD151/GSE8507) reproduces within tolerance (1.81 vs 1.7, direction up-in-Job's holds). C2 (CEACAM1/GSE88796) is bracketed by one of three probes (1.68–1.74 vs 1.69) but the paper pins neither the sample subset nor probe selection, so it is derivable yet not uniquely reproducible. Deviations are small and sit on the input/grouping side (our self-chosen cohort + the paper's underspecification), not in core computation; no fabrication, overall a solid but not clean-1:1 reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.