Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Predicting favorable landing pads for targeted integrations in Chinese hamster ovary cell lines by learning stability characteristics from random transgene inte

Comput Struct Biotechnol J · 2020
not yet assessed 2/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

salvaged by watchdog from agreement.json (agent omitted ROOM_RESULT.json)

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-18 ⛓ ecbfed85b09a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-06-30

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study investigates the molecular causes of transgene expression instability in CHO production cell lines by comparing genomic, transcriptomic and epigenetic profiles around random transgene integration sites in stably vs unstably expressing clones, in order to identify favorable genomic landing pads for targeted transgene integration.

Core claims
  • Expression stability in CHO cell lines is controlled at three levels: choice of integration site, integrity/concatemerization pattern of the transgene, and stress-related cellular processes. finding
  • Genome-wide favorable and unfavorable genomic loci for targeted transgene integration can be predicted by learning stability characteristics from random integration events. resource
  • Unfavorable (high copy number, unstable) cell lines show upregulation of gene sets associated with apoptosis, cell signaling and extracellular matrix components. finding
  • Genes associated with glucose metabolism promoting Akt signaling are enriched in all expressing cell lines compared to the host. mechanism
  • Exact transgene integration sites can be identified by targeted locus amplification sequencing (TLA-Seq) and characterized for surrounding genomic, transcriptomic and epigenetic patterns. method
  • Genomic variability across the genome can be scored using an M-value based on median absolute deviation of variant counts in 2 kb bins to screen for landing pads. method
  • Stable-expression integration sites are located within transcriptionally active regions (intergenic or intronic regions of expressed genes), consistent with prior lentiviral studies. finding
Experimental setups
Assay System Perturbation Readout Platform
Whole genome sequencing (WGS) Horizon Discovery CHO-K1 (HD-BIOP3) GS-/- derived cell lines plus ATCC and Horizon host cell lines random transgene integration (mAb + GS expression plasmid); host controls none genomic variants (SNPs, INDELs), structural variations, genomic variability Illumina HiSeq 2x150bp; NEBNext Ultra DNA Library Prep Kit
RNA-Seq (bulk, stranded total RNA) Six phenotype CHO cell lines (low/med/high copy, stable/unstable) and three host cell lines, biological triplicates random transgene integration vs host transcriptome / gene expression profiles Illumina HiSeq 2x150bp; Illumina TruSeq Stranded Total RNA kit, Ribo-Zero rRNA depletion
Targeted locus amplification sequencing (TLA-Seq) 13 Horizon CHO cell lines (6 phenotyped at P1 and P10, plus 7 additional) random transgene integration transgene integration sites, estimated copy number, structural changes, transgene genetic alterations Cergentis TLA pipeline (NGS)
Digital Droplet PCR (ddPCR) copy number CHO-K1 GS-/- clones expressing >1 g/L random transgene integration transgene (heavy/light chain) copy number normalized to GcgR housekeeping gene Bio-Rad QX200; DNeasy Blood & Tissue Kit
Titer / expression stability assay CHO clones over 10 passages without selective pressure none (passaging without selection) antibody titer change (ΔTiter %) from passage 1 to passage 10 Octet (ForteBio) Protein A
Fluorescence-based colony screening CHO-K1 GS-/- transfectant single colonies in methylcellulose random transgene integration secreted antibody (protein G Alexa Fluor 488 fluorescence), colony size/shape ClonePix FL (Molecular Devices)
Key results
  • Upregulation of apoptosis, cell signaling and extracellular matrix gene sets in unfavorable high-copy unstable cell lines
  • Enrichment of glucose metabolism genes promoting Akt signaling in all expressing cell lines vs host
  • Stable-expression transgene integration sites located within transcriptionally active intergenic or intronic regions
  • Genomic rearrangements at integration sites associated with loss of transgene copies over time in unstable lines
Key statistics
  • other >25% titer change = unstable; <25% = stable (threshold defining expression stability over passages 1-10)
  • count Low = 1-3 copies, Medium = 4-15 copies, High >= 15 copies (transgene copy number categories by ddPCR)
  • other M-value > 0.7 (threshold for high genomic variability regions in 2 kb bins when screening landing pads)
  • count 13 (Horizon CHO cell lines analyzed by TLA-Seq)
  • count over 70% (proportion of recombinant biopharmaceuticals produced in CHO cells)
  • count 84% (proportion of monoclonal antibodies produced in CHO expression systems)
  • other $140 billion (2013), $188 billion (2017) (estimated total biopharmaceutical sales)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper compares Chinese hamster ovary (CHO) cell line clones grouped by transgene copy number (low/medium/high) and by expression stability (stable vs. unstable, defined by a fixed percentage drop in titer over 10 passages) using whole-genome sequencing, RNA-seq, and targeted locus amplification sequencing (TLA-Seq). Results are reported primarily as classifications (stable/unstable, high/low genomic variability) and as bioinformatic pipeline outputs (variant calls, structural variants, integration site mapping) rather than through a described classical hypothesis-testing framework. The excerpted methods do not include a dedicated 'statistical analysis' section covering group-comparison tests, effect sizes, or multiple-testing correction for the RNA-seq or genomic comparisons.

Replicationmixed Sample sizeOne subclone per phenotype category (six phenotypes: low/medium/high copy number x stable/unstable) was selected for WGS and TLA-Seq; RNA-seq was performed in biological triplicate per phenotype and per host cell line. No formal sample-size or power calculation is described in the provided text. GroupsStable vs unstable transgene expression across low/medium/high copy-number phenotypes, plus comparisons to host (non-producing) cell lines Pairingunclear Randomization/blindingnot stated Dispersionnone
Statistical tests used
Test Applied to n Assumptions
Fixed-threshold classification of titer change (ΔTiter = [Avg(P1,P2)-Avg(P9,P10)]/Avg(P1,P2) x100; <25% = stable, >25% = unstable) Classifying each clone as stable vs unstable transgene expression not stated
Median Absolute Deviation (MAD)-based M-value scoring with a threshold of 0.7 Identifying high vs low genomic variability regions across 2 kb genome-wide bins not stated
GATK hard-filtering thresholds (QD, FS, MQ, MQRankSum, ReadPosRankSum, SOR) SNP/INDEL variant calling from WGS data na not stated
Approaches that could also have been used
  • Cell lines were classified as 'stable' or 'unstable' using a fixed 25% titer-change cutoff between early and late passages.
    Could also: Treating percentage titer change as a continuous outcome in a regression or mixed-effects model (with copy number as a covariate) rather than dichotomizing at a single threshold. — A continuous-outcome approach preserves statistical power that dichotomization can lose and allows assessment of whether downstream findings are sensitive to the specific cutoff chosen.
  • Genomic variability regions were defined using a MAD-based M-value with a threshold of 0.7, without a stated formal significance test.
    Could also: A permutation-based or Benjamini-Hochberg false-discovery-rate framework applied to the bin-level variability scores. — This would provide an explicit multiple-testing-aware significance criterion for calling 'high variability' bins across the many thousands of genome-wide bins being screened simultaneously.
  • Outcomes were compared descriptively across six phenotype categories defined by two crossed factors (copy number: low/medium/high; stability: stable/unstable).
    Could also: A two-way ANOVA or generalized linear model with copy number and stability (and their interaction) as factors. — A factorial model would formally test main effects and any interaction between copy number and stability on measured outcomes, which a purely descriptive category-by-category comparison does not directly quantify.
  • RNA-seq was generated in biological triplicate per phenotype, but the excerpted methods do not describe a differential-expression statistical model.
    Could also: Standard RNA-seq differential expression tools such as DESeq2 (Wald test) or edgeR (negative-binomial GLM), paired with FDR correction (e.g., Benjamini-Hochberg). — These methods are designed for small-n triplicate RNA-seq designs and provide variance-stabilized effect size estimates, p-values, and multiple-testing correction for gene-level comparisons between stable and unstable phenotypes.
  • A single subclone was sequenced (WGS, TLA-Seq) per phenotype category rather than multiple independent subclones within each category.
    Could also: Sequencing several independent subclones per phenotype group. — This would allow estimation of clone-to-clone (biological) variance within each stability/copy-number group, supporting formal inferential statistical comparisons in addition to single-sample profiling.
Software: Trimmomatic 0.36 · FastQC 0.11.5 · BWA (mem) 0.7.17 · Picard tools 2.3.0 · Qualimap 2.2.1 · GATK HaplotypeCaller 3.8 · Manta 1.6.0 · Delly 0.7.9 · Gviz (R package) · bcl2fastq 2.17

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-33304461

Paper: Dhiman et al. 2020, Comput Struct Biotechnol J 18:3632. "Predicting favorable landing pads for targeted integrations in CHO cell lines by learning stability characteristics from random transgene integrations." DOI 10.1016/j.csbj.2020.11.008. Code: https://github.com/hd4git/StabilityAnalysis (own code; master @ 2020-07-13). Data: ENA PRJEB39258 — 27 RNA-Seq + 8 WGS runs (paired Illumina HiSeq).

Study design

6 phenotyped IgG-producer CHO clones (G9, 5G10, 1E3, 1C11, 6A6, 6H1) spanning low/medium/high copy × stable/unstable, plus 3 host lines (C1835A=ATCC CHO-K1, C3234A=Horizon HD-BIOP3, PE24=Horizon process-evolved). Each profiled by WGS and RNA-Seq (triplicate). Reference genome CriGri-PICR (GCF_003668045.1).

Pipeline-derived results (IN SCOPE)

# Result Pipeline Reproducibility
R1 RNA-seq DE expressing-vs-host: 806 up / 485 down (§3.4) Trimmomatic→Hisat2→htseq-count→DESeq2 lfcShrink(apeglm), |LFC|>0.585 & padj<0.01 data+code shipped → ATTEMPTED
R2 RNA-seq DE favorable-vs-unfavorable: 7 up fav / 50 up unfav (§3.5) same counts matrix, favorability_unfav_vs_fav ATTEMPTED
R3 WGS structural variants total: 250 INV / 1261 INS / 2403 DEL / 256 DUP / 366 TRANS (§3.2) BWA→GATK + Manta/Delly/Lumpy consensus HEAVY — attempted after R1/R2
R4 Variability regions (M-value>0.7 over 2kb bins) wgs_analysis + mval scripts depends on R3 alignments
R5 Landing pads: 7,166 favorable / 16,048 unfavorable regions (§3.6) combine DE peaks + variability + histone marks partially out of scope (needs external histone-mark data)

OUT OF SCOPE (not attempted)

  • TLA-Seq integration-site profiling (§3.3, Table 3): the TLA-Seq raw data is NOT in PRJEB39258 (deposit holds only RNA-Seq+WGS). Integration sites were called by Cergentis (commercial TLA service); the TLA/*.R repo scripts consume Cergentis output tables that are not shipped. → data_unavailable for this sub-result.
  • Wet-lab measurements: titres (mg/L), copy number (ddPCR), passage stability (%titre change). Manual/experimental, not pipeline-derived.
  • Histone-mark chromatin states (six marks used for landing-pad annotation) come from an external CHO epigenome resource, not this deposit → R5 only partially reproducible.

Primary target

R1 + R2 (RNA-seq DE) — fully specified, data+code shipped, one counts matrix drives both. This is the quick-minimum. R3 (WGS SV) pursued next as the harder result.

Known deviations from the authors' exact pipeline

  • The authors pre-filter reads against a confidential vector sequence (VectorA) before trimming (rnaseq/trim1.sh). That sequence is not public → step omitted. Impact on endogenous-gene counts is negligible (vector reads do not map to the CHO genome anyway).
  • htseq-count run with -r pos (coordinate-sorted BAM) vs the script's default; counts are identical, only mate-pairing buffering differs.
  • Tool minor versions pinned to match the paper where available (Hisat2 2.1.0, htseq 0.11.0, Trimmomatic 0.36, samtools 1.9); DESeq2/apeglm at current Bioconductor (paper used 1.24.0).
Figures / tables: Fig 4ATable

No individual results have been recorded for this entry yet.

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 44/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

Reproduction incomplete, our-side truncation. The agent downloaded the full RNA-seq dataset and validated the alignment->counts pipeline end-to-end (clean 27,873-gene counts), but hit its session limit at 18:00 with only 16/27 samples counted and never ran the final merge+DESeq2 — so the primary claims (806 up / 485 down; 7 fav / 50 unfav DEGs) have no reproduced value, outputs/ is empty, and ROOM_RESULT.json was watchdog-salvaged. This is resumable our-side incompletion, not an authors' defect or fabrication; the RNA-seq data+code are shipped and identical. Secondary results are genuinely limited by the deposit (TLA-seq raw data and histone marks absent, one WGS run missing), and there is a documented our-method tool substitution (STAR for hisat2+htseq-count). q8=red flags it as a critical non-result that should simply be re-run/resumed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

822 k
tokens (I/O) · 131.2 M incl. cache
176 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.