Therapeutic Stress-Induced Remodeling of Transposable Elements and TE-Gene Chimeras in KYSE150 Esophageal Squamous Cell Carcinoma Cells.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🔴A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL, honest reproduction. The paper is described well enough to reproduce a standard, low-ambiguity pipeline (Trimmomatic 0.39 -> STAR 2.7.11b on GRCh38/GENCODE v44 -> TEtranscripts 2.2.3 multi/DESeq2) on public RNA-seq (3 treated 125I+carfilzomib vs 3 control KYSE150). I ran the WHOLE primary pipeline end-to-end on «our HPC» («job» COMPLETED): downloaded all 6 ENA runs (24-26M pairs each), trimmed, aligned (6 BAMs ~4GB), and ran TEtranscripts+DESeq2 to a TE differential-expression table. RESULT: C1b (the qualitative headline 'ERV1 LTR enrichment') reproduces EXACTLY - among the significantly deregulated TEs the LTR class dominates (78%) and ERV1 is the single most enriched family (63%), and the treatment predominantly induces TEs (35 up vs 11 down), matching 'stress-induced remodeling'. C1 (the exact count) reproduces the phenomenon and order of magnitude but is ~3x lower than reported: 46 deregulated TEs vs 148. This gap is explained, not fabricated: the paper does not pin the DESeq2 version (we got a newer DESeq2 via r-base>=4.3, which changes log2FC shrinkage and hence how many of the 322 FDR<0.05 TEs clear |log2FC|>1), the GENCODE release, exact STAR parameters, or the TE-GTF version - all of which materially move TE DE counts. KEY AUDIT FINDING (run-independent): the paper's Data-Availability cites BioProject PRJNA936875, which is an UNRELATED KYSE150/KYSE410 hypoxia study (SRR23558695-706); the runs the paper actually lists in-text (SRR29055194-199) belong to PRJNA1111689/SRP508124. Verifiable wrong-accession citation error; the real data is public and was used here, so it is not a data-availability drop. NOT attempted: C2 TE-gene chimeras (ChimeraTE, 80/20 secondary), plus FIMO/clusterProfiler/cCRE tertiary analyses and all wet-lab steps (out of scope). All grades are provisional for human sign-off.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-24no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors hypothesize that combined genotoxic (125I radiation) and proteotoxic (carfilzomib) stress catalyzes structured reconfiguration of transposable element (TE) transcription and promotes TE-gene chimera formation, thereby driving stress-induced transcriptome reprogramming in KYSE150 ESCC cells.
- ★ Combined 125I radiation and carfilzomib treatment causes structured, non-random remodeling of the TE transcriptome in KYSE150 ESCC cells. finding
- ★ 148 TEs are significantly dysregulated (FDR<0.05, |log2FC|>1), with ERV1 LTR elements as the most affected subclass. finding
- ★ 301 significant TE-gene chimeric events were identified, with increased TE-initiated and TE-exonic chimeras but decreased TE-terminal events. finding
- ★ The TE families most transcriptionally altered are not the same families driving chimeric events, indicating global TE activation does not passively cause chimera remodeling. mechanism
- ★ Gene repression is strongly associated with chimeric transcript formation; gene expression change is negatively correlated with chimerism frequency. finding
- ★ SPANXN1, IL1RL1, and RSAD2 are strongly downregulated genes that produce novel TE-derived isoforms and represent high-potential functional candidates. finding
- Exonized TE-gene chimeras substantially overlap candidate cis-regulatory elements (cCREs), suggesting occurrence in regulatory-active genomic regions. finding
- Pathway enrichment shows upregulated genes linked to cell cycle progression and DNA repair, while downregulated genes are linked to autophagy inhibition. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (TE differential expression) | KYSE150 ESCC cell line | combined 125I radiation + carfilzomib | TE transcript expression (log2FC, FDR) and PCA/hierarchical clustering | — |
| RNA-seq-based TE-gene chimeric transcript detection | KYSE150 ESCC cell line | combined 125I radiation + carfilzomib | counts/categories of TE-initiated, TE-exonic, TE-terminal chimeric transcripts | — |
| differential gene expression analysis (RNA-seq) | KYSE150 ESCC cell line | combined 125I radiation + carfilzomib | DEGs associated with each chimeric transcript category | — |
| transcription factor motif enrichment analysis | KYSE150 (upstream TE sequences of TE-initiated chimeras) | combined 125I radiation + carfilzomib | motif density/co-occurrence (KLF, ZNF, PRDM9, POU, FOXC2) | — |
| epigenetic overlap analysis (cCRE intersection) | KYSE150 ESCC cell line | combined 125I radiation + carfilzomib | intersections between exonized chimeras and candidate cis-regulatory elements | — |
| GO enrichment analysis | KYSE150 ESCC cell line | combined 125I radiation + carfilzomib | biological process/cellular component/molecular function enrichment of chimera-associated genes | — |
| KEGG pathway analysis | KYSE150 ESCC cell line | combined 125I radiation + carfilzomib | pathway enrichment (DNA replication, cell cycle, base excision repair, Fanconi anemia, p53 signaling, metabolic pathways) | — |
| chromosomal distribution analysis | KYSE150 ESCC cell line | combined 125I radiation + carfilzomib | genomic/chromosomal location of TE-gene chimeric events | — |
- – 148 TEs significantly dysregulated (FDR<0.05, |log2FC|>1)
- – ERV1 LTR elements were the most affected TE subclass n=27
- – 301 significant TE-gene chimeric events identified across categories (FDR<0.05)
- ▼ TE-terminal transcripts decreased after treatment 1192 to 729 transcripts
- ▲ TE-initiated transcripts increased after treatment 170 to 211 transcripts
- ▲ TE-exonic transcripts increased after treatment 35,367 to 38,535 transcripts
- ▼ SPANXN1, IL1RL1, and RSAD2 strongly downregulated with novel TE-derived isoforms appearing only after treatment log2FC = -6.70, -6.15, -5.63 respectively
- ▼ Negative correlation between gene expression change and TE-gene chimeric frequency Pearson r=-0.556
- correlation Pearson r = -0.556, p = 4.69 × 10^-119 (gene expression change vs. TE-gene chimeric frequency)
- correlation Spearman r = -0.601, p = 7.34 × 10^-144 (gene expression change vs. TE-gene chimeric frequency)
- count 148 dysregulated TEs (FDR<0.05, |log2FC|>1 criterion)
- count 301 significant TE-gene chimeric events (FDR<0.05 across TE-initiated, TE-exonic, TE-terminal categories)
- fold_change log2FC = -6.70 (SPANXN1 downregulation)
- fold_change log2FC = -6.15, adjusted p<0.001 (IL1RL1 downregulation)
- fold_change log2FC = -5.63, adjusted p<0.001 (RSAD2 downregulation)
- other χ2 = 2.08, d.f. = 23, p = 1.00; Mann-Whitney U = 221.5, p = 0.173 (chromosomal distribution of TE-gene chimeras not significantly altered by treatment)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used RNA-seq differential expression analysis (FDR < 0.05, |log2FC| > 1 thresholds, with variance-stabilized data for PCA and clustering) to compare transposable element (TE) expression and TE-gene chimeric transcripts between control and combined 125I-radiation/carfilzomib-treated KYSE150 cells. TE-gene chimera formation and its relationship to gene expression were further examined using Pearson and Spearman correlation, chi-square and Mann-Whitney U tests for chromosomal distribution, and GO/KEGG pathway enrichment. Results were reported mainly as log2 fold-changes with FDR-adjusted p-values, or as exact p-values for the correlation/distribution tests, without confidence intervals or SD/SEM-type dispersion measures.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression testing with FDR correction and log2FC threshold (test/model not explicitly named; variance-stabilized transformation mentioned, consistent with a DESeq2/edgeR-type negative-binomial framework) | TE expression, Figure 1A–D | — | not stated |
| FDR-based significance calling for TE-gene chimeric events | 301 significant TE-chimeric events across categories, Figure 2, Figure S4 | — | not stated |
| Differential expression analysis (log2FC, adjusted p) for genes linked to TE-initiated, TE-terminal, and TE-exonic chimeras | DEG analyses, Figure 5, Figures S6–S7 | — | not stated |
| Pearson and Spearman correlation | Gene expression change vs. chimeric frequency, Figure 7A | — | not stated |
| Chi-square test | Chromosomal distribution of TE-gene chimeric events, Figure S10 | 23 degrees of freedom (24 categories implied) | not stated |
| Mann-Whitney U test | Chromosomal distribution of TE-gene chimeric events, Figure S10 | — | not stated |
-
TE and gene differential expression were thresholded with FDR < 0.05 and |log2FC| > 1, and variance-stabilized data were used for PCA/clustering, but the underlying statistical model (e.g., a negative-binomial generalized linear model) is not named.↳ Could also: Explicitly naming and describing the count-based DE model (e.g., a Wald or likelihood-ratio test within a negative-binomial GLM framework such as DESeq2 or edgeR) — Stating the model and its dispersion-estimation approach would let readers evaluate how variance was handled across replicates and how the reported thresholds map onto model-based test statistics.
-
The number of biological replicates per condition is not explicitly stated in the text provided.↳ Could also: Reporting the exact replicate count per group, and where feasible a power or precision justification for that number — Explicit replicate counts help readers gauge the precision of variance estimates in RNA-seq differential expression, which is particularly informative when replicate numbers are small.
-
The relationship between gene expression change and chimeric frequency was summarized with Pearson and Spearman correlation coefficients accompanied by very small p-values (e.g., p = 4.69×10⁻¹¹⁹).↳ Could also: Reporting a bootstrap or analytic confidence interval around the correlation coefficient in addition to the p-value — With very large sample sizes, p-values can become extremely small even for modest correlations; a confidence interval around r would additionally convey the magnitude and precision of the association.
-
Chromosomal distribution of chimeric events was assessed with a chi-square test and a Mann-Whitney U test treating chromosomal counts as independent categories.↳ Could also: A permutation-based or genomic-interval bootstrap test that accounts for spatial clustering of events along chromosomes — Genomic co-localization data can show local clustering that violates the independence assumption of chi-square/rank-based tests; a permutation approach tailored to genomic coordinates could complement these results.
-
Several distinct families of tests (TE differential expression, chimera significance calling, category-specific DEG lists, correlation analysis, chromosomal distribution tests) each apply their own significance threshold.↳ Could also: An explicit statement of whether multiplicity correction was performed within each family only, or across all families of tests performed on the same dataset — Clarifying the scope of correction helps readers assess the overall false-discovery rate across the full set of analyses conducted on overlapping data.
-
A custom "functional impact scoring framework" was used to prioritize chimeric events, with candidates above a score of 0.8 termed high-impact, without a stated null distribution for the score.↳ Could also: Deriving an empirical null distribution for the composite score (e.g., via permutation of TE-gene pairings) to translate the 0.8 cutoff into an estimated false-positive rate — Tying the threshold to a null distribution would let readers relate the chosen cutoff to a quantifiable error rate rather than an a priori numeric threshold.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-42074115
Paper: Majid et al. 2026, Int J Mol Sci 27(8):3471. DOI 10.3390/ijms27083471. "Therapeutic Stress-Induced Remodeling of Transposable Elements and TE-Gene Chimeras in KYSE150 Esophageal Squamous Cell Carcinoma Cells."
Data discrepancy (flagged for human audit)
- The paper's Data Availability cites BioProject PRJNA936875 (also the RU metadata). That accession is an unrelated KYSE150/KYSE410 hypoxia/normoxia study (runs SRR23558695–706), NOT this paper's data.
- The SRR runs the paper actually lists (SRR29055194–199) belong to PRJNA1111689 / SRP508124: "RNA seq of KYSE150 cells treated with 125I seed radiation and carfilzomib". These are the real data and ARE downloadable.
- → Citation error in the paper (wrong BioProject), not a drop. We use the real runs. Design: TREATED (125I+carfilzomib) = SRR29055194/195/196; CONTROL (untreated) = SRR29055197/198/199. Paired-end RNA-seq, 3 vs 3.
In scope (pipeline-derived, attempted)
| ID | Result | Pipeline | Priority |
|---|---|---|---|
| C1 | "148 TEs significantly deregulated" (FDR<0.05, | log2FC | >1), ERV1-LTR enriched |
| C2 | "301 significant TE-gene chimeric events" (FDR<0.05); AluSx/AluJb/MIRb top | ChimeraTE v1.0 Mode 1 (GRCh38.p14) | SECONDARY (20%, attempt-if-time) |
Out of scope / not attempted
- Wet-lab: cell culture, 125I irradiation, carfilzomib dosing (experimental, not computational).
- Secondary numeric claims tied to figures (per-family counts e.g. ERV1 n=27, Pearson r=−0.556, cCRE 10,393 intersections, individual gene log2FC like SPANXN1 −6.70): downstream of C1/C2 and figure-derived; not independently reproduced in this pass (80/20). Recorded as reported values only.
- FIMO motif, clusterProfiler GO/KEGG, ENCODE cCRE overlap: tertiary, not attempted.
Tools / refs pinned
- STAR 2.7.11b, Trimmomatic 0.39, TEtranscripts 2.2.3, DESeq2 (bioconductor), sra-tools.
- Genome: GENCODE v44 GRCh38 primary assembly + gene GTF (assumption — paper says "GRCh38" without release). TE GTF: Hammell-lab canonical GRCh38_GENCODE_rmsk_TE.gtf.
- STAR multimapper params per TEtranscripts docs: --outFilterMultimapNmax 100 --winAnchorMultimapNmax 100.
- "Code" repo in metadata is github.com/ncbi/sra-tools (download tool only); the actual analysis tools are third-party (TEtranscripts, ChimeraTE) applied to the paper's data — valid per brief rule P16.
Reproducibility caveats (why exact 148 is uncertain)
The paper does not pin: GENCODE/Ensembl release, exact STAR params beyond "TEtranscripts docs", TE GTF version, or the "preprocessed FASTQ" provenance. TE DE counts are sensitive to gene-GTF release and multimapper settings, so an exact 148 match is not expected; closeness + ERV1-LTR enrichment is the honest target.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The primary TE differential-expression pipeline ran end-to-end on the paper's real public data, and the central conclusion reproduces exactly: ERV1-LTR enrichment (LTR 78.3%, ERV1 63%) with TEs predominantly induced (35 up / 11 down). The only quantitative deviation is the absolute count of deregulated TEs (148 reported vs 46 reproduced, ~3x lower), which sits in the DESeq2 log2FC step and is best explained by version drift / unpinned software — our-method-and-underspecification, not fabrication or non-derivability. Two documented caveats lower confidence without breaking the claim: a verifiable wrong-accession citation (PRJNA936875 vs the correct PRJNA1111689/SRP508124) and the un-attempted C2 chimera analysis. Overall a solid yellow: explainable deviation, headline confirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.