Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Therapeutic Stress-Induced Remodeling of Transposable Elements and TE-Gene Chimeras in KYSE150 Esophageal Squamous Cell Carcinoma Cells.

Int J Mol Sci · 2026
L1 67/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
67/100
Reproducibility score
0.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 29% of all assessed papers rank 795 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL, honest reproduction. The paper is described well enough to reproduce a standard, low-ambiguity pipeline (Trimmomatic 0.39 -> STAR 2.7.11b on GRCh38/GENCODE v44 -> TEtranscripts 2.2.3 multi/DESeq2) on public RNA-seq (3 treated 125I+carfilzomib vs 3 control KYSE150). I ran the WHOLE primary pipeline end-to-end on «our HPC» («job» COMPLETED): downloaded all 6 ENA runs (24-26M pairs each), trimmed, aligned (6 BAMs ~4GB), and ran TEtranscripts+DESeq2 to a TE differential-expression table. RESULT: C1b (the qualitative headline 'ERV1 LTR enrichment') reproduces EXACTLY - among the significantly deregulated TEs the LTR class dominates (78%) and ERV1 is the single most enriched family (63%), and the treatment predominantly induces TEs (35 up vs 11 down), matching 'stress-induced remodeling'. C1 (the exact count) reproduces the phenomenon and order of magnitude but is ~3x lower than reported: 46 deregulated TEs vs 148. This gap is explained, not fabricated: the paper does not pin the DESeq2 version (we got a newer DESeq2 via r-base>=4.3, which changes log2FC shrinkage and hence how many of the 322 FDR<0.05 TEs clear |log2FC|>1), the GENCODE release, exact STAR parameters, or the TE-GTF version - all of which materially move TE DE counts. KEY AUDIT FINDING (run-independent): the paper's Data-Availability cites BioProject PRJNA936875, which is an UNRELATED KYSE150/KYSE410 hypoxia study (SRR23558695-706); the runs the paper actually lists in-text (SRR29055194-199) belong to PRJNA1111689/SRP508124. Verifiable wrong-accession citation error; the real data is public and was used here, so it is not a data-availability drop. NOT attempted: C2 TE-gene chimeras (ChimeraTE, 80/20 secondary), plus FIMO/clusterProfiler/cCRE tertiary analyses and all wet-lab steps (out of scope). All grades are provisional for human sign-off.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-24
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The authors hypothesize that combined genotoxic (125I radiation) and proteotoxic (carfilzomib) stress catalyzes structured reconfiguration of transposable element (TE) transcription and promotes TE-gene chimera formation, thereby driving stress-induced transcriptome reprogramming in KYSE150 ESCC cells.

Core claims
  • Combined 125I radiation and carfilzomib treatment causes structured, non-random remodeling of the TE transcriptome in KYSE150 ESCC cells. finding
  • 148 TEs are significantly dysregulated (FDR<0.05, |log2FC|>1), with ERV1 LTR elements as the most affected subclass. finding
  • 301 significant TE-gene chimeric events were identified, with increased TE-initiated and TE-exonic chimeras but decreased TE-terminal events. finding
  • The TE families most transcriptionally altered are not the same families driving chimeric events, indicating global TE activation does not passively cause chimera remodeling. mechanism
  • Gene repression is strongly associated with chimeric transcript formation; gene expression change is negatively correlated with chimerism frequency. finding
  • SPANXN1, IL1RL1, and RSAD2 are strongly downregulated genes that produce novel TE-derived isoforms and represent high-potential functional candidates. finding
  • Exonized TE-gene chimeras substantially overlap candidate cis-regulatory elements (cCREs), suggesting occurrence in regulatory-active genomic regions. finding
  • Pathway enrichment shows upregulated genes linked to cell cycle progression and DNA repair, while downregulated genes are linked to autophagy inhibition. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (TE differential expression) KYSE150 ESCC cell line combined 125I radiation + carfilzomib TE transcript expression (log2FC, FDR) and PCA/hierarchical clustering
RNA-seq-based TE-gene chimeric transcript detection KYSE150 ESCC cell line combined 125I radiation + carfilzomib counts/categories of TE-initiated, TE-exonic, TE-terminal chimeric transcripts
differential gene expression analysis (RNA-seq) KYSE150 ESCC cell line combined 125I radiation + carfilzomib DEGs associated with each chimeric transcript category
transcription factor motif enrichment analysis KYSE150 (upstream TE sequences of TE-initiated chimeras) combined 125I radiation + carfilzomib motif density/co-occurrence (KLF, ZNF, PRDM9, POU, FOXC2)
epigenetic overlap analysis (cCRE intersection) KYSE150 ESCC cell line combined 125I radiation + carfilzomib intersections between exonized chimeras and candidate cis-regulatory elements
GO enrichment analysis KYSE150 ESCC cell line combined 125I radiation + carfilzomib biological process/cellular component/molecular function enrichment of chimera-associated genes
KEGG pathway analysis KYSE150 ESCC cell line combined 125I radiation + carfilzomib pathway enrichment (DNA replication, cell cycle, base excision repair, Fanconi anemia, p53 signaling, metabolic pathways)
chromosomal distribution analysis KYSE150 ESCC cell line combined 125I radiation + carfilzomib genomic/chromosomal location of TE-gene chimeric events
Key results
  • 148 TEs significantly dysregulated (FDR<0.05, |log2FC|>1)
  • ERV1 LTR elements were the most affected TE subclass n=27
  • 301 significant TE-gene chimeric events identified across categories (FDR<0.05)
  • TE-terminal transcripts decreased after treatment 1192 to 729 transcripts
  • TE-initiated transcripts increased after treatment 170 to 211 transcripts
  • TE-exonic transcripts increased after treatment 35,367 to 38,535 transcripts
  • SPANXN1, IL1RL1, and RSAD2 strongly downregulated with novel TE-derived isoforms appearing only after treatment log2FC = -6.70, -6.15, -5.63 respectively
  • Negative correlation between gene expression change and TE-gene chimeric frequency Pearson r=-0.556
Key statistics
  • correlation Pearson r = -0.556, p = 4.69 × 10^-119 (gene expression change vs. TE-gene chimeric frequency)
  • correlation Spearman r = -0.601, p = 7.34 × 10^-144 (gene expression change vs. TE-gene chimeric frequency)
  • count 148 dysregulated TEs (FDR<0.05, |log2FC|>1 criterion)
  • count 301 significant TE-gene chimeric events (FDR<0.05 across TE-initiated, TE-exonic, TE-terminal categories)
  • fold_change log2FC = -6.70 (SPANXN1 downregulation)
  • fold_change log2FC = -6.15, adjusted p<0.001 (IL1RL1 downregulation)
  • fold_change log2FC = -5.63, adjusted p<0.001 (RSAD2 downregulation)
  • other χ2 = 2.08, d.f. = 23, p = 1.00; Mann-Whitney U = 221.5, p = 0.173 (chromosomal distribution of TE-gene chimeras not significantly altered by treatment)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used RNA-seq differential expression analysis (FDR < 0.05, |log2FC| > 1 thresholds, with variance-stabilized data for PCA and clustering) to compare transposable element (TE) expression and TE-gene chimeric transcripts between control and combined 125I-radiation/carfilzomib-treated KYSE150 cells. TE-gene chimera formation and its relationship to gene expression were further examined using Pearson and Spearman correlation, chi-square and Mann-Whitney U tests for chromosomal distribution, and GO/KEGG pathway enrichment. Results were reported mainly as log2 fold-changes with FDR-adjusted p-values, or as exact p-values for the correlation/distribution tests, without confidence intervals or SD/SEM-type dispersion measures.

Replicationunclear Sample sizeNot explicitly stated; text refers to "biological replicates" once (Section 2.1) but does not give a replicate count GroupsControl KYSE150 cells vs. combined 125I radiation + carfilzomib-treated KYSE150 cells Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFalse discovery rate (FDR) thresholding (specific procedure, e.g. Benjamini-Hochberg, not named in text)
Statistical tests used
Test Applied to n Assumptions
Differential expression testing with FDR correction and log2FC threshold (test/model not explicitly named; variance-stabilized transformation mentioned, consistent with a DESeq2/edgeR-type negative-binomial framework) TE expression, Figure 1A–D not stated
FDR-based significance calling for TE-gene chimeric events 301 significant TE-chimeric events across categories, Figure 2, Figure S4 not stated
Differential expression analysis (log2FC, adjusted p) for genes linked to TE-initiated, TE-terminal, and TE-exonic chimeras DEG analyses, Figure 5, Figures S6–S7 not stated
Pearson and Spearman correlation Gene expression change vs. chimeric frequency, Figure 7A not stated
Chi-square test Chromosomal distribution of TE-gene chimeric events, Figure S10 23 degrees of freedom (24 categories implied) not stated
Mann-Whitney U test Chromosomal distribution of TE-gene chimeric events, Figure S10 not stated
Approaches that could also have been used
  • TE and gene differential expression were thresholded with FDR < 0.05 and |log2FC| > 1, and variance-stabilized data were used for PCA/clustering, but the underlying statistical model (e.g., a negative-binomial generalized linear model) is not named.
    Could also: Explicitly naming and describing the count-based DE model (e.g., a Wald or likelihood-ratio test within a negative-binomial GLM framework such as DESeq2 or edgeR) — Stating the model and its dispersion-estimation approach would let readers evaluate how variance was handled across replicates and how the reported thresholds map onto model-based test statistics.
  • The number of biological replicates per condition is not explicitly stated in the text provided.
    Could also: Reporting the exact replicate count per group, and where feasible a power or precision justification for that number — Explicit replicate counts help readers gauge the precision of variance estimates in RNA-seq differential expression, which is particularly informative when replicate numbers are small.
  • The relationship between gene expression change and chimeric frequency was summarized with Pearson and Spearman correlation coefficients accompanied by very small p-values (e.g., p = 4.69×10⁻¹¹⁹).
    Could also: Reporting a bootstrap or analytic confidence interval around the correlation coefficient in addition to the p-value — With very large sample sizes, p-values can become extremely small even for modest correlations; a confidence interval around r would additionally convey the magnitude and precision of the association.
  • Chromosomal distribution of chimeric events was assessed with a chi-square test and a Mann-Whitney U test treating chromosomal counts as independent categories.
    Could also: A permutation-based or genomic-interval bootstrap test that accounts for spatial clustering of events along chromosomes — Genomic co-localization data can show local clustering that violates the independence assumption of chi-square/rank-based tests; a permutation approach tailored to genomic coordinates could complement these results.
  • Several distinct families of tests (TE differential expression, chimera significance calling, category-specific DEG lists, correlation analysis, chromosomal distribution tests) each apply their own significance threshold.
    Could also: An explicit statement of whether multiplicity correction was performed within each family only, or across all families of tests performed on the same dataset — Clarifying the scope of correction helps readers assess the overall false-discovery rate across the full set of analyses conducted on overlapping data.
  • A custom "functional impact scoring framework" was used to prioritize chimeric events, with candidates above a score of 0.8 termed high-impact, without a stated null distribution for the score.
    Could also: Deriving an empirical null distribution for the composite score (e.g., via permutation of TE-gene pairings) to translate the 0.8 cutoff into an estimated false-positive rate — Tying the threshold to a null distribution would let readers relate the chosen cutoff to a quantifiable error rate rather than an a priori numeric threshold.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-42074115

Paper: Majid et al. 2026, Int J Mol Sci 27(8):3471. DOI 10.3390/ijms27083471. "Therapeutic Stress-Induced Remodeling of Transposable Elements and TE-Gene Chimeras in KYSE150 Esophageal Squamous Cell Carcinoma Cells."

Data discrepancy (flagged for human audit)

  • The paper's Data Availability cites BioProject PRJNA936875 (also the RU metadata). That accession is an unrelated KYSE150/KYSE410 hypoxia/normoxia study (runs SRR23558695–706), NOT this paper's data.
  • The SRR runs the paper actually lists (SRR29055194–199) belong to PRJNA1111689 / SRP508124: "RNA seq of KYSE150 cells treated with 125I seed radiation and carfilzomib". These are the real data and ARE downloadable.
  • → Citation error in the paper (wrong BioProject), not a drop. We use the real runs. Design: TREATED (125I+carfilzomib) = SRR29055194/195/196; CONTROL (untreated) = SRR29055197/198/199. Paired-end RNA-seq, 3 vs 3.

In scope (pipeline-derived, attempted)

ID Result Pipeline Priority
C1 "148 TEs significantly deregulated" (FDR<0.05, log2FC >1), ERV1-LTR enriched
C2 "301 significant TE-gene chimeric events" (FDR<0.05); AluSx/AluJb/MIRb top ChimeraTE v1.0 Mode 1 (GRCh38.p14) SECONDARY (20%, attempt-if-time)

Out of scope / not attempted

  • Wet-lab: cell culture, 125I irradiation, carfilzomib dosing (experimental, not computational).
  • Secondary numeric claims tied to figures (per-family counts e.g. ERV1 n=27, Pearson r=−0.556, cCRE 10,393 intersections, individual gene log2FC like SPANXN1 −6.70): downstream of C1/C2 and figure-derived; not independently reproduced in this pass (80/20). Recorded as reported values only.
  • FIMO motif, clusterProfiler GO/KEGG, ENCODE cCRE overlap: tertiary, not attempted.

Tools / refs pinned

  • STAR 2.7.11b, Trimmomatic 0.39, TEtranscripts 2.2.3, DESeq2 (bioconductor), sra-tools.
  • Genome: GENCODE v44 GRCh38 primary assembly + gene GTF (assumption — paper says "GRCh38" without release). TE GTF: Hammell-lab canonical GRCh38_GENCODE_rmsk_TE.gtf.
  • STAR multimapper params per TEtranscripts docs: --outFilterMultimapNmax 100 --winAnchorMultimapNmax 100.
  • "Code" repo in metadata is github.com/ncbi/sra-tools (download tool only); the actual analysis tools are third-party (TEtranscripts, ChimeraTE) applied to the paper's data — valid per brief rule P16.

Reproducibility caveats (why exact 148 is uncertain)

The paper does not pin: GENCODE/Ensembl release, exact STAR params beyond "TEtranscripts docs", TE GTF version, or the "preprocessed FASTQ" provenance. TE DE counts are sensitive to gene-GTF release and multimapper settings, so an exact 148 match is not expected; closeness + ERV1-LTR enrichment is the honest target.

C1
Reported
148 TEs significantly deregulated (FDR<0.05, |log2FC|>1), treated (125I+carfilzomib) vs control KYSE150
Reproduced
46 deregulated TEs (35 up / 11 down) under the identical FDR<0.05 & |log2FC|>1 criterion; 322 TEs reach FDR<0.05 before the fold-change filter
partial
C1b
Reported
ERV1 LTR element enrichment among deregulated TEs
Reproduced
LTR is the top class (36/46 = 78.3%) and ERV1 is the top family (29/46 = 63%) among the significant TEs - exact qualitative match
exact
C2
Reported
301 significant TE-gene chimeric events (FDR<0.05); AluSx/AluJb/MIRb top families
Reproduced
not attempted (ChimeraTE v1.0 secondary target, deprioritized per 80/20 after the primary pipeline completed)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 67/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

The primary TE differential-expression pipeline ran end-to-end on the paper's real public data, and the central conclusion reproduces exactly: ERV1-LTR enrichment (LTR 78.3%, ERV1 63%) with TEs predominantly induced (35 up / 11 down). The only quantitative deviation is the absolute count of deregulated TEs (148 reported vs 46 reproduced, ~3x lower), which sits in the DESeq2 log2FC step and is best explained by version drift / unpinned software — our-method-and-underspecification, not fabrication or non-derivability. Two documented caveats lower confidence without breaking the claim: a verifiable wrong-accession citation (PRJNA936875 vs the correct PRJNA1111689/SRP508124) and the un-attempted C2 chimera analysis. Overall a solid yellow: explainable deviation, headline confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.