Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Cell type- and species-specific regulation of hepatic lncRNAs by TCDD-activated aryl hydrocarbon receptor.

Sci Rep · 2025
L1 71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the bulk RNA-seq DE analysis 1:1 from SHIPPED inputs. The repo commits the STAR ReadsPerGene count matrices (mouse 40 smp / GSE203302, rat 24 smp / GSE110293), so alignment was correctly skipped; we ran DESeq2 on them with the documented DE rule (|FC|>=1.5 & padj<=0.05, union over dose-vs-0 contrasts). The only non-shipped input, the sample->dose metadata, was reconstructed and VERIFIED exact against GEO. MOUSE reproduces near-exactly: the method-pure union DE-gene set is 8460 vs the paper's 2386+6071=8457 (within 3 genes, 0.04%), and the lncRNA/mRNA split (2349/5944) is within ~2% of 2386/6071 — the small residual is from reconstructing the biotype map (authors' custom gene->Type annotation is NOT committed) using RefSeq instead. RAT is partial: total DEGs 4784 vs reported 3972 (+20%; lncRNA 1014 vs 916, mRNA 3536 vs 3056), most plausibly because the committed rat counts are n=3/group while the paper text says n=5, plus weaker rn7 symbol->biotype matching. NOT attempted (hard 20% / not low-hanging): re-alignment from FASTQ (unnecessary), exact custom-GTF/biotype reconstruction (liftOver+MGI+bedtools; map file not shipped), snRNAseq cell-type DE counts (Fig 4/6), AHR ChIP-seq / pDRE enrichment, and the cross-species shared-gene overlap (203 lncRNA / 1492 mRNA). No fabrication signal for the mouse claims; rat gap reads as a reproducibility (n + annotation) issue for a human to confirm.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-15 ⛓ 031323962faf
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether TCDD-activated aryl hydrocarbon receptor (AHR) mediates dose-dependent, cell type- and species-specific differential expression of hepatic long non-coding RNAs (lncRNAs) that may contribute to the progression of steatosis to steatohepatitis with fibrosis.

Core claims
  • TCDD elicits dose-dependent AHR-mediated differential expression of hepatic lncRNAs in both mouse and rat liver, with more DE lncRNAs in mice than rats and a subset conserved across species. finding
  • DE lncRNAs show AHR genomic enrichment and putative dioxin response element (pDRE) patterns similar to mRNA-coding genes, with greater frequency proximal to the transcription start site. finding
  • lncRNA differential expression by TCDD is cell-type- and zone-specific, with pericentral and periportal hepatocytes and macrophages showing the most changes. finding
  • 52 previously annotated hepatocyte lncRNAs differentially expressed by TCDD are associated with human liver diseases including steatosis, fibrosis, and hepatocellular carcinoma. finding
  • AHR-mediated lncRNA dysregulation may play a significant role in TCDD-elicited progression of steatosis to steatohepatitis with fibrosis (MASLD-like pathology). mechanism
  • Integration of bulk and single-nuclei RNAseq with AHR ChIPseq and pDRE data provides a framework to identify lncRNAs regulated via canonical and non-canonical AHR pathways. method
  • A custom GTF gene model (20,973 mRNAs and 55,038 lncRNAs) was lifted from mm10 to mm39 and to rat rn7 to enable cross-species conserved lncRNA identification. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (re-analysis) male C57BL/6Crl mouse liver TCDD oral gavage (0.03–30 µg/kg) vs sesame oil vehicle dose-dependent differential lncRNA and mRNA expression FASTQC v0.12.1, Trimmomatic v0.39, STAR v2.7.1, DESeq2 v1.34.0 (GSE203302)
bulk RNA-seq (re-analysis) female Sprague-Dawley rat liver TCDD oral gavage (0.01–10 µg/kg) vs sesame oil vehicle dose-dependent differential lncRNA and mRNA expression STAR v2.7.1, DESeq2 v1.34.0 (GSE110293)
single-nuclei RNA-seq (re-analysis) mouse liver, cell-type resolved (PC/PP hepatocytes, macrophages, LSECs, HSCs) TCDD oral gavage (0.01–30 µg/kg) every 4 days for 28 days vs vehicle cell-type- and zone-specific DE lncRNA/mRNA (pseudobulk) 10X Genomics Chromium Single Cell 3' v3.1, CellRanger v8.0.1, Scater v1.22.0, DESeq2 v1.34.0, ScanPy v1.10.0 (GSE148339)
AHR ChIP-seq (re-analysis) mouse liver TCDD 30 µg/kg, 2-hour exposure (AHR vs IgG antibody) AHR genomic enrichment peaks (FDR ≤ 0.05) within 10 kb of TSS Bowtie/CisGenome, liftOver to mm39 (GSE97634)
AHR ChIP-seq (re-analysis) rat liver TCDD 10 µg/kg, 2-hour exposure (AHR vs IgG antibody) AHR genomic enrichment peaks (FDR ≤ 0.05) within 10 kb of TSS Bowtie/CisGenome, liftOver to rn7 (GSE110282)
pDRE computational identification mouse (mm39) and rat (rn7) genomes none pDRE core 5'-GCGTG-3' matrix similarity score distribution relative to TSS position weight matrix (PWM)
K-means clustering mouse and rat bulk RNAseq DE genes none clustering of pseudobulk log2 fold-change into 5 groups scikit-learn v0.24.2
Key results
  • Mouse had 2,386 DE lncRNAs and rat had 916 DE lncRNAs, with 203 common to both species 2,386 mouse; 916 rat; 203 common
  • Mouse had 6,071 DE mRNAs and rat had 3,056 DE mRNAs, with 1,492 in common 6,071 mouse; 3,056 rat; 1,492 common
  • snRNAseq revealed 5,495 DE lncRNAs across all liver cell subtypes 5,495
  • Periportal hepatocytes exhibited 3,654 DE lncRNAs and 5,722 DE mRNAs; pericentral hepatocytes 3,463 DE lncRNAs and 5,372 DE mRNAs PP 3,654/5,722; PC 3,463/5,372
  • Macrophages exhibited 2,161 DE lncRNAs (213, ~9.9%, with pDREs in AHR-enriched regions) and 5,890 DE mRNAs (1,135, 19.3%, with pDREs in AHR-enriched regions) 2,161 lncRNAs; 213 (9.9%); 5,890 mRNAs; 1,135 (19.3%)
  • 52 previously annotated hepatocyte lncRNAs were differentially expressed by TCDD, many associated with steatosis, fibrosis, and HCC 52
  • pDREs and AHR genomic enrichment frequency increased proximal to the TSS for both lncRNAs and mRNAs in mouse and rat
  • TCDD generally repressed lncRNA expression at the 30 µg/kg dose, with most DE changes occurring at 10 and 30 µg/kg
Key statistics
  • count 2,386 mouse DE lncRNAs vs 916 rat DE lncRNAs, 203 common (cross-species bulk RNAseq DE lncRNAs)
  • count 6,071 mouse DE mRNAs vs 3,056 rat DE mRNAs, 1,492 common (cross-species bulk RNAseq DE mRNAs)
  • count 5,495 (DE lncRNAs across all liver cell subtypes (snRNAseq))
  • count PP 3,654 / PC 3,463 DE lncRNAs; PP 5,722 / PC 5,372 DE mRNAs (periportal and pericentral hepatocyte DE genes)
  • count 2,161 DE lncRNAs; 213 (~9.9%) with pDRE+AHR; 5,890 DE mRNAs; 1,135 (19.3%) with pDRE+AHR (macrophage DE genes and AHR enrichment overlap)
  • fold_change |fold-change| ≥ 1.5, adjusted p-value ≤ 0.05 (differential expression cutoff criteria (bulk n=5/group))
  • other MSS ≥ 0.861 (mm39); MSS ≥ 0.825 (rn7) (pDRE matrix similarity score thresholds by species)
  • other macrophages comprise ~30% of liver cell population after TCDD (immune cell infiltration)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper re-analyzes three previously published in vivo transcriptomic datasets (bulk RNAseq from mouse and rat liver; snRNAseq from mouse liver) to characterize dose-dependent, cell-type-specific lncRNA differential expression following TCDD treatment. Differential expression was assessed with DESeq2 (using a combined fold-change and adjusted p-value threshold), and dose-response patterns were summarized by K-means clustering into five groups. Results were reported primarily as log2 fold-change heatmaps, cluster trend lines, and Venn diagram overlaps between species and cell types.

Replicationbiological Sample sizeBulk RNAseq: n=5 per group explicitly stated. snRNAseq: n=2–3 biological replicates per dose group as stated. No formal power analysis described. GroupsVehicle control vs. 7–8 TCDD dose levels per species; mouse vs. rat species comparison; cross-cell-type comparison within snRNAseq Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionAdjusted p-value as produced by DESeq2 (Benjamini-Hochberg FDR by default); method name not explicitly stated in text
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test (negative binomial GLM) Bulk RNAseq differential expression: vehicle vs. each TCDD dose, mouse liver (GSE203302) n=5 biological replicates per group not stated
DESeq2 Wald test (negative binomial GLM) Bulk RNAseq differential expression: vehicle vs. each TCDD dose, rat liver (GSE110293) n=5 biological replicates per group not stated
DESeq2 Wald test on pseudobulk aggregates (via Scater) snRNAseq differential expression per cell type: vehicle vs. each TCDD dose (GSE148339) n=2–3 biological replicates per dose group not stated
K-means clustering (scikit-learn) Grouping dose-response log2FC trajectories of DE lncRNAs and mRNAs in mouse and rat bulk RNAseq; k fixed at 5 na
FDR threshold (source: CisGenome peak caller, lifted-over coordinates) AHR ChIPseq enriched peak selection: FDR ≤ 0.05 (mouse GSE97634, rat GSE110282) n=5 per group (AHR and IgG antibodies) not stated
Position weight matrix scoring (matrix similarity score threshold) pDRE identification genome-wide: MSS ≥ 0.861 (mm39), MSS ≥ 0.825 (rn7) not stated
Approaches that could also have been used
  • The number of K-means clusters was fixed at k=5 to facilitate comparison with a prior mRNA analysis, without a data-driven selection of k.
    Could also: The optimal number of clusters could also be determined empirically using the silhouette score, gap statistic, or within-cluster sum of squares elbow plot before fixing k. — Data-driven k selection would provide an independent check on whether five clusters is the most parsimonious grouping for the lncRNA dose-response trajectories, and could reveal whether the lncRNA data structure differs from the mRNA data structure used to motivate k=5.
  • Species overlap of differentially expressed lncRNAs and mRNAs was summarized descriptively via Venn diagrams (counts of shared DE genes).
    Could also: A hypergeometric test or Fisher's exact test applied to the overlap could also quantify whether the observed sharing exceeds chance expectation given the sizes of the DE gene lists and the total annotated gene universe. — Formal testing of overlap significance would distinguish biologically meaningful conservation of response from overlap that could arise by chance when large numbers of genes are DE in both species.
  • Differential expression in snRNAseq was assessed using a pseudobulk approach (Scater + DESeq2) with n=2–3 replicates per dose group.
    Could also: Mixed-effects models (e.g., NEBULA, glmmTMB) or alternative pseudobulk frameworks (e.g., edgeR with quasi-likelihood) could also handle the small and unequal replicate numbers per dose. — With only n=2–3 replicates, variance estimation in DESeq2 relies heavily on empirical Bayes shrinkage; mixed-effects approaches that explicitly model both within- and between-replicate variance can be compared as a sensitivity check, particularly for cell types with sparse counts.
  • Differential expression across all dose levels was assessed by treating each dose independently against vehicle rather than fitting a single dose-response model.
    Could also: A monotone trend test (e.g., likelihood-ratio test across ordered doses in DESeq2, or ImpulseDE2) could also model the full dose-response curve simultaneously. — A model-based dose-response approach borrows information across dose levels, can improve sensitivity for genes with gradual responses, and directly characterizes the shape (monotone, non-monotone, threshold) of the response without requiring post-hoc clustering.
  • The multiplicity correction was applied within each pairwise dose-vs-vehicle comparison separately; no correction was described for the family of tests spanning multiple doses or multiple cell types.
    Could also: A global FDR correction applied across all doses within a species (or across all cell types within snRNAseq) could also be used to control the experiment-wide false discovery rate. — When hundreds of comparisons are run (e.g., 8 dose levels × multiple cell types), treating each comparison's FDR independently can result in a higher aggregate false discovery rate than the nominal 5%; a joint correction would make the aggregate error rate explicit.
  • Bar plots of log2 fold-change are shown without any measure of variability (no error bars, SEM, SD, or CI) across the n=5 bulk replicates.
    Could also: Displaying 95% confidence intervals or SD around the mean log2FC, or showing individual replicate values as overlaid points, would also convey within-group variability. — Variability estimates alongside effect sizes allow readers to gauge the consistency of the fold-change across replicates and the precision of the estimate, which is especially informative for genes highlighted as biologically important.
Software: DESeq2 1.34.0 · STAR 2.7.1 · Trimmomatic 0.39 · FASTQC 0.12.1 · CellRanger (10X Genomics) 8.0.1 · ScanPy 1.10.0 · Scater 1.22.0 · scikit-learn 0.24.2 · UCSC liftOver

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Authors · 4
Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41136526

Paper: Cell type- and species-specific regulation of hepatic lncRNAs by TCDD-activated AhR (Post et al., Sci Rep 2025) Repo: github.com/zacharewskilab/publication_analyses (Analyses/TCDD_lncRNA_Repo) Screened_utc: 2026-06-15T23:20:53Z

Pipeline-derived results (candidate targets)

The paper reuses previously-published GEO datasets and re-quantifies them against a custom GTF (RefSeq mm39/rn7 + Karri et al. lncRNAs + MGI lncRNAs; 20,973 mRNAs + 55,038 lncRNAs). Pipeline: FASTQC->Trimmomatic->STAR(mm39/rn7)->DESeq2 v1.34.0. DE criterion: |fold-change| >= 1.5 AND adjusted p <= 0.05 (union across dose-vs-0 contrasts).

IN SCOPE (low-hanging, shipped inputs)

The repo COMMITS the STAR ReadsPerGene.out.tab count matrices (mouse 40 samples, GSE203302; rat 24 samples, GSE110293). So alignment is bypassed and the DESeq2 -> DEG-count step is directly reproducible. Targets (Cross-species comparison, bulk RNA-seq): C1 Mouse DE lncRNAs = 2,386 C2 Mouse DE mRNAs = 6,071 C3 Rat DE lncRNAs = 916 C4 Rat DE mRNAs = 3,056 Biotype split source: lncRNA Type in {lncRNA,antisense,lincRNA,NR,lncOfInterest}; mRNA Type in {NM,NM#NR,mitochondrial protein-coding gene} (from 05a notebooks).

OUT OF SCOPE (hard 20% / not attempted, with reason)

  • Re-alignment from FASTQ (STAR/Trimmomatic): unnecessary, counts shipped.
  • Exact custom-GTF / biotype-map reconstruction (liftOver mm10->mm39, MGI merge, bedtools): the gene_id->Type map file (MGI_and_Karri_Annotations_Restructured_mm39.txt) is NOT committed. Biotype is reconstructed approximately (Karri lnc_ prefix + RefSeq biotype for symbol genes) -> mRNA/lncRNA split graded honestly.
  • snRNAseq cell-type DE counts (Fig 4/6), AHR ChIP-seq / pDRE enrichment (Sec 99): depend on the unshipped annotation + multi-step external builds. Not attempted now.
  • Harvard Dataverse pDRE set: separate, not needed for C1-C4.

Metadata reconstruction — VERIFIED against GEO (not an assumption)

Sample->dose from GEO matches the authors' desired_order grouping exactly. Mouse (GSE203302, Prj161-L3-NN -> L-id): 8 doses x n=5. Rat (GSE110293, title=L-id): 8 doses x n=3. (paper text says n=5 for rat; shipped data + GEO are n=3 -> flagged.)

C0
Reported
8457 (mouse lncRNA+mRNA DEGs)
Reproduced
8460
within tolerance
C1
Reported
2386 (mouse DE lncRNAs)
Reproduced
2349
within tolerance
C2
Reported
6071 (mouse DE mRNAs)
Reproduced
5944
within tolerance
C3
Reported
916 (rat DE lncRNAs)
Reproduced
1014
partial
C4
Reported
3056 (rat DE mRNAs)
Reproduced
3536
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Mouse bulk RNA-seq DE reproduces near-exactly — union DEG set 8460 vs the paper's 8457 (within 3 genes), and the lncRNA/mRNA split (2349/5944) lands within ~2% of 2386/6071, the residual fully explained by a RefSeq biotype proxy substituting for the authors' unshipped custom annotation. Rat is partial, overcounting ~20% (total 4784 vs 3972; lncRNA 1014 vs 916, mRNA 3536 vs 3056), most plausibly because the deposited rat counts are n=3/group while the paper text states n=5 — a deposit/cohort inconsistency on the authors' side, compounded by weaker rn7 symbol→biotype matching. The deviation is moderate (magnitude and direction hold) and the central species-specific lncRNA-regulation conclusion stands; no fabrication signal — the mouse magnitudes are cleanly regenerable from shipped counts. Net: a solid, explainable partial reproduction (q8 yellow).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

202.5 k
tokens (I/O) · 15.4 M incl. cache
21 min
runtime · 0.03 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine