Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Systematic Analysis of Long Non-Coding RNA Genes in Nonalcoholic Fatty Liver Disease.

Noncoding RNA · 2022
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (strong). Pipeline (STAR GRCh38.103 + edgeR GLM, threshold 2-fold & FDR<0.05) reconstructed from Methods + repo; the shipped code itself does NOT apply the threshold/biotype/intersection, so those headline-producing steps were reconstructed. From the deposited HGNC-filtered app_data.rds: C1 (>1000 DEGs) and C3 (24/14 cross-dataset protein-coding) reproduce EXACTLY (NASH markers TREM2/CIDEC/LPL/AKR1B10 in the set). Full STAR+edgeR re-alignment of all 57 SRA runs reproduces C2up within-tol (166 vs 163) and recovers ALL SIX named consistent lncRNAs of C4 (incl. clone-named AJ009632.2/AC110995.1 that the app's biomaRt HGNC filter drops) -- confirming the deposited-data undershoot is purely that filter. Sole discrepancy: C2dn (32 vs 12 down lncRNAs; the 12 are a subset of our 32, excess plausibly newer-edgeR sensitivity). No fabrication indicators. NOT attempted: wet-lab LINC01639 knockdown/qPCR (out of scope), GSE135251 full re-alignment (low marginal value).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-22 ⛓ ad9599e09e1e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Secondary analysis of published NAFLD RNA-seq data can reveal long non-coding RNAs (lncRNAs) that are dysregulated in NAFLD and may contribute to its disease mechanism, complementing prior protein-coding gene analyses.

Core claims
  • Many lncRNAs, like protein-coding genes, are differentially expressed in NAFLD patients compared to healthy normal-weight and obese individuals. finding
  • LINC01639 regulates genes involved in apoptosis, TNF/TGF and cytokine signaling, growth factors, and genes upregulated in NAFLD. finding
  • Four lncRNAs (AJ009632.2, LINC01639, MIRLET7BHG, SNAP25-AS1) are consistently up-regulated and two (AC110995.1, MANCR) consistently down-regulated across independent NAFLD RNA-seq datasets and conditions. finding
  • Overlap of differentially expressed genes (lncRNA and protein-coding) between two independent NAFLD RNA-seq studies is low, likely reflecting biopsy sample heterogeneity. finding
  • LiverDB, a web-based database cataloging expression (CPM/RPKM/TPM) of protein-coding and lncRNA genes across NAFLD studies, was built to facilitate hepatic lncRNA research. resource
  • A Nextflow-based RNA-seq reprocessing pipeline (SRA Toolkit, fastp, STAR, edgeR) was used to reanalyze public NAFLD datasets for lncRNA quantification not provided in original studies. method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (secondary analysis) human liver biopsy (NAFL n=15, NASH n=16, healthy normal weight n=14, obese n=12) none (disease state comparison) differential expression of protein-coding genes and lncRNAs GSE126848; Illumina NextSeq 500; STAR/edgeR
bulk RNA-seq (secondary analysis) human liver biopsy, European NAFLD cohort (early NASH n=138, moderate NASH n=68 vs healthy obese controls n=10) none (fibrosis stage/disease progression comparison) differential expression of protein-coding genes and lncRNAs GSE135251; STAR/edgeR
siRNA knockdown + gene expression analysis Huh-7 hepatocyte-derived carcinoma cell line siRNA silencing of LINC01639 vs scrambled control siRNA expression changes of TYMS, DUSP2, SRF, CXCL3, TNFRSF10D, BAD, FGF21, IGF1 Lipofectamine RNAiMAX
siRNA knockdown + fatty acid treatment (NAFLD cell model) + gene expression analysis Huh-7 cells LINC01639 siRNA silencing combined with oleic acid + palmitic acid (FA) treatment expression of CCL2, IL-6, RAC2
lipid accumulation assay Huh-7 cells fatty acid mixture (oleic acid + palmitic acid) treatment vs BSA control cellular lipid accumulation 96-well plate assay
Key results
  • Over 1000 genes differentially expressed in both NAFL and NASH vs healthy/obese controls (2-fold, FDR<0.05) 2-fold
  • 163 up- and 12 down-regulated lncRNAs shared across all NAFLD-vs-control comparisons in GSE126848
  • Only very few lncRNAs retained the same up/down trend when the GSE126848-derived list was compared against the independent GSE135251 dataset
  • 24 protein-coding genes commonly up-regulated and 14 commonly down-regulated across both datasets/conditions
  • 4 lncRNAs (AJ009632.2, LINC01639, MIRLET7BHG, SNAP25-AS1) commonly up-regulated and 2 (AC110995.1, MANCR) commonly down-regulated across all conditions/datasets
  • LINC01639 knockdown downregulated TYMS, DUSP2, SRF, CXCL3, TNFRSF10D, BAD, FGF21, IGF1 in Huh-7 cells
  • LINC01639 knockdown combined with fatty acid treatment upregulated CCL2, IL-6, and RAC2
Key statistics
  • fold_change 2-fold change (Threshold used to define differentially expressed genes)
  • pvalue FDR-adjusted p-value < 0.05 (Threshold used to define differentially expressed genes)
  • count 163 up-regulated and 12 down-regulated lncRNAs (lncRNAs shared among all NAFLD-vs-control comparisons in GSE126848)
  • count 24 up-regulated and 14 down-regulated protein-coding genes (Genes commonly dysregulated across both NAFLD datasets)
  • count 4 up-regulated and 2 down-regulated lncRNAs (lncRNAs commonly dysregulated across both NAFLD datasets/conditions)
  • count NAFL n=15, NASH n=16, healthy normal weight n=14, obese n=12 (GSE126848 cohort sample sizes)
  • count NAFLD patients n=206 (early n=138, moderate n=68) vs healthy obese controls n=10 (GSE135251 cohort sample sizes)
  • other ~25% of the world population (Global prevalence of NAFLD)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper performed secondary bioinformatic analysis of two published bulk liver RNA-seq datasets (GSE126848 and GSE135251), using a STAR-edgeR pipeline to call differentially expressed genes and lncRNAs with a combined threshold of 2-fold change and FDR-adjusted p-value < 0.05. Overlaps of differentially expressed genes across patient subgroups and between the two independent studies were examined descriptively (Venn/UpSet plots), and pathway/gene-ontology enrichment was performed with Enrichr and DAVID. A follow-up wet-lab loss-of-function experiment (siRNA knockdown of LINC01639 in Huh-7 cells, with and without a fatty-acid treatment) reported gene expression changes described as up- or down-regulated, without a named inferential test or exact p-values in the provided text.

Replicationunclear Sample sizePatient cohort sizes per group are stated for both RNA-seq datasets (e.g., NAFL n=15, NASH n=16, healthy n=14, obese n=12; early NASH n=138, moderate n=68, controls n=10); replicate numbers for the siRNA/cell-culture experiments are not given in the provided text GroupsNAFL/NASH vs healthy/obese individuals; early/moderate NASH vs healthy obese controls; LINC01639 siRNA-silenced vs scrambled-control Huh-7 cells, with/without fatty-acid treatment Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Multiplicity correctionFalse discovery rate (FDR)-adjusted p-values (as implemented in edgeR)
Statistical tests used
Test Applied to n Assumptions
edgeR differential expression analysis (fold-change + FDR-adjusted p-value threshold) NAFL (n=15) and NASH (n=16) vs healthy normal weight (n=14) and obese (n=12) individuals, GSE126848 n=15 (NAFL), n=16 (NASH), n=14 (healthy), n=12 (obese) not stated
edgeR differential expression analysis (fold-change + FDR-adjusted p-value threshold) early (n=138) and moderate (n=68) NASH vs healthy obese controls (n=10), GSE135251 n=138 (early), n=68 (moderate), n=10 (controls) not stated
Pathway/gene-ontology enrichment (Enrichr; DAVID v6.8) differentially expressed protein-coding/lncRNA gene sets not stated
Unspecified test for 'significant' up/down-regulation qPCR-based gene expression changes after LINC01639 siRNA silencing, Figure 4B/4D not stated
Approaches that could also have been used
  • Differentially expressed genes/lncRNAs were called using edgeR with a combined 2-fold change and FDR-adjusted p<0.05 threshold.
    Could also: DESeq2 (with shrinkage estimators such as apeglm) or limma-voom — These tools model count dispersion differently and are commonly used as complementary or cross-validating pipelines, which can be informative for consistency checks, especially with smaller group sizes.
  • Overlap of differentially expressed genes/lncRNAs between the two independent NAFLD RNA-seq datasets was assessed descriptively using Venn diagrams and UpSet plots.
    Could also: A formal meta-analytic approach, such as combining effect sizes across studies with a random-effects model or rank-aggregation method — This would provide a quantitative summary of cross-study consistency while explicitly accounting for between-study heterogeneity, complementing the descriptive overlap counts.
  • Pathway and gene-ontology enrichment of differentially expressed gene sets was performed with Enrichr and DAVID, which typically use hypergeometric/Fisher's exact tests on threshold-selected gene lists.
    Could also: Gene set enrichment analysis (GSEA) using the full ranked list of fold changes or test statistics — Rank-based enrichment can detect coordinated pathway-level shifts even when individual genes fall short of a hard significance cutoff.
  • Expression changes following LINC01639 siRNA silencing were described as significant up- or down-regulation without a named statistical test, exact p-value, or stated replicate number in the excerpted text.
    Could also: A explicitly named test such as a two-tailed Student's t-test or one-way ANOVA with post-hoc correction, reported with exact p-values and stated biological replicate n — Naming the test and reporting exact p-values with replicate numbers lets readers gauge the statistical strength and reproducibility of the knockdown effect.
  • Significance of differential expression was communicated as a fold-change/FDR threshold call rather than with an accompanying measure of estimate precision.
    Could also: Reporting log fold-change with its standard error or a 95% confidence interval — This would convey the precision of each effect-size estimate alongside its significance status, which can be useful when comparing effect magnitude across genes or studies.
  • Multiple hypothesis testing across genome-wide comparisons was addressed via FDR correction as implemented in edgeR.
    Could also: Alternative multiplicity-control approaches such as Bonferroni correction or independent hypothesis weighting — Different correction strategies offer different trade-offs between statistical power and stringency depending on the total number of tests and prior expectations about the proportion of true effects.
Software: Nextflow 21.10.6.5660 · SRA Toolkit 2.10.0 · fastp 0.23.2 · STAR 2.7.10a · R/edgeR 3.38.1 · R/GenomicFeatures 1.48.1 · Enrichr · DAVID v6.8

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 35893239 (LiverDB / Ilieva et al. 2022, Noncoding RNA)

Title: Systematic Analysis of Long Non-Coding RNA Genes in Nonalcoholic Fatty Liver Disease. Code: https://github.com/Bishop-Laboratory/LiverDB (preprocess/ pipeline) Data: GEO GSE126848 (primary, 57 liver-biopsy RNA-seq samples) + GSE135251 (secondary, validation)

Pipeline (from Methods + repo preprocess/)

SRA-toolkit (2.10.0) download → fastp 0.23.2 trim → STAR 2.7.10a align to GRCh38.103 (--quantMode GeneCounts, ReadsPerGene.out.tab) → edgeR 3.38.1 GLM: DGEList → calcNormFactors (TMM) → estimateGLMCommonDisp/Trended/Tagwise(design=~0+groups) → glmFit → glmLRT(contrast=c(1,-1)) → topTags(n=Inf). Output GSE126848_degs.csv.gz = (gene_id, numerator, denominator, logFC, FDR) for all genes, unfiltered. Contrasts for GSE126848 (metadata/contrasts.csv): NAFL/Healthy, NASH/Healthy, Obese/Healthy, Obese/NAFL, Obese/NASH, NAFL/NASH. Reported DEG threshold (Methods §4): 2-fold change AND FDR-adjusted p < 0.05 (i.e. |log2FC| ≥ 1 & FDR < 0.05).

KEY AUDIT NOTE — what the deposited code does NOT do

The shipped pipeline (process_raw_counts.R, prepare_app_data.R) computes edgeR logFC/FDR and adds HGNC symbols via biomaRt only. It does NOT (a) apply the 2-fold/FDR<0.05 threshold, (b) classify genes as lncRNA vs protein-coding, or (c) compute the cross-contrast intersections. Therefore the paper's HEADLINE numbers — 163/12 shared lncRNAs, 24/14 protein-coding, the 6 consistent lncRNAs — come from undeposited downstream steps described in Methods text only. Reproduction = take the authors' deposited DEG table, apply the described threshold, biotype-annotate via GRCh38.103 Ensembl gene_biotype, and intersect. A match confirms the headline numbers are derivable from the deposited data; a mismatch is a flag.

IN SCOPE (pipeline-derived, reproducible)

id result pipeline tier
C1 >1000 DE genes in both NAFL & NASH vs Healthy (Fig 1A) edgeR + threshold on deposited degs Tier-1 (light)
C2 163 shared up- + 12 down-regulated lncRNAs (Fig 1C, Supp S7/S8) + biotype + intersect Tier-1
C3 24 up- + 14 down-regulated common protein-coding genes (§2.2) + biotype + intersect Tier-1
C4 6 consistent lncRNAs: up AJ009632.2, LINC01639, MIRLET7BHG, SNAP25-AS1; down AC110995.1, MANCR (Fig 3) cross-dataset intersect (needs GSE135251) Tier-2
C5 Full regeneration of degs from raw SRA via STAR+edgeR (57 samples) whole pipeline on «our HPC» Tier-2 (heavy)

OUT OF SCOPE (wet-lab / manual / not a pipeline result)

  • LINC01639 siRNA knockdown in Huh-7, qPCR validation, fatty-acid NAFLD cell model (CCL2/IL-6/RAC2) — wet-lab.
  • Enrichr pathway enrichment narrative interpretation (the enrichr step is in repo but the biological conclusions are manual) — enrichr table reproducible but low-value, deferred.
  • LiverDB Shiny app itself (a data browser, not a reported numeric result).

Strategy

  1. Tier-1 (no heavy compute): download authors' GSE126848_degs.csv.gz (+ GTF for biotype) to «infra», apply threshold + biotype + intersections → reproduce C1, C2, C3. This is the auditability core.
  2. Tier-2 (heavy, «our HPC»/SLURM): regenerate degs from raw SRA counts (STAR+edgeR) and add GSE135251 for C4. Pursued after Tier-1 lands.

All data + GTF live on «infra»; «host» holds only small result tables. Heavy compute via SLURM on «our HPC».

Figures / tables: Figure 1AFigure 1CFigure 3
C1
Reported
>1000 DE genes in both NAFL & NASH vs Healthy
Reproduced
T1 NAFL=2036/NASH=2662; T2 NAFL=2656/NASH=3575
exact
C2up
Reported
163 shared up lncRNAs (4 comparisons, GSE126848)
Reproduced
166 (full STAR+edgeR re-align)
within tolerance
C2dn
Reported
12 shared down lncRNAs
Reproduced
32 (edgeR 4.8.2; 12 are a subset)
did not match
C3up
Reported
24 cross-dataset up protein-coding
Reproduced
24
exact
C3dn
Reported
14 cross-dataset down protein-coding
Reproduced
14
exact
C4up
Reported
4 up lncRNAs (AJ009632.2, LINC01639, MIRLET7BHG, SNAP25-AS1)
Reproduced
all 4 recovered in GSE126848 shared set
exact
C4dn
Reported
2 down lncRNAs (AC110995.1, MANCR)
Reproduced
both recovered in GSE126848 shared set
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

Strong reproduction. Using public GEO data (GSE126848/GSE135251) re-aligned from raw SRA, C1 (>1000 DEGs → 2036/2662), C3 (24 up / 14 down cross-dataset protein-coding) reproduce exactly, and full STAR+edgeR re-alignment recovers all six named consistent lncRNAs of C4, confirming the deposited-data undershoot is purely the app's HGNC filter. The single discrepancy is C2dn (32 vs 12 down lncRNAs) — an over-count where the reported 12 are a strict subset of our 32, plausibly newer edgeR-version sensitivity; direction and magnitude hold and the central conclusions are unaffected. Deviations sit on our methodology/tooling side (version + the authors' deposited-data filter), with no fabrication indicators — every headline number is derivable from the shared data.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

429.7 k
tokens (I/O) · 43.2 M incl. cache
155 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.