Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

An atlas of the human liver diurnal transcriptome and its perturbation by hepatitis C virus infection.

Nat Commun · 2024
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED the named third-party pipeline (dryR::dryseq, naef-lab/dryR 1.0.0, period 24) control-vs-HCV on GSE200811 (36 samples, 17663 expressed genes) on «our HPC» (SLURM «job»). The dryR 5-model framework reproduces EXACTLY: model mapping validated against fitted coefficients (m4 conserved 2675/2675 coeff-equal; m5 altered 0/119 equal; control-rhythmic<=>{2,4,5}, HCV-rhythmic<=>{3,4,5}). The qualitative claim reproduces (HCV perturbs a substantial fraction of rhythmic genes via loss/gain/altered rhythmicity; the majority of rhythmic genes stay conserved). The HEADLINE percentage does NOT match 1:1: paper ~22% vs reproduced 45.5% over all rhythmic genes (37.1% at a BICW threshold where the rhythmic-gene count matches the paper's ~1700). The ~1.7-2x gap is NOT explained by biotype (protein-coding: 44.5%) or BICW-confidence filtering (both tested); it tracks the paper's restriction to protein-coding ORTHOLOGOUS rhythmic genes (enriched for conserved core-clock/metabolic genes -> larger model-4 share), which is out of scope (needs the separate WT mouse liver time course + ortholog map). Honest verdict: faithful pipeline + qualitative reproduction, partial quantitative match with a documented methodological cause. This replaces the prior 'error' (VPN-blocked, no compute) run.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-20 ⛓ d3c2a3ab0d47
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Using human liver chimeric mice as a surrogate for human liver, the study investigates whether the human hepatic circadian clock generates a diurnal transcriptome/epigenome and whether chronic hepatitis C virus (HCV) infection perturbs this rhythmicity in ways relevant to liver disease and hepatocellular carcinoma (HCC) development.

Core claims
  • Human hepatocytes engrafted in liver chimeric mice display a large rhythmic transcriptome of ~1700 protein-coding orthologous genes, including transcription factors, chromatin modifiers, and metabolic enzymes. finding
  • dryR algorithm classifies rhythmic genes into distinct models (human-only cycling, mouse-only cycling, unaltered rhythm in both species, altered rhythm between species), enabling cross-species comparison of rhythmicity. method
  • ~140 transcription factors (~8% of rhythmic genes) show rhythmic expression in human hepatocytes, with some (IRF2, NCOR2, JUNB, RELB, IRF1) uniquely rhythmic in human vs. mouse cells. finding
  • Diurnal H3K27ac ChIP-seq reveals rhythmic epigenetic remodeling of promoters/enhancers in human hepatocytes, including human-specific rhythmicity at the IRF2 locus not seen in mouse. finding
  • Chronic HCV infection perturbs the rhythmicity of expression of more than 1000 genes in human hepatocytes in vivo. finding
  • HCV-perturbed rhythmic pathways activate processes mediating metabolic alterations, fibrosis, and cancer, and remain dysregulated in patients with advanced liver disease. finding
  • The human liver chimeric mouse (HLCM) model recapitulates key aspects of human liver disease biology, including chronic viral infection, and is a viable surrogate for studying human hepatic diurnal biology. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq human liver chimeric mice (primary human hepatocytes engrafted, male) none (baseline diurnal timecourse, ZT0-ZT24 every 4h) rhythmic transcript abundance / gene expression rhythmicity
ChIP-seq (H3K27ac) human liver chimeric mice liver tissue vs. wild-type mouse liver none (diurnal timecourse) diurnal histone H3K27ac occupancy at promoters/enhancers
bulk RNA-seq HCV-infected human liver chimeric mice liver tissue HCV infection (patient-derived HCV, 10 weeks) perturbation of diurnal transcriptome rhythmicity
serum albumin quantification human liver chimeric mice serum none / HCV infection (comparison) degree of hepatocyte engraftment (humanization)
immunostaining (CK8-18) human liver chimeric mice liver tissue none human hepatocyte-specific staining confirming humanization
histology (H&E staining) human liver chimeric mice liver tissue none hepatic lobular architecture
viral load quantification serum of HCV-infected human liver chimeric mice HCV infection circulating HCV viral load
Key results
  • dryR identified ~1700 rhythmic protein-coding orthologous genes in human hepatocytes of HLCM liver ~1700 genes
  • 824 genes were uniquely rhythmic in human hepatocytes (model 2) 824 genes
  • 749 genes showed unaltered rhythmic expression shared between human and mouse hepatocytes (model 4) 749 genes
  • 103 genes showed altered rhythmicity (phase/amplitude) between human and mouse hepatocytes (model 5) 103 genes
  • ~140 transcription factors identified as rhythmically expressed in human hepatocytes, representing ~8% of all rhythmic genes ~140 genes (~8%)
  • H3K27ac levels around the IRF2 promoter-enhancer were rhythmic only in human hepatocytes, not in mouse liver
  • HCV infection altered the rhythmicity of expression of more than 1000 genes in human hepatocytes >1000 genes
  • Chimeric livers showed approximately 65-70% humanization confirmed by CK8-18 immunostaining ~65-70%
Key statistics
  • count ~1700 rhythmic protein-coding orthologous genes (dryR-classified rhythmic genes in human hepatocytes per timepoint)
  • count 824 genes (model 2, human-only cycling) (dryR rhythmicity model classification)
  • count 749 genes (model 4, unaltered rhythm in both species) (dryR rhythmicity model classification)
  • count 103 genes (model 5, altered rhythm) (dryR rhythmicity model classification)
  • count ~140 rhythmic transcription factors (~8% of rhythmic genes) (TF analysis using dataset of 1600 human TFs)
  • count >1000 genes with perturbed rhythmicity (HCV infection effect on diurnal transcriptome)
  • mean Series 1: ~14,203 μg/mL; Series 2: ~14,973 μg/mL (human serum albumin levels indicating hepatocyte engraftment)
  • other ~35 million reads per sample (average) (RNA-seq sequencing depth)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a human liver chimeric mouse (HLCM) model, sacrificing mice every 4 hours across a 24-hour cycle in two independent experimental series to profile the diurnal liver transcriptome (RNA-seq) and epigenome (H3K27ac ChIP-seq) in human vs. residual mouse hepatocytes, and to compare control vs. HCV-infected animals. Rhythmicity and differential rhythmicity of gene expression were classified using the dryR algorithm across multiple models (species-specific, shared, or altered rhythmicity). Pathway-level enrichment of rhythmic genes was tested against MSigDB HALLMARK gene sets with an FDR<0.05 threshold, and overlap between rhythmic gene sets from different datasets/species was tested with a hypergeometric test. Based on the excerpted text, results were reported mainly through expression plots (means with SD) and enrichment scores rather than extensive tables of exact p-values.

Replicationbiological Sample sizeNumber of chimeric mice per timepoint stated (Series 1: n=2/timepoint, Series 2: n=3/timepoint; merged/combined analyses used n=5 HLCM/timepoint); no formal a priori power or sample-size calculation described in the excerpted text Groupshuman vs. mouse hepatocyte rhythmic transcriptome/epigenome; control vs. HCV-infected HLCM across circadian (ZT) timepoints Pairingunclear Randomization/blindingnot stated DispersionSD Multiplicity correctionFDR (specific correction procedure, e.g., Benjamini-Hochberg, not stated)
Statistical tests used
Test Applied to n Assumptions
dryR (harmonic regression-based classification of rhythmic/differentially rhythmic genes) classification of rhythmic genes into models 1-5 comparing human and mouse hepatocytes and conditions (Fig. 1c, e; Supplementary Data 3) n=5 HLCM per timepoint (merged Series 1 and 2) not stated
Hypergeometric test assessing significance of overlap of shared rhythmic genes between HLCM, WT mouse liver, and post-mortem human liver datasets (Supplementary Fig. 4a, b) not stated (gene set sizes not specified in text) not stated
Pathway enrichment analysis (MSigDB HALLMARK gene sets, FDR<0.05) identifying rhythmic pathways in human hepatocytes and comparison with WT mice (Fig. 1d, Supplementary Fig. 5a, b) n=5 HLCM per timepoint not stated
DESeq2 normalization of RNA-seq counts expression quantification for transcription factor expression patterns (Fig. 1f) n=5 HLCM per timepoint not stated
Approaches that could also have been used
  • Rhythmicity of gene expression was classified using dryR, a harmonic-regression-based model-selection algorithm comparing rhythmicity across species/conditions.
    Could also: Algorithms such as JTK_CYCLE, RAIN, or MetaCycle could also be applied to time-course expression data. — These are widely used alternative circadian-detection methods with different underlying assumptions (e.g., non-parametric rank-based testing or ensemble approaches), and applying more than one can serve as a cross-check on which genes are called rhythmic.
  • Overlap between rhythmic gene sets from different datasets/species was tested for significance using a hypergeometric test.
    Could also: A permutation-based overlap test (e.g., resampling-based methods) or Fisher's exact test could also be used. — Permutation approaches can be useful when the appropriate background gene universe is uncertain, offering an alternative way to estimate the null distribution of overlap counts.
  • Pathway enrichment of rhythmic genes was determined using a threshold-based approach (FDR<0.05) on MSigDB HALLMARK gene sets.
    Could also: A rank-based gene set enrichment analysis (GSEA) using the full ranked gene list rather than a significance-threshold-selected subset could also be used. — Rank-based GSEA can detect coordinated pathway-level shifts even when individual genes do not each cross a significance threshold.
  • Group-level expression patterns (e.g., transcription factor expression in Fig. 1f) were summarized with mean and SD across n=5 biological replicates per timepoint.
    Could also: Reporting SEM, a 95% confidence interval, or showing individual data points alongside the summary statistic could also be used. — For small per-group sample sizes, individual data points or CIs can convey the underlying variability and precision of the estimate more directly than SD alone.
  • Sample sizes per timepoint (2, 3, or 5 animals) were arrived at via post-hoc quality-control clustering and exclusion of low-read samples rather than a pre-specified power calculation.
    Could also: An a priori power analysis based on anticipated effect sizes from pilot or published circadian amplitude data could also be used to set target sample sizes. — Pre-specified power calculations help document the sensitivity of the design to detect a target effect size before data collection.
  • Two independent experimental series (Series 1 and 2, with differing n per timepoint) were merged for combined analyses.
    Could also: A mixed-effects or linear model explicitly including 'Series' as a batch covariate could also be used when pooling the two experiments. — Explicit batch modeling can help account for between-experiment variability when combining independent series into a single analysis.
Software: dryR · DESeq2 · MSigDB (gene set database for HALLMARK pathway analysis)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39209804

Paper: Mukherji et al., An atlas of the human liver diurnal transcriptome and its perturbation by hepatitis C virus infection. Nat Commun 2024. PMID 39209804 · PMC11362569 · DOI 10.1038/s41467-024-51698-8.

Code link is a text-mining FALSE POSITIVE

The registry code_url = github.com/genepattern/gparc-module-docs is an archive of GenePattern module HTML docs frozen 2021-12-09 (description: "an archive of gparc module docs as of 21.12.09"). It is not this paper's analysis code and contains nothing specific to GSE200812. The paper ships no authors' code repo; it names standard tools + one specific rhythmicity package (dryR).

Per brief rule P16, applying the named third-party tool to the paper's own data is an equally valid reproduction. We do exactly that.

Pipeline (from Methods)

  • RNA-seq: HISAT2 (GRCh38/GRCm38) → htseq-count → DESeq2 v1.28.1 (size-factor norm).
  • Rhythmicity / differential rhythmicity: dryR (naef-lab/dryR) — the 5-model framework the paper's "model 2 / model 4 / model 5" language comes from.
  • (Out of scope: ChIP-seq Bowtie/MACS3/HOMER/JASPAR; GSEA/GSVA enrichment; wet-lab albumin/Sirius-red/immunostaining; the 216-patient cirrhosis PLS cohort.)

Data

GSE200812 SuperSeries = GSE200809 (RNA-Seq I, HCV patient pre/post-DAA biopsies), GSE200810 (H3K27ac ChIP), GSE200811 (RNA-Seq II — the circadian time course). GSE200811 ships a processed raw gene count matrix (GSE200811_gene_counts_matrix.tsv.gz, 1.3 MB) + a DESeq2-normalized matrix. Design (from GEO): 36 samples = 6 Zeitgeber times {ZT0,4,8,12,16,20} × {untreated control, HCV-infected} × 3 replicates (humanized uPA/SCID chimeric liver). Sample titles are internal IDs (e.g. S76781); ZT + treatment are in the series-matrix characteristics.

IN SCOPE — what we reproduce (one clean pipeline output)

Run dryR's 2-condition differential-rhythmicity (dryseq, period 24 h) on the GSE200811 count matrix, control vs HCV, and recover the dryR model partition:

  • # rhythmic genes (rhythmic in ≥1 condition) — paper context: ~1,700 rhythmic orthologous genes (note: paper's headline ~1,700 is the human-vs-mouse ortholog comparison, a different contrast; our number is the control-vs-HCV human-hepatocyte rhythmic set, reported as our own derived value, not asserted equal).
  • % of rhythmic genes perturbed by HCV = (lost + gained + altered)/(all rhythmic) = dryR models {2,3,5}/{2,3,4,5}. Paper claim: ~22% (loss, gain, altered rhythmicity). THIS is the primary 1:1 comparison.

OUT OF SCOPE (the hard last ~20%, not attempted)

  • Human-vs-mouse ortholog dryR (824/749/103 model 2/4/5) — needs the separate wild-type mouse liver time course + ortholog mapping; not a single-matrix run.
  • ChIP-seq, enrichment, PLS/186-gene signature, all wet-lab quantitations.

Drop-reason check

Not a drop: data is public + processed matrix shipped; the analysis method (dryR) is public, installable, and explicitly named; an expected value (~22%) is pinnable.

Figures / tables: Fig.2bFig.1cFig.2
C1
Reported
~22% of all rhythmic genes perturbed by HCV (dryR models 2,3,5 of 2,3,4,5)
Reproduced
45.5% over all 4907 rhythmic genes (m2=1112 lost, m3=1001 gained, m4=2675 conserved, m5=119 altered); 37.1% at BICW>=0.7 where the protein-coding rhythmic count (~1614) matches the paper's ~1700
partial
C2
Reported
~1,700 rhythmic protein-coding orthologous genes
Reproduced
NOT ATTEMPTED (human-vs-mouse ortholog contrast; needs mouse data + ortholog map)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

The named third-party pipeline (dryR::dryseq, period 24) reproduces exactly on the authors' own GSE200811 matrix — the 5-model partition and lost/gained/conserved/altered semantics are validated 1:1, and the qualitative conclusion (substantial HCV perturbation, majority of rhythmic genes conserved) holds. The deviation is confined to the headline percentage: reproduced 45.5% (37.1% at a BICW threshold where n~1614 matches the paper's ~1700) vs the reported ~22%, a ~1.7–2x gap. The cause is on our methodology/scope side, not the authors': the paper's denominator is ~1,700 protein-coding orthologous rhythmic genes, and reproducing that restriction requires the separate WT mouse liver time course + ortholog map (out of scope); biotype and BICW filters were tested and do not close the gap. No fabrication suspicion — overall a solid reproduction with a documented, explainable denominator deviation.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

444.2 k
tokens (I/O) · 34 M incl. cache
86 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.