Utilizing the codon adaptation index to evaluate the susceptibility to HIV-1 and SARS-CoV-2 related coronaviruses in possible target cells in humans.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the CORE method: the CAI computation is fully specified and self-contained in the authors' CAI.R. Reproduced 1:1 by running their VERBATIM RSCU()+CAI() functions (commit bac2f29) on public reference viral ORFs (SARS-CoV-2 NC_045512.2; HIV-1 HXB2 K03455.1, +vpu) against a transparent human highly-expressed-gene background. The paper's central biological orderings reproduce: SARS-CoV-2 N-highest and E/ORF10/ORF6-lowest are EXACT; SARS-CoV-2 below endogenous holds; HIV-1 vpu-lowest reproduces exactly (value 0.455 vs reported avg 0.500); tat is among the highest (rev edges it by 0.01 in single-reference HXB2 -> partial). DID NOT attempt the hard 20%: the full upstream RNA-seq alignment pipeline (SRA->HISAT2->featureCounts->top200 genes for ~21 GEO datasets incl GSE159249) and scRNA-seq Seurat clustering -- those scripts hardcode «path» paths and are unparameterized -- so the EXACT per-cell-type CAI averages were not regenerated. One text/code discrepancy flagged: pseudocount 0.01 (text) vs 0.1 (code), numerically negligible here. Verdict: partial -- a faithful, clean reproduction of the paper's core method and its qualitative conclusions; the per-cell-type numeric averages remain unverified-but-plausible.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-15 ⛓ dc809f5a40c5
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether the codon adaptation index (CAI), as a measure of viral ORF translational efficiency at the translational elongation level, can be used at the cell-type level to evaluate the susceptibility of different human cell types to HIV-1 and SARS-CoV-2.
- ★ CAI is positively correlated with translational efficiency, supporting its use as a proxy for viral mRNA translation in cell types. finding
- ★ Compared to high-expression endogenous genes, the CAIs of viral ORFs are relatively low, implying HIV-1 and SARS-CoV-2 are not well adapted to human cell translational machinery. finding
- ★ Presumptive susceptibility to viruses based on CAI is usually consistent with experimental results, but with some exceptions. finding
- ★ HIV-1 and SARS-CoV-2 have different effects on cellular translational mechanisms: HIV-1 decouples CAI and translational efficiency of endogenous genes, while SARS-CoV-2 exhibits increased CAI for its ORFs in infected cells. mechanism
- ★ CAI should be constructed separately per cell type using cell-type-specific high-expression background gene sets rather than a single species-level set, refining analysis to the cell-type level. method
- ★ CAI can serve as an auxiliary index to assess cell susceptibility to viruses but cannot be the sole evidence to identify viral target cells. finding
- A curated resource of gene-level expression matrices and per-cell-type high-expression gene sets for the analyzed RNA-seq datasets is provided (GitHub CAIvirus). resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk RNA-seq (reanalysis) | Human monocyte-macrophage system cell types (monocytes, dermal macrophages, dendritic cells, Langerhans cells, osteoclasts, Kupffer cells, colonic macrophages, microglia) | none (unstimulated) | Gene-level expression (FPKM) used to build high-expression gene sets for CAI | — |
| Bulk RNA-seq (reanalysis) | Human CD4+ T lymphocyte subtypes (naive, nonnaive, Tfh, Treg) | none (unstimulated) | Gene-level expression (FPKM) for CAI | — |
| Bulk RNA-seq (reanalysis) | Human kidney cells (podocytes, mesangial cells) | none (unstimulated) | Gene-level expression (FPKM) for CAI | — |
| Bulk RNA-seq (reanalysis) | Human metabolic organ cells (hepatocytes, cholangiocytes, hepatic satellite cells, adipocytes) | none (unstimulated) | Gene-level expression (FPKM) for CAI | — |
| Single-cell RNA-seq (Smart-seq2) | Human lung (16 cell types incl. AT1, AT2, immune cells) and PBMCs (7 cell types) | none | Per-cell-type expression for CAI; Seurat clustering/annotation | Smart-seq2; Seurat v4.1.1 |
| Bulk RNA-seq (reanalysis) | Human cell lines/primary cells (HEK293T-hACE2, A549, A549-hACE2, Calu3, NHBE, lung organoid, HPPT) | SARS-CoV-2 infection or cytokine stimulation (IFNα/β/γ, IL-1β) vs control | Gene expression and viral ORF CAI in control vs infected/stimulated | — |
| Bulk RNA-seq + paired Ribo-seq (reanalysis) | Volunteer-derived primary CD4+ T cells (HIV-1) and HBEC cells (SARS-CoV-2) | HIV-1 or SARS-CoV-2 infection vs control at multiple timepoints (4–96h) | Translational efficiency (Ribo-seq/RNA-seq) vs CAI correlation | — |
| Codon adaptation index (CAI) computation | Viral ORFs (HIV-1, SARS-CoV-2 and related coronaviruses) across human cell-type-specific background gene sets | none | CAI of viral ORFs (top 200 high-expression genes as background, isoform-resolved, +0.01 correction) | — |
- ▲ CAI is positively correlated with measured translational efficiency, validating the method.
- ▼ Viral ORFs of HIV-1 and SARS-CoV-2 show relatively low CAIs compared to high-expression endogenous genes.
- – CAI-based predicted susceptibility largely matches experimental susceptibility data, with exceptions.
- – HIV-1 decouples CAI from translational efficiency of endogenous genes in host cells.
- ▲ SARS-CoV-2 exhibits increased CAI for its ORFs in infected cells.
- count 19 bulk RNA-seq datasets (including 2 with paired Ribo-seq) and 2 single-cell RNA-seq datasets (Datasets selected from NCBI GEO for constructing cell-type background gene sets)
- count top 200 protein-coding genes (Highest mean FPKM genes used to construct each cell type's high-expression gene set)
- count 16 cell types (lung scRNA-seq); 7 cell types (PBMC scRNA-seq) (Cell types annotated in single-cell datasets)
- other pk+0.01 and qk+0.01 correction; minimum >10 cells per cell type (CAI calculation corrections and cell-type inclusion threshold)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This computational study uses the Codon Adaptation Index (CAI) — calculated from cell-type-specific high-expression gene sets (top 200 protein-coding genes by mean FPKM) derived from 19 bulk and 2 single-cell RNA-seq datasets — to predict translational efficiency of HIV-1 and SARS-CoV-2 ORFs across dozens of human cell types. CAI values are compared descriptively across cell types and between infected and control conditions. A positive correlation between CAI and translational efficiency was verified using paired RNA-seq and Ribo-seq data, and GO-BP and KEGG enrichment analyses were performed to characterize the high-expression reference gene sets.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Correlation analysis (type not specified) between CAI and translational efficiency | Validation of CAI as a proxy for translational efficiency using paired RNA-seq and Ribo-seq datasets (GSE158930) | — | not stated |
| GO-BP and KEGG enrichment analysis via R/clusterProfiler | Functional characterization of high-expression gene sets in three representative cell types (blood monocytes, CD4+ Tfh, scRNA-seq dendritic cells) | — | not stated |
| Descriptive comparison of CAI point estimates across cell types and viral ORFs | All cell-type-level CAI comparisons for HIV-1 and SARS-CoV-2 ORFs | — | na |
| Descriptive comparison of CAI between control and SARS-CoV-2- or HIV-1-infected cells | Analysis of viral infection effect on high-expression gene set codon usage and viral ORF CAI (multiple cell line datasets from GSE158930, GSE147507, GSE169158, GSE160435, GSE161916) | — | not stated |
-
CAI was used as the sole codon-adaptation indicator for predicting translational efficiency of viral ORFs↳ Could also: The tRNA Adaptation Index (tAI) — which uses cellular tRNA gene copy numbers as a proxy for tRNA pool availability — could also be used alongside or instead of CAI — tAI captures the supply side of codon-anticodon matching more directly; comparing CAI and tAI results would reveal whether conclusions are robust across complementary measures of codon adaptation
-
CAI values across cell types and ORFs were compared descriptively, without inferential tests or uncertainty quantification↳ Could also: Bootstrapping over gene-set composition (resampling the top-200 reference genes) or permutation tests could also be used to derive confidence intervals for CAI estimates and assess whether differences between cell types exceed chance variation — Point estimates of CAI carry no formal uncertainty; bootstrapping would show how sensitive cell-type rankings are to the specific genes that happen to meet the top-200 threshold, supporting stronger inferential claims about differential susceptibility
-
The reference high-expression gene set was defined as the top 200 protein-coding genes by mean FPKM, with a fixed cutoff of SD > mean for exclusion↳ Could also: Sensitivity analyses varying the reference set size (e.g., top 100, 500) or using an alternative threshold for variability exclusion (e.g., coefficient of variation) could also be reported — CAI is directly determined by the composition of the reference set; demonstrating stability of cell-type CAI rankings across plausible reference-set definitions would strengthen confidence in conclusions about relative susceptibility
-
The correlation between CAI and translational efficiency was verified using paired RNA-seq and Ribo-seq data, but the correlation method is not named in the main text↳ Could also: Both Pearson and Spearman rank correlation could be reported and compared, with Spearman being robust to the non-normal, heavy-tailed distributions typical of gene expression data — Specifying and justifying the correlation method — and reporting it with a confidence interval — would make the strength of the CAI-translational efficiency relationship precisely interpretable and reproducible
-
Cross-study integration of 19 bulk RNA-seq datasets from different GEO accessions was performed using FPKM normalization alone↳ Could also: Explicit batch-correction methods (e.g., ComBat, limma removeBatchEffect, or quantile normalization across studies) could also be applied before constructing high-expression gene sets — Technical variation between sequencing runs and laboratories can shift FPKM distributions in ways that affect which genes rank in the top 200; batch correction would reduce the risk that cross-study technical differences influence CAI estimates and cell-type comparisons
-
Single-cell RNA-seq data were clustered with Seurat and cell types annotated by marker genes from the corresponding literature↳ Could also: Cluster stability metrics (e.g., silhouette scores, bootstrapped reproducibility, or resolution sweeps) could also be reported, and an independent tool (e.g., Scanpy) could be used to verify major cluster assignments — Cell-type assignment at the scRNA-seq level directly determines which cells contribute to the high-expression gene set; documenting cluster robustness would clarify how sensitive lung and PBMC CAI estimates are to clustering choices
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36760235
Paper: Zhou H, Ren R, Yau SS. Utilizing the codon adaptation index to evaluate the susceptibility to HIV-1 and SARS-CoV-2 related coronaviruses in possible target cells in humans. Front Cell Infect Microbiol 2023. PMID 36760235 / PMC9905242 / DOI 10.3389/fcimb.2022.1085397.
Code: https://github.com/Renruohan/CAIvirus @ commit bac2f29 (cloned to «infra»).
Repo ships code only (9 R/shell scripts, ~9 KB) — NO data, NO intermediate files,
NO README beyond a one-line file list. The viral ORF sequences (HIVsequence/*.csv,
SARSCOV2sequence/*.fasta) and the per-cell-type top-200 CDS backgrounds
(CAImaxgeneCDS/.../top200codinggenemaxtranscripts.csv) are read from local paths
that are not in the repo and must be regenerated.
Pipeline map (what produces each reported result)
| Stage | Script | In scope? |
|---|---|---|
| SRA download → trim → HISAT2 align → featureCounts → FPKM (21 datasets incl. GSE159249) | upstream.sh, gettop200maxgenelist.R |
NO — hard 20%. Heavy alignment of dozens of SRA runs; scripts hardcode «path», unparameterized. |
| Pick top-200 highly-expressed coding genes per cell type; extract their max-expressed-transcript CDS | gettop200maxgenelist.R, gettop200CDSlist.R |
NO (depends on stage above) |
| scRNA-seq lung/PBMC Seurat clustering + annotation | scRNAseq-lung.R, scRNAseq-PBMC.R |
NO (heavy, depends on raw scRNA data) |
| CAI of viral ORFs = geometric mean of pseudocount-adjusted relative adaptiveness (w) over codons, background = top-200 gene RSCU | CAI.R (functions Generatecodon, RSCU, CAI) |
YES — core, fully specified, self-contained. |
| CAI vs translational-efficiency regression; HIV downstream | CAIandTEregression.R, HIVanalysis.R |
NO (needs Ribo-seq TE data + full backgrounds) |
What we attempt (in scope)
Run the authors' verbatim RSCU + CAI functions (the paper's core method) on
public reference viral ORFs (SARS-CoV-2 NC_045512.2; HIV-1 HXB2 K03455.1, both
via NCBI fasta_cds_na) against a transparent, reproducible human
highly-expressed-gene background (30 canonical highly-expressed human RefSeq CDS).
We test the paper's central, background-robust qualitative claims and value ranges:
- HIV-1 per-ORF ordering: vpu lowest, tat highest; single-ORF range 0.346–0.765.
- SARS-CoV-2 per-ORF ordering: N highest; E / ORF10 / ORF6 lowest.
- SARS-CoV-2 overall CAI lower than endogenous genes.
- Endogenous CAI in ~0.6–0.85.
What we do NOT attempt (out of scope / hard 20%)
- The full upstream RNA-seq alignment for 21 datasets (incl. GSE159249) → so we do NOT reproduce the exact per-cell-type top-200 background or the per-cell-type CAI averages (e.g. "vpu 0.500 / tat 0.658", "Kupffer-cell highest"). We substitute a transparent human highly-expressed-gene background; absolute values are expected to differ in the 2nd decimal while orderings/ranges are the testable claims.
- scRNA-seq Seurat re-clustering; CAI–TE regression coefficients (ρ=0.102 etc.).
Notable discrepancy (possible-fabrication / spec note)
Paper text states pseudocount 0.01; the shipped RSCU() code adds 0.1
((counts+0.1)/(maxcounts+0.1*familysize)). We run the code's value (0.1) and
flag the mismatch.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Running the authors' verbatim CAI.R on public reference viral ORFs reproduces the paper's central qualitative results cleanly: SARS-CoV-2 N-highest and E/ORF10/ORF6-lowest are exact, SARS-CoV-2 < endogenous holds, and HIV-1 vpu-lowest reproduces (0.455 vs reported 0.500). The remaining gaps sit on our side (self-chosen 30-gene background, single HXB2 reference, aggregation differs from the authors' 58-strain × per-cell-type averages), explaining the C5 tat/rev 0.01 swap and the C7 range shift. One genuine authors'-side flag — pseudocount 0.01 (text) vs 0.1 (code) (C8) — is numerically negligible and shows no fabrication signal. The exact per-cell-type averages remain unverified only because the upstream pipeline hardcodes paths and was not rerun, so this is a solid reproduction with explainable, input-side deviations rather than a substantive discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.