Detection of Virus-Related Sequences Associated With Potential Etiologies of Hepatitis in Liver Tissue Samples From Rats, Mice, Shrews, and Bats.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
COMPLETE (15/15 libraries). Viral-metagenomics paper He et al 2021 (Front Microbiol). The deterministic, well-specified FRONT END of the pipeline reproduced cleanly on the EXACT deposited data (PRJNA695121 11 runs + PRJNA701687 4 runs), every fastq.gz md5-verified against ENA: raw read counts match ENA EXACTLY for all 15 (R2 exact); Sickle 1.33 Q20/L50 survival 86.5-97.6% mean 94.75% (R3); IDBA-UD + 264bp filter, 706-588705 contigs/lib (R4); DIAMOND vs RefSeq viral confirms abundant, retrovirus-dominated viral signal across ALL animals plus large DNA viruses/NCLDVs/phages (R5 qualitative, 37,844 viral contigs). Serum libraries are sparser than liver, the expected tropism difference. The HEADLINE taxonomy numbers (T1-T3: no-result-read %s, virus family/genus counts, per-genus abundances) depend on full NT/NR BLAST at an unpinned 2020-21 DB version and are therefore non-deterministic / not exactly reproducible by design -> honestly out of scope, NOT scored as mismatch. Notable deposit finding: 11 deposited libraries vs 'nine pooled samples' stated (2 extra codes WSM/SRN), and 2 runs use a 3-file ENA layout (paired+orphan singles) that reconciles to read_count exactly. Overall: front-end 1:1 reproducible; biological headline numbers irreproducible by DB-version design, not by data/code defect.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 57assessed: 2026-06-18 ⛓ 20b0bf229c35
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetBecause the etiology of 10–20% of human hepatitis cases remains unclear and several hepatitis-associated viruses are zoonotic, the study asks what viral communities (particularly hepatitis-associated and potentially hepatitis-causing viruses) are present in liver tissue of rats, mice, house shrews, and bats, and whether these differ across species and between liver and serum.
- ★ Viral metagenomics of liver tissue from rats, mice, shrews, and bats revealed a diverse set of sequences related to herpesviruses, orthomyxoviruses, anelloviruses, hepeviruses, hepadnaviruses, flaviviruses, parvoviruses, and picornaviruses finding
- ★ First PCR detection of hepatovirus sequences in Hipposideros larvatus (3.85%) finding
- ★ First detection of Zika virus-related sequences in urban rats and house shrews finding
- ★ Influenza A virus and herpesvirus sequences were detected only in liver tissue samples (not serum) finding
- ★ Pegivirus detection rates were higher in liver and serum of rats than in house shrews finding
- ★ Torque teno virus (TTV) detection rates were higher in serum than in liver for both rats and house shrews finding
- ★ Near-full-length genomes of pegivirus and torque teno virus were amplified resource
- ★ Liver and serum viral community compositions differ, with some hepatitis-associated viruses (e.g., herpesviruses) showing apparent liver tropism finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| viral metagenomic sequencing (Illumina HiSeq, 2x150bp paired-end) | pooled liver tissue from urban rats, house shrews, bats, laboratory rats/mice | none | viral sequence reads classified via BLASTN/BLASTP against NCBI NT/NR | Illumina HiSeq |
| PCR screening | individual liver tissue samples from bats, urban rats, house shrews | none | presence/prevalence of hepatovirus, hepacivirus, HDV, parechovirus, herpesvirus, influenza A virus, adenovirus, DENV, CHIKV, ZIKV, pegivirus, TTV, LV, KIs-V, SEN-V sequences | — |
| viral metagenomic sequencing (paired comparison) | liver tissue vs serum from Rattus norvegicus (24 animals, pooled) | none | relative abundance and PCA-based comparison of viral community composition | Illumina HiSeq |
| strand-specific RT-PCR (minus-strand detection) | serum and liver tissue from pegivirus-positive animals | none | presence of minus-strand pegivirus RNA as evidence of replication | — |
| Sanger sequencing and phylogenetic analysis (MAFFT alignment, MrBayes tree) | PCR-positive amplicons/near-full-length genomes (pegivirus, TTV, others) | none | phylogenetic placement/genetic diversity | MEGA v6.0, MrBayes v3.2 |
| cytB gene sequencing | trapped rats, mice, house shrews, bats | none | species identification | — |
- – Hepatovirus sequences detected in Hipposideros larvatus liver 3.85%
- – Zika virus-related sequences detected in urban rats and house shrews 0.64% (urban rats), 1.13% (house shrews)
- ▲ Pegivirus detection rate higher in rat liver and serum than house shrews 7.85% (liver), 15.79% (serum) in rats
- ▲ TTV detection rate higher in serum than liver for rats and house shrews 52.72% (rat serum), 5.26% (shrew serum)
- – Mammalian virus-related sequence diversity differed by species 17 families/26 genera (urban rats), 15 families/18 genera (lab animals), 12 families/18 genera (house shrews), 14 families/24 genera (bats)
- – In bats, only hepatovirus, influenza A virus, and herpesvirus sequences detected by PCR 0.99% (hepatovirus), 2.97% (influenza A), 20.79% (herpesvirus)
- – Serum samples had higher relative abundance of anelloviruses and flaviviruses; liver samples had higher relative abundance of retroviruses and herpesviruses
- – Rhinovirus C-related sequences detected only in liver, not serum, of R. norvegicus 94% nucleotide identity
- other 1,003 wild animals trapped (177 bats, 624 urban rats, 202 house shrews) (total sampled animal population)
- count 247 serum samples and 1,003 liver tissue samples collected (sample collection)
- other hepatovirus detection rate 3.85% (Hipposideros larvatus liver, first detection)
- other pegivirus detection 7.85% (liver) vs 15.79% (serum) (rat pegivirus prevalence)
- other TTV detection 52.72% (rat serum) vs 5.26% (shrew serum) (TTV prevalence in serum)
- other nucleotide identities: herpesvirus 66%, influenza A virus 100%, HEV 78%, TTV 89%, pegivirus 93%, rhinovirus C 94% (sequence identity of detected hepatitis-associated viral sequences)
- other reads with no BLAST hits: 85.24% (bats), 86.67% (urban rats), 74.27% (house shrews) (unannotated sequencing reads)
- pvalue P < 0.05 (significance threshold for chi-square tests of viral prevalence differences)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This cross-sectional surveillance study combined viral metagenomics and targeted PCR to characterize liver viral communities in four small-mammal taxa (bats, urban rats, house shrews, laboratory rodents) trapped in southern China. Viral community composition was explored with alpha-diversity indices and Euclidean-distance PCA; prevalence of individual viruses was estimated from PCR screening and compared across host species with chi-square tests. Results were reported primarily as percent detection rates, with phylogenetic relationships inferred by Bayesian and maximum-likelihood methods.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson chi-square test | Differences in viral prevalence (PCR detection rates) between animal species/sample types | Up to 1,003 individual animals (177 bats, 624 urban rats, 202 house shrews, 10 laboratory rodents); specific per-test n not stated | not stated |
| Principal component analysis (PCA, Euclidean distance) | Beta diversity of viral community composition among pooled metagenomic samples (liver) and between liver vs. serum in R. norvegicus | 9 pooled liver samples; 4 additional pooled samples (R. norvegicus liver vs. serum) | na |
| Alpha-diversity indices: Chao, ACE, Shannon, Simpson | Richness and evenness of viral communities per pooled metagenomic sample at genus level | 9–13 pooled samples; individual-level read counts within pools | na |
| Bayesian phylogenetic inference (MrBayes, GTR+G+I model, 2,000,000 generations, 25% burn-in) | Phylogenetic trees for detected virus sequences relative to GenBank references | Number of sequences per tree not stated | not stated |
| Maximum-likelihood phylogenetic analysis (MEGA 6.0) | Additional phylogenetic placement of virus sequences | Not stated | not stated |
-
Chi-square tests were used to compare PCR prevalence rates across host species, some of which had small group sizes (e.g., 1 Bandicota indica, 6 Pipistrellus abramus)↳ Could also: Fisher's exact test (or its generalized form for r×c tables) could also be used for comparisons involving small expected cell counts — Fisher's exact test does not rely on the large-sample approximation underlying the chi-square and is preferred when any expected cell count falls below 5; it produces valid p-values in the same small-n scenarios that make chi-square unreliable
-
Multiple chi-square tests were conducted comparing prevalence of up to 15 different viruses across several host groups with no stated multiplicity correction↳ Could also: A Bonferroni correction, Benjamini-Hochberg false discovery rate, or similar procedure could also be applied across the family of prevalence comparisons — When many tests are conducted simultaneously, the probability of at least one false positive rises; a correction procedure makes the family-wise or per-comparison error rate explicit and allows readers to calibrate findings accordingly
-
Beta diversity was visualized with PCA based on Euclidean distance on relative-abundance data↳ Could also: Non-metric multidimensional scaling (NMDS) on Bray-Curtis dissimilarity could also be used, accompanied by a permutation-based PERMANOVA (adonis) test — Bray-Curtis dissimilarity is widely used for sparse compositional data (many zeros are expected in virome tables); NMDS does not assume linearity, and PERMANOVA provides a formal significance test for group separation rather than purely visual interpretation
-
Viral prevalence estimates were reported as percentages without accompanying uncertainty measures↳ Could also: Wilson score or Clopper-Pearson 95% confidence intervals for each prevalence estimate could also be reported — Confidence intervals communicate both the point estimate and its precision, which is especially informative when sample sizes differ substantially across host species (as here, ranging from 1 to 624 animals)
-
Metagenomics was performed on pooled samples (up to multiple animals per pool) rather than individual samples↳ Could also: Individual-sample sequencing combined with mixed-effects or zero-inflated count models could also be used to estimate prevalence and test for host-group differences on a per-animal basis — Pooling is a practical and cost-effective strategy; individual sequencing would additionally allow within-group variability to be quantified and formal statistical comparisons of community composition to be made with appropriate error structure
-
Alpha-diversity indices (Chao, ACE, Shannon, Simpson) were computed but no statistical comparisons between host groups are described for these indices↳ Could also: Rarefaction to a common sequencing depth followed by permutation tests (e.g., Kruskal-Wallis with Dunn post-hoc) could also formally compare diversity between groups — Rarefaction equalizes sampling effort before comparing richness estimators; formal tests make the comparison interpretable beyond visual inspection, which matters when pool sizes and read depths differ across groups
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34177835
Paper: He W, Gao Y, Wen Y, Ke X, Ou Z, Li Y, He H, Chen Q. Detection of Virus-Related Sequences Associated With Potential Etiologies of Hepatitis in Liver Tissue Samples From Rats, Mice, Shrews, and Bats. Front Microbiol 2021. PMID 34177835 · PMC8221242 · doi:10.3389/fmicb.2021.653873
Type: Viral metagenomics (virome) of pooled liver/serum tissue from wild + lab rodents, shrews, bats. Sequenced on Illumina HiSeq 2×150 bp.
Named code / tools (Methods → "Metagenomic Sequencing and Analysis")
The RU "Code" pointer is github.com/najoshi/sickle — the read-trimmer named in Methods. Per BRIEF P16, applying this third-party tool to the paper's own data is fully valid reproduction. The full pipeline (each tool named in the paper):
- Sickle — quality cutoff 20; drop reads with >10 bp N; drop reads <50 bp.
- BWA — align to host genome, remove host-similar reads.
- BLASTN vs NCBI NT — taxonomic assignment of reads.
- IDBA-UD — de Bruijn assembly of short reads.
- Contigs filtered to min length 264 bp.
- Metagene — gene prediction; keep protein-coding genes ≥100 bp.
- CD-HIT — cluster genes (95% identity, 90% overlap) → non-redundant catalog.
- BLASTP vs NCBI NR — taxonomic annotation (E ≤ 1e-5).
- Contamination filtering: remove lab-component viral sequences (Asplund et al. 2019 Supp Tables 2–5); remove index-hopping (read count <0.1% of max library count).
Data
- PRJNA695121 — 11 metagenomic runs (SRR13569749–759), cross-animal liver viromes.
- PRJNA701687 — 4 runs (SRR13717604–607), R. norvegicus liver vs serum comparison.
- Deposited representative viral genomes: MW055869–MW055884, MW389532–MW389537 (GenBank).
IN SCOPE (pipeline-derived, deterministic, well-specified) — ATTEMPTED
| # | Reported result | Pipeline step | Feasibility |
|---|---|---|---|
| R1 | Per-library raw read counts (Supp Table 3) | SRA deposit | EASY — verify N runs + read_count vs ENA |
| R2 | QC-passing reads after Sickle (Q20, N>10, <50bp) | Sickle | EASY — exact named tool + params on exact data |
| R3 | Assembly: contigs, contigs ≥264 bp | IDBA-UD | MEDIUM — deterministic given reads |
| R4 | Non-redundant gene catalog size | Metagene + CD-HIT | MEDIUM |
| R5 | Virus-related sequences ARE detectable in liver viromes (central qualitative claim) | DIAMOND/BLAST vs viral DB | MEDIUM (qualitative, viral-subset DB) |
OUT OF SCOPE — NOT attempted (with reason)
- Exact taxonomy percentages ("85.24% / 86.67% / 74.27% no-result reads"; viral
family/genus counts 17/26, 15/18, 12/18, 14/24; per-genus relative abundances; %
nucleotide identities). These require BLASTN vs full NCBI NT + BLASTP vs full NR at the
2020–2021 database version. The NT/NR DB version is not pinned and not recoverable, so
these numbers are non-deterministic / not exactly reproducible even with a full run.
Recorded as
partial/uncheckable, not attempted for exact match. - PCR screening (Table 2) — wet-lab PCR prevalence across 1003 animals. Wet-lab.
- Phylogenetics (Figs 4–6), pegivirus minus-strand replication assay — wet-lab/manual.
- Deposited representative genomes (MW...) — Sanger-sequenced amplicons, wet-lab.
- PCA plots (Figs 2–3) — derived from the full NT/NR taxonomy table (out of scope input).
Honest reproducibility assessment
The front end of the pipeline (Sickle QC → IDBA-UD assembly) is clearly specified with the exact named tool + parameters and is reproducible on the exact deposited data. The headline biological numbers depend on full-NT/NR BLAST at an unrecoverable DB version and are therefore not exactly reproducible by design — at best qualitatively confirmable (viruses present, broad families). No integrated pipeline script is shipped; only the trimmer repo is named. Expect: front-end reproduced, taxonomy partial/uncheckable.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a preliminary, honestly-documented partial reproduction of a viral-metagenomics paper. The deterministic front-end (Sickle QC, IDBA-UD assembly, qualitative viral detection) is being reproduced on the exact public SRA data, and the central qualitative claim — virus-related sequences detectable in liver viromes — is supported. The headline quantitative results (no-result-read % such as 85.24/86.67/74.27 and family/genus counts 17/26, 15/18, 12/18, 14/24) are not comparable by design, because they require a full NCBI NT/NR database at an unpinned, unrecoverable 2020-21 version — a data-availability/version limitation on the authors'/external side, not a fabrication signal. The only concrete anomaly is a minor 11-vs-9 deposited-library discrepancy. Net: solid partial reproduction with explainable deviations → yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.