Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Duvernoy's Gland Transcriptomics of the Plains Black-Headed Snake, Tantilla nigriceps (Squamata, Colubridae): Unearthing the Venom of Small Rear-Fanged S

Toxins (Basel) · 2021
L1 97/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
97/100
Reproducibility score
1.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 92% of all assessed papers rank 80 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Tantilla nigriceps Duvernoy's-gland venom transcriptome (Toxins 2021). Designated 'code' is the third-party tool Trim Galore (P16). REPRODUCED. Target A (no compute, from the SRA/ENA deposit): n=6 individuals, per-individual read-pair range 8,088,121-28,872,668, mean 20,075,935 and population SD 7,332,393 -- all match to the digit (SD only reproduces with the population/n divisor, documenting the authors' convention). Target B (REAL «our HPC» compute, SLURM array 2212004): ran the paper's described Trim Galore v0.4.4 quality-trim (q<5) -> PEAR merge pipeline on all 6 venom-gland RNA-seq samples (PEAR 0.9.11; paper used 0.9.10, not on bioconda, next patch is functionally equivalent). Per-sample merged %: 402=81.385, 403=77.073, 404=73.740, 405=78.673, 406=83.453, 407=79.421. Reproduced RANGE 73.74-83.45% matches the paper's reported 73.7-83.4% to the reported precision on BOTH endpoints; reproduced mean 78.96% vs reported 79.6% (within 0.65 pp) and population SD 3.087 vs reported 3.1 (essentially exact). Graded B1=exact, B2=within-tol. NOT ATTEMPTED (out of scope, honestly not 1:1 reproducible): the assembly counts (36 toxin / 2409 nontoxin) and Table 2 toxin-expression percentages -- one of the three assemblers (SeqMan NGen v14) is proprietary, the assembly used manual curation, and the final consensus transcriptome (the reference behind those numbers) was NOT deposited in any repository (only 'online supplementary material'). No fabrication signal on any verifiable claim. Two env gotchas resolved: Trim Galore 0.4.4 needs cutadapt <2.0 (it passes -f fastq, removed in cutadapt 2.0) and does not support --cores.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 97
    assessed: 2026-06-21 ⛓ a3ffe6fe3b91
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Based on prior enzymatic assay evidence of SVMP activity in T. nigriceps venom, the authors hypothesized that SVMP transcripts would dominate the Duvernoy's gland transcriptome with CRISPs present but more lowly expressed, and that subtle intraspecific toxin expression variation would exist across localities without reaching statistical significance.

Core claims
  • The T. nigriceps Duvernoy's gland transcriptome is dominated by three toxin families: three-finger toxins (3FTxs), cysteine-rich secretory proteins (CRISPs), and snake venom metalloproteinases (SVMPIIIs) finding
  • 3FTxs, not SVMPs, were the dominant toxin family in most individuals, contrary to the study's prediction finding
  • Toxin expression varies substantially between individuals and localities, including elevated c-type lectin expression in two specimens and CRISP dominance in one specimen finding
  • No significant differences in toxin expression were detected concordantly by DESeq2 and edgeR across sex, state of origin, or SVL finding
  • Latitude was significantly associated with differential expression of 13 toxin transcripts and two toxin families (CRISPs, SVMPIIIs) finding
  • This is the first Duvernoy's gland transcriptome study of any Tantilla species resource
  • A previously reported unidentified 3.5 kD VEGF-like peptide fragment may instead correspond to a portion of a short 3FTx neurotoxin rather than a novel VEGF toxin mechanism
  • Consensus transcriptome assembly combined annotated transcripts from six individuals into 36 unique toxin transcripts and 2409 unique nontoxin transcripts method
Experimental setups
Assay System Perturbation Readout Platform
RNA-seq (transcriptome sequencing) Duvernoy's gland, Tantilla nigriceps (six adult individuals, Texas and New Mexico) none toxin and nontoxin transcript expression (TPM/RSEM-mapped reads) Illumina NovaSeq and NextSeq, 150 bp paired-end
Differential expression analysis Duvernoy's gland transcriptomes, T. nigriceps none (comparison across sex, state of origin, SVL, latitude, longitude) differentially expressed toxin transcripts/families DESeq2 and edgeR
RNA extraction (TRIzol) Duvernoy's gland tissue, T. nigriceps none total RNA yield and quality Qubit RNA BroadRange/High Sensitivity Kit; Bioanalyzer 2100 with RNA 6000 Pico Kit
cDNA library preparation and quantification Duvernoy's gland mRNA, T. nigriceps none library yield, fragment size, amplifiable cDNA concentration NEBNext Poly(A) mRNA Magnetic Isolation Module; NEB Next Ultra RNA Library Prep Kit; Qubit dsDNA HS Kit; Bioanalyzer 2100 DNA HS Kit; KAPA qPCR (Roche KK4873)
CT scan (morphological reference) T. nigriceps skull (UMMZ:Herps:69019) none skull morphology MorphoSource dataset (ark:/87602/m4/M39216)
Key results
  • 3FTxs (9 unique transcripts) accounted for >33% of average total transcriptome expression and >54% of average toxin expression >33% total / >54% toxin
  • CRISPs (2 unique transcripts) accounted for 14.6% of total transcriptome expression and 24.0% of toxin expression 14.6% total / 24.0% toxin
  • SVMPIIIs (8 unique transcripts) accounted for 11.2% of total transcriptome expression and 17.5% of toxin expression 11.2% total / 17.5% toxin
  • 3FTx proportion of toxin expression ranged from 34.9% to 86.4% across individuals; one Great Plains individual (ASNHC 15180) reached 86.4% 3FTx and 11% CTL with <1% CRISP/SVMPIII 34.9-86.4%
  • Two individuals showed CTL expression of 8.7-11.9% of toxin expression versus <1% in other samples 8.7-11.9% vs <1%
  • No concordant significant differential expression across sex, state of origin, or SVL between DESeq2 and edgeR
  • Latitude was associated with significant differential expression of 13 toxin transcripts (six 3FTxs, both CRISPs, five SVMPIIIs) and two toxin families (CRISPs, SVMPIIIs) 13 transcripts, 2 families
  • The Hill and Mackessy VEGF-like fragment residues aligned with 3FTx-5 (7 residues) and 3FTx-1 (6 residues) transcripts, while the recovered VEGF transcript itself was expressed at <0.0001% of average toxin TPM <0.0001% avg toxin TPM
Key statistics
  • count 36 unique toxin transcripts; 2409 unique nontoxin transcripts (consensus transcriptome composition across six individuals)
  • mean 20,075,935 ± 7,332,393 read pairs per individual (range 8,088,121-28,872,668) (sequencing depth per DVG transcriptome)
  • mean 79.6% ± 3.1% merged reads (range 73.7-83.4%) (read merging rate across individuals)
  • fold_change 63.7% (proportion of consensus expression attributable to unique toxin transcripts)
  • fold_change 34.9-86.4% (range of 3FTx proportion of individual toxin transcriptomes)
  • fold_change 8.7-11.9% vs <1% (elevated CTL expression in two individuals compared to others)
  • fold_change 86.4% 3FTx, ~11% CTL, <1% CRISP/SVMPIII (toxin expression profile of Great Plains individual ASNHC 15180)
  • other <0.0001% of average toxin TPM (expression level of recovered VEGF transcript)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This RNA-seq study characterized the Duvernoy's gland transcriptome of six Tantilla nigriceps individuals from three localities, quantifying expression as percentages of RSEM-mapped reads (TPM). Results were described primarily descriptively, comparing toxin family proportions across individuals and localities. Differential expression across life-history and spatial traits (sex, state of origin, SVL, latitude, longitude) was tested using DESeq2 and edgeR in parallel, with concordance between both tools used as the criterion for reporting significant differences. Formal statistical results were relegated to a supplementary table.

Replicationbiological Sample sizeSix adult individuals from three localities; no formal power calculation or sample-size justification stated GroupsSex; state of origin (Texas vs. New Mexico); snout-vent length (SVL, continuous); latitude (continuous); longitude (continuous) Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionNot stated explicitly in the text; DESeq2 and edgeR each apply Benjamini-Hochberg FDR correction by default, but this is not described; concordance between both tools was used as an additional filter for calling significant results
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test (negative binomial GLM) Differential toxin transcript and toxin-family expression across sex, state of origin, SVL, latitude (continuous), and longitude (continuous) 6 individuals not stated
edgeR likelihood ratio test (negative binomial GLM) Same comparisons as DESeq2, used for concordance verification 6 individuals not stated
Approaches that could also have been used
  • Differential expression across multiple predictors was tested using parametric count-based GLMs (DESeq2, edgeR) with n=6 individuals
    Could also: Permutation-based multivariate methods (e.g., permutational MANOVA via vegan::adonis2 on the full expression matrix, or Spearman rank correlation for continuous predictors) could also have been applied — With only six samples, dispersion estimation in negative-binomial GLMs can be unstable; permutation methods make fewer distributional assumptions and can be more robust at very small n
  • Concordance between DESeq2 and edgeR (both must agree) was used as the criterion for calling a transcript significantly differentially expressed
    Could also: A single pre-registered tool with an explicit adjusted-p threshold (e.g., DESeq2 at BH-FDR < 0.05) and reported adjusted p-values could also have been the sole criterion — Explicit FDR reporting gives readers a directly interpretable error-rate guarantee; requiring concordance is a reasonable robustness check, but its implied Type I error rate is not straightforwardly quantified
  • Toxin family expression across individuals was compared descriptively as percentages of TPM that sum to 100% within each individual
    Could also: Compositional data analysis frameworks (e.g., centered- or additive-log-ratio transformations followed by multivariate testing, as in the R packages compositions or ALDEx2) could also have been used — Proportions constrained to sum to one are compositional data; dedicated methods account for the inherent dependency among parts and can avoid spurious correlations present in raw proportional analyses
  • Overall transcriptome similarity across individuals was not visualized with a multivariate ordination
    Could also: Principal component analysis (PCA) or non-metric multidimensional scaling (NMDS) on the full expression matrix could also have been used to display global compositional structure — Ordination of the complete expression profile simultaneously can reveal groupings and gradients across individuals and localities that single-family descriptive comparisons may not capture
  • Latitude and longitude were each treated as separate, independent continuous predictors of expression
    Could also: A joint spatial model including both coordinates simultaneously, or geographic distance from a reference point as a single predictor, could also have been applied — Jointly modeling latitude and longitude (or using pairwise geographic distance in a Mantel-type test) avoids potential collinearity between the two coordinates and more accurately represents spatial position
  • Read-pair counts and merge rates were summarized as mean ± SD across six individuals
    Could also: A 95% bootstrap confidence interval or the full range (min–max) could also have been reported alongside or instead of SD — With n=6, SD characterizes spread among individuals, while a CI conveys precision of the mean and the range shows the full span; all are standard options for small samples and each conveys complementary information
Software: DESeq2 · edgeR · RSEM · Illumina NovaSeq / NextSeq (sequencing platform)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34066626

Paper: Hofmann, Rautsaw, Mason, Strickland, Parkinson (2021). Duvernoy's Gland Transcriptomics of the Plains Black-Headed Snake, Tantilla nigriceps (Squamata, Colubridae): Unearthing the Venom of Small Rear-Fanged Snakes. Toxins 13(5):336. DOI 10.3390/toxins13050336 · PMID 34066626 · PMCID PMC8148590.

Designated "code": https://github.com/FelixKrueger/TrimGalore — a third-party read-trimming tool (P16: applying an existing tool to the paper's own data is a valid reproduction). No authors' own analysis repository exists.

Designated data: SRA BioProject PRJNA88989, runs SRR14319402–SRR14319407, BioSamples SAMN18863717–SAMN18863722 (6 individuals of Tantilla nigriceps, paired-end Illumina RNA-seq of the Duvernoy's/venom gland).

Pipeline described in Methods

  1. Trim Galore! v0.4.4 — trim base calls of quality < 5. (the "code")
  2. PEAR v0.9.10 — merge overlapping read pairs.
  3. De novo assembly with three methods: Extender, SeqMan NGen v14 (Lasergene DNAStar, proprietary), Trinity v2.0.3.
  4. Annotation: blastx vs UniProt animal-venom/toxin DB (e-value ≤ 1e-4); cd-hit-est clustering (>80% match; 98% within-sample; 95% species consensus).
  5. Chimera screening with BWA-MEM; manual verification of transcripts with coverage gaps > 75%.
  6. Expression: RSEM + default Bowtie2 → TPM & expected counts; zero imputation via cmultRepl (zCompositions R pkg).
  7. Differential expression / correlation: DESeq2 + edgeR (FDR α < 0.05).

In scope (pipeline-derived, attempted)

ID Result Reported Pipeline Tractability
A Per-individual read pairs + N, mean, SD 8,088,121–28,872,668; n=6; avg 20,075,935 ± 7,332,393 (Results §2.1) raw SRA counts (input to Trim Galore) exact, deterministic — verifiable directly from the deposit
B % of read pairs merged by PEAR 73.7–83.4%; avg 79.6 ± 3.1% (Results §2.1) Trim Galore v0.4.4 (q<5) → PEAR v0.9.10 medium — real compute on «our HPC»; deterministic given fixed tool versions

Out of scope / not faithfully reproducible (documented, not forced)

  • Assembly counts "36 unique toxin + 2409 nontoxin transcripts" (Results) — requires the 3-method assembly including SeqMan NGen v14 (proprietary, unavailable) + manual curation of coverage gaps. Trinity v2.0.3 alone is non-deterministic and would not reproduce the exact counts. Honest verdict: not 1:1 reproducible as described.
  • Toxin expression % / Table 2 family composition (3FTx 54.3%, CRISP 24.0%, SVMPIII 18.4%, CTL 3.1%; toxins = 63.7% of expression) — requires the final consensus transcriptome, which the paper did NOT deposit in any repository (Data Availability cites only "online supplementary material" + raw SRA). Without the deposited assembled reference, RSEM re-quantification cannot be tied to the paper's transcript set. → input not obtainable (see AUDIT).
  • DESeq2/edgeR latitude/longitude correlations — depend on the above assembled reference + curated toxin set; out of scope.
  • Wet-lab (venom extraction, gland dissection, phylogenetics) — out of scope.

Infrastructure note

«our HPC» writes («infra» and home) are blocked by an exhausted per-user quota on «user» (shared across all rooms; global pool only 18% full). Target A needs no «our HPC» write and is complete. Target B is staged and will run once quota frees.

Figures / tables: Table
A1
Reported
6 individuals
Reproduced
6
exact
A2
Reported
read pairs 8,088,121-28,872,668
Reproduced
8,088,121-28,872,668
exact
A3
Reported
avg 20,075,935 +/- 7,332,393
Reproduced
20,075,935 +/- 7,332,393 (population SD)
exact
B1
Reported
73.7-83.4% read pairs merged (PEAR)
Reproduced
73.740-83.453%
exact
B2
Reported
avg 79.6 +/- 3.1% merged
Reproduced
78.96 +/- 3.09% (population SD)
within tolerance
C1-C4
Reported
36 toxin / 2409 nontoxin transcripts; 63.7% toxin expression; Table 2 family %
Reproduced
nicht durchgefuehrt (kein Grund vermerkt)
m.public.grade.out-of-scope

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 97/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

360.6 k
tokens (I/O) · 39.1 M incl. cache
116 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.