Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Responses of peritubular macrophages and the testis transcriptome profiles of peripubertal and adult rodents exposed to an acute dose of MEHP.

Toxicol Sci · 2024
66/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
66/100
Reproducibility score
0.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 28% of all assessed papers rank 830 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL 1:1 reproduction via the linked third-party tool GO_MWU (P16). CLEAN RE-RUN after the 2026-06-22 requeue (prior «infra» workdir reclaimed): conda env rebuilt and GO_MWU genuinely re-run from scratch on «our HPC». Described well enough? Partly - Methods name the pipeline (FastQC->Bowtie2 Rnor_6.0/mm10->Samtools->DESeq2->GO_MWU) but omit DEG thresholds and the GO_MWU input metric. GEO GSE229734 deposits the authors' own DESeq2 + GO_MWU tables, enabling a genuine GO_MWU re-run on the paper's own data. COMPUTE: «our HPC» SLURM «job» (prep, COMPLETED) + 2221781 (GO_MWU, COMPLETED), GO_MWU commit 67b0007, go.obo releases/2020-12-08, env R 4.3.3 + org.Rn/Mm.eg.db. RESULT vs PAPER: the GO_MWU step reproduces 1:1 in theme across ALL THREE ontologies - adult rat is immune-dominated (C3 within-tol: 4/5 paper top-5 BP terms in re-run top-4, FDR ~1e-14; deposit matches paper exactly; Spearman -log10FDR vs deposit 0.85-0.91 for BP/MF/CC) and mouse is weaker + metabolic/non-immune (C4 within-tol: 84 vs 729 sig BP terms). Provenance is exact (PROV): GO_MWU input = DESeq2 Wald 'stat' (r=1.0, diff=0.0); adult-rat analysis = the treatment_MEHP_vs_CO contrast. MISMATCH/FLAG: the paper's '>13,000 genes differentially expressed vs control' in peripubertal rats is NOT supported by the deposited DESeq2 - the MEHP-vs-control contrast yields only 140-1248 DEGs; ~13,000 corresponds to the AGE (pubertal vs adult) comparison. Likely mislabel/overstatement, flagged for human review (caveat: pubertal-only per-group model not deposited, so not fully refutable). NOT ATTEMPTED: from-FASTQ rebuild (counts/FASTQ absent, params underspecified); independent pubertal-rat GO_MWU re-run (input ranking not deposited); wet-lab assays (out of scope). Verdict: partial - strong genuine-compute 1:1 on the GO_MWU step (C3/C4 across BP/MF/CC) + exact provenance, one flagged headline mismatch (C1), one deposit-only confirmation (C2).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 66
    assessed: 2026-06-20 ⛓ ffe0f40ac770
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether the extent of MEHP-induced testicular injury correlates with peritubular macrophage (PTMφ) numbers, hypothesizing that rodent age/species groups known to be less sensitive to MEHP toxicity (adult rats, peripubertal mice) would not show the significant PTMφ increase seen in MEHP-sensitive peripubertal rats.

Core claims
  • MEHP exposure previously increased PTMφ numbers 6-fold in peripubertal Fischer rats finding
  • Adult Fischer rats have a 2-fold higher basal PTMφ level than peripubertal rats and show no significant PTMφ increase after MEHP exposure finding
  • Peripubertal C57BJ/6 mice have a 5-fold higher basal PTMφ level than peripubertal rats and show no significant PTMφ increase after MEHP exposure finding
  • Peripubertal rats show significant testis transcriptome changes after MEHP, adult rats show lesser changes, and peripubertal mice show only minor changes finding
  • PTMφ numbers are associated with rodent sensitivity to MEHP in an age- and species-dependent manner mechanism
  • MEHP causes acute spermatocyte apoptosis (increased apoptotic index) in peripubertal PND26 rats but not in adult PND75 rats finding
  • Increases in PLZF+ undifferentiated/differentiating spermatogonia correlate with PTMφ increases in MEHP-treated peripubertal rats, but not in less-sensitive groups finding
  • 3' Tag sequencing with DESeq2 differential expression and GO-MWU enrichment analysis was used to profile testis transcriptomes across ages/species method
Experimental setups
Assay System Perturbation Readout Platform
Immunofluorescence whole-mount seminiferous tubule staining (MHCII/PLZF) PND26 peripubertal Fischer CDF344 rat testis MEHP 700 mg/kg oral gavage vs vehicle (corn oil) PTMφ and PLZF+ spermatogonia counts per 10^5 pixel area Zeiss LSM 710 confocal microscope; ImageJ 1.49T; NIS.ai GA3
Immunofluorescence whole-mount seminiferous tubule staining (MHCII/PLZF) PND75 adult Fischer CDF344 rat testis MEHP 700 mg/kg oral gavage vs vehicle (corn oil) PTMφ and PLZF+ spermatogonia counts per 10^5 pixel area Zeiss LSM 710 confocal microscope; ImageJ 1.49T; NIS.ai GA3
Immunofluorescence whole-mount seminiferous tubule staining (F4/80/PLZF) PND26 peripubertal C57BJ/6 mouse testis MEHP 700 mg/kg oral gavage vs vehicle (corn oil) PTMφ and PLZF+ spermatogonia counts per 10^5 pixel area Zeiss LSM 710 confocal microscope; ImageJ 1.49T; NIS.ai GA3
Immunofluorescent staining of testis cross sections (Caspase-3) PND26 peripubertal Fischer CDF344 rat testis MEHP 700 mg/kg oral gavage vs vehicle Apoptotic index (% tubules with >3 Caspase-3+ germ cells)
Immunofluorescent staining of testis cross sections (Caspase-3) PND75 adult Fischer CDF344 rat testis MEHP 700 mg/kg oral gavage vs vehicle Apoptotic index (% tubules with >3 Caspase-3+ germ cells)
3' Tag RNA sequencing PND26 peripubertal and PND75 adult Fischer CDF344 rat testis MEHP 700 mg/kg oral gavage vs vehicle Differential gene expression, PCA clustering, GO enrichment Bowtie2 alignment (Rnor_6.0), Samtools, DESeq2, GO-MWU
3' Tag RNA sequencing PND26 peripubertal C57BJ/6 mouse testis MEHP 700 mg/kg oral gavage vs vehicle Differential gene expression, PCA clustering, GO enrichment Bowtie2 alignment (mm10), Samtools, DESeq2, GO-MWU
Key results
  • Apoptotic index in peripubertal rats increased significantly from 15 to 28 after MEHP exposure ~1.87-fold
  • Apoptotic index did not significantly increase in adult PND75 rats after MEHP exposure
  • PTMφ numbers previously shown to increase 6-fold in peripubertal rats after MEHP 6-fold
  • Adult rats had ~1.8 PTMφ per 10^5 pixel area basally (2-fold higher than peripubertal rats) with no significant increase after MEHP 2-fold (basal)
  • Peripubertal mice had ~13 PTMφ per 10^5 pixel area basally (5-fold higher than peripubertal rats) with no significant increase after MEHP 5-fold (basal)
  • PCA showed distinct MEHP vs control clusters in peripubertal and adult rats, but randomly distributed (non-separated) clusters in mice
  • PLZF+ spermatogonia increased with MEHP in peripubertal rats correlating with PTMφ increase, but not significantly changed in adult rats or mice
Key statistics
  • fold_change 6-fold (PTMφ increase in peripubertal rats after MEHP (prior study))
  • fold_change 2-fold (Adult rat basal PTMφ level vs peripubertal rat basal level)
  • fold_change 5-fold (Peripubertal mouse basal PTMφ level vs peripubertal rat basal level)
  • mean 1.8 PTMφ per 10^5 pixel area (Adult rat basal PTMφ level)
  • mean 13 PTMφ per 10^5 pixel area (Peripubertal mouse basal PTMφ level)
  • other AI increased from 15 to 28 (Apoptotic index in peripubertal rats after MEHP)
  • count n=7 per treatment group (Sample size per group across experiments)
  • pvalue p<.05 (Significance threshold for all statistical tests)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study combined immunofluorescence quantification of peritubular macrophages and spermatogonia in whole seminiferous tubule preparations with 3′ Tag RNA sequencing of testicular tissue across three rodent groups (peripubertal rats, adult rats, peripubertal mice) exposed to a single dose of MEHP or vehicle (n = 7 biological replicates per group). Bench-level comparisons (cell counts, apoptotic index) were tested with Student's t-test or one-way ANOVA with Tukey post-hoc; transcriptomic differential expression was performed with DESeq2 and GO enrichment with the GO-MWU package. Results were presented as individual data points with means ± SEM, with p < .05 as the significance threshold.

Replicationbiological Sample sizen = 7 animals per treatment group stated in figure legends and methods; no a priori power calculation described GroupsControl vs MEHP-treated within each of three rodent groups: PND 26 Fischer rats, PND 75 Fischer rats, PND 26 C57BJ/6 mice Pairingunpaired Randomization/blindingstated DispersionSEM Effect sizesno Confidence intervalsno Multiplicity correctionTukey HSD (for ANOVA post-hoc comparisons of bench data); DESeq2 default multiple testing correction not explicitly named in the statistics section
Statistical tests used
Test Applied to n Assumptions
Student's t-test (two-group, parametric) Pairwise comparisons of apoptotic index, PTMφ counts, and PLZF+ spermatogonial counts between control and MEHP-treated animals within each age/species group n = 7 biological replicates per group not stated
One-way ANOVA followed by Tukey post-hoc test Multi-group comparisons of cell count outcomes across age and species groups n = 7 biological replicates per group not stated
DESeq2 (negative binomial Wald test) Differential gene expression analysis of 3′ Tag sequencing data, control vs MEHP-treated within each rodent group n = 7 biological replicates per group (2 peripubertal MEHP-treated rats noted as potential non-respondents by PCA) not stated
GO-MWU (Mann-Whitney U-based GO enrichment) Gene Ontology enrichment analysis of DESeq2 results to identify age- and species-specific differences null not stated
Principal component analysis (PCA) Exploratory visualization of transcriptomic sample clustering for rats and mice separately n = 7 per group na
Approaches that could also have been used
  • Dispersion around means was reported as SEM throughout, with n = 7 per group
    Could also: SD or 95% confidence intervals could also be used to summarize spread — With small group sizes (n = 7), SD conveys the actual variability of individual animals, and a 95% CI conveys estimation uncertainty; SEM shrinks with larger n and can visually minimize spread, so SD or CI is often preferred when communicating biological variability in small-n preclinical studies
  • Multiple pairwise Student's t-tests were used alongside one-way ANOVA and Tukey post-hoc for the bench outcomes
    Could also: A single two-way ANOVA (age/species × treatment) with interaction term could also be applied to the full design — A two-way ANOVA would formally test whether the treatment effect differs by age or species (the interaction), which is the central biological question, and would control the family-wise error rate across all group comparisons in a single model
  • Normality of the cell-count and apoptotic-index outcomes was not assessed or stated before applying Student's t-test and ANOVA
    Could also: Non-parametric alternatives such as the Mann-Whitney U test (for two groups) or Kruskal-Wallis with Dunn post-hoc (for multiple groups) could also be used — With n = 7, parametric assumptions (normality, homoscedasticity) are difficult to verify; non-parametric tests make fewer distributional assumptions and are commonly applied to immunofluorescence count data in small preclinical studies
  • Two potential non-responders in the peripubertal MEHP-treated rat group were identified by PCA but their handling in downstream differential expression analysis is not described
    Could also: Sensitivity analyses including and excluding the flagged samples, or outlier-robust differential expression methods, could also be reported — Documenting how apparent non-responders affect the DESeq2 results allows readers to gauge the robustness of the transcriptomic findings to this source of variability
  • DESeq2 was used for differential expression; the multiple-testing correction method applied is not explicitly stated in the statistics section
    Could also: edgeR or limma-voom are also standard tools for count-based RNA-seq differential expression and are widely used as comparators — Explicitly stating the FDR method and threshold used in DESeq2, and optionally cross-checking with a second tool, is common practice in transcriptomic studies to assess robustness of the gene lists
  • Cell counts were normalized to tubule area and averaged across 10 tubules per animal, with the animal treated as the statistical unit
    Could also: A mixed-effects model with tubule nested within animal could also be used to account for within-animal correlation while retaining tubule-level information — Mixed models use all tubule-level observations without collapsing to per-animal means, potentially increasing power while still correctly treating the animal as the unit of inference; the authors' per-animal averaging approach is a valid and common simplification
Software: GraphPad Prism 5.0 · R 4.2.0 · Python · DESeq2 · GO-MWU · Bowtie2 · Samtools · ImageJ 1.49T · FastQC

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38113427

Paper: Fang X, Tiwary R, Nguyen VP, Richburg JH. Responses of peritubular macrophages and the testis transcriptome profiles of peripubertal and adult rodents exposed to an acute dose of MEHP. Toxicol Sci. 2024. DOI 10.1093/toxsci/kfad128 · PMCID PMC10901151 · GEO GSE229734.

Linked code (P16, third-party tool — equally valid): GO_MWU (https://github.com/z0on/GO_MWU), GO enrichment by Mann-Whitney U on a continuous gene ranking. Author: M. Matz.

Experimental design

3 rodent groups × (MEHP 700 mg/kg single oral gavage vs corn-oil vehicle), testis sampled 48 h post-dose, Tag-seq, n≈7/group:

  • Peripubertal (PND26) Fischer CDF344 rat
  • Adult (PND75) Fischer CDF344 rat
  • Peripubertal (PND26) C57BL/6 mouse

Reported computational pipeline (Methods)

FastQC → adapter/quality trim (drop <50 bp) → Bowtie2 align (Rnor_6.0 rat, mm10 mouse) → Samtools → DESeq2 (differential expression) → GO_MWU (GO enrichment on the DE ranking, BP). GO-MWU display threshold: adj p ≤ 1e-04.

In scope (pipeline-derived, attempted)

# Reported result Pipeline Approach
C1 Peripubertal rat: ">13,000 genes differentially expressed vs control" DESeq2 Count DEGs from deposited GSE229734_deseq_results_rat.xlsx at the implied threshold; sanity-check magnitude.
C2 GO_MWU top enriched GO terms, peripubertal rat (top-5: regulation of locomotion; regulation of cellular component movement; regulation of immune system process; neg/pos regulation of multicellular organismal process) GO_MWU Re-run the linked GO_MWU tool on the deposited DESeq2 rat result (signed −log10 p ranking), compare top terms vs paper + vs deposited GSE229734_gomwu_pubertal_rat.xlsx.
C3 GO_MWU top terms, adult rat (top-5 immune-dominated) GO_MWU Same, vs GSE229734_gomwu_adult_rat.xlsx.
C4 GO_MWU mouse: weaker enrichment (log10FC all <10) GO_MWU Same, vs GSE229734_gomwu_mouse.xlsx.

GEO deposits the authors' own DESeq2 results (deseq_results_rat/mouse.xlsx) and GO_MWU files (gomwu_*_rat/mouse.xlsx), so the GO_MWU step can be reproduced 1:1 on the paper's own data — exactly the P16 case (run the linked tool on the paper's data).

Out of scope (not attempted) — and why

  • Raw read alignment / DESeq2 from FASTQ. Bowtie2 on raw Tag-seq + re-deriving the count matrix is the hard last ~20%; the count matrix is not deposited on GEO (only DESeq2 result tables + GO_MWU files), so a from-FASTQ rebuild is not cleanly specified (no exact trim params, no count step detail). Skipped per 80/20.
  • Wet-lab: immunostaining of peritubular macrophages, flow cytometry, qPCR — manual/bench, not pipeline.
  • Exact DEG threshold for ">13,000": the paper gives no padj/FC cutoff, so C1 is a magnitude check, graded partial unless a single threshold reproduces it.
Figures / tables: figurefigures
C1
Reported
peripubertal Fischer rat: 'more than 13 000 genes ... was differentially expressed from that of the control group' (Results)
Reproduced
Recounted from deposited DESeq2 (GSE229734_deseq_results_rat.xlsx): treatment_MEHP_vs_CO = 140 (padj<0.05) / 190 (padj<0.1) / 1248 (p<0.05); interaction = 17/40/1382. The ~13,000 figure matches ONLY the age_pubertal_vs_adult contrast (12,897/13,682/13,313). FLAG: the headline DEG count attributed to MEHP-vs-control corresponds in the deposit to the AGE comparison, not the treatment effect. Caveat: a pubertal-only per-group DESeq2 is not deposited, so not fully refutable.
did not match
C2
Reported
peripubertal rat GO-MWU top BP terms (regulation of locomotion; regulation of cellular component movement; regulation of immune system process; +/- regulation of multicellular organismal process)
Reproduced
Deposited GSE229734_gomwu_pubertal_rat BP top rows match the paper. INDEPENDENT RE-RUN NOT POSSIBLE: the pubertal-rat GO_MWU input (per-group DESeq2 stat) is not deposited (matches no deposited contrast; Pearson r vs deposited treatment stat ~0.57, max_abs_diff 6.0).
partial
C3
Reported
adult rat GO enrichment immune-dominated (top BP: immune system process; regulation of immune system process; defense response; immune response; response to biotic stimulus)
Reproduced
GENUINE independent GO_MWU re-run («our HPC» «job», COMPLETED) on the deposited treatment_MEHP_vs_CO Wald stat reproduces an immune-dominated top set: regulation of immune system process (FDR 1.5e-14); response to biotic stimulus; immune response; defense response; ... 4 of the paper's 5 top-5 terms appear in the re-run top-4; 729 BP terms at FDR<0.1 (deposit 619). Spearman of -log10(FDR) over shared GO ids = 0.887. Also reproduced for MF (99 vs 73, rho 0.895) and CC (121 vs 89, rho 0.911). Deposited adult-rat GO_MWU top-5 match the paper exactly.
within tolerance
C4
Reported
mice show substantially weaker GO enrichment; top-15 GO log10 fold-change: pubertal rat 20-30, adult rat 10-20, mouse all below 10
Reproduced
GENUINE re-run («job»): mouse BP = 84 terms at FDR<0.1 (vs 729 adult rat); top terms metabolic/biosynthetic (organonitrogen compound biosynthetic process; small molecule metabolic process; ...), NOT immune; Spearman vs deposit 0.846. Mouse MF=11, CC=48 (also much fewer than rat). Deposited log10(FDR) bands: pubertal rat 18-30, adult rat 9-19, mouse 3-9 (matches the 20-30 / 10-20 / <10 structure).
within tolerance
PROV
Reported
(not stated in paper) GO_MWU input ranking metric unspecified in Methods
Reproduced
GO_MWU input measure = DESeq2 Wald 'stat'. The independent re-run on the deposited treatment_MEHP_vs_CO stat reproduces the deposited adult-rat GO_MWU output (input==stat: Pearson r=1.0, max_abs_diff=0.0 over 3493 shared genes), confirming the adult-rat analysis = the combined-model treatment contrast; mouse input==deposited mouse mehp-vs-co stat (r=1.0); pubertal-rat input = non-deposited per-group model (r~0.57).
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

749.6 k
tokens (I/O) · 53 M incl. cache
268 min
runtime · 1.83 CPU-h
19.9 GB
peak RAM
2
HPC jobs
hummel
machine