The butyrophilin 1a1 knockout mouse revisited: Ablation of Btn1a1 leads to concurrent cell death and renewal in the mammary epithelium during lactation.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. Jeong et al. 2021 FASEB BioAdv -- Btn1a1 KO vs WT mouse mammary RNA-seq at lactation (GSE182075, 5 WT + 7 KO). In-scope pipeline = STAR->RSEM->voom/limma DE, DEGs at |FC|>=2 & adj.p<=0.05; listed code CCBR/Pipeliner is a generic third-party tool (P16). EVIDENCE at increasing independence: (L1) applying the paper's exact thresholds to the deposited limma table -- which uniquely ships the per-sample log2CPM matrix for all 12 samples -- yields EXACTLY 113 up / 53 down and all 7 Table-1 top genes to every printed digit, incl. Btn1a1 as the top down-regulated gene at -43.21 (adj.p 4.95E-11). (L2) an independent limma re-run on the deposited log2CPM matrix reproduces per-gene logFC at r=0.999. (L3) a FULLY INDEPENDENT end-to-end re-quantification from raw FASTQ (12 ENA runs -> Salmon selective-alignment vs GENCODE M18, ISR, gene-level -> voom/limma) recovers the same biology: Btn1a1 is the #1 most-significant gene and top down-regulated transcript at -44.6 fold (~the reported -43.2), all 7 top genes are direction-concordant, 12,319 genes tested (vs 11,834), DEG count 187/32 (same up>down asymmetry as 113/53; exact count differs as expected for a different quantifier). FABRICATION CHECK: none -- every pinnable pipeline-derived value is supported by the deposited data and independently re-derivable from raw reads. DEVIATION: the paper's STAR+RSEM were unusable on this cluster (STAR index deadlock; broken bioconda RSEM) so L3 quantified with Salmon (documented). NOT ATTEMPTED (the hard ~20%): IPA canonical pathways (proprietary), GSEA 721-GO-set clustering (figure-level, underspecified), exact PCA/heatmap/volcano rendering, all wet-lab assays. All heavy compute on «our HPC»/«infra»; «host» holds only small result files + pointers. Grades provisional; a human reviewer decides (see AUDIT.md).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-15 ⛓ 7fa10b902ec6
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates whether ablation of Btn1a1 (which the authors hypothesize forms a membrane-based secretion complex with xanthine oxidoreductase, Xdh, to drive lipid droplet secretion) disrupts lipid secretion and produces broader downstream consequences for mammary epithelial gene/protein expression, cell survival, and renewal during lactation.
- ★ BTN1A1 functions as an apical membrane receptor that forms a lipid-droplet secretion complex with the redox enzyme xanthine oxidoreductase (XDH) mechanism
- ★ Ablation of Btn1a1 disrupts lipid droplet secretion, producing large unstable droplets released into alveolar spaces with fragmented surface membranes finding
- Original Btn1a1-/- KO lines showed reduced litter weights (60-80% of wild type) and up to half of pups dying before weaning finding
- ★ Disruption of Btn1a1 causes cytoplasmic build-up of Xdh, induction of acute phase response genes, and Lif-activated Stat3 phosphorylation finding
- ★ Approximately 10% of mammary epithelial cells are dying at peak lactation in Btn1a1-/- mice, assessed by TUNEL finding
- ★ Cell death in Btn1a1-/- epithelium proceeds via multiple pathways including caspase 8/activated caspase 3, autophagy, Slc5a8-mediated survivin (Birc5) inactivation, and pStat3-mediated lysosomal lysis mechanism
- ★ Milk secretion is prolonged in Btn1a1-/- mice through compensatory renewal of the secretory epithelium, marked by Ki67 upregulation in ~10% of nuclei and expression of cyclins and Fos/Jun finding
- ★ RNAseq and proteomic analysis of the Btn1a1-/- KO2 mouse line was used to characterize gene/protein expression changes upon Btn1a1 loss method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-seq | mammary gland tissue, Btn1a1+/+ (n=5) vs Btn1a1-/- (n=7) mice, days 9-11 lactation | Btn1a1 knockout | differential gene expression | Illumina HiSeq 4000 with TruSeq mRNA library prep, STAR alignment to mm10 |
| Proteomic analysis | mammary gland tissue, Btn1a1+/+ vs Btn1a1-/- mice | Btn1a1 knockout | differential protein expression | LC-MS with MS-grade trypsin digestion |
| qRT-PCR | mammary gland tissue, day 5 lactation, Btn1a1+/+ vs Btn1a1-/- mice (n=6 each) | Btn1a1 knockout | mRNA expression of selected genes, normalized to Slc44a3/SPG21 | iCycler iQ Real-time PCR Detection System (Bio-Rad) |
| Immunoblot (Western blot) | mammary gland tissue fractions (total homogenate, post-nuclear supernatant, microsomal membrane), Btn1a1+/+ vs Btn1a1-/- mice | Btn1a1 knockout | relative protein levels (e.g., Xdh, Stat3/pStat3, Stat5/pStat5, cathepsin B, oxidative damage markers) | ChemiDoc Imaging System with QuantityOne densitometry |
| Immunofluorescence on unfixed frozen sections | mammary gland tissue, Btn1a1+/+ and Btn1a1-/- mice | Btn1a1 knockout; digitonin permeabilization | Xdh localization, lipid droplets (Nile red), apical membranes (WGA) | Leica SP5X confocal laser scanning microscope |
| Immunohistochemistry (paraffin-embedded, DAB) | mammary gland tissue, Btn1a1+/+ vs Btn1a1-/- mice | Btn1a1 knockout | pStat3, pStat5, and Ki67 nuclear staining | ABC VectorStain, light microscopy |
| TUNEL assay | mammary gland cryosections, Btn1a1+/+ and Btn1a1-/- mice at day 10 lactation and Btn1a1+/+ at day 2 involution | Btn1a1 knockout / involution | percentage of apoptotic (fluorescein-positive) nuclei | Leica SP5X confocal microscope, ImagePro software |
| Electron microscopy | mammary gland tissue | Btn1a1 knockout | ultrastructure of lipid droplet secretion and membrane integrity | Zeiss EM10CA electron microscope |
- ▲ Cytoplasmic accumulation of Xdh in Btn1a1-/- mammary epithelium
- ▲ Induction of acute phase response genes in Btn1a1-/- mammary gland
- ▲ Lif-activated Stat3 phosphorylation in Btn1a1-/- mammary gland
- ▲ Approx. 10% of mammary epithelial cells undergoing cell death at peak lactation, assessed by TUNEL ~10%
- ▲ Ki67 upregulation in approx. 10% of cell nuclei, indicating compensatory epithelial renewal ~10%
- ▼ Litter weights reduced in original Btn1a1-/- KO lines 60%-80% of wild type
- ▼ Up to half of pups died before weaning in original Btn1a1-/- KO characterization
- – RNAseq samples met high quality thresholds for integrity and sequencing depth
- count 5 Btn1a1+/+ and 7 Btn1a1-/- mice (RNAseq sample size (days 9-11 lactation))
- other RIN 8.2-9.6 (RNA integrity numbers for RNAseq samples)
- count 95-167 million pass filter reads (RNAseq sequencing depth per sample)
- other >92% of bases above quality score Q30 (RNAseq base call quality)
- other >97% average mapping rate; >80% unique alignment; 2.09%-2.63% unmapped reads (RNAseq alignment to mm10 reference)
- fold_change 60%-80% of wild type litter weight (Original Btn1a1-/- KO line litter weight reduction)
- count 6 Btn1a1+/+ and 6 Btn1a1-/- mice, analyzed in duplicate (qRT-PCR validation sample size)
- other 9%-16% non-duplicate reads (RNAseq library complexity (Picard MarkDuplicate))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper combined RNAseq (n=5 WT, 7 KO), qRT-PCR (n=6 per group), proteomics, immunoblot densitometry, and histological cell-counting assays (TUNEL, Ki67 immunohistochemistry) to characterize the mammary epithelial phenotype of Btn1a1−/− mice at peak lactation versus wild-type controls. RNAseq reads were aligned with STAR v2.4.2a via the CCBR Pipeliner pipeline; the specific differential-expression statistical model is not named in the provided text. Cell-death and proliferation endpoints were reported as approximate percentile ratios of positive nuclei (~10% for both TUNEL and Ki67), while protein levels were assessed by densitometry normalized to total Ponceau S staining and a reference antigen.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Not stated (RNAseq differential expression; alignment via STAR/CCBR Pipeliner; DE model not named in provided text) | Genome-wide mRNA differential expression, Btn1a1−/− vs Btn1a1+/+ mammary gland (days 9–11 of lactation) | 5 Btn1a1+/+ and 7 Btn1a1−/− mice | not stated |
| qRT-PCR relative quantification with geometric mean normalization to two reference genes (Slc44a3, SPG21); between-group inferential test not named | Selected gene expression validation, day 5 lactation mammary gland | 6 Btn1a1+/+ and 6 Btn1a1−/− mice, each run in duplicate | not stated |
| TUNEL assay (fluorescein-positive nucleus counting reported as percentile ratio, aided by ImagePro software) | Apoptotic cell frequency in mammary gland cryosections at day 10 lactation (Btn1a1+/+ and Btn1a1−/−) and day 2 involution (Btn1a1+/+) | not stated | not stated |
| Ki67 immunohistochemistry (positive nucleus counting, approximate percentage reported descriptively) | Proliferating-cell frequency in mammary gland paraffin sections | not stated | not stated |
| Immunoblot densitometry (relative protein quantification normalized to Ponceau S total-protein staining and a reference antigen set to 100%) | Xdh, pStat3, pStat5, and other proteins in total homogenate or subcellular fractions | not stated | not stated |
-
The differential-expression statistical model applied to RNAseq count data is not named in the provided text, though STAR alignment and the CCBR Pipeliner are described↳ Could also: DESeq2 (negative-binomial Wald test with shrinkage estimation) or edgeR (quasi-likelihood F-test) are standard, widely-cited methods for two-group RNAseq comparisons with unequal sample sizes — Naming the DE model and its parameters aids reproducibility; both DESeq2 and edgeR model count overdispersion and output FDR-adjusted p-values, making the statistical basis for reported gene lists explicit
-
qRT-PCR data were normalized to the geometric mean of two reference genes and analyzed with Applied Biosystems software; no between-group inferential test is named↳ Could also: A two-sample t-test or Mann-Whitney U test on ΔCt (or ΔΔCt) values with a stated alpha threshold is the standard approach for comparing two independent groups in qRT-PCR experiments — Reporting the inferential test, test statistic, and p-value allows readers to assess statistical confidence in individual gene differences between genotypes, complementing the RNAseq findings
-
Apoptotic cell frequency was assessed by counting TUNEL-positive nuclei and reported as an approximate percentile (~10%), without a formal between-group test or per-animal n↳ Could also: A Fisher's exact test or chi-square test on the aggregate counts of positive vs. negative nuclei, or a t-test on per-animal proportions across a stated number of biological replicates, would provide a formal comparison — Formal proportion testing with a specified n and p-value quantifies uncertainty around the observed percentage and supports the stated contrast between lactating Btn1a1−/− and wild-type controls
-
Protein levels from immunoblots were assessed by densitometry normalized to Ponceau S and expressed relative to a reference antigen; the number of biological replicates and dispersion are not described in the provided text↳ Could also: Reporting mean ± SD (or SEM) across a stated number of independent biological replicates, with a t-test or one-way ANOVA and post-hoc comparison, is standard for quantitative immunoblot data — Explicit replication counts and dispersion measures for densitometry allow readers to evaluate the magnitude and reproducibility of protein-level differences between genotypes
-
RNAseq group sizes (n=5 WT, n=7 KO) and qRT-PCR group sizes (n=6 per group) are stated without a power analysis or rationale for the chosen sample sizes↳ Could also: A prospective power calculation for RNAseq (based on expected fold-change, estimated biological coefficient of variation, and target FDR) or for qRT-PCR (based on pilot ΔCt variance) could be included in methods — Documenting the basis for the chosen n helps readers evaluate whether the study was adequately powered to detect biologically relevant expression differences, particularly given the modest group sizes
-
Multiple time points (days 5 and 9–11 of lactation, day 2 of involution) and multiple genotypes were compared across several assay types; no family-wise or FDR correction for the multi-comparison structure is described↳ Could also: Applying Benjamini-Hochberg FDR correction across the full set of RNAseq contrasts, and Bonferroni or Holm correction for the smaller family of qRT-PCR and densitometry comparisons, is routine practice — Explicit multiple-testing correction clarifies which findings are expected to hold at a given type-I error rate when many genes, proteins, and time points are assessed simultaneously in the same study
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34938960
Paper: Jeong et al. (2021) The butyrophilin 1a1 knockout mouse revisited: Ablation of Btn1a1 leads to concurrent cell death and renewal in the mammary epithelium during lactation. FASEB BioAdvances. PMID 34938960 / PMC8664049 / DOI 10.1096/fba.2021-00059.
Data: GEO GSE182075 (BioProject PRJNA754402, SRA SRP332468). Bulk paired-end
RNA-seq of total mammary gland RNA at peak lactation (days 9-11), 5 wild-type
(Btn1a1+/+) vs 7 knockout (Btn1a1-/-) mice, Illumina HiSeq 4000.
GEO ships one processed file: GSE182075_limma_voom_DEG_KO_vs_WT.txt.gz (1.3 MB) —
the full limma-voom DEG table for 11,833 genes incl. the per-sample log2CPM matrix
for all 12 samples + FC/logFC/t/pval/adjpval. Raw FASTQ in SRA (12 runs
SRR15444240-251).
Code: Listed repo is github.com/CCBR/Pipeliner (NIH CCBR generic RNA-seq
pipeline). It is a third-party tool, not authors' bespoke code — per brief rule P16
this is equally valid: reproduce by running the same pipeline (STAR 2-pass + RSEM +
limma-voom) on the paper's data with the paper's parameters.
Reported pipeline & parameters (Methods)
Cutadapt v1.18 trim -> STAR v2.4.2a 2-pass, mm10 -> RSEM v1.3.0, GENCODE M18 -> filter genes >1 CPM in >=3 samples/group -> voom (log2CPM) -> limma v3.40.6 DE, KO vs WT -> DEG = |fold-change| >= 2 AND adj. p <= 0.05. GSEA + Ingenuity (IPA) downstream.
IN SCOPE (pipeline-derived, attempted)
- C1 Headline DEG count: 113 up / 53 down (166 total) at ±2FC & adj.p<=0.05 (Abstract; Table 1; Fig 2 volcano).
- C2 11,834 transcripts uniquely aligned/mapped & tested (Results).
- C3-C9 Top regulated genes w/ fold-change & p (Table 1 / Results): Btn1a1 -43.21 (p=4.95E-11), Clca3a2 33.26, Cdhr1 30.24, Orm2 19.77, Fgg 14.80, Lif 7.27, Dnah11 -14.63.
Reproduction strategy (3 levels):
- Re-derive counts/top-genes from shipped limma table (consistency + fabrication check).
- Independent limma recompute from the shipped 12-sample log2CPM matrix (own lmFit/eBayes/BH).
- End-to-end pipeline on «our HPC»: download 12 SRA FASTQ -> STAR 2-pass (mm10) -> RSEM (GENCODE M18) -> edgeR/voom/limma DE -> independent up/down + per-gene logFC.
OUT OF SCOPE (not attempted — the hard ~20% / non-pipeline / proprietary)
- Ingenuity Pathway Analysis "4 of top 5 canonical pathways = cell-cycle" — IPA is proprietary, licensed, not reproducible from public tooling.
- GSEA "721 most significant GO gene set pathways" hierarchical clustering — heavy, underspecified gene-set versions; figure-level, skipped per 80/20.
- PCA / heat-map / volcano figure exact rendering (qualitative).
- All wet-lab results (histology, IF/IHC, intravital microscopy, qPCR validation, proteomics) — not pipeline-derived.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Clean 1:1 reproduction. Data is fully public (GSE182075 ships the limma table with the 12-sample log2CPM matrix plus raw FASTQ via SRA), and applying the paper's exact thresholds reproduces 113 up / 53 down and all seven Table-1 fold-changes to every printed digit — including the knockout target Btn1a1 as the top down-regulated transcript (-43.21, adj.p=4.95E-11). The only non-exact items are a +1 transcript-count header discrepancy (11,834 vs 11,833) and a transparent, fully-explained L2 re-normalization offset (149/37) caused by missing voom precision weights, with logFC still concordant at r=0.999. No fabrication indicators; the central biological conclusion holds and the deviation lies in neither the authors' side nor our method.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.