Global and precise identification of functional miRNA targets in mESCs by integrative analysis.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH; reproduces 1:1. Integrative bioinformatics study (Ciaudo lab) calling functional miRNA targets in mESCs. Note: the screening brief mis-attributed code (github.com/alexdobin/STAR) and data (GSE122627) -- the paper deposits its FULL own pipeline (github.com/moritzschaefer/mesc-regulation_pub @ revisions1) and its results as EMBO Datasets EV1-EV7. Per the repo README, the central integrative analysis is reproducible WITHOUT FASTQ re-alignment by importing the deposited supp tables. Result: 8 of 9 in-scope numeric claims reproduce EXACTLY from the deposited data using the repo's exact logic -- C3=3609 and C4=2956 with IDENTICAL gene membership (0 diff) to the deposited lists; C5=360, C6=203, C7=728, C8=106, C9=260(72%) all exact; C1=759 confirmed as the all-filters intersection in deposited Dataset EV3. C2 is partial: the abstract's '<10% of expressed genes' holds robustly (4.45%) but the exact '7%' depends on the authors' specific ~10843 expressed-gene denominator not pinned in the deterministic path run. No fabrication indicators: every count is independently re-derivable from the shipped tables and the deposit is internally consistent. NOT attempted (out of scope): raw FASTQ re-alignment (STAR/snakePipes), wet-lab (CRISPR/qPCR/WB/IF), proteomics raw .wiff processing, web tools (ClueGO/PWMScan). Optional stretch in progress: fresh end-to-end regeneration of the 759 set from TargetScan + AGO2-HEAP to independently confirm Dataset EV3 (env build «job»).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-19 ⛓ d1531eb406de
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-26
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors hypothesize that current miRNA target prediction models fail to accurately identify functional interactions because they lack context-specific factors, and test whether integrating miRNA-depletion transcriptomics, AGO-binding profiles, AGO-loading data, and sequence-based predictions can identify the direct and functional miRNA targets in mouse embryonic stem cells (mESCs).
- ★ Fewer than 10% of expressed genes are directly and functionally regulated by miRNAs in mESCs finding
- ★ Integrative multi-OMICs scoring (AGO2-binding, AGO-loading, target upregulation in miRNA_KO lines, TargetScan predictions) identifies 759 high-confidence candidate miRNA target genes method
- ★ TFAP4 is identified as a miR-290-295 cluster target and an important transcription factor for early development, linked to Wnt signaling regulation finding
- ★ Gene expression changes in miRNA_KO mESCs alone are not sufficient for accurate miRNA target prediction due to mutant-specific and secondary regulatory effects finding
- ★ Ribosome profiling in miRNA_KO mESCs reinforces the validity of the integrative candidate target list by showing increased ribosome occupancy for candidates finding
- A subset of mESC-predicted miRNA interactions show functional conservation in human ESCs finding
- CIC is identified as a novel functional miR-290-295 target not previously catalogued in miRTarBase finding
- Number of predicted MREs/interactions per gene positively correlates with differential ribosome occupancy, suggesting cooperative miRNA targeting finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq | WT and miRNA_KO (Drosha_KO, Dicer_KO, Ago2&1_KO) mESCs | KO | differential gene expression / upregulated & downregulated genes | — |
| Western blot / immunoblot | WT and KO mESCs | KO | DROSHA, DICER, AGO1, AGO2 protein levels (TUBULIN loading control) | — |
| AGO2 CLIP-seq (HEAP) binding profiles | mESCs (published dataset, Li et al 2020) | none | AGO2-binding peaks at miRNA response elements | — |
| AGO RNA immunoprecipitation and sequencing (RIP-seq) / small RNA-seq | WT mESCs (Ngondo et al 2018) | none | miRNA loading into AGO1/AGO2 complexes | — |
| SWATH-MS proteomics | WT and miRNA_KO mESCs | KO | differential protein abundance of candidate miRNA targets | SWATH-MS |
| Ribosome profiling (Ribo-seq) | WT and miRNA_KO mESCs | KO | differential ribosome occupancy of candidate miRNA targets | — |
| Western blot / immunoblot | WT, Drosha_KO, Dicer_KO, Ago2&1_KO mESCs | KO | CIC protein levels normalized to TUBULIN | — |
| CLIP-seq / AGO-binding and seed-match conservation analysis | human ESCs (Lipchina et al 2011) | none | conservation of MREs and correlation with miRNA expression/AGO-binding | — |
- – Only a small fraction of upregulated genes overlapped across all three miRNA_KO mutants 1,109 of 6,408 upregulated genes (adj. P<0.2)
- – Integrative scoring narrowed candidate miRNA target genes to a restrictive set 759 genes (~7% of expressed genes)
- ▲ SWATH-MS showed significantly enriched positive log2FC in protein abundance for candidate targets in all KO mutants ≥60% of 203/759 detectable genes with positive log2FC, P<0.0051 per mutant
- ▲ Ribo-seq showed stronger and more significant enrichment of positive log2FC in ribosome occupancy than proteomics 76-79% of 728/759 detectable genes with positive log2FC, P<1e-51 per mutant
- ▲ Number of predicted interactions per gene positively correlated with differential ribosome occupancy Pearson r=0.2-0.24, P<2.8e-8
- – Only about half of mouse miRTarBase-annotated targets (filtered for mESC-loaded miRNAs) were recapitulated by the integrative approach 362 of 6,422 targets
- – A subset of mESC-predicted interactions showed conserved MREs with associated AGO-binding in human ESCs 10% of interactions, >2.5-fold over scrambled control
- ▲ Conserved MREs correlated with hESC miRNA expression and were more likely to show AGO-binding when the miRNA was expressed 3-fold more likely when expressed; Pearson r=0.22, P<1e-50
- count 1,109 of 6,408 upregulated genes (overlap of upregulated genes across three miRNA_KO mutants, adj. P<0.2)
- count 759 candidate miRNA target genes (7% of expressed genes) (final integrative candidate target list)
- pvalue P<0.0051 for every mutant (SWATH-MS significance of increased protein abundance for candidate targets)
- pvalue P<1e-51 for every mutant (Ribo-seq significance of increased ribosome occupancy for candidate targets)
- correlation Pearson r=0.2-0.24, P<2.8e-8 (MRE/interaction count per gene vs differential ribosome occupancy)
- count 27% (203/759) detectable in SWATH-MS (proteomics detection sensitivity for candidate targets)
- count 96% (728/759) detectable in Ribo-seq (ribosome profiling detection sensitivity for candidate targets)
- correlation Pearson r=0.22, P<1e-50 (hESC miRNA expression vs CLIP-seq read correlation for conserved interactions)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combines RNA-seq (DESeq2 differential expression, adjusted P-value thresholds), published AGO2 CLIP/HITS-CLIP and small-RNA/AGO-loading sequencing, and sequence-based TargetScan predictions into a per-interaction and per-gene integrative scoring scheme to nominate functional miRNA targets in mouse embryonic stem cells. Validation of candidate targets uses SWATH-MS proteomics and ribosome profiling, compared to WT with Student's t-tests on cumulative distributions of log2 fold-changes, Pearson correlation between interaction counts and ribosome occupancy, GO/GSEA enrichment (via Enrichr), and immunoblot quantification (Student's t-test vs WT) for a selected target (CIC). Results are reported largely as P-value thresholds/inequalities (e.g., P < 0.0051, P < 1e-51) alongside fold-change and correlation coefficients, with bar graphs summarized as mean ± SD.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Student's t-test | Differential protein abundance (SWATH-MS) of candidate miRNA targets vs control distribution (Fig EV2E) | 203 of 759 candidate targets | not stated |
| Student's t-test | Differential ribosome occupancy of candidate miRNA targets vs control distribution (Fig 2A) | 728 of 759 candidate targets | not stated |
| Student's t-test | CIC immunoblot intensity in each miRNA_KO line vs WT (Fig 2F) | three biological replicates | not stated |
| DESeq2 differential expression analysis (adjusted P-value threshold) | RNA-seq comparisons of miRNA_KO lines vs WT to call up/downregulated genes (Fig EV1C–E, Fig 1) | not explicitly stated in provided text (replicate RNA-seq per genotype) | not stated |
| Pearson correlation | Number of predicted interactions per gene vs differential ribosome occupancy (Fig 2D); miRNA expression vs AGO-binding signal in hESCs (Fig EV2K) | not explicitly stated | not stated |
| GO term / gene set enrichment analysis (Enrichr) | Comparison of candidate miRNA target genes vs commonly up- and downregulated gene sets (Fig 1E) | 759 target genes; 3,609 upregulated; 2,956 downregulated | not stated |
-
Student's t-tests were used to compare cumulative distributions of log2 fold-changes (protein abundance, ribosome occupancy) between mutants and a control distribution↳ Could also: A nonparametric test such as the Kolmogorov-Smirnov test or Mann-Whitney U test — These tests compare full distributions/ranks without assuming normality, which can be a natural fit when the comparison is visualized as a CDF, as done in Fig EV2E and Fig 2A
-
Separate Student's t-tests were run for each of the three miRNA_KO mutants against WT in several validation analyses (Fig EV2E, Fig 2A, Fig 2F)↳ Could also: A correction for multiple comparisons (e.g. Bonferroni or Holm) across the three parallel mutant-vs-WT tests, or a mixed-effects/ANOVA framework treating mutant as a factor — This would jointly account for the family-wise error rate introduced by testing the same hypothesis in three independent mutant backgrounds
-
Correlation between number of predicted interactions and ribosome occupancy (and miRNA expression vs AGO-binding) was assessed with Pearson's r↳ Could also: Spearman's rank correlation — Spearman correlation is robust to outliers and does not assume a linear relationship or normally distributed variables, which can be useful for count-based or skewed genomic metrics
-
Differential expression significance was based on a DESeq2 adjusted P-value threshold (0.1 or 0.2) without stating the specific multiple-testing correction method↳ Could also: Explicitly reporting the correction method (e.g. Benjamini-Hochberg FDR) alongside the threshold, or comparing results across DE tools such as edgeR or limma-voom — Naming the correction method aids reproducibility, and cross-tool comparison can show whether target calls are robust to the specific DE modeling approach
-
Bar graphs (e.g. Fig 2F) summarize replicate measurements as mean ± SD↳ Could also: Reporting as mean with 95% confidence interval, or showing individual data points alongside the summary statistic — With small replicate numbers (e.g. n=3), a CI or dot plot can more directly convey the precision of the estimate than SD alone
-
GO term enrichment for candidate miRNA targets was assessed via Enrichr without specifying the underlying statistical test↳ Could also: A hypergeometric/Fisher's exact test with explicit multiple-testing correction (e.g. FDR) reported alongside gene set size and background set — Making the underlying enrichment test and background gene set explicit allows readers to evaluate the specificity of the reported GO terms
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 35899551
Title: Global and precise identification of functional miRNA targets in mESCs by integrative analysis. Schaefer, Nabih, Spies et al., Ciaudo lab. EMBO Reports 2022;23(9):e54762. PMCID PMC9442311. DOI 10.15252/embr.202254762.
What the paper is
An integrative bioinformatics study. It combines several omics layers in mouse embryonic stem cells (mESCs) to call functional, direct miRNA target genes:
- transcriptome of miRNA-depleted mESC lines (Drosha_KO, Dicer_KO, Ago2&1_KO) vs WT (RNA-seq → DESeq2)
- AGO RIP-seq (miRNA loading into Argonautes)
- AGO2-binding profiles (HEAP / iCLIP from external GEO series)
- sequence-based predictions (TargetScan mouse 7.2 context++ scores)
- validation layers: full proteome (SWATH-MS), ribosome profiling, QuantSeq of miR-290-295_KO + siPOOLs.
The headline result: integrating these layers narrows candidate miRNA targets to 759 genes (≈7% of expressed genes → "<10% of expressed genes are functionally and directly regulated by miRNAs in mESCs").
Code / data availability (RICHER than the screening brief)
The screening BRIEF lists only github.com/alexdobin/STAR as "Code" and GSE122627 as "Data".
That is incomplete / misattributed. The paper actually deposits its full own pipeline:
github.com/moritzschaefer/mesc-regulation_pub— "Main pipeline producing all data and figures" (snakemake; branchrevisions1). This is the primary reproduction asset.github.com/moritzschaefer/srna-seq— RIP-seq pipelinegithub.com/moritzschaefer/snakepipelines_pub— RNA-seq & QuantSeq snakePipes configsgithub.com/moritzschaefer/moritzsphd_pub— auxiliary library (moritzsphd, pip-installable; used to fetch/annotate HEAP data)
GSE122627 is not this paper's own deposit — it is one reused input (Drosha_KO mESC RNA-seq,
Cirera-Salinas et al). The paper's OWN deposits: GSE110942 (Ago2&1_KO RNA-seq), GSE181393 (QuantSeq),
GSE135577 (Ribo-seq), ProteomeXchange PXD014484 (proteome). Reused inputs: GSE78971/GSE78973/GSE122627
(WT/Dicer/Drosha RNA-seq), GSE80454 (RIP-seq), GSE139345 (HEAP AGO2), GSE61348 (iCLIP), SRR359787 (hESC PAR-CLIP).
IN SCOPE (pipeline-derived, reproducible)
The README states the integrative analysis can be re-run without re-aligning any FASTQ by importing
the published Supplementary Tables (DESeq2 results, RIP-seq, etc.) via import_supptables.snakefile,
then running the integration scripts. This makes the central claims a deterministic, light-weight
reproduction (pandas / gffutils / bedtools / TargetScan parsing — no read alignment).
Primary target — the integrative target-gene set and its derived counts:
- 759 candidate miRNA target genes (Fig 1B / Dataset EV3) —
mirna_target_genes.txtfromrules/direct_mirna_targets.smk(combine_direct_mirna_interactions.py→rank_direct_mirna_interactions.py). - 3609 commonly upregulated (
miRNA_KOUP) and 2956 commonly downregulated genes (Fig 1) — computed inrank_direct_mirna_interactions.pyfrom TableS1 (padj<0.2, log2FC sign, ≥2 mutants). - 360 / 759 genes that are miR-290-295 targets (Fig EV3A).
- Overlap-with-validation counts: 203/759 in proteome (SWATH-MS), 728/759 in ribo-seq, 106/360 sig. up in miR-290-295_KO, 260/360 positive log2FC.
Inputs the integration needs (all downloadable / deposited):
supp/TableS1_RNA-seq.xlsx(DESeq2 per-mutant log2FC/padj/tpm),supp/TableS2_RIP-seq.xlsx(miRNA loading),supp/quant_seq_counts.tsv(GSE181393);- TargetScan mouse 7.2 (Conserved + Nonconserved Site Context Scores, HTTP download);
- AGO2 HEAP peaks GSE139345 (via
moritzsphd); - FANTOM5 mouse small-RNA CPMs (mesc_mirnas); GENCODE GRCm38.98 + vM3 annotations; pyensembl release 98.
OUT OF SCOPE (not attempted / not pipeline-derived here)
- Re-alignment of raw FASTQ (STAR/snakePipes from scratch) — heavy and unnecessary; the DESeq2 outputs are deposited as supp tables. May be attempted as a secon
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.