Integrated omics in Drosophila uncover a circadian kinome.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH + REPRODUCED 1:1. The paper's own pipeline (iCMod, Java + R/MetaCycle 1.2.0, commit d40b32f) ships its derived input tables in test_data, so the pipeline-derived results are fully reproducible without re-running the wet-lab MS or the external GPS GUI. A FRESH end-to-end «our HPC» run («job»: clone repo @d40b32f, download Ensembl-110 BDGP6 pep fasta, port Windows paths, compile Java, normalize -> MetaCycle meta2d ARS/JTK/LS minper=20 noIntegration -> GPS-based hypergeometric kinase enrichment) reproduced the core results essentially exactly and BYTE-IDENTICALLY to the earlier run (matching sha256 on KAA_wt_16, my_circadian, ARSresult, and wt_enrichout): WT background M = 4686 p-sites (== shipped id_all.txt); MetaCycle ARSER p<0.01 yields 789 oscillating NCPs whose SET is byte-identical to the authors' shipped 789-site id.txt (Jaccard 1.0, 0 differences); the GPS circadian ssKSR network = 153 kinases / 778 substrate-sites (exact) and 36,800 ssKSRs (vs 36,522, +0.76%); the hypergeometric enrichment recovers exactly 27 potential circadian kinases (p<0.05 & E>1), INCLUDING all six kinases the paper later validated by RNAi (hep, Dsor1, gish, CKIalpha, gskt, bsk). The faithful run (authors' id.txt) and the fully end-to-end run (our own MetaCycle set) produced byte-identical enrichment output (same sha256). No fabrication indicators: every value is deterministically derivable from the shipped data + code; the 789-site set equals the authors' and the wet-lab-validated kinases all reappear. NOT ATTEMPTED (out of scope): upstream MaxQuant MS search, cuffdiff/FPKM transcriptome filtering, GPS 2.1 prediction itself, per0 mutant arm & cycling counts, wet-lab RNAi/behavioral validation. This room was re-run because the prior ROOM_RESULT predated the required qc_room/datasets schema; the science verdict is unchanged. All grades PROVISIONAL pending human audit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 98assessed: 2026-06-16 ⛓ 9229403c473f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates how circadian clocks drive integrated molecular oscillations by asking which protein kinases regulate rhythmic phosphorylation events (a 'circadian kinome') that shape circadian rhythms in Drosophila.
- ★ iCMod, a computational pipeline integrating transcriptomic, proteomic, and phosphoproteomic circadian data, was developed to accurately identify normalized circadian p-sites (NCPs) method
- ★ 789 (~17%) phosphorylation sites in fly heads show circadian oscillation (NCPs), identified from 431 proteins finding
- ★ 27 potential circadian kinases were computationally predicted to phosphorylate NCPs, including 7 kinases already known to function in the core clock finding
- ★ Screening of the remaining 20 predicted kinases identified gskt, Dsor1, bsk, gish, hep, and CKIalpha as involved or potentially involved in modulating locomotor rhythm finding
- ★ GASKET (GSKT) was identified as a potentially important regulator within a reconstructed circadian kinase signaling web and acts to reduce TIM protein but not mRNA level mechanism
- ★ The majority of circadian oscillations at mRNA, protein, and phosphorylation levels are abolished in per0 mutants, indicating they are driven by the molecular clock finding
- Temporal correlation between proteome and phosphoproteome is much higher than between transcriptome and proteome, suggesting phosphorylation has a major role in regulating temporal protein-level changes finding
- SGG (SHAGGY10 isoform) tyrosine 214 phosphorylation shows significant circadian variation that is eliminated in per0 mutants, validating an NCP by Western blot finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| TMT-based quantitative phosphoproteomics (LC-MS/MS) | Drosophila (WT and per0) fly heads | per0 mutation; samples collected at 3h intervals over 2 days in constant darkness (DD) | phosphopeptide/p-site abundance and circadian oscillation (NCPs) | LC-MS/MS with Tandem Mass Tag (TMT) labeling |
| TMT-based quantitative proteomics (LC-MS/MS) | Drosophila (WT and per0) fly heads | per0 mutation; DD, 3h intervals, 2 days | protein abundance and circadian oscillation | LC-MS/MS with TMT labeling |
| RNA sequencing (RNA-seq) | Drosophila (WT and per0) fly heads | per0 mutation; DD, 3h intervals, 2 days | mRNA expression (FPKM) and circadian oscillation | — |
| Western blot | Drosophila (WT and per0) whole-head extracts | per0 mutation | SGG pY214, total SGG protein, normalized SGG pY214 levels across circadian time | — |
| Locomotor rhythm behavioral assay | Drosophila, genetic manipulation of candidate kinase genes | genetic knockdown/mutation of 20 candidate kinases | period length and power of locomotor rhythm | — |
| Computational kinase-substrate prediction (GPS 2.1) and hypergeometric enrichment test | in silico analysis of NCPs from fly head phosphoproteome | none | site-specific kinase-substrate relations (ssKSRs) and enrichment of NCPs per kinase | GPS 2.1; ARSER for rhythmicity detection |
- – 789 (16.84%) p-sites identified as NCPs with circadian oscillation in WT fly heads, from 4686 high-confidence quantified p-sites 16.84%
- – 27 potential circadian kinases predicted via hypergeometric enrichment of NCPs among kinase substrates; 7 already known clock regulators
- ▼ per0 mutation abolishes cycling of 93.95% of mRNAs, 84.52% of proteins, and 87.96% of p-sites compared to WT 93.95%/84.52%/87.96%
- – Six kinases (gskt, Dsor1, bsk, gish, hep, CKIalpha) altered locomotor rhythm period by ≥1h or reduced power by ≥50% when genetically manipulated period change ≥1h or power reduction ≥50%
- – SGG pY214 shows significant temporal variation in WT that is not significant in per0
- – 98.10% of identified NCPs were predicted with ≥2 kinases 98.10%
- ▲ Spearman correlation between phosphoproteome and proteome temporal variation is higher than between transcriptome and proteome
- – 661 (6.67%) mRNAs and 620 (16.42%) proteins identified as significantly cycling in WT fly heads 6.67%/16.42%
- correlation Spearman's rank correlation coefficients of 0.99 (mRNA), 0.95 (protein), 0.85 (phosphorylation) between two monitored cycles (reproducibility of multi-omics measurements across two circadian cycles)
- count 789 (16.84%) NCPs (circadian phosphorylation sites identified from total quantified high-confidence p-sites in WT)
- count 661 (6.67%) cycling mRNAs; 620 (16.42%) cycling proteins (circadian oscillation identified by iCMod in WT fly heads)
- pvalue p value = 0.00001, 0.0614, 0.1299, 0.0013, 0.0305, 0.1572 (hypergeometric enrichment of CGDB genes and translatome-oscillating mRNAs against cyclic mRNA/protein/phosphoprotein sets)
- pvalue p value of mRNA < 2.2 × 10−5, Pro. < 1.4 × 10−4, Phos. < 3.1 × 10−7 (GO-based enrichment analysis of cycling mRNAs, proteins, and NCPs)
- pvalue one-way ANOVA p value = 0.00033 (SGG pY214 WT); p value = 0.00002 (normalized SGG pY214 WT) (statistical significance of SGG Y214 phosphorylation rhythm in WT vs per0)
- count 27 potential circadian kinases predicted; 7 previously known, 20 newly tested, 6 validated as involved/potentially involved (kinase prediction and validation pipeline outcome)
- count 36,522 potential ssKSRs between 153 protein kinases and 778 phosphorylated substrates (GPS-predicted kinase-substrate relations for NCPs)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study is an integrative multi-omics circadian profiling in Drosophila heads (transcriptome via RNA-seq, proteome and phosphoproteome via TMT LC-MS/MS) comparing wild-type and per0 flies sampled at 3 h intervals over two days in constant darkness. Rhythmicity in mRNAs, proteins, and normalized phosphorylation sites was called computationally with the ARSER periodicity algorithm after a custom normalization/filtering pipeline (iCMod), kinase–substrate relations were predicted with GPS 2.1, and enrichment of cycling sites/kinases was assessed with two-sided hypergeometric tests. Targeted molecular validations (e.g., SGG phosphorylation Western blots) were summarized as means with SEM and compared with one-way ANOVA.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| ARSER periodicity detection (cycling mRNA/protein/p-site calling) | Identification of cycling mRNAs (661), proteins (620) and NCPs (789) across the omics data | 16 samples per genotype (3 h intervals over 2 days in DD) | not stated |
| Two-sided hypergeometric test | GO term enrichment; enrichment of CGDB/translatome gene sets among cycling data sets (p = 0.00001, 0.0614, 0.1299, 0.0013, 0.0305, 0.1572); enrichment of NCPs among predicted kinase substrates to call 27 circadian kinases | — | na |
| One-way ANOVA | Quantification of SGG pY214 and normalized SGG pY214 temporal variation in Western blots (Fig. 3i; p = 0.00033, p = 0.00002) | n = 7 for SGG pY214 and normalized pY214; n = 5 for SGG protein | not stated |
| Spearman's rank correlation | Reproducibility between Day 1 and Day 2 of transcriptome, proteome, phosphoproteome (r = 0.99, 0.95, 0.85) | — | na |
-
Rhythmicity was detected with ARSER, a single periodicity-detection algorithm.↳ Could also: Complementary rhythm-detection methods such as JTK_CYCLE, RAIN, MetaCycle (which combines several algorithms), or empirical JTK could also be applied. — Running more than one detector and reporting concordance can convey robustness of cycling calls across methods that make different waveform assumptions, which is often informative for circadian time-series data.
-
Genome/proteome-wide cycling and enrichment were assessed with periodicity and hypergeometric tests; the text does not explicitly describe a multiple-testing correction.↳ Could also: A stated FDR control (e.g., Benjamini–Hochberg q-values) across the families of rhythmicity p-values and enrichment p-values could also be reported. — Explicit FDR reporting for large test families helps readers gauge the expected proportion of false positives among the many simultaneous comparisons.
-
Targeted Western blot quantifications were summarized using SEM.↳ Could also: Standard deviation or a 95% confidence interval could also be shown alongside or instead of SEM. — SD or a CI conveys the spread of the underlying observations directly, which is often preferred when n is small so readers can see variability rather than only the precision of the mean.
-
Temporal variation across circadian time points was assessed with one-way ANOVA.↳ Could also: A method that explicitly models the cyclic/time structure (e.g., cosinor regression) or, if distributional assumptions are a concern, a nonparametric Kruskal–Wallis test could also be used. — Cosinor or rhythmicity-aware models directly estimate amplitude/phase of a 24 h oscillation, while a nonparametric test relaxes normality assumptions for small samples.
-
ANOVA results were reported with overall p-values for the time effect.↳ Could also: Reporting accompanying effect sizes (e.g., eta-squared) and the specific post-hoc comparisons with their correction could also be included. — Effect sizes and named post-hoc procedures help readers interpret the magnitude of temporal differences and see which time points drive the overall effect.
-
Kinase–substrate relations were predicted with GPS 2.1 and prioritized via hypergeometric enrichment.↳ Could also: Alternative or complementary kinase-enrichment frameworks (e.g., KSEA, NetworKIN-based scoring) could also be applied. — Cross-referencing predictions from multiple kinase-inference tools can convey how stable the prioritized circadian kinase set is to the choice of prediction method.
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — pmid-32483184
Paper: Wang et al. 2020, Integrated omics in Drosophila uncover a circadian kinome. Nat Commun 11:2710. PMID 32483184 · PMC7264355 · DOI 10.1038/s41467-020-16514-z.
Code (authors' own, P-not-needed since it is theirs): https://github.com/CuckooWang/iCMod
pinned commit d40b32fb26f81fe07f9e9b4d0451c756fb0b982c (2019-07-13, only commit on master).
Pipeline iCMod = 7 Java programs + 1 R script (MetaCycle), Java 17.
Shipped test_data (decisive): the repo ships the actual derived input tables used
by the paper — proteome & phosphoproteome MaxQuant intensity matrices
(Proteome/new_data1..4.txt, Phosphoproteome/new_data1..4.txt), the GPS 2.1
kinase–substrate prediction output (gps_out.txt, 218,665 ssKSRs), and the
circadian-site lists id.txt (789 lines) and id_all.txt (4686 lines). This makes
a self-contained reproduction of the core computational results possible without
re-running the upstream wet-lab MS or the external GPS GUI.
In scope (pipeline-derived, attempted)
| # | Result | Pipeline step | Inputs (all shipped) | Reproduces claim |
|---|---|---|---|---|
| R1 | Cross-batch WT_3 normalization + phospho-by-protein normalization → KAA_wt_16.txt (WT time course, 16 tp) |
CorrectByWT3.java (+ phospho block) → PhosNormByPro.java |
Proteome/Phosphoproteome new_data1..4.txt | background p-site set M (≈ id_all 4686) |
| R2 | 789 oscillating p-sites (NCPs) in WT by ARSER p<0.01 | MetaCycleRun.R (meta2d ARS/JTK/LS, minper=20) on KAA_wt_16.txt |
KAA_wt_16.txt | "789 (16.84%) p-sites significantly oscillating" (Fig 3a); compare set to shipped id.txt |
| R3 | ssKSR network restricted to NCPs (36,522 ssKSRs / 153 kinases / 778 substrates) | counting on gps_out.txt restricted to NCP sites |
gps_out.txt + id lists | Results / Fig 4 |
| R4 | 27 potential circadian kinases (hypergeometric enrichment, p<0.05 & E-ratio>1) | CalCricadianSiteKa.java → enrichment.java |
gps_out.txt, BDGP6 pep fasta, id.txt, KAA_wt_16.txt | "27 potential circadian kinases" (Fig 4c,d) |
Out of scope (not attempted, with reason)
- Upstream transcriptome / FPKM filtering (
FPKM1.java): requiresgene_exp.diffcuffdiff outputs which the repo ships only partially ("some files are partially present due to upload size limit"). Cannot regenerate the FPKM≥1 reference DB faithfully. - MaxQuant MS/MS search (raw → peptide/p-site intensities): wet-lab + proprietary MaxQuant; the new_data*.txt intensity tables are taken as given (shipped).
- GPS 2.1 prediction itself: external Windows GUI tool (gps.biocuckoo.org); its
output
gps_out.txtis shipped and used as given. - per0 mutant arm, mRNA/protein cycling counts (661 mRNAs, 620 proteins), wet-lab RNAi validation (6 kinases), behavioral assays: not pipeline-derivable from the shipped data (mRNA cuffdiff partial; RNAi/behavior are wet-lab).
Reference data fetched
- Ensembl BDGP6 peptide fasta (
Drosophila_melanogaster.BDGP6.pep.all.fa, release-110, 30,799 FBpp) — used only to attach gene-symbol labels to kinase FBpp IDs inCalCricadianSiteKa/enrichment; does not affect counts.
All compute on «our HPC» (SLURM, account kubisch_std, partition std); all data on «infra»
«path».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Near-perfect 1:1 reproduction using the authors' own iCMod pipeline on their shipped test_data: C1–C5 match exactly except C3c (36,800 vs 36,522 ssKSRs, +0.76%), a negligible boundary effect on the externally-shipped GPS output. Our independently re-run MetaCycle reproduces the authors' 789-site circadian list byte-identically (Jaccard 1.0) and the enrichment recovers exactly 27 kinases including all 6 RNAi-validated ones. No deviation on any side and no fabrication indicators; the only caveat is that upstream raw MS/RNA-seq processing was out of scope, so this confirms pipeline determinism on shipped derived inputs.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.