Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Systematic and computational identification of Androctonus crassicauda long non-coding RNAs.

Sci Rep · 2021
L1 71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to PARTIALLY reproduce on the paper's own deposited data (P16). The paper's FINAL lncRNA set is public as GenBank TLS KEPY01000000 (13,399 sequences, Trinity v2.0.3). On that deposit: transcript count (13,399 = Table 1) and mean length (747.3 vs weighted 747.7 bp) reproduce 1:1; CPC2 calls 100% non-coding and PLEK 94.87% non-coding (the paper's own tools confirm the deposit is genuinely non-coding RNA). Gene count differs (+6.2%, likely a gene-collapse definition gap) and GC content MISMATCHES (36.46% observed, robust clean ACGT, vs Fig 7 ~42.6% -> flagged for human review, since the same dataset's N and length match exactly). NOT attempted (honest blocker): the full 952,725-transcript Trinity assembly and all Table 1 per-tool funnel counts (CPC2 47,982/904,743; PLEK 40,503/911,471; Annocript 122,421/5,955) -- the raw assembly is not deposited and the raw reads (PRJNA687110/SAMN17133090) are not retrievable from SRA/ENA (no linked runs), so re-assembly is impossible (and Trinity de novo is non-deterministic anyway). Annocript itself not run (legacy MySQL/BLAST, un-deposited input). All grades provisional; human signs off.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-18 ⛓ d344b440a23d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a stringent step-by-step experimental and computational filtering (ECF) pipeline combined with machine-learning tools identify and characterize species-specific long non-coding RNAs in the scorpion Androctonus crassicauda, an organism lacking a reference genome?

Core claims
  • A custom ECF pipeline identified 13,401 lncRNAs in the A. crassicauda transcriptome (12,642 novel, 759 known). finding
  • The majority of predicted scorpion lncRNAs (12,642/13,401) have no identifiable orthologs even in closely related species and are considered novel. finding
  • The ECF pipeline distinguishes coding from non-coding transcripts in species without reference genomes better than standalone tools, achieving balanced high accuracy. method
  • A. crassicauda lncRNAs are characterized by lower protein-coding potential, lower GC content, shorter transcript length, and fewer isoforms per gene than protein-coding transcripts. finding
  • This is the first comprehensive analysis and characterization of lncRNAs in scorpions, providing a resource (TLS project accession KEPY00000000). resource
  • Among standalone tools, Annocript produced the best lncRNA prediction results, while PLEK misclassified many transcripts. finding
  • Using prediction tools built on distantly related organisms increases false-positive rates in species lacking close relatives. finding
Experimental setups
Assay System Perturbation Readout Platform
paired-end RNA-seq (de novo transcriptome assembly) Androctonus crassicauda venom gland, six male/female scorpions of mature and immature age none assembled transcripts/genes; lncRNA identification Trinity (assembly), default parameters
coding potential scoring / lncRNA prediction A. crassicauda assembled transcripts none coding vs non-coding classification (CP threshold 0.4) CPC2 web server
protein homology/domain search A. crassicauda candidate transcripts none removal of transcripts with protein hits BLASTX vs Nr/Swissprot/Uniprot/toxin DB and Pfam, E-value 1e-3
ncRNA family classification A. crassicauda candidate ncRNAs none removal of housekeeping/small ncRNAs INFERNAL, Rfam, RNACentral v14
known lncRNA alignment A. crassicauda lncRNAs none identification of known lncRNAs by homology blastn vs RNACentral v14 and NONCODE v3.0, E-value <1e-5
expression quantification A. crassicauda venom gland lncRNAs none FPKM expression (cutoff FPKM<1 dropped) RSEM
alignment-free lncRNA validation / benchmarking A. crassicauda and Drosophila melanogaster datasets none sensitivity, specificity, accuracy, PPV, NPV, AUC/ROC PLEK, CNIT, CPC2, Annocript
Key results
  • 472 million clean reads assembled into 952,725 transcripts (585,177 genes) 952,725 transcripts
  • Final set of 12,642 novel lncRNA transcripts (11,039 genes) plus 759 known, totaling 13,401 lncRNAs 13,401 transcripts
  • 12,642 of 13,401 lncRNAs have no identifiable orthologs; 759 (5.7%) have homologs in other species 5.7%
  • ECF pipeline achieved highest accuracy on fruit fly dataset (acc 0.99, spec 1, sens 0.91, PPV 1, NPV 0.99) accuracy 0.99
  • ECF pipeline correctly predicted 92.38% (3673/3976) lncRNAs and 100% (30,588/30,588) mRNAs of fruit fly 92.38%
  • On scorpion dataset, CPC2, PLEK, CNIT accuracies were 0.53, 0.49, 0.52 vs ECF pipeline 1 0.49-0.53 vs 1
  • Known and novel lncRNAs averaged 1.1 isoforms per gene vs >2 for protein-coding genes 1.1 isoforms/gene
  • 131,311 putative scorpion-specific lncRNAs retained after FPKM<1 filtering before PLEK validation 131,311 transcripts
Key statistics
  • count 472 million clean reads; 952,725 transcripts; 585,177 genes (de novo Trinity assembly input)
  • count 13,401 total lncRNAs (12,642 novel transcripts / 11,039 genes; 759 known) (final ECF pipeline lncRNA set)
  • other sensitivity 0.91, specificity 1, accuracy 0.99, PPV 1, NPV 0.99 (ECF pipeline on Drosophila dataset)
  • other CPC2 sensitivity 0.94, specificity 0.95, accuracy 0.95 (fruit fly) (benchmark Table 2)
  • other scorpion accuracies: CPC2 0.53, PLEK 0.49, CNIT 0.52, ECF 1 (benchmark Table 3)
  • count misclassified as coding: scorpion CNIT 1.03%, CPC 0%, PLEK 15.91% (false positive non-coding transcripts)
  • count misclassified as coding: fruit fly CPC 6.19%, CNIT 8.07%, PLEK 9.45% (false positive non-coding transcripts)
  • mean 1.1 isoforms per gene for lncRNAs vs >2 for protein-coding (lncRNA characterization)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study developed and benchmarked a multi-step experimental and computational filtering (ECF) pipeline for de novo identification of lncRNAs in the Androctonus crassicauda transcriptome, assembled from RNA-seq data of six scorpion specimens. Pipeline performance was evaluated using standard binary classification metrics (sensitivity, specificity, accuracy, PPV, NPV) and ROC/AUC analysis on both a scorpion-derived dataset and a Drosophila melanogaster reference dataset with known annotations. Predicted lncRNAs were characterized by comparing feature distributions (coding probability, GC content, transcript length, isoforms per gene) against protein-coding transcripts, and sequence novelty was assessed via BLAST-based ortholog searches.

Replicationbiological Sample sizeSix A. crassicauda specimens (male and female, mature and immature); no formal power analysis stated GroupslncRNA vs. protein-coding transcripts; ECF pipeline vs. PLEK, CPC2, CNIT, Annocript; scorpion dataset vs. D. melanogaster reference dataset Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Binary classification performance metrics (sensitivity, specificity, accuracy, PPV, NPV) Benchmarking of ECF pipeline and competing tools (PLEK, CPC2, CNIT) on D. melanogaster and A. crassicauda datasets (Tables 2 and 3) Fruit fly: 3,976 lncRNAs + 30,588 mRNAs; Scorpion: 131,311 lncRNAs + 202,064 mRNAs not stated
ROC curve / AUC analysis Visual and quantitative comparison of classifier performance across tools on fruit fly and scorpion datasets (Figure 6) Same labeled sets as Tables 2 and 3 not stated
BLAST E-value thresholding (blastx, blastn) Sequential filtering of transcripts against Swissprot, Nr, Pfam, Uniprot, manually curated toxin database, RNACentral v14, and NONCODE v3.0 (ECF pipeline steps) 952,725 de novo assembled transcripts at pipeline entry na
FPKM threshold filter (FPKM < 1 excluded) Removal of low-expression transcripts from lncRNA candidate set prior to PLEK validation step 367,332 lncRNA candidates entering this filter not stated
Venn diagram overlap analysis (descriptive set intersection) Comparison of mRNA and lncRNA classification sets across PLEK, CPC2, CNIT, and Annocript (Figure 3) 952,725 total assembled transcripts na
Approaches that could also have been used
  • AUC values from ROC curves were compared descriptively across tools without uncertainty quantification
    Could also: DeLong's method or bootstrap resampling could be used to construct confidence intervals around each AUC and formally test whether differences between classifiers on the same labeled dataset are statistically meaningful — Point-estimate AUC comparisons on class-imbalanced datasets (30,588 mRNAs vs. 3,976 lncRNAs in the fruit fly set) can be misleading without uncertainty bounds; CIs and pairwise significance tests would help readers judge whether observed performance differences are reliable or within sampling variability
  • Classifier error rates on the same test set were compared via tabulated point estimates (Tables 2 and 3)
    Could also: McNemar's test could formally compare the misclassification rates of two classifiers evaluated on the same labeled examples, using the discordant prediction pairs as the test statistic — Because all tools are applied to identical transcripts, the predictions are correlated; McNemar's test accounts for this dependency and provides a p-value for whether two classifiers' error rates differ beyond chance
  • A fixed coding-potential probability threshold of 0.4 was applied to discard putatively coding transcripts, stated without a derivation
    Could also: The threshold could be selected empirically on a labeled validation subset by plotting Youden's J (sensitivity + specificity − 1) or F1 score across the range of CPC2 output scores and choosing the value that optimizes the chosen criterion — A data-driven threshold makes the cutoff choice reproducible and explicitly balances sensitivity against specificity according to the study's priorities, rather than relying on a fixed value borrowed from prior pipelines
  • A fixed FPKM threshold of 1 was used to exclude low-expression transcripts, without a stated basis for the cutoff
    Could also: TPM (transcripts per million) is an increasingly preferred metric over FPKM for cross-sample comparisons, and the expression cutoff could be selected using a mixture-model or kernel-density approach to distinguish signal from background noise empirically — TPM normalizes more consistently across samples with different sequencing depths; a model-based threshold would reduce the arbitrariness of the cutoff and is more straightforward to reproduce across different RNA-seq experiments
  • Venn diagrams were used to display four-way set intersections of mRNA and lncRNA classifications across tools (Figure 3)
    Could also: UpSet plots could represent the same multi-set intersections as ranked bar charts of intersection sizes — Four-set Venn diagrams produce overlapping ellipses that are visually difficult to parse quantitatively; UpSet plots scale well to many sets and make intersection sizes directly readable, which is particularly useful when some intersections are very small relative to others
  • Distributions of lncRNA vs. protein-coding transcript features (GC content, length, coding probability, isoforms per gene) appear to have been compared, though the reporting detail is truncated in the supplied text
    Could also: Non-parametric tests (Mann-Whitney U) accompanied by effect-size measures such as rank-biserial correlation or common-language effect size could formally quantify differences between the two transcript classes — Given the very large and likely non-normally distributed transcript populations, effect-size reporting alongside any significance test would allow readers to assess the practical magnitude of the observed differences, not just their statistical detectability
Software: Trinity · RSEM · CPC2 (Coding Potential Calculator 2) · PLEK · CNIT (Coding-Non-Coding Identifying Tool) · Annocript · BLAST (blastx, blastn, rpsblast) · INFERNAL (used via Rfam annotation)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 33633149

Title: Systematic and computational identification of Androctonus crassicauda long non-coding RNAs. Venue: Sci Rep 2021 · DOI 10.1038/s41598-021-83815-8 · PMCID PMC7907363 Brief code pointer: github.com/frankMusacchia/Annocript (one of several tools used). Brief data pointer: SRA / BioProject PRJNA687110.

What the paper did (pipeline)

  1. New RNA-seq: venom-gland tissue, 6 scorpions, Illumina HiSeq 2000, 150 bp PE, ~472 M clean reads (BioProject PRJNA687110, BioSample SAMN17133090).
  2. De novo assembly with Trinity v2.0.3952,725 transcripts (585,177 genes).
  3. A multi-tool "ECF" lncRNA-identification pipeline applied to the assembly: coding-potential (CPC2, PLEK, CNIT), protein-domain removal (Pfam/UniProt/ Annocript), length/ORF filters (>300 nt, ORF <300 nt → non-coding), housekeeping-RNA removal (Rfam/RNACentral v14), FPKM>1 expression filter, and known-lncRNA matching (NONCODE v3.0, RNACentral) → 13,401 final lncRNAs (11,039 genes; 12,642 novel + 759 known).

In scope (pipeline-derived, attempted)

The paper's final deposited lncRNA set is the GenBank/INSDC TLS accession KEPY01000000 (records KEPY01000001–KEPY01013399 = 13,399 sequences, Trinity v2.0.3). This is the paper's own output data. Per brief rule P16, we reproduce by running the paper's own third-party tools (CPC2, PLEK) on this deposited data and by recomputing the sequence-level statistics the paper reports:

  • R1 Final lncRNA transcript count (Results / Table 1).
  • R2 Final lncRNA gene count (Results) — from Trinity gene IDs in FASTA headers.
  • R3 Mean lncRNA length (Fig 7).
  • R4 lncRNA GC content (Fig 7).
  • R5 Non-coding classification by CPC2 (independent coding-potential tool the authors also used) — does the deposit deliver non-coding RNA as promised?
  • R6 Non-coding classification by PLEK (second independent tool from Table 1).
  • R7 Length filter (>300 nt) consistency.

Out of scope / NOT attempted (with reason)

  • The full assembly funnel (952,725 → CPC2 47,982 coding / 904,743 nc → 745,889 → 387,637 → 367,332 → 131,311) and the Table 1 per-tool counts on the full assembly (CPC2 47,982/904,743; PLEK 40,503/911,471; Annocript 122,421/ 5,955): these operate on the 952,725-transcript raw Trinity assembly, which is NOT deposited. Only the final 13,399 lncRNAs were deposited.
  • Re-assembly from raw reads: BioProject PRJNA687110 / BioSample SAMN17133090 exist, but no SRA run is linked or retrievable (NCBI esearch sra=0; elink biosample→sra empty). Raw reads are not publicly obtainable, so the assembly cannot be regenerated. (And Trinity de novo on ~472 M reads is non-deterministic — exact 952,725 would not reproduce even with reads.)
  • Annocript itself (MySQL + full-UniProt BLAST pipeline, last release 2016/18): not run — heavy/legacy, and its paper numbers are on the un-deposited full assembly. Out of reach with the deposited data only.
  • Wet-lab (RNA extraction, sequencing) — non-computational, out of scope.

Approach

Download KEPY01 (13,399 seqs) to «infra» via «our HPC» front1; recompute N / gene count / length / GC; run CPC2 and PLEK on «our HPC». Compare to Results + Fig 7 + Table 1. All grades provisional — human reviewer decides.

Figures / tables: TableFig 7
R1
Reported
13,401 transcripts (Results) / 13,399 (Table 1)
Reproduced
13,399 deposited (KEPY01000001-013399)
exact
R2
Reported
11,039 lncRNA genes
Reproduced
11,724 unique Trinity genes
partial
R3
Reported
mean length ~747.7 bp (weighted Fig 7: novel 762.2 / known 504.15)
Reproduced
747.3 bp
exact
R4
Reported
GC ~42.6% (Fig 7: novel 42.6 / known 43.4)
Reproduced
36.46%
did not match
R5
Reported
lncRNAs non-coding (CPC2 filter)
Reproduced
100.0% non-coding (13,399/13,399, CPC2)
exact
R6
Reported
lncRNAs non-coding (PLEK filter)
Reproduced
94.87% non-coding (12,711/13,399, PLEK)
within tolerance
R7
Reported
length filter >300 nt
Reproduced
min 282 nt; 60 seqs <=300
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

On the paper's own deposited final lncRNA set (GenBank TLS KEPY01, 13,399 seqs) the central claim reproduces 1:1: transcript count = Table 1, mean length 747.3 vs 747.7 bp, and both authors' tools confirm non-coding nature (CPC2 100%, PLEK 94.87%). The notable defects are on the authors'/data side: the Fig 7 GC content (~42.6%) is robustly not derivable from the identical deposited sequences (observed 36.46%), the gene count differs +6.2% from a definition gap, and 60 seqs violate the stated >300 nt filter. The entire upstream funnel (952,725-transcript assembly, Table 1 per-tool splits) is unreachable because the raw reads and full assembly are not deposited/retrievable. Net: solid partial reproduction with one genuine unexplained discrepancy (GC) and major coverage limited by data availability — moderate, not fabrication-level.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

176.7 k
tokens (I/O) · 10.3 M incl. cache
22 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.