Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genome-wide identification and characterization of germin-like protein family in Brassica juncea reveals their role against biotic stress.

BMC Plant Biol · 2025
L1 58/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
58/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 18% of all assessed papers rank 950 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Plant gene-family identification paper; code = third-party TBtools (P16). Genome pinned to BRAD Braju_tum_V2.0 (BjuT84V2); the paper's BjuVA/BjuVB gene IDs resolve via the BRAD per-gene API (coords match the supplement exactly). Reproduced ProtParam (==Biopython) + cupin_1 PF00190 (==hmmer) on 98/101 family proteins on «our HPC» («job»). NOT a clean 1:1: the deposited Supplementary Table 1 physicochemical values do NOT map to their gene IDs -- only 17/98 protein lengths (15/98 MW, 18/98 pI) match the gene printed beside them. The numbers are genuine ProtParam outputs but row-misaligned (demonstrable -1 row shift across blocks, e.g. BjuGLP03-14 where gene N's real value == gene N-1's printed value), and ~9 large values (1800-1900 aa / 196-205 kDa) match no listed locus -> DATA-INTEGRITY / possible-fabrication flag (provisional, for human audit). What DOES reproduce: cupin_1 membership 98/98 (31 double) = exact; chromosomal max/min (B02=9, B08=1) = matches 'Chr02=9/Chr08=1'. Count discrepancy: text says 102 GLPs, the table lists 101 (40+/61- vs 41+/61-). NOT attempted: full BLASTp re-identification (no bulk proteome in mirror) and the RNA-seq DESeq2 expression pipeline on PRJNA860050 (heavy 12-library, the hard 20%).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 58
    assessed: 2026-06-16 ⛓ c5707e6767d9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Germin-like proteins (GLPs) were poorly characterized in Brassica juncea and its progenitor species; this study aimed to identify and characterize the GLP family in B. juncea, B. nigra, and B. rapa and to determine candidate B. juncea GLP genes conferring resistance against biotic stress, using Alternaria leaf spot disease (A. alternata) as the model biotic stress.

Core claims
  • 102 GLPs were identified in B. juncea, 51 in B. nigra, and 48 in B. rapa via genome-wide in-silico analysis finding
  • This is the first identification and characterization of the GLP family in B. juncea and its parental species B. nigra and B. rapa resource
  • Ten differentially expressed BjuGLPs were validated by RT-qPCR: 5 upregulated (BjuGLP06, BjuGLP23, BjuGLP34, BjuGLP70, BjuGLP97) and 5 downregulated (BjuGLP04, BjuGLP33, BjuGLP71, BjuGLP72, BjuGLP91) under A. alternata infection finding
  • In-silico RNA-seq expression trends of B. juncea GLPs were confirmed by in-vitro RT-qPCR, indicating GLPs contribute to defense against Alternaria leaf spot finding
  • Phylogenetic analysis of GLPs across B. juncea, B. nigra, B. rapa and A. thaliana revealed 4 major clades divided into 21 subclades/members finding
  • Synteny analysis showed GLPs are conserved and transferred from both progenitor genomes into B. juncea, with gene duplication events observed in B. juncea and B. nigra finding
  • All GLPs possess a conserved cupin_1 domain and conserved germin box motif, with 20 conserved motifs predicted across the three species finding
  • Ka/Ks analysis revealed mixed purifying and positive selection acting on GLP gene pairs in B. juncea and B. nigra finding
Experimental setups
Assay System Perturbation Readout Platform
Genome-wide in-silico GLP identification (orthology to AtGLPs) B. juncea, B. nigra, B. rapa genomes none number and physiochemical properties of GLP genes/proteins
Phylogenetic analysis B. juncea, B. nigra, B. rapa, A. thaliana GLPs none clade/subclade structure IQ-TREE v2.1.2 (1000 bootstrap)
Motif/domain/gene structure analysis BjuGLPs, BniGLPs, BraGLPs none conserved motifs, cupin_1 domain, exon-intron structure TBtools v2.210, MEME (E=1e-5), Pfam
Synteny analysis B. juncea (AABB), B. nigra (BB), B. rapa (AA) none collinear/syntenic gene relationships MCScanX in TBtools v2.210 (E-value 1e-10)
Cis-regulatory element (CRE) analysis BjuGLP/BniGLP/BraGLP promoters (1500 bp upstream) none CRE types in promoters PlantCARE
Ka/Ks selection analysis GLP paralog pairs in B. juncea and B. nigra none Ka/Ks ratios (selection pressure)
In-silico RNA-Seq differential expression analysis B. juncea under Alternaria brassicae infection (2 DPI, 4 DPI) Alternaria infection differentially expressed GLPs (|log2FC|=1) Galaxy (usegalaxy.org); SRA PRJNA860050
RT-qPCR expression validation Super raya (B. juncea) plants, diseased vs control A. alternata inoculation at 2-3 leaf stage relative expression of 10 selected BjuGLPs (ref gene UBQ9)
Key results
  • 102 GLPs identified in B. juncea, 51 in B. nigra, 48 in B. rapa 102/51/48
  • 5 BjuGLPs upregulated (BjuGLP06, BjuGLP23, BjuGLP34, BjuGLP70, BjuGLP97) under A. alternata infection ≥2-fold
  • 5 BjuGLPs downregulated (BjuGLP04, BjuGLP33, BjuGLP71, BjuGLP72, BjuGLP91) under A. alternata infection ≥2-fold
  • 25 GLPs at 2 DPI and 30 GLPs at 4 DPI identified as differentially expressed in RNA-seq |log2FC|=1
  • Phylogenetic tree resolved 4 major clades (GLP1, GLP3, GLP4, GLP5) and 21 subclades 4 clades / 21 members
  • Highest number of GLPs predicted in extracellular space (66 BjuGLPs, 31 BniGLPs, 34 BraGLPs)
  • In-silico expression trends confirmed in in-vitro RT-qPCR for all 10 validated B. juncea GLPs
  • No duplicated GLP genes found in B. rapa; duplications found in B. nigra and B. juncea
Key statistics
  • count 102 GLPs in B. juncea, 51 in B. nigra, 48 in B. rapa (genome-wide GLP identification)
  • fold_change 2-fold (|log2FC| = 1) (threshold for DEG selection and RT-qPCR validation)
  • count 25 DEG GLPs at 2 DPI, 30 DEG GLPs at 4 DPI (RNA-seq differential expression (45 detected at 2 DPI, 42 at 4 DPI))
  • other isoelectric point 4.77 to 10.2 (physiochemical properties of GLPs)
  • other protein length 160–2000 aa; molecular weight 12–206 kDa (physiochemical properties)
  • count 47 distinct CRE types (light 7, auxin 6, defense/stress 2) (cis-regulatory element analysis of promoters)
  • count B. juncea strand: 41 positive, 61 negative (45% positive, 55% negative) (strand position analysis)
  • other up to 47% yield losses worldwide (Alternaria leaf spot disease impact)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This genome-wide study combined in-silico bioinformatics analyses (physiochemical characterization, maximum-likelihood phylogenetics, synteny, cis-regulatory element prediction, and Ka/Ks evolutionary analysis) with RNA-Seq differential expression screening to identify candidate GLP genes in Brassica juncea and its progenitor species responsive to biotic stress. Differentially expressed genes were filtered by a 2-fold change threshold (|log2FC| ≥ 1) from publicly available RNA-Seq data processed via Galaxy, and 10 representative genes were selected for RT-qPCR validation in plants inoculated with Alternaria alternata. Results were reported primarily as descriptive gene-family characterizations, fold-change categories, and qualitative expression trends (upregulated/downregulated), without explicit reporting of p-values or dispersion measures for most analyses.

Replicationmixed Sample sizeNot described with formal power analysis; wet-lab RT-qPCR used 6 of 8 collected plant RNA samples (4 diseased, remainder control); RNA-Seq replication structure of public dataset PRJNA860050 not described in the paper GroupsDiseased (A. alternata-inoculated) vs. control B. juncea plants; cross-species comparison of B. juncea, B. nigra, and B. rapa gene families Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Maximum-likelihood phylogenetic inference (IQ-TREE) with bootstrap resampling Phylogenetic tree of BjuGLPs, BniGLPs, BraGLPs, and AtGLPs (Fig. 4) 102 BjuGLPs + 51 BniGLPs + 48 BraGLPs + AtGLPs (exact AtGLP n not stated) not stated
Log2 fold-change threshold filter (|log2FC| ≥ 1, i.e., ≥2-fold change) for RNA-Seq differential expression In-silico expression analysis of BjuGLPs at 2 DPI and 4 DPI (PRJNA860050, Fig. 7 heat map) 45 BjuGLPs detected at 2 DPI; 42 at 4 DPI; no biological replicate n stated not stated
Ka/Ks (nonsynonymous-to-synonymous substitution rate) ratio analysis Evolutionary selection analysis of duplicated BjuGLP and BniGLP gene pairs 6 BjuGLP pairs; multiple BniGLP pairs (exact n not stated) not stated
MEME motif discovery (E-value threshold 1e-5) Conserved motif identification in BjuGLPs, BniGLPs, BraGLPs (Figs. 1–3) 102 BjuGLPs, 51 BniGLPs, 48 BraGLPs not stated
MCScanX collinearity/synteny analysis (E-value cut-off 1e-10; similarity threshold >90%) Synteny analysis between B. juncea, B. nigra, and B. rapa (Fig. 5) Whole-genome gene complements of three species not stated
RT-qPCR relative quantification (statistical test not explicitly stated in available text) Validation of 10 selected BjuGLPs in A. alternata-inoculated vs. control Super raya plants 6 RNA samples used (2 of 8 excluded due to faint bands); exact biological replicate n not stated not stated
Approaches that could also have been used
  • Differentially expressed BjuGLPs were identified using a fold-change threshold (|log2FC| ≥ 1) without a reported statistical test or false discovery rate correction
    Could also: A negative binomial model-based DEG caller such as DESeq2 or edgeR, with Benjamini-Hochberg FDR correction, could also have been applied to the same Galaxy-processed read counts — Combining a fold-change threshold with an adjusted p-value (e.g., FDR < 0.05) controls the expected proportion of false positives across the family of ~100 tested genes, which is a widely recommended practice when the number of simultaneous comparisons is large
  • RT-qPCR results were described qualitatively as 'upregulated' or 'downregulated' without a stated statistical test or measure of dispersion across biological replicates
    Could also: A Student's t-test or one-way ANOVA (with Tukey or Bonferroni post-hoc correction for multiple genes) on ΔΔCt values across biological replicates, reported with mean ± SD and individual replicate n, could also be used — Formal testing and dispersion reporting allow readers to assess whether observed expression differences exceed random biological variation and are standard for RT-qPCR validation studies in plant biology
  • Phylogenetic confidence was assessed using 1000 bootstrap replicates with IQ-TREE (maximum-likelihood framework)
    Could also: Bayesian inference (e.g., MrBayes or BEAST) with posterior probability support could also be applied to the same aligned sequences — Bayesian posterior probabilities and ML bootstrap values both quantify node support but differ in interpretation; reporting both or choosing one with an explicit model-selection step (e.g., ModelTest-NG) is a common complementary approach in plant gene-family studies
  • Ka/Ks ratios for duplicated gene pairs were interpreted descriptively by comparing the ratio to thresholds (< 1, = 1, > 1)
    Could also: Branch-site or site models in PAML (codeml) could also be used to test statistically whether positive selection (Ka/Ks > 1) is significant via a likelihood ratio test — A likelihood ratio test against a null model adds a formal probability statement to the selective-pressure inference, complementing the ratio-threshold description used here
  • Cis-regulatory elements were identified in 1500 bp upstream sequences using PlantCARE and counted descriptively
    Could also: Motif enrichment analysis (e.g., AME from the MEME suite or Homer) comparing GLP promoters to a background set of non-GLP promoters could also be applied — Enrichment testing identifies which CRE types are over-represented specifically in GLP promoters relative to genome-wide expectation, adding statistical context to the observed CRE catalogue
  • Two RNA samples (JD3, JC4) were excluded from RT-qPCR analysis due to faint RNA bands, reducing the effective n without a stated imputation or sensitivity analysis
    Could also: Re-extraction of those two samples, or a sensitivity analysis comparing results with and without the excluded samples, could also be performed — Documenting whether conclusions hold after exclusion of low-quality samples, or attempting re-extraction, helps readers assess whether the exclusion affected the directional findings
Software: Galaxy (usegalaxy.org) · IQ-TREE 2.1.2 · TBtools 2.210 · MEME · MCScanX (via TBtools comparative genomics module) · PlantCARE (online)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PF00190 Pfam in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet
PRJNA860050 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

1 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.
PRJNA860050 BioProject reused by 2 papers in the literature
Most-cited downstream papers:

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41327045

Paper: Genome-wide identification and characterization of germin-like protein (GLP) family in Brassica juncea reveals their role against biotic stress. Abdul Sattar G, Rana IA, Atif RM, Niazi AK. BMC Plant Biol 2025. PMC12763953 · DOI 10.1186/s12870-025-07805-y.

"Code" (P16): TBtools v2.210 (third-party GUI) + ExPASy ProtParam, CDD/Pfam, MEME v5.5.0, IQ-TREE v2.1.2, MCScanX, BUSCA, PlantCARE, Galaxy (HISAT2/StringTie/ DESeq2). No authors' own scripts — applying standard tools to the paper's data is P16-valid and reproducible.

Genome (pinned): BRAD V3.0 (http://brassicadb.cn/). B. juncea assembly Braju_tum_V2.0 (server internal name BjuT84V2; var. tumida T84-66). Gene-ID format BjuVA##G##### (AA subgenome), BjuVB##G##### (BB subgenome), Contig##G#####. The BRAD download mirror («ip»:82) only ships juncea V1.1/V1.5; the V2 proteins are served per-gene by the API http://«ip»:8001/api/search-gene-test/?geneID=<ID> (returns protein_fa, cds_fa, coords). NOTE: paper Methods says "Bju_tum_V2"; the gene IDs are the Varuna-style BjuV* tags but the API confirms species=BjuT84V2/Braju_tum_V2.0, coordinates matching the supplement exactly → genome correctly identified.

In scope (pipeline-derived, attempted)

ID Result Source Pipeline Plan
C1 Physicochemical properties of the BjuGLP family (protein length PL, MW kDa, pI) per gene Suppl. Table 1 (MOESM2 .docx), 101 rows BjuGLP01–101 ExPASy ProtParam Fetch each gene's protein from BRAD API; recompute length+MW+pI with Biopython ProteinAnalysis (== ExPASy Bjellqvist/average-mass); compare per gene
C2 Family membership criterion: cupin_1 (PF00190) domain present in all members Methods + Results CDD/Pfam (hmmer) hmmsearch PF00190 against the 101 proteins; count members with domain; single vs double cupin
C3 Chromosomal + strand distribution Results: "Chr02=9 highest, Chr08=1 lowest"; strand 41+/61− TBtools gene-location Count per (sub)chromosome and strand from coords; compare
C4 Family size = 102 BjuGLPs (abstract/text) Abstract/Results BLASTp(AtGLP×36)+CDD Consistency check vs Suppl. Table 1 (which lists 101); full re-BLAST deferred (bulk V2 proteome not in mirror) — method-sensitive 20%

Out of scope / not attempted (with reason)

  • RNA-seq expression (DESeq2 on PRJNA860050) — heavy multi-sample align/quant/DE (12 libraries, HISAT2+StringTie+DESeq2 via Galaxy). The hard 20%; deferred for budget. Reported numbers (45/42 genes detected, 25/30 DEGs) not regenerated.
  • Phylogeny (IQ-TREE), conserved motifs (MEME, "20 motifs"), gene structure, synteny/MCScanX, cis-elements (PlantCARE) — figure-level/visual, no single pinnable scalar to grade 1:1; not attempted.
  • RT-qPCR validation — wet-lab, out of scope.

Auditability notes / possible-fabrication flags (provisional)

  • F1 (count): Abstract/text say 102 BjuGLPs; Suppl. Table 1 enumerates only 101 (BjuGLP01–BjuGLP101). Strand split in the table is 40+/61− (=101); the paper states 41+/61− (=102). Off-by-one between prose and the deposited table.
Figures / tables: Table
C1_physchem_table
Reported
Suppl. Table 1: per-gene protein length / MW(kDa) / pI for the BjuGLP family
Reproduced
as-printed agreement length 17/98, MW 15/98, pI 18/98; values are genuine ProtParam outputs but ROW-MISALIGNED to gene IDs (block -1 shift, BjuGLP03-14 gene N = reported N-1); 9 large values (1808-1897 aa) match no listed gene
did not match
C1_count
Reported
102 BjuGLP genes (abstract/text)
Reproduced
Suppl. Table 1 lists 101 (BjuGLP01-101); strand 40+/61- not 41+/61-
did not match
C2_cupin_PF00190
Reported
all members possess cupin_1 domain; single & double
Reproduced
98/98 fetched proteins contain cupin_1 (PF00190); 31 double
exact
C3_chrom_max
Reported
Chr02 = 9 (highest)
Reproduced
B02=9 (B07=9 tie)
within tolerance
C3_chrom_min
Reported
Chr08 = 1 (lowest)
Reproduced
B08=1
exact
C3_strand
Reported
41+ / 61- (=102)
Reproduced
40+ / 61- (=101)
partial
C4_reidentify_family
Reported
102 via BLASTp(36 AtGLP,E<1e-5)+cupin
Reproduced
not attempted (bulk Braju_tum_V2 proteome not in BRAD download mirror)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 58/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🟡7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The deterministic, scriptable parts of the pipeline reproduce cleanly — cupin_1/PF00190 membership is exact (98/98, 31 double) and chromosomal max/min match (Chr02=9, Chr08=1) — and the method is validated by exact anchors (BjuGLP01/02 length/MW/pI 1:1), so input-data identity and comparability are sound. The defect is on the authors' side: Suppl. Table 1's physicochemical columns are demonstrably row-shifted against the gene IDs (only ~17/98 match as printed) and include ~9 large orphan values (up to 1897 aa) that map to no listed gene, plus a 102-vs-101 family-count discrepancy. The underlying numbers are genuine ProtParam outputs but mis-tabulated, making per-gene lookups unreliable and triggering a possible-fabrication flag. The core family/domain conclusion survives, but the characterization table's integrity failure makes this a critical reproduction outcome.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

376.7 k
tokens (I/O) · 34.3 M incl. cache
52 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine