Regulatory Noncoding Small RNAs Are Diverse and Abundant in an Extremophilic Microbial Community.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to attempt, but the headline pipeline result is NOT a clean 1:1 against the public artifacts. SnapT (github.com/ursky/SnapT @a51df26, the authors' own tool) was read in full; the authors deposited its actual output as GEO GSE137164 supplementary files (GFFs + per-library TPM tables, SHA256 recorded). Comparing those deposited outputs to the paper's Table 1 yields a robust, auditable discrepancy: the deposited final sRNA set has 1750 sRNAs (1037 antisense + 713 intergenic), ~14% MORE than the reported 1538 (925 + 613). The deposited TPM tables (1003/712) sit between. No single-sample TPM cutoff on the GFF recovers 1538; the reported counts shrink monotonically (GFF 1750 -> TPM-table 1715 -> Table1 1538), implying an additional, undocumented cross-library expression filter (SnapT's final manual TPM cutoff is never pinned to a value). The paper's Table 1 is internally self-consistent (all percentages match the 1538 base), so the deposited data is a looser superset of the headline numbers rather than evidence of fabrication. REPRODUCED cleanly: sRNA size range 50-500 nt (deposited 51-499, exact), antisense/intergenic split ~60/40% (deposited 59.3/40.7%, within-tol), n=45 libraries (ENA, exact). NOT ATTEMPTED / BLOCKED: full end-to-end SnapT rerun (HISAT2->StringTie->Prodigal->DIAMOND/nr->Infernal/Rfam) because the reference co-assembled metagenome (JGI IMG taxon 3300027982, Ga0224629) is login-gated and not anonymously downloadable (files.jgi.doe.gov=0 hits; IMG TaxonDetail HTTP 403; absent from NCBI PRJNA484015 and the ursky/timeline_paper repo). The 45 raw reads are public but cannot reproduce the paper's exact contigs/coordinates without the identical reference; a surrogate assembly would be a different analysis, so it was not fabricated. Consequently the Rfam(79), conserved(155), and DE(109) claims are blocked (they need sRNA sequences from the gated reference and/or raw counts not deposited). Net: a strong, fully hand-auditable PARTIAL reproduction whose central finding is a deposited-vs-published count discrepancy / reproducibility gap, plus a confirmed login-gated reference blocker for the full rerun.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 57assessed: 2026-06-16 ⛓ 3884678f7e83
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDo regulatory noncoding small RNAs (intergenic and antisense sRNAs) exist, are they diverse and abundant, and do they participate in environmental adaptive gene regulation in a natural extremophilic microbial community inhabiting halite nodules in the Atacama Desert, as detectable at the community level via metatranscriptomics?
- ★ Hundreds of intergenic (itsRNAs) and antisense (asRNAs) sRNAs are diverse and abundant in the halite endolithic microbial community, with 1,538 total ncRNAs discovered across Archaea and Bacteria. finding
- ★ SnapT, a new bioinformatic pipeline, enables sRNA discovery and annotation in any microbial community using strand-specific metatranscriptomics aligned to a reference metagenome. resource
- ★ asRNA expression levels are negatively correlated with that of their overlapping putative target genes, indicating cis-antisense repression at the community level. mechanism
- ★ A subset of itsRNAs are conserved and significantly differentially expressed between two sampling time points (2016 vs 2017), showing sRNAs can be modulated in the natural environment. finding
- ★ Target prediction and structure modeling of conserved sRNAs link putative mRNA targets to environmental challenges such as osmotic adjustment and nutrient competition. method
- ★ itsRNAs and asRNAs are expressed ~2-fold higher than protein-encoding genes (normalized by contig abundance), suggesting functional relevance. finding
- 79 ncRNAs match 6 known Rfam families and 155 additional ncRNAs are conserved in NCBI nt, validating the experimental and computational approach. finding
- Several sRNAs were experimentally validated by RT-PCR in environmental and enrichment cultures. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| strand-specific metatranscriptomics (stranded RNA-seq) | halite endolithic microbial community, Atacama Desert (field replicates 2016 n=21, 2017 n=24) | none (natural environmental variation between sampling time points) | sRNA/transcript expression levels (TPM), differential expression | — |
| genome-resolved metagenomics (coassembly) | halite microbial community metagenome | none | taxonomic distribution, gene annotation, metagenome-assembled genomes (MAGs) | — |
| sRNA annotation pipeline (SnapT, bioinformatic) | halite community metatranscriptome aligned to coassembled metagenome | none | itsRNA/asRNA identification with coverage (5x/10x) and size (50-500 nt) thresholds | SnapT (https://github.com/ursky/SnapT) |
| homology/conservation analysis (Rfam search, blastn) | discovered ncRNAs vs Rfam and NCBI nt databases | none | number of conserved ncRNA families and hits | blastn (E<=1E-3, >=70% similarity, >=50% coverage); Rfam database |
| RT-PCR experimental validation | halite environmental samples and enrichment cultures (haloarchaea, Cyanobacteria) | culture media with high (25%) and low (18%) salt, various carbon sources | confirmation of sRNA transcript presence via amplicon sequencing | — |
| sRNA structure and target prediction | conserved/differentially expressed itsRNAs with high-quality MAGs | none | consensus secondary structures and predicted mRNA targets | LocARNA-P (STAR profiles); IntaRNA |
| differential expression analysis | itsRNAs across 2016 vs 2017 field samples | natural environmental variation (post rain event) | significantly differentially expressed itsRNAs (FDR<5%) | — |
- – 1,538 total ncRNAs discovered; 925 (60%) antisense sRNAs and 613 (40%) intergenic sRNAs 1,538 total
- – ncRNAs distributed 54% Archaea / 46% Bacteria, similar to total metatranscriptomic reads 54% vs 46%
- – 3 times more itsRNAs in Archaea than Bacteria; asRNAs more abundant in Bacteria (Cyanobacteria 38%, Bacteroidetes 15%) 3-fold
- ▼ Highly expressed asRNAs (>100 TPM) associated with lowly expressed target genes (<0.1 TPM); 9 statistically significant negatively correlated asRNA-gene pairs P<0.01; 9 pairs
- – 109 (18%) of regulatory itsRNAs significantly differentially expressed between 2016 and 2017 (FDR<5%); 72% archaea, 28% bacteria, 16 conserved 109 (18%)
- ▲ itsRNAs and asRNAs expressed 2-fold higher than protein-encoding genes (normalized by contig abundance) in both years 2-fold
- – 79 ncRNAs (5%) match 6 Rfam families (incl. RNaseP, SRP RNAs, tRNAs, cobalamin riboswitch, CyVA-1); 155 (10%) conserved in NCBI nt (60% archaea, 40% bacteria) 79 Rfam; 155 conserved
- – Of asRNA-target pairs with asRNA>100 TPM and gene<0.1 TPM, 77% haloarchaea, 12% Cyanobacteria, 11% other bacteria; functions enriched for transport (16%) and cell membrane/wall (5%), 44% hypothetical 77%/12%/11%
- count 1,538 (100%) total ncRNA; 925 (60%) antisense; 613 (40%) intergenic; 79 (5%) Rfam; 155 (10%) conserved (Summary of ncRNAs discovered in halite community (Table 1))
- count 21 and 24 replicates for 2016 and 2017 respectively (metatranscriptomic field replicate samples)
- fold_change 2-fold higher (itsRNA/asRNA expression vs protein-encoding genes (normalized by contig abundance))
- fold_change 3 times more (itsRNAs in Archaea vs Bacteria)
- count 109 (18%) itsRNAs differentially expressed, FDR<5% (DE between 2016 and 2017 samples)
- pvalue P<0.01 (9 significant negatively correlated asRNA-gene pairs (Pearson))
- count taxonomic split 54% Archaea / 46% Bacteria (ncRNA assignment across domains)
- other blastn maximum E value 1E-3, >=70% similarity, >=50% coverage (thresholds for conserved ncRNA detection in NCBI nt)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This metatranscriptomic study identified and characterized sRNAs in a halite microbial community using a custom bioinformatic pipeline (SnapT) applied to strand-specific RNA-seq data from field samples collected in two years (21 replicates in 2016, 24 in 2017). Differential expression of intergenic sRNAs between the two time points was assessed using FDR-controlled analysis, and co-expression relationships between antisense sRNAs and their putative mRNA targets were evaluated via Pearson correlation. Results were reported primarily as TPM-normalized expression values, log2 fold changes, PCA plots, and heat maps, with significance thresholds of FDR < 5% (differential expression) and P < 0.01 (correlations).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression analysis with FDR control (method not specified; likely a negative-binomial-based test such as DESeq2 Wald or edgeR likelihood ratio) | itsRNA expression levels compared between 2016 and 2017 field samples (Fig. 3, Data Set S1) | 21 replicates (2016) and 24 replicates (2017) | not stated |
| Pearson correlation | Expression levels of asRNAs versus their putative mRNA target genes across all replicates (Fig. 2B, Fig. S5A) | 45 total replicate samples (21 + 24), exact per-pair n not stated | not stated |
| Principal component analysis (PCA) | itsRNA expression levels clustered by year (Fig. 3A); metatranscriptomic annotated gene expression across samples (Fig. S5B) | 45 total replicate samples | na |
| Coverage threshold filtering (5× for itsRNAs, 10× for asRNAs) | SnapT pipeline for sRNA candidate selection (Table 1, Fig. S2A) | — | not stated |
| blastn similarity search (E ≤ 1E−3, ≥70% identity, ≥50% coverage) | Identification of conserved sRNAs in NCBI nt database (Table 1) | — | na |
-
Differential expression was assessed with an FDR threshold of 5%, but the underlying statistical model and software are not named↳ Could also: Explicitly named tools such as DESeq2, edgeR, or limma-voom with stated dispersion estimation method could also be applied and reported — Naming the model makes the analysis reproducible and allows readers to assess normalization strategy, dispersion shrinkage, and model assumptions — particularly important for metatranscriptomic data where sequencing depth varies across community members
-
Pearson correlation was used to assess co-expression between asRNAs and mRNA targets across replicates, with a P < 0.01 threshold↳ Could also: Spearman rank correlation could also be used; additionally, FDR adjustment across all tested asRNA-mRNA pairs would also control the family-wise error rate for the set of correlations — TPM values are right-skewed and non-normal; Spearman is robust to this. With 925 asRNAs each tested against one or more targets, an FDR adjustment across correlations would also provide a formal multiplicity control complementing the per-test P threshold
-
Expression levels were summarized and plotted as mean TPM across replicates (Fig. 2A, C, D) without dispersion estimates↳ Could also: Reporting SD, SEM, or 95% CIs alongside mean expression would also convey within-group variability — With 21–24 biological replicates per year, variability estimates are well-powered and would help readers gauge the consistency of expression patterns across the community
-
PCA was used to visualize separation of 2016 vs. 2017 itsRNA expression profiles↳ Could also: Permutational MANOVA (e.g., adonis/PERMANOVA in the vegan R package) could also formally test whether year explains a significant proportion of multivariate expression variance — PCA is a visualization tool; a permutation-based multivariate test would add a significance statement to the observed clustering and is common in community-level omics analyses
-
Coverage-based thresholds (5× for itsRNAs, 10× for asRNAs) were used to filter candidate sRNAs↳ Could also: Expression-based filtering using a counts-per-million (CPM) threshold implemented within the chosen DE framework (e.g., edgeR's filterByExpr) could also be applied — Library-size-normalized expression filters adapt to sequencing depth per sample and are integrated with DE pipelines, which can reduce threshold sensitivity to total read counts across heterogeneous metatranscriptomic libraries
-
The two time points (2016, 2017) were treated as independent groups in the differential expression design↳ Could also: A design that also models individual halite nodule identity (if paired or nested samples were available) or includes continuous environmental covariates (e.g., time since rain event in days) could also be used — Incorporating the 3-month vs. 15-month gradient as a continuous predictor, or blocking on nodule if the same nodules were sampled repeatedly, would increase statistical power and model the biological gradient more precisely
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
79 (5%) of discovered ncRNAs match known Rfam families including RNase P, SRP RNA, tRNAs, and cobalamin riboswitch; 155 (10%) are conserved in NCBI nt, with 60% archaeal and 40% bacterial hits.other atacama desert halite endolithic community 2020×1papers★ This paper is the founder (earliest)
-
Highly expressed antisense sRNAs (>100 TPM) are significantly negatively correlated with lowly expressed target genes (<0.1 TPM) in the halite community, with 9 statistically significant asRNA-gene pairs (P<0.01).RNA-seq atacama desert halite endolithic community down 2020×1papers★ This paper is the founder (earliest)
-
asRNA targets in the halite community are enriched for transport (16%) and cell membrane/wall (5%) functions; 77% of high-confidence asRNA-target pairs originate from haloarchaea and 44% involve hypothetical proteins.RNA-seq atacama desert halite endolithic community 2020×1papers★ This paper is the founder (earliest)
-
109 (18%) of intergenic sRNAs are significantly differentially expressed between 2016 and 2017 sampling years (FDR<5%), with 72% archaeal origin, indicating environmental responsiveness of the ncRNA repertoire.RNA-seq atacama desert halite endolithic community mixed 2020×1papers★ This paper is the founder (earliest)
-
Intergenic sRNAs are 3-fold more abundant in Archaea than Bacteria in the halite community, while antisense sRNAs are enriched in Bacteria, particularly Cyanobacteria (38%).RNA-seq atacama desert halite endolithic community 2020×1papers★ This paper is the founder (earliest)
-
ncRNAs are distributed proportionally between Archaea (54%) and Bacteria (46%) in the halite endolithic community, mirroring overall metatranscriptomic composition.RNA-seq atacama desert halite endolithic community 2020×1papers★ This paper is the founder (earliest)
-
1,538 small regulatory RNAs are identified in the halite endolithic microbial community (60% antisense, 40% intergenic), demonstrating high ncRNA diversity and abundance in an extremophilic environment.RNA-seq atacama desert halite endolithic community 2020×1papers★ This paper is the founder (earliest)
-
sRNAs (itsRNAs and asRNAs) are expressed 2-fold higher than protein-encoding genes in the halite endolithic community when normalized by contig abundance.RNA-seq atacama desert halite endolithic community up 2020×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32019831
Paper: Gelsinger DR, Uritskiy G, Reddy R, Munn A, Farney K, DiRuggiero J. "Regulatory Noncoding Small RNAs Are Diverse and Abundant in an Extremophilic Microbial Community." mSystems 5:e00584-19 (2020). PMID 32019831 / PMC7002113 / DOI 10.1128/msystems.00584-19.
Tool / repo: SnapT — Small ncRNA Annotation Pipeline for Transcriptomes.
https://github.com/ursky/SnapT (HEAD commit a51df26f3b66617059ba5130d0da312d353f6994,
v0.4). This is the authors' own tool (lead/2nd author = Uritskiy/Gelsinger), so it
is both "own code" and a general third-party tool (P16-eligible either way).
Data:
- Metatranscriptome RNA-seq: GEO GSE137164 = SRA SRP221175 = BioProject PRJNA564677. 45 paired-end RNA-Seq runs (verified against ENA, exact).
- Reference: co-assembled halite metagenome, JGI IMG taxon OID 3300027982
(metaSPAdes via metaWRAP; contigs
NODE_*, IMG gene prefixGa0224629). Deposited only at the JGI Genome Portal — login-gated (see blocker below). BioProject for the metagenome reads: PRJNA484015. Previous study describing the co-assembly: Uritskiy et al. 2019, PMC6794293; analysis repo https://github.com/ursky/timeline_paper (ships IMG annotation tables but not the assembly FASTA).
SnapT pipeline (per the repo bin/snapt, v0.4)
- HISAT2 align stranded RNA-seq to the (meta)genome (
--no-spliced-alignment,--rna-strandness FR,--score-min C,-12). - StringTie reference-guided assembly, run twice:
-c 10(antisense) and-c 5(intergenic);-m 50. - Prodigal ORFs; intersect to drop coding transcripts (
snapt_intersect_gff.py,snapt_consolidate_transcripts.py); keep intergenic (≥30 nt from genes) and antisense (≥10 nt overlap opposite strand). - Re-Prodigal on transcripts; drop those with an ORF > 1/2 length; edge-trim;
size-select 50–500 nt (
snapt_size_select.py). - DIAMOND blastx vs NCBI nr (drop protein-coding: bitscore>50, e<1e-4, pid>30, qcov>30).
- Infernal
cmscanvs Rfam 14.1 (--cut_ga --FZ 5 --nohmmonly); drop tRNA/RNaseP (v0.4) hits. - TPM expression curve (
snapt_tpm_curve.py) → manual low-expression cutoff recommended.
In scope (pipeline-derived) — what we attempt
| result | claim | pipeline | status |
|---|---|---|---|
| Total sRNAs | 1,538 | SnapT full | mismatch vs deposited (deposited=1750) |
| Antisense / intergenic split | 925 / 613 (60/40%) | SnapT | counts mismatch; % within-tol |
| sRNA size range | 50–500 nt | SnapT size-select | exact (deposited 51–499) |
| n libraries | 45 | — | exact (ENA: 45 RNA-Seq runs) |
| Rfam matches | 79 (5%) | Infernal/Rfam | blocked (needs reference FASTA) |
| Conserved sRNAs | 155 (10%) | clustering | blocked (needs sequences) |
| DE itsRNAs | 109 (18%) | featureCounts+DESeq2 | blocked (needs raw counts; only TPM deposited) |
Out of scope (not pipeline / not attempted)
- IntaRNA target prediction, LocARNA-P structure, Northern-blot validation, qPCR, taxonomic 54/46% split (derived from the gated metagenome annotation): wet-lab or dependent on the gated reference.
Hard blocker (honest)
The exact reference co-assembly (JGI IMG 3300027982) is not anonymously
downloadable — files.jgi.doe.gov search returns 0 files; IMG TaxonDetail returns
HTTP 403. Every sequence-level SnapT step (HISAT2 index, StringTie, getfasta of
sRNAs, DIAMOND, Infernal) requires this FASTA. The 45 raw-read runs are public but
cannot reproduce the paper's contigs/coordinates without the identical reference,
and substituting a self-made assembly would be a different analysis, not a
reproduction. → full end-to-end rerun = data_restricted (reference login-gated).
What we therefore reproduce (achievable, fully auditable)
A deposited-output vs published-Table-1 audit: the authors deposited the actual SnapT output (GFFs + TPM tables) as GEO supplementary files. We compare those artifacts directly to the paper's reported numbers. This
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The headline SnapT counts do not reproduce 1:1: the authors' own deposited GEO GSE137164 output holds 1750 sRNAs (1037 antisense + 713 intergenic) versus Table 1's 1538 (925 + 613), a consistent ~14% superset, and the 1538 is not recoverable from any public file because the final manual TPM cutoff is never specified. This sits on the authors'/data-availability side (underspecified final filter + login-gated JGI reference blocking the full rerun and the Rfam/conserved/DE claims), not on our method. Severity is moderate: direction, the ~60/40 antisense split, the 50–500 nt size range and n=45 libraries all reproduce, so the central "diverse and abundant antisense-dominated sRNA" conclusion holds with limited confirmation. No fabrication signature — an explainable deposited-vs-published auditability gap.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.