Blood Transcriptome Analysis of Septic Patients Reveals a Long Non-Coding Alu-RNA in the Complement C5a Receptor 1 Gene.
The main results reproduced, with only marginal, non-material deviations.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction of PMID 35447887 (Emblem 2022) for its SECONDARY/validation accession PRJNA607653 (Taizhou PBMC). Compute ran fully on «our HPC» (STAR 2.7.9a + samtools 1.12, EXACT versions; 44/44 runs aligned 92-94% uniquely; featureCounts over a bedtools-derived Alu_genes.gtf; DESeq2-style median-of-ratios normalization). GENOME-REF claims reproduce well: A1 total Alu within-tol (1218045 vs 1209364, 0.72%; 'without alt loci'=exclude _alt contigs), A2 families within-tol (rounded), A3 within-tol (760343 vs ~800000). A4 PARTIAL+FLAGGED: 863426 gene-pairs vs reported 993146 (-13%); not reproducible by any natural bedtools parameterization -> possible underspecification, human-review flag. DATA-DERIVED validation claims: C2 (the paper's headline validation) REPRODUCES - C5aR1 3'UTR AluSx1 is upregulated in Taizhou sepsis (FC 2.82), with the whole C5aR1 3'UTR Alu cluster trending up. C1 over-counts: 25/26 immune genes recur (only NFKBIZ absent) vs the reported 15/26, because the open substitute is more permissive than the proprietary StrandNGS pipeline run on an unspecified 18+18 subset - validation direction confirmed but exact count differs. D1 mismatch: the deposit holds 44 runs (24 sepsis+20 healthy), not the 36 the paper says it used, and the exact subset is unspecified. NOT attempted: PRJNA647880 primary headline numbers (different accession), wet-lab, on-request whole-blood Illumina/Ion data, StrandNGS internals.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 10assessed: 2026-06-19 ⛓ b9167a4be9e6
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates whether immune genes contain embedded Alu elements that give rise to functional long non-coding Alu-RNAs, and whether such Alu-lncRNAs are differentially expressed in the blood transcriptome of septic patients versus healthy controls, using the complement receptor gene C5aR1 as a candidate target.
- ★ A computational pipeline intersecting immune gene coordinates with Alu element coordinates can identify candidate Alu-lncRNAs method
- ★ 48 Alu insertions were identified in 26 immune genes as robust candidates in sepsis blood transcriptome data finding
- ★ The complement receptor gene C5aR1 contains a novel Alu-lncRNA with an independent transcriptional start site finding
- ★ Alu-containing transcripts cluster by health condition (healthy vs. inflammation) by hierarchical clustering, and by all four conditions by UMAP finding
- ★ Findings were reproduced in an independent sepsis cohort (PRJNA607653) and validated with RNA-seq from an ex vivo S. aureus-activated whole blood model finding
- Alu elements are predominantly located in introns (~90%) with fewer than 5% in 3'UTRs of genes finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (external dataset PRJNA647880) | peripheral blood, human sepsis/septic shock/infection patients and healthy controls | disease state (sepsis/septic shock/infection) vs. healthy | Alu transcript expression levels (raw reads, fold change) | — |
| bulk RNA-seq (external dataset PRJNA607653) | peripheral blood mononuclear cells (PBMC), human sepsis patients and healthy controls | sepsis vs. healthy | Alu transcript expression levels (validation cohort) | — |
| Illumina RNA sequencing | ex vivo whole blood model, healthy human donors (n=6, pooled) | heat-inactivated S. aureus (strain Cowan) vs. PBS, 0/30/60/120 min | transcriptome expression to verify whole blood model as imitation of blood infection | HiSeq 2000, 100 bp paired-end |
| Ion S5 stranded RNA sequencing | ex vivo whole blood model, healthy human donors | heat-inactivated S. aureus vs. PBS | transcriptional direction/strand of Alu-RNAs | Ion 540 chip, Ion GeneStudio S5 System |
| computational genome-wide intersection (bedtools/UCSC RepeatMasker/Ensembl) | human genome (Hg38) | none | number and location of Alu element insertions within genes | — |
| Homer read quantification (analyzeRepeats.pl) | human blood transcriptome BAM files | none | raw/normalized read counts per Alu-gene coordinate | Homer v4.11 |
| manual visual inspection of aligned reads | human blood RNA-seq BAM files | none | confirmation of candidate Alu-lncRNA transcript structure | IGV v2.11 |
- – 1.2 million Alu element insertions identified in the human genome, divided into AluJ (320,000), AluS (727,000), and AluY (149,000)
- – Approximately 800,000 Alu insertions located within genes; ~90% in introns and <5% in 3'UTRs ~90% introns, <5% 3'UTR
- – 1173 Alu transcripts (400 raw read cutoff) intersected in 726 genes, including 324 immune-related genes
- – Hierarchical clustering separated healthy controls from inflammation patients but not by inflammatory subtype; UMAP additionally separated all four conditions
- – 58 genes met fold change (1.3) and raw read (>1000) thresholds; 26 of these (45%) were classified as immune genes fold change 1.3, 45%
- – 48 Alu elements embedded in the 26 immune genes were characterized in Table 1, all differentially expressed transcripts significant at p ≤ 0.001 p ≤ 0.001
- – One of the 48 candidates, located within C5aR1, identified as a novel Alu-lncRNA and validated using whole blood model RNA-seq
- count 1,209,364 Alu coordinates (Alu coordinates parsed from UCSC RepeatMasker track)
- count 993,146 intersections (bedtools intersect of Alu elements with reference gene transcripts)
- count 1173 Alu transcripts in 726 genes (Homer quantification with 400 raw read cutoff)
- count 324 immune-related genes (subset of the 726 genes intersected with Alu transcripts)
- fold_change 1.3 fold change threshold, ≥1000 raw reads (filtering criteria yielding 58 candidate genes)
- count 26 of 58 genes (45%) (genes classified as immune genes among top candidates)
- count 48 Alu insertions in 26 immune genes (final candidate Alu-lncRNA list (Table 1))
- pvalue p ≤ 0.001 (significance threshold for differentially expressed Alu transcripts, two-way t-test with Benjamini-Hochberg correction)
- count 88 samples (40 healthy, 18 sepsis, 18 septic shock, 12 infection) (primary sepsis cohort (PRJNA647880) used for analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study applied a bioinformatic pipeline to bulk RNA-seq data from 88 peripheral blood samples (40 healthy controls, 48 sepsis/septic shock/severe infection patients) to identify Alu-containing lncRNAs within immune genes. Differential expression across the 1,173 quantified Alu transcripts was tested with a two-tailed t-test (described as 'two-way t-test') with Benjamini–Hochberg FDR correction, using a significance threshold of p ≤ 0.001 combined with fold-change (≥1.3) and raw-read (≥1000) filters. Results were visualized via hierarchical clustering (Ward's linkage) and UMAP, and key findings were validated in a second external cohort and an ex vivo whole-blood S. aureus model.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Two-tailed t-test (paper terms: 'two-way t-test'), implemented via Homer/StrandNGS | All 1,173 differentially expressed normalised Alu transcripts; comparison of sepsis/septic shock/infection versus healthy controls | 88 samples (40 healthy controls, 18 sepsis, 18 septic shock, 12 infection) | not stated |
| Benjamini–Hochberg false discovery rate correction | Applied to p-values from all 1,173 Alu transcript comparisons | 1173 tests | na |
| Hierarchical clustering (Euclidean distance, Ward's linkage) | Heatmap of normalised Alu expression values across all 88 participants and 1,173 transcripts (Figure 2a) | 88 samples, 1173 transcripts | na |
| UMAP (Uniform Manifold Approximation and Projection) | Dimensionality reduction and condition-level separation of all 1,173 transcripts across 88 samples (Figure 2b) | 88 samples, 1173 transcripts | na |
| DESeq normalisation (variance stabilisation) | Normalisation of raw Homer read counts prior to clustering and group comparisons | 88 samples | not stated |
-
Differential expression of Alu-containing transcripts was assessed with a two-tailed t-test (labelled 'two-way t-test') on DESeq-normalised counts↳ Could also: A negative-binomial model implemented in DESeq2 (Wald test) or edgeR (likelihood-ratio test) could also be applied directly to raw count data — Negative-binomial models are specifically designed for RNA-seq count data, which are overdispersed; applying them directly to raw counts rather than normalised values is a common alternative that explicitly models count variability and does not require a normality assumption
-
The 1,173 candidate Alu transcripts were filtered by both a statistical threshold (p ≤ 0.001 after BH-FDR) and two additional non-statistical thresholds (fold change ≥1.3 and raw reads ≥1000)↳ Could also: Candidates could also be selected using statistical criteria alone (e.g., adjusted p-value plus a minimum mean normalised count), or by reporting the full ranked list with effect sizes — Combining hard raw-read and fold-change cutoffs with a p-value threshold is one common approach; relying on statistically derived thresholds alone (e.g., an FDR-controlled fold-change estimate with a credible interval) would make the selection criterion fully quantitative and easier to reproduce across datasets with different sequencing depths
-
Four conditions (healthy, infection, sepsis, septic shock) were compared, but the t-test framework appears to have been applied as pairwise or pooled contrasts rather than a single omnibus multi-group model↳ Could also: A one-way ANOVA (or its RNA-seq analogue, a likelihood-ratio test with a multi-level group factor) could also test all four conditions simultaneously before post-hoc pairwise contrasts — An omnibus test first controls the family-wise error rate across all group comparisons in a single model and provides a natural framework for ordered contrasts (e.g., a severity gradient: healthy → infection → sepsis → septic shock)
-
UMAP was used for dimensionality reduction and sample-level visualisation↳ Could also: PCA or t-SNE could also be used for the same visualisation purpose — PCA is linear and directly interpretable in terms of variance explained; t-SNE preserves local structure; the authors cite evidence favouring UMAP for bulk RNA-seq, but all three are widely used and each has contexts where it performs differently — reporting PCA alongside UMAP would show how much variance is captured by the top components
-
The whole-blood ex vivo model used RNA pooled from 6 donors for RNA-seq, resulting in a single pooled library per time point↳ Could also: Individual (non-pooled) libraries from each of the 6 donors could also be sequenced and analysed, preserving donor-level replication — Pooling RNA before library preparation collapses biological variability into a single observation per condition, making it impossible to estimate inter-donor variance or apply inferential statistics; individual libraries would support formal statistical testing of time-course changes
-
Dispersion in expression values is displayed as box plots (Figure 2c) without an explicitly labelled dispersion statistic↳ Could also: Reporting SD, SEM, or 95% CI alongside or instead of box plots would also convey spread in a labelled, reproducible form — Box plots show median and IQR but the specific dispersion metric is often inferred rather than stated; explicitly labelling the measure (and, for small n, preferring SD or 95% CI over SEM to avoid overstating precision) is a common recommendation in reporting guidelines such as those from Nature Methods and the MIQE guidelines for RNA data
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35447887
Paper: Emblem et al. 2022, Noncoding RNA 8(2):24. "Blood Transcriptome Analysis of Septic Patients Reveals a Long Non-Coding Alu-RNA in the Complement C5a Receptor 1 Gene." DOI 10.3390/ncrna8020024 · PMCID PMC9027897.
The paper's pipeline (the "in-house Alu-lncRNA pipeline", Methods §2.1–2.9)
- Alu coordinates from UCSC RepeatMasker track (hg38), Table Browser → GTF.
- Gene coordinates from Ensembl GRCh38.105 → BED.
- bedtools intersect Alu vs genes (
-wa -wb -f 0.009) →Alu_genes.gtf. - STAR v2.7.9a unguided alignment (no GTF) of FASTQs to hg38 (no alt loci).
- samtools v1.12 sort.
- Homer v4.11
analyzeRepeats.pl Alu_genes.gtf hg38 -strand both -min 400 -raw→ count matrix. - StrandNGS v4 (proprietary): DESeq, hierarchical clustering, UMAP; filter fold-change ≥1.3 AND raw reads ≥1000 in ≥1 sample → candidate list.
- Manual IGV inspection; FANTOM5 CAGE TSS annotation; t-test + Benjamini-Hochberg (p<0.001).
Datasets the paper uses (and which one this room is assigned)
| accession | role in paper | N (paper) | access |
|---|---|---|---|
| PRJNA647880 | PRIMARY dataset → all headline numbers (Fig 2, Table 1, Table S2): 1173 transcripts, 726 genes, 58 genes/48 Alu, 26 immune genes | 88 (sepsis 18, septic shock 18, infection 12, healthy 40) | open |
| PRJNA607653 | ← THIS ROOM. SECONDARY validation dataset (Taizhou Hospital PBMC), Figure S3 / Table S3 | "Sepsis n=18, Healthy n=18" = 36 | open |
| Illumina whole-blood model | own data, Fig 3 / Table S4 | 6 donors pooled | on request (restricted) |
| Ion Torrent whole-blood model | own data, Fig 3, direction check | 6 donors pooled | on request (restricted) |
This room's accession (brief) = PRJNA607653. It is the paper's validation set, NOT the primary set. Therefore the room's directly-reproducible, data-derived claims are the ones the paper states for PRJNA607653 (Fig S3 / Table S3):
- C5aR1 3′UTR AluSx1 upregulated in the Taizhou sepsis data.
- 15 of the 26 Fig 2d immune genes also appear in the Taizhou figure (Fig S3b).
In scope (pipeline-derived → attempted)
A. Genome-reference results (no patient data; public refs; cheap). Upstream of BOTH datasets, fully specified, high-value quick wins:
- A1: total Alu insertions in hg38 RepeatMasker = 1,209,364 (Methods §2.8).
- A2: family split AluJ ≈320,000 / AluS ≈727,000 / AluY ≈149,000 (Results §3.1).
- A3: ≈800,000 Alu within genes (Results §3.1).
- A4: bedtools intersections with genes (
-f 0.009) = 993,146 (Methods §2.8).
B. PRJNA607653 data-derived («our HPC» compute).
- B1: download 36/44 FASTQs, STAR→hg38, Homer count matrix over
Alu_genes.gtf. - B2: DE (sepsis vs healthy), fold-change ≥1.3 & raw ≥1000 filter → candidate Alu list.
- B3: overlap of those candidates' immune genes with the paper's 26 → expect 15 (Fig S3b).
- B4: C5aR1 AluSx1 direction/up-regulation in PRJNA607653.
Out of scope (not attempted; reason)
- Wet-lab: RNA isolation, S. aureus whole-blood model, RIN, library prep — not computational.
- Illumina + Ion Torrent whole-blood data (Fig 3, Tables S4): "available on request" → restricted.
- StrandNGS proprietary software: license-gated. Substituted with open tools (DESeq2/edgeR + explicit fold-change/raw filter) per BRIEF P16 (third-party-tool reproduction is valid). Substitution recorded as a caveat, not a 1:1 of StrandNGS internals.
- UMAP / hierarchical-cluster figures (Fig 2a/b, Fig S3a): qualitative cluster figures from the proprietary tool; the derived counts (overlap, direction) are the auditable targets.
- PRJNA647880 headline numbers: that is a different accession (the primary set), not the one assigned to this room. Reported here for context only; reproducing it is a separate RU.
Key reproducibility caveats found in the paper
- Sample selection unspecified: paper used "18 sepsis + 18 healthy" from PRJNA607653, but the
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.