Does excitatory fronto-extracerebral tDCS lead to improved working memory performance?
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reconstructed the group assignment (Active n=11, Sham n=10) via exhaustive combinatorial search since it is absent from the shared data, then reproduced Table 1 descriptives (38/40 cells within ~0.01-0.03, 2 cells flagged as likely paper transcription errors), the baseline d' t-test (within-tol), the Day-1 tDCS exploratory ANCOVA contrast (within-tol), and the RM-ANOVA main effect of time (within-tol) and group (partial) -- but found a genuine mismatch on the group x time interaction (we find it significant, p=0.038; paper reports non-significant, p=0.277), likely driven by our necessarily simpler complete-case, no-covariate ANOVA versus the paper's SPSS MIXED heterogeneous-AR1 model with Satterthwaite df, since no R/scipy/statsmodels was available on the HPC compute nodes and pure-Python distribution functions were implemented and validated instead.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-31
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDoes excitatory (anodal) 1 mA tDCS applied to the left dorsolateral prefrontal cortex with a contralateral extracerebral (cheek) reference electrode improve working memory (3-back) performance relative to sham stimulation? The authors hypothesized that active stimulation recipients would show greater task performance improvement relative to baseline than sham recipients across two stimulation days.
- ★ Active anodal left DLPFC tDCS with a contralateral cheek reference did not significantly enhance 3-back working memory performance over sham across the two-day experiment (no main effect of group, no group x time interaction). finding
- ★ Exploratory Bonferroni-corrected comparisons revealed a significant advantage of active over sham stimulation on d' during the first stimulation phase (day 1 tDCS) only, with a large effect size. finding
- ★ The results raise the possibility that tDCS effects on working memory may be greatest during early learning stages of the task. finding
- Task performance improved across both groups with increasing exposure to the 3-back (significant main effect of time). finding
- During day 1 stimulation, hit rate was significantly better and correct rejection rate was better at trend level in the active group, while reaction times did not differ. finding
- ★ A fronto-extracerebral montage (anode F3, cathode on contralateral cheek) avoids the interpretational confound of a scalp reference electrode, where effects could arise from excitation, inhibition, or their combination. method
- A double-blind, between-subjects design using the device 'study mode' kept stimulation-administering researchers blinded to condition. method
- The 3-back task code (Matlab/Cogent) and the participant performance dataset are made freely and permanently available (DOI 10.5281/zenodo.7148; data under CC0). resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| 3-back working memory task (behavioural, d' signal detection) at baseline | Healthy human volunteers (N=21, 14 female, mean age 23.09 years, SD 3.95), right-handed, UCL subject pool | none (5-minute pre-stimulation baseline, day 1) | d' (Z(hit rate) - Z(false alarm rate)), hit rate, correct rejection rate, reaction time | Task coded in Matlab release 2008b for Windows (Mathworks, Natick, MA, USA) with the Cogent Toolbox |
| 3-back working memory task performed concurrently with tDCS (D1 tDCS, D2 tDCS) | Healthy human volunteers randomized to active (N=10) or sham (N=11) stimulation | 1 mA anodal tDCS over F3 (left DLPFC) with cathodal reference on contralateral cheek, 10 minutes with 15-s fade-in/fade-out, vs sham (100-200 µA pulses every 400-550 ms between 15-s ramps) | d', hit rate, correct rejection rate, hit and correct-rejection reaction times during stimulation | Neuroconn DC-Stimulator (Neuroconn, Germany); 7 cm × 5 cm rubber electrodes in saline-dampened sponges; F3 located with a 10-20 EEG cap |
| 3-back working memory task immediately following stimulation (D1 post-tDCS, D2 post-tDCS) | Healthy human volunteers, active vs sham groups | no stimulation (10-minute post-tDCS run); monetary performance incentive (£10 bonus on day 2 for beating day 1 test phase) | d', hit rate, correct rejection rate, reaction time post-stimulation | Matlab (2008b) with Cogent Toolbox; Neuroconn DC-Stimulator used in preceding phase |
| Linear mixed model analysis of repeated 3-back performance (and follow-up general linear models, independent-samples t-tests, Bonferroni-corrected linear contrasts) | Data from 21 healthy volunteers across five testing sessions over two days (24-48 h apart) | Fixed effects of time (4 post-baseline sessions), group (active/sham) and time-by-group; participant as random effect; baseline performance as covariate; heterogeneous first-order autoregressive covariance structure | d' as dependent variable; follow-up models on hit rate, correct rejection rate and reaction time | SPSS version 21 (IBM Corp New York 2012) |
- – No main effect of stimulation group on d' performance across the experiment F(1,16) = 2.228, P = 0.155
- – No group × time interaction on d' performance F(3,36) = 1.339, P = 0.277
- ▲ Significant main effect of time on d', reflecting improvement across both groups with increasing task exposure F(3,36) = 7.669, P < 0.001
- ▲ Active stimulation outperformed sham on d' at the day 1 tDCS time point (Bonferroni corrected, controlling for baseline) F(1,13.373) = 10.747, P = 0.006; Cohen's d = 1.427, r2 = 0.337
- – No group differences in d' at any other post-baseline time point all F < 1.2, P > 0.3
- ▲ Hit rate during day 1 stimulation was significantly higher in the active group than sham F(1,18) = 4.454, P = 0.049, ηp2 = 0.198
- ▲ Correct rejection rate during day 1 stimulation was better at trend level in the active group F(1,18) = 3.680, P = 0.071, ηp2 = 0.170
- – No reaction time differences at the day 1 stimulation time point for hits or correct rejections hits F(1,18) = 0.010, P = 0.923, ηp2 = 0.001; correct rejections F(1,18) = 0.202, P = 0.659, ηp2 = 0.011
- pvalue F(1,13.373) = 10.747, P = 0.006 (Active vs sham d' difference at day 1 tDCS time point, Bonferroni corrected, baseline as covariate)
- other Cohen's d = 1.427, r2 = 0.337 (Effect size for the day 1 during-stimulation group difference)
- pvalue F(3,36) = 7.669, P < 0.001 (Main effect of time on d' in the linear mixed model)
- pvalue F(1,16) = 2.228, P = 0.155 (Non-significant main effect of stimulation group on d')
- pvalue t(19) = 1.044, P = 0.309 (No baseline d' difference between active and sham groups)
- count active N = 10, sham N = 11; 21 total (14 females) (Group allocation of right-handed participants, mean age 23.09 years (SD 3.95))
- mean Day 1 tDCS hit rate: anodal 0.5250 (SD 0.1336) vs sham 0.4030 (SD 0.1197) (Table 1 hit rate during day 1 stimulation)
- other 80% power to detect a large effect size (d = 1.3) at P = 0.05 two-tailed (A priori power for between-group comparison given the sample size)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This double-blind, between-subjects study (active tDCS N=10, sham N=11) analyzed 3-back working memory performance (d') using a linear mixed model in SPSS (time, group, and time×group as fixed effects; participant as random effect; baseline performance as a covariate; heterogeneous first-order autoregressive covariance structure). The omnibus model showed a significant main effect of time but no significant main effect of group or group×time interaction. Exploratory, Bonferroni-corrected pairwise contrasts at each of four post-baseline time points, followed up with general linear models, identified a significant group difference only at the day 1 stimulation time point, which was also examined for hit rate, correct rejection rate, and reaction time.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Independent samples t-test | baseline d' performance between groups; age between groups | N=21 (10 active, 11 sham) | not stated |
| Chi-square test | gender distribution between groups | N=21 | not stated |
| Linear mixed model (fixed effects of time, group, time×group; random effect of participant; baseline as covariate; heterogeneous AR1 covariance structure) | d' across four post-baseline testing sessions (D1 tDCS, D1 post-tDCS, D2 tDCS, D2 post-tDCS) | N=21, with one participant's incomplete/missing session data included in the model | not stated |
| Bonferroni-corrected pairwise linear contrasts | group differences in d' at each of the four post-baseline time points | not stated per contrast | not stated |
| General linear model with baseline as covariate | follow-up group comparison at the D1 tDCS time point for d', hit rate, correct rejection rate, and reaction time | denominator df of 18 reported for hit rate/correct rejection rate/RT models; 13.373 for the d' contrast | not stated |
-
Exploratory pairwise group comparisons at each of the four post-baseline time points were corrected for multiplicity using the Bonferroni method.↳ Could also: A false discovery rate procedure (e.g., Benjamini-Hochberg) or a set of pre-planned contrasts embedded directly in the mixed-model framework — FDR-based correction can offer greater power than Bonferroni when testing a modest number of related comparisons, which may be useful when comparisons are of similar exploratory interest.
-
The primary time-course analysis used a linear mixed model with a heterogeneous first-order autoregressive covariance structure.↳ Could also: Comparing alternative covariance structures (e.g., compound symmetry, unstructured) via information criteria (AIC/BIC), or a repeated-measures ANOVA with a Greenhouse-Geisser correction — Comparing covariance structures can show how sensitive the time-by-group effect is to the specific structure chosen, and repeated-measures ANOVA is a widely used alternative framework for this type of design.
-
Spread was reported as SD in the summary table and as SEM in the main results figure.↳ Could also: Reporting 95% confidence intervals for the group difference estimates — CIs convey the precision and plausible range of an estimated effect directly, which can be a useful complement to p-values, particularly with per-group sample sizes of 10 and 11.
-
Sample size was justified by a stated 80% power to detect a large effect (d=1.3) at P=0.05, described in relation to the sample obtained.↳ Could also: An a priori power analysis specifying the minimum sample size needed to detect the smallest effect of theoretical interest before data collection — A prospective power calculation ties the planned sample size directly to a pre-specified effect of interest, which can be a useful complement to a post hoc power statement.
-
Baseline and follow-up group comparisons relied on parametric tests (t-tests, GLM, LMM) assuming approximately normal residuals.↳ Could also: Non-parametric alternatives such as the Mann-Whitney U test for baseline group comparisons, or permutation-based tests for the mixed-model contrasts — Non-parametric or permutation approaches do not depend on distributional assumptions, which can be a helpful check with modest per-group sample sizes (N=10-11).
-
d', hit rate, correct rejection rate, and reaction time were each analyzed with separate GLMs at the D1 tDCS time point.↳ Could also: A multivariate analysis of variance (MANOVA) or a joint model across the correlated outcome measures — Modeling correlated outcomes jointly can account for their shared variance and control the family-wise error rate across the set of related dependent measures.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The deposited data are genuinely raw and near-complete (124/126 files, the 2 gaps matching the stated day-2 dropout), but they omit the Active/Sham group labels, so the grouping had to be recovered by exhaustive search over 705,432 splits — a best-fit estimate, not ground truth, that propagates into every group statistic. Under that reconstruction the paper holds up well: 38/40 Table 1 cells match within ~0.01–0.03, baseline t(19)=1.044→1.180, the key Day-1 ANCOVA F=10.747, P=0.006→F=11.515, P=0.0032, and the time effect F=7.669→8.971 all confirm. Two deviations are on the authors' side: the Sham CRR SDs of 0.2757/0.2821 are unreachable from per-subject CRRs spanning only 0.817–0.983 (transcription error, not fabrication), and the reported non-significant group×time interaction F(3,36)=1.339, P=0.277 became F(3,54)=3.004, P=0.038 on reproduction. That flip is confounded with our own forced simplification (no R/scipy/statsmodels on the compute nodes → OLS instead of heterogeneous-AR1 MIXED, complete-case N=20), so it is a limitation of the reproduction as much as a challenge to the paper — hence yellow overall rather than red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.