Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Does excitatory fronto-extracerebral tDCS lead to improved working memory performance?

F1000Res · 2013
L1 61/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
61/100
Reproducibility score
0.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 22% of all assessed papers rank 906 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reconstructed the group assignment (Active n=11, Sham n=10) via exhaustive combinatorial search since it is absent from the shared data, then reproduced Table 1 descriptives (38/40 cells within ~0.01-0.03, 2 cells flagged as likely paper transcription errors), the baseline d' t-test (within-tol), the Day-1 tDCS exploratory ANCOVA contrast (within-tol), and the RM-ANOVA main effect of time (within-tol) and group (partial) -- but found a genuine mismatch on the group x time interaction (we find it significant, p=0.038; paper reports non-significant, p=0.277), likely driven by our necessarily simpler complete-case, no-covariate ANOVA versus the paper's SPSS MIXED heterogeneous-AR1 model with Satterthwaite df, since no R/scipy/statsmodels was available on the HPC compute nodes and pure-Python distribution functions were implemented and validated instead.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.7148

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-31
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Does excitatory (anodal) 1 mA tDCS applied to the left dorsolateral prefrontal cortex with a contralateral extracerebral (cheek) reference electrode improve working memory (3-back) performance relative to sham stimulation? The authors hypothesized that active stimulation recipients would show greater task performance improvement relative to baseline than sham recipients across two stimulation days.

Core claims
  • Active anodal left DLPFC tDCS with a contralateral cheek reference did not significantly enhance 3-back working memory performance over sham across the two-day experiment (no main effect of group, no group x time interaction). finding
  • Exploratory Bonferroni-corrected comparisons revealed a significant advantage of active over sham stimulation on d' during the first stimulation phase (day 1 tDCS) only, with a large effect size. finding
  • The results raise the possibility that tDCS effects on working memory may be greatest during early learning stages of the task. finding
  • Task performance improved across both groups with increasing exposure to the 3-back (significant main effect of time). finding
  • During day 1 stimulation, hit rate was significantly better and correct rejection rate was better at trend level in the active group, while reaction times did not differ. finding
  • A fronto-extracerebral montage (anode F3, cathode on contralateral cheek) avoids the interpretational confound of a scalp reference electrode, where effects could arise from excitation, inhibition, or their combination. method
  • A double-blind, between-subjects design using the device 'study mode' kept stimulation-administering researchers blinded to condition. method
  • The 3-back task code (Matlab/Cogent) and the participant performance dataset are made freely and permanently available (DOI 10.5281/zenodo.7148; data under CC0). resource
Experimental setups
Assay System Perturbation Readout Platform
3-back working memory task (behavioural, d' signal detection) at baseline Healthy human volunteers (N=21, 14 female, mean age 23.09 years, SD 3.95), right-handed, UCL subject pool none (5-minute pre-stimulation baseline, day 1) d' (Z(hit rate) - Z(false alarm rate)), hit rate, correct rejection rate, reaction time Task coded in Matlab release 2008b for Windows (Mathworks, Natick, MA, USA) with the Cogent Toolbox
3-back working memory task performed concurrently with tDCS (D1 tDCS, D2 tDCS) Healthy human volunteers randomized to active (N=10) or sham (N=11) stimulation 1 mA anodal tDCS over F3 (left DLPFC) with cathodal reference on contralateral cheek, 10 minutes with 15-s fade-in/fade-out, vs sham (100-200 µA pulses every 400-550 ms between 15-s ramps) d', hit rate, correct rejection rate, hit and correct-rejection reaction times during stimulation Neuroconn DC-Stimulator (Neuroconn, Germany); 7 cm × 5 cm rubber electrodes in saline-dampened sponges; F3 located with a 10-20 EEG cap
3-back working memory task immediately following stimulation (D1 post-tDCS, D2 post-tDCS) Healthy human volunteers, active vs sham groups no stimulation (10-minute post-tDCS run); monetary performance incentive (£10 bonus on day 2 for beating day 1 test phase) d', hit rate, correct rejection rate, reaction time post-stimulation Matlab (2008b) with Cogent Toolbox; Neuroconn DC-Stimulator used in preceding phase
Linear mixed model analysis of repeated 3-back performance (and follow-up general linear models, independent-samples t-tests, Bonferroni-corrected linear contrasts) Data from 21 healthy volunteers across five testing sessions over two days (24-48 h apart) Fixed effects of time (4 post-baseline sessions), group (active/sham) and time-by-group; participant as random effect; baseline performance as covariate; heterogeneous first-order autoregressive covariance structure d' as dependent variable; follow-up models on hit rate, correct rejection rate and reaction time SPSS version 21 (IBM Corp New York 2012)
Key results
  • No main effect of stimulation group on d' performance across the experiment F(1,16) = 2.228, P = 0.155
  • No group × time interaction on d' performance F(3,36) = 1.339, P = 0.277
  • Significant main effect of time on d', reflecting improvement across both groups with increasing task exposure F(3,36) = 7.669, P < 0.001
  • Active stimulation outperformed sham on d' at the day 1 tDCS time point (Bonferroni corrected, controlling for baseline) F(1,13.373) = 10.747, P = 0.006; Cohen's d = 1.427, r2 = 0.337
  • No group differences in d' at any other post-baseline time point all F < 1.2, P > 0.3
  • Hit rate during day 1 stimulation was significantly higher in the active group than sham F(1,18) = 4.454, P = 0.049, ηp2 = 0.198
  • Correct rejection rate during day 1 stimulation was better at trend level in the active group F(1,18) = 3.680, P = 0.071, ηp2 = 0.170
  • No reaction time differences at the day 1 stimulation time point for hits or correct rejections hits F(1,18) = 0.010, P = 0.923, ηp2 = 0.001; correct rejections F(1,18) = 0.202, P = 0.659, ηp2 = 0.011
Key statistics
  • pvalue F(1,13.373) = 10.747, P = 0.006 (Active vs sham d' difference at day 1 tDCS time point, Bonferroni corrected, baseline as covariate)
  • other Cohen's d = 1.427, r2 = 0.337 (Effect size for the day 1 during-stimulation group difference)
  • pvalue F(3,36) = 7.669, P < 0.001 (Main effect of time on d' in the linear mixed model)
  • pvalue F(1,16) = 2.228, P = 0.155 (Non-significant main effect of stimulation group on d')
  • pvalue t(19) = 1.044, P = 0.309 (No baseline d' difference between active and sham groups)
  • count active N = 10, sham N = 11; 21 total (14 females) (Group allocation of right-handed participants, mean age 23.09 years (SD 3.95))
  • mean Day 1 tDCS hit rate: anodal 0.5250 (SD 0.1336) vs sham 0.4030 (SD 0.1197) (Table 1 hit rate during day 1 stimulation)
  • other 80% power to detect a large effect size (d = 1.3) at P = 0.05 two-tailed (A priori power for between-group comparison given the sample size)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This double-blind, between-subjects study (active tDCS N=10, sham N=11) analyzed 3-back working memory performance (d') using a linear mixed model in SPSS (time, group, and time×group as fixed effects; participant as random effect; baseline performance as a covariate; heterogeneous first-order autoregressive covariance structure). The omnibus model showed a significant main effect of time but no significant main effect of group or group×time interaction. Exploratory, Bonferroni-corrected pairwise contrasts at each of four post-baseline time points, followed up with general linear models, identified a significant group difference only at the day 1 stimulation time point, which was also examined for hit rate, correct rejection rate, and reaction time.

Replicationbiological Sample sizeAuthors state that, based on their sample size, they had 80% power to detect a large effect size (d=1.3) at P=0.05, two-tailed, between the stimulation groups. Groupsactive (anodal) DLPFC tDCS vs sham stimulation, between-subjects (N=10 vs N=11) Pairingmixed Randomization/blindingstated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionBonferroni correction
Statistical tests used
Test Applied to n Assumptions
Independent samples t-test baseline d' performance between groups; age between groups N=21 (10 active, 11 sham) not stated
Chi-square test gender distribution between groups N=21 not stated
Linear mixed model (fixed effects of time, group, time×group; random effect of participant; baseline as covariate; heterogeneous AR1 covariance structure) d' across four post-baseline testing sessions (D1 tDCS, D1 post-tDCS, D2 tDCS, D2 post-tDCS) N=21, with one participant's incomplete/missing session data included in the model not stated
Bonferroni-corrected pairwise linear contrasts group differences in d' at each of the four post-baseline time points not stated per contrast not stated
General linear model with baseline as covariate follow-up group comparison at the D1 tDCS time point for d', hit rate, correct rejection rate, and reaction time denominator df of 18 reported for hit rate/correct rejection rate/RT models; 13.373 for the d' contrast not stated
Approaches that could also have been used
  • Exploratory pairwise group comparisons at each of the four post-baseline time points were corrected for multiplicity using the Bonferroni method.
    Could also: A false discovery rate procedure (e.g., Benjamini-Hochberg) or a set of pre-planned contrasts embedded directly in the mixed-model framework — FDR-based correction can offer greater power than Bonferroni when testing a modest number of related comparisons, which may be useful when comparisons are of similar exploratory interest.
  • The primary time-course analysis used a linear mixed model with a heterogeneous first-order autoregressive covariance structure.
    Could also: Comparing alternative covariance structures (e.g., compound symmetry, unstructured) via information criteria (AIC/BIC), or a repeated-measures ANOVA with a Greenhouse-Geisser correction — Comparing covariance structures can show how sensitive the time-by-group effect is to the specific structure chosen, and repeated-measures ANOVA is a widely used alternative framework for this type of design.
  • Spread was reported as SD in the summary table and as SEM in the main results figure.
    Could also: Reporting 95% confidence intervals for the group difference estimates — CIs convey the precision and plausible range of an estimated effect directly, which can be a useful complement to p-values, particularly with per-group sample sizes of 10 and 11.
  • Sample size was justified by a stated 80% power to detect a large effect (d=1.3) at P=0.05, described in relation to the sample obtained.
    Could also: An a priori power analysis specifying the minimum sample size needed to detect the smallest effect of theoretical interest before data collection — A prospective power calculation ties the planned sample size directly to a pre-specified effect of interest, which can be a useful complement to a post hoc power statement.
  • Baseline and follow-up group comparisons relied on parametric tests (t-tests, GLM, LMM) assuming approximately normal residuals.
    Could also: Non-parametric alternatives such as the Mann-Whitney U test for baseline group comparisons, or permutation-based tests for the mixed-model contrasts — Non-parametric or permutation approaches do not depend on distributional assumptions, which can be a helpful check with modest per-group sample sizes (N=10-11).
  • d', hit rate, correct rejection rate, and reaction time were each analyzed with separate GLMs at the D1 tDCS time point.
    Could also: A multivariate analysis of variance (MANOVA) or a joint model across the correlated outcome measures — Modeling correlated outcomes jointly can account for their shared variance and control the family-wise error rate across the set of related dependent measures.
Software: SPSS 21 (IBM Corp, New York, 2012)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

table1_descriptives
Reported
Table 1: per-group (Active/Sham) means and SDs of HR, CRR, Hit RT, CR RT across 5 sessions (Baseline, D1 tDCS, D1 post-tDCS, D2 tDCS, D2 post-tDCS)
Reproduced
partial
baseline_ttest_dprime
Reported
No significant difference in d' between groups at baseline
Reproduced
within tolerance
day1_tdcs_exploratory_contrast
Reported
Exploratory Bonferroni-corrected pairwise comparison: active vs sham d' at Day 1 tDCS timepoint, controlling for baseline
Reproduced
within tolerance
main_effect_time
Reported
Significant main effect of time on d' across the 4 post-baseline sessions
Reproduced
within tolerance
main_effect_group
Reported
No significant main effect of stimulation group on d'
Reproduced
partial
group_time_interaction
Reported
No significant group x time interaction on d'
Reproduced
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 61/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The deposited data are genuinely raw and near-complete (124/126 files, the 2 gaps matching the stated day-2 dropout), but they omit the Active/Sham group labels, so the grouping had to be recovered by exhaustive search over 705,432 splits — a best-fit estimate, not ground truth, that propagates into every group statistic. Under that reconstruction the paper holds up well: 38/40 Table 1 cells match within ~0.01–0.03, baseline t(19)=1.044→1.180, the key Day-1 ANCOVA F=10.747, P=0.006→F=11.515, P=0.0032, and the time effect F=7.669→8.971 all confirm. Two deviations are on the authors' side: the Sham CRR SDs of 0.2757/0.2821 are unreachable from per-subject CRRs spanning only 0.817–0.983 (transcription error, not fabrication), and the reported non-significant group×time interaction F(3,36)=1.339, P=0.277 became F(3,54)=3.004, P=0.038 on reproduction. That flip is confounded with our own forced simplification (no R/scipy/statsmodels on the compute nodes → OLS instead of heterogeneous-AR1 MIXED, complete-case N=20), so it is a limitation of the reproduction as much as a challenge to the paper — hence yellow overall rather than red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.