EEG × Representation learning

NeuroLLM

An experimental EEG transformer pipeline for masked reconstruction and motor-imagery classification. The present work concerns data boundaries, model components and evaluation design; real-data gains are not yet established.

PyTorch · NumPy/SciPy · EEG · Transformers

Problem and scope

Can reconstruction of unlabeled EEG provide a useful starting point for supervised motor-imagery decoding? NeuroLLM implements patch embeddings, masked pre-training and a classification head to investigate that question. Its current artifacts do not establish improved decoding accuracy over classical or neural baselines.

Data and evaluation boundaries

InputIntended roleEvidence status
TUH EEGUnlabeled pre-trainingNo verified real-data pre-training result presented
BCI Competition IV 2aSeparate training and evaluation sessionsReal-data access and a clean held-out protocol required
Synthetic arraysExercise loaders, shapes and training codeSoftware smoke tests, not EEG decoding evidence

Synthetic data must be explicitly requested and distinguishable from real recordings. Missing files must fail visibly. Training and evaluation sessions must map to different input files, and model selection must use validation data rather than the held-out test labels.

Implemented model and separate experiments

  • Representation: channel/time patches with learned channel and temporal encodings.
  • Pre-training objective: masked reconstruction from the surrounding signal; its downstream benefit requires an ablation against the same model without pre-training.
  • Encoder: the reviewed default uses PyTorch TransformerEncoder, not the separate frequency-band attention kernel.
  • Model size: the default 22-channel classifier contains 3,265,284 parameters; alternative configurations need their own count.
  • Experimental kernel: frequency-band attention has separate source. Its integration, numerical behavior and speedup require independent validation.

Model source · Dataset loaders · Attention experiments

What can be concluded

No real-EEG performance result is established here. Approximate literature scores and synthetic smoke-test outputs are not comparable benchmarks for this implementation. The reviewed training script also selects a checkpoint using the evaluation loader, so its reported test score is not an untouched held-out estimate.

A defensible comparison requires disjoint sessions, train-only preprocessing and checkpoint selection, identical evaluation subjects, documented seeds, per-subject results and a frozen test protocol. Compare CSP+SVM, EEGNet and the transformer with and without pre-training under that same protocol.

Implemented model path

flowchart TB EEG["EEG channel × time array"] --> Patch["Patch + channel/time embeddings"] Patch --> Encoder["PyTorch TransformerEncoder"] Encoder --> Reconstruction["Masked reconstruction head"] Encoder --> Classification["CLS token + classification head"] Separate["Frequency-band kernel: separate experiment"]

Validation update · 26 September 2026

Explicit data boundaries. The loader repair maps train/test to distinct T/E files, rejects missing real data and requires an explicit synthetic mode with distinct deterministic partitions. On Python 3.11.14 and PyTorch 2.8.0, 54 focused data/preprocessing tests passed without training. Test-based checkpoint selection remains unresolved.

Inspect the tested revision and validation details. This repair is available on a review branch; the repository’s main branch has not yet been updated.

Inspect the protocol

Fine-tuning and checkpoint selection · Evaluation code · Tests. Loader regression tests can establish separation and explicit input handling; they do not establish transfer learning performance.

Read the technical note →

Building a similar system?

Get in touch about your project.