EEG × Representation learning

Neural signal decoding via transformer pre-training

An EEG representation-learning prototype, with emphasis on disjoint data, explicit synthetic inputs and the difference between an implemented objective and a demonstrated transfer-learning result.

On this page

A hypothesis to test

Masked reconstruction may provide a useful representation before supervised motor-imagery classification. NeuroLLM implements components for testing that hypothesis. It has not established a real-data improvement over CSP+SVM, EEGNet or a transformer trained from scratch.

What is wired into the model

Channel/time patches receive learned channel and temporal embeddings before a PyTorch TransformerEncoder. A reconstruction head supports masked pre-training; a classification head uses the CLS representation. The reviewed default 22-channel classifier has 3,265,284 parameters; this count excludes the masked-reconstruction head.

Inspect the model. The frequency-band attention code is a separate experiment, not the attention module used by this encoder.

Reconstruction is an objective, not a result

Mask part of the input representation and predict the withheld signal from context. Check masking, target construction and loss accounting before running a large experiment. A lower reconstruction loss alone does not establish better motor-imagery decoding.

Synthetic arrays can exercise shapes and optimization plumbing, but they cannot substitute for real recordings. Dataset identity, preprocessing, channel mappings and training provenance must accompany any transfer-learning claim.

Keep train, validation and test separate

BCI session labels require an explicit mapping: the training file uses T and the evaluation file uses E. Taking the first letter of the words “train” and “test” maps both to T and leaks the training session into evaluation. Missing real files must raise an error rather than silently generate data.

Fixing file selection is only one boundary. The reviewed fine-tuning loop selects its checkpoint using the test loader; that protocol must change before reporting a held-out test score. Select checkpoints using a validation subset of training data, freeze that choice, and evaluate the test session once.

Dataset source · Training and selection source

Frequency-band attention remains separate

The repository includes an attention experiment motivated by frequency structure. It must be assessed on its own numerical behavior and then deliberately integrated before an ablation can attribute any model change to it. Kernel source, model integration and demonstrated benefit are three different milestones.

No comparable accuracy table yet

A table of approximate literature ranges alongside synthetic outputs would combine different datasets and protocols. This note therefore does not present it as a benchmark. A valid next comparison should use the same subjects and sessions, train-only preprocessing, reproducible seeds and identical evaluation rules for every baseline.

Report per-subject accuracy and uncertainty, checkpoint provenance and a same-architecture pre-training ablation. A synthetic four-class task’s chance expectation is not a measured real-EEG result.

Visualizations need a defined interpretation

Attention visualizations can help inspect indexing and model behavior, but attractive channel maps alone do not validate neurophysiological interpretation. Record the input, layer, head, aggregation method and data split. Avoid presenting a synthetic visualization as evidence from real motor-imagery recordings.

What is needed next

The immediate technical priority is evaluation integrity: explicit sources, distinct partitions and model selection that never sees held-out test labels. Once that protocol is fixed and reviewed, a small real-data baseline is more informative than an unsupported claim about large-scale pre-training.

Project scope, evidence and source links →

Keep exploring

Working on a similar problem? Let’s discuss it.

Get new articles via RSS