Problem and scope
Can reconstruction of unlabeled EEG provide a useful starting point for supervised motor-imagery decoding? NeuroLLM implements patch embeddings, masked pre-training and a classification head to investigate that question. Its current artifacts do not establish improved decoding accuracy over classical or neural baselines.
Data and evaluation boundaries
| Input | Intended role | Evidence status |
|---|---|---|
| TUH EEG | Unlabeled pre-training | No verified real-data pre-training result presented |
| BCI Competition IV 2a | Separate training and evaluation sessions | Real-data access and a clean held-out protocol required |
| Synthetic arrays | Exercise loaders, shapes and training code | Software smoke tests, not EEG decoding evidence |
Synthetic data must be explicitly requested and distinguishable from real recordings. Missing files must fail visibly. Training and evaluation sessions must map to different input files, and model selection must use validation data rather than the held-out test labels.
Implemented model and separate experiments
- Representation: channel/time patches with learned channel and temporal encodings.
- Pre-training objective: masked reconstruction from the surrounding signal; its downstream benefit requires an ablation against the same model without pre-training.
- Encoder: the reviewed default uses PyTorch TransformerEncoder, not the separate frequency-band attention kernel.
- Model size: the default 22-channel classifier contains 3,265,284 parameters; alternative configurations need their own count.
- Experimental kernel: frequency-band attention has separate source. Its integration, numerical behavior and speedup require independent validation.
What can be concluded
No real-EEG performance result is established here. Approximate literature scores and synthetic smoke-test outputs are not comparable benchmarks for this implementation. The reviewed training script also selects a checkpoint using the evaluation loader, so its reported test score is not an untouched held-out estimate.
A defensible comparison requires disjoint sessions, train-only preprocessing and checkpoint selection, identical evaluation subjects, documented seeds, per-subject results and a frozen test protocol. Compare CSP+SVM, EEGNet and the transformer with and without pre-training under that same protocol.
Implemented model path
Validation update · 26 September 2026
Explicit data boundaries. The loader repair maps train/test to distinct T/E files, rejects missing real data and requires an explicit synthetic mode with distinct deterministic partitions. On Python 3.11.14 and PyTorch 2.8.0, 54 focused data/preprocessing tests passed without training. Test-based checkpoint selection remains unresolved.
Inspect the tested revision and validation details. This repair is available on a review branch; the repository’s main branch has not yet been updated.
Inspect the protocol
Fine-tuning and checkpoint selection · Evaluation code · Tests. Loader regression tests can establish separation and explicit input handling; they do not establish transfer learning performance.