Interpretable Deep Neural Networks for Single-Trial EEG Classification
Irene Sturm, Sebastian Bach, Wojciech Samek, Klaus-Robert Müller
I Introduction
Deep Neural Networks (DNNs) are powerful methods for solving complex classification tasks in fields such as computer vision , natural language processing , video analysis and physics . Although researchers have recently started introducing this promising technology into the domain of cognitive neuroscience and Brain-Computer Interfacing (BCI) , most of the current techniques in these fields are still based on linear methods . A limiting factor for the applicability of DNN in these fields is the notion of a DNN as a black box. In the domain of cognitive neuroscience this is a particular drawback because obtaining neurophysiological insights is of utmost importance beyond the classification performance of a system.
Recently, the interpretability aspect of deep neural networks has been addressed by the Layer-wise Relevance Propagation (LRP) method. LRP explains individual classification decisions of a DNN by decomposing its output in terms of input variables. It is a principled method which has close relation to Taylor decomposition and is applicable to arbitrary DNN architectures. From a practitioners perspective LRP adds a new dimension to the application of DNNs (e.g., in computer vision ) by making the prediction transparent. Within the scope of cognitive neuroscience this means that DNN with LRP, may provide not only a highly effective (non-linear) classification technique that is suitable for complex high-dimensional data, but also yield detailed single-trial accounts of the distribution of decision-relevant information, a feature that is lacking in commonly applied DNN techniques and also in other state-of-the art methods (such as those discussed below).
Here we propose using DNN with LRP for the first time for EEG analysis. For that we train a DNN to solve a classification task related to motor-imaginery BCI. On two example data sets we compare the classification performance of DNN to that of CSP-LDA, a standard technique . We then apply LRP to produce heatmaps that indicate the relevance of each data point of a spatio-temporal EEG epoch for the classifier’s decision in single trial. We present several examples of such heatmaps and demonstrate their neurophysiological plausibility. Critically, we point out that the spatio-temporal heatmaps represent a new quality of explanatory resolution that allows to explain why the classifier reaches a certain decision in a single instance. Note that such information can not be derived from CSP-LDA. Finally, we provide a range of future applications of this technique in neuroscience. We discuss why equipping the extremely powerful non-linear technology of DNN with the diagnostic power of LRP may contribute to extending the scope of DNN techniques.
II A Deep Neural Network for EEG Classification
The network applied here consists of two linear sum-pooling layers with bias-inputs, followed by an activation or normalization step each. The first linear layer accepts an input of the dimensionality 301 time points 118 channels EEG features (301 time point 58 channels for subjects od-obx) as a 33518 (od-obx: 17458) dimensional input vector and produces a 500-dimensional tanh-activated output vector. The next layer reduces the 500-dimensional space to a 2-dimensional output space followed by a softmax layer for activation in order to produce output probabilities for each class. The network was trained using a standard error back-propagation algorithm using batches of 5 randomly drawn training samples. The above prediction accuracy was achieved after terminating the training procedure after 3000 iterations .
II-B Interpretability
The DNN assigns a classification score to every input data sample at prediction time. Layer-wise Relevance Propagation decomposes the classifier output in terms of relevances attributing to each input component its share with which it contributes to the classification decision
These relevance values are backpropaged from the network output to the input layer using a local redistribution rule
where indexes a neuron at a particular layer , where runs over all lower-layer neurons connected to neuron , and where are parameters specific to pairs of adjacent neurons and learned from the data. This redistribution rule has been showed to fulfill the layer-wise conservation property and to be closely related to a deep variant of Taylor decomposition .
III Evaluation
The application of DNN with LRP on EEG data was demonstrated on dataset IVa from BCI competition III (cued motor imagery data with classes right hand vs. foot from 5 subjects ) and on a subset of 5 subjects from where subjects had to perform left and right hand motor imaginery while dealing with different types of distractions. Here, we only analyzed data obtained in the condition ‘no distraction’, a standard motor imaginery BCI setting. As in the competition, we did not use test data for training for dataset IVa (subjects aa, al, av, aw, ay). For the other data set (subjects od, njy, njk, nko, obx) a leave-on-out cross-validation was performed. For both data sets the potential of DNN for subject-to-subject transfer was evaluated: for each subject a DNN was trained on all available data of the other four subjects and evaluated on its own test data. This was was done in a sequential fashion, so that the network was once initialized and then trained on the data of each of the four subjects successively. The entire process of training and testing was repeated five times for different orders of the four subjects and the classification performance on the test data was averaged.
All data sets were downsampled to 100Hz and bandpass filtered in the range of 9-13 Hz. The CSP algorithm was performed on a ms epoch after the cue and 3 pairs of spatial filters were selected. On the extracted features a regularized LDA classifier with analytically determined shrinkage parameter was trained. For training and evaluating the DNN the envelope of each epoch ( ms after cue) was calculated and an epochwise baseline of ms before the cue was subtracted. Each epoch’s spatio-temporal features (301 time points 118 channels for aa-ay, 301 time point 58 channels for subject od-obx) were vectorized into one vector with 33518 (17458) dimensions. Relevance maps were calculated for each trial from the two-valued DNN output according to Equation 2.
III-B Results
Classification results for the different methods are summarized in Table I. Overall, classification performance of DNN is lower than that of CSP-LDA. Subjects ay and njy, the subjects with the lowest performance, represent an exception: here DNN effects an increase in classification accuracy. The performance of inter-subject DNN is inferior to that of single-subject DNN in 6/10 subjects. In the remaining four subjects inter-subject DNN effects a substantial increase in classification accuracy.
Fig. 1 (a) gives an example of relevance maps obtained with LRP for two single trials of subject od. The matrices depict the relevance of each EEG channel at each time point of the epoch. Note that these relevance maps differ from CSP patterns where the absolute magnitude of a weight determines its relevance and its sign the polarity. In LRP-derived heatmaps positive and negative values refer to the relevance and non-relevance with respect to the specific decision of the DNN. For instance, in a trial assigned to class ‘right hand’ with high confidence positive values may be understood as speaking for class ‘right hand’ membership and negative values as speaking against class ‘right hand’ membership. For a given time point the relevance information can be plotted as a scalp topography. The example scalp maps at the bottom show typical lateralized motor activation patterns that can be related to a single time point in a single trial. The average of the spatio-temporal relevance matrix across the entire epoch (top) reveals similar scalp patterns. The average of all time-averaged relevance maps of one class (Fig. 1 (b)) is highly similar in topographical distribution to the patterns of the first pair of CSP filters. Fig. 1 (c) shows examples of time-averaged relevance maps for a selection of correctly/incorrectly classified trials. In those trials that were classified correctly and with high confidence (classifier output 0 or 1), relevant information is confined to small regions with neurophysiologically highly plausible distribution. In incorrectly or with less confidence classified trials influences outside the sensorimotor areas seem to have influenced the network’s decision. These are located in occipital and frontal regions and may indicate the influence of visual activity and of eye movements.
IV Discussion
We have provided the first application of DNN with LRP on EEG data. In terms of classification performance, our relatively simple DNN network does not outperform the benchmark methodology of CSP-LDA. However, we provide some examples that training a network successively on several other subjects is advantageous. For instance, this substantially increased classification accuracy in a subject with particularly low accuracy. This is a first hint that DNN technology may be beneficial for subject-to-subject transfer of learned neural representations, and, ultimately, may advance subject-independent zero training strategies in BCI .
The most important and novel contribution in this work is the application of LRP. We have demonstrated that LRP produces neurophysiologically highly plausible explanations of how a DNN reaches a decision. More specifically, LRP produced textbook-like motor imaginery patterns in single instants of single trials. These represent accounts of neural activity at an unprecedented level of specificity and detail. In contrast, CSP-LDA (and also other methods) only allow to examine discriminating information at the level of the whole ensemble of samples of one class. In a direct application in the BCI context LRP helped to diagnose influences that led to low-confidence or erroneous decisions of the network.
Outside BCI DNN with LRP may add a new dimension of explanation in any setting where detailed single-trial information is valued. In clinical applications it may represent a sensitive tool for neurophysiological interpretation of anomalies or differences between populations. Here, the opportunity to integrate prior knowledge about clinical populations through inter-subject DNN analyses may be a further advantage. In contexts where the trial-to-trial variability of EEG is not viewed as a notorious obstacle for analysis, but as a source of information, LRP can contribute high-resolving spatio-temporal representations of underlying neurophysiological phenomena. In particular, this might be interesting for linking brain indices to single instances of behavioral measures , for understanding subtle aspects of complex perceptual processes, such as perception of video or audio quality , and of dynamic cognitive processes, such as decision making . Finally, a trained network produces relevance maps for any (even artificially generated) DNN decision. This means that LRP can derive a representation of what a network has learned, e.g., by performing LRP on a ‘ideal’ specimen of a given class or even by systematically exploring the space of possible decisions. This might be an interesting alternative to network visualization techniques .
V Conclusion
In summary, we have provided a showcase of how LRP can add an explanatory layer to the highly effective technique of DNN in the EEG/BCI domain. Our results show that LRP provides highly detailed accounts of relevant information in high-dimensional EEG data that may be useful in analysis scenarios where single trials need to be considered individually.
Acknowledgement This work was supported by the Brain Korea 21 Plus Program and by the Deutsche Forschungsgemeinschaft (DFG). This publication only reflects the authors views. Funding agencies are not liable for any use that may be made of the information contained herein. Correspondence to WS and KRM.