Vowel Speech Recognition From Rat Electroencephalography Using Long Short-term Memory Neural Network Part 3

Dec 28, 2023

Machine learning classifiers

The performance of BiLSTM was compared with conventional machine learning classifiers: SVM with linear kernel (SVM_lin), SVM with radial basis function kernel (SVM_rbf), random forests (RF), NB, and KNN. 

Random forest is a machine learning algorithm that is currently widely used in various data analysis and prediction fields. Compared with other machine learning algorithms, it has better robustness and accuracy, while also effectively reducing overfitting. In recent years, the application scope of random forests has been expanding, and it can even be used to predict certain human memory abilities.

In the field of cognitive psychology, memory is a very important research direction. Scientists have been looking for a simple and effective way to assess human memory levels. In recent years, the emergence of random forests has brought new ideas and methods to this field.

Random Forest can train a model to predict a variable, which can be anything you want to predict, including memory ability. Scientists can feed relevant factors into a random forest model to predict how well a person will score on a memory test. These factors can be factors such as age, education level, gender, weight, etc., or biological indicators such as brain structure. According to research, there is a certain relationship between these factors and human memory ability.

By collecting and analyzing large amounts of test data from subjects, scientists can build a model that predicts memory ability in a random forest model. Predicting results can provide valuable information about a subject's future performance on a certain memory test.

In summary, the random forest algorithm provides scientists with a new way to evaluate human memory levels. In the future, its application may play an important role in cognitive psychology, neuroscience, and other fields. We have reason to believe that the combination of random forests and other techniques can provide a broader perspective and deeper understanding of human brain function research. It can be seen that we need to improve memory, and Cistanche deserticola can significantly improve memory, because Cistanche deserticola can also regulate the balance of neurotransmitters, such as increasing the levels of acetylcholine and growth factors. These substances are very important for memory and learning. In addition, Meat can also improve blood flow and promote oxygen delivery, which can ensure that the brain receives sufficient nutrients and energy, thereby improving brain vitality and endurance.

memory enhancement

Click know ways to improve brain function

SVM [74] aims to determine the optimally separated hyperplane by maximizing the margin, which is the distance between the support vectors. By using the kernel trick, SVM is capable of mapping feature space from low to high dimensions; therefore, it can efficiently perform linear classification and non-linear classification. 

RF [75] operates by constructing multiple decision trees during the training phase and generating the final class that combines the results of each decision tree. NB [76, 77] is a probabilistic classifier based on Bayes' theorem and conditional probability which usually assumes that all features are independent of each other. 

KNN [78] is a non-parametric approach that classifies the input based on the majority class of its k-nearest neighbors in the feature space. Usually, the k value is selected as an odd number to avoid tied classes. 

To train and evaluate the above machine learning models, the same 10-CV was used as in BiLSTM. All machine learning models were implemented using the Scikit-Learn library [73] in Python.

Statistical analyses

All statistical analyses were performed using SPSS software (SPSS version 20.0, SPSS Inc., Armonk, NY, USA) and MATLAB software version 2017b (Mathworks, Inc., MA, USA). 

The data was analyzed with parametric statistics since all the data in the study showed a normal distribution in the Shapiro–Wilk test (p > 0.05). ANOVA was used to analyze the statistical significance of the TFRs according to the different vowel stimuli.
In addition, a repeated measures ANOVA was conducted to compare the performance of each classifier. Subsequently, pairwise comparisons using paired t-tests were performed between the BiLSTM network and other classical machine-learning classifiers, and a Bonferroni correction was performed to adjust for the type I error rate inflation. 

The statistical significance of the p-value was set at 0.01 when comparing the TFR of EEG responses, while the significance level of the p-value was set at 0.05 when comparing the performance between the BiLSTM network and other machine learning classifiers.

Results

Auditory evoked potentials in response to vowel sounds

A total of 19 Sprague-Dawley rats underwent epidural electrode implantation surgery, and all rats survived the surgical procedure. As a result, EEG responses to five English vowel sounds were recorded from 19 isoflurane-anesthetized rats. To extract the mean AEP waveforms, all the neural responses were averaged over the subjects for each stimulus. Fig 4 presents the averaged AEP waveforms for each vowel sound from bilateral AAF.

As expected, each categorical vowel sound evoked distinct neural activities in the bilateral AAF with varying peak amplitudes and latencies. The peak amplitude of AEPs, defined as the highest recorded voltage after the vowel stimuli, was smallest for /i/ (61.74 ㎶ in left AAF and 61.27 ㎶ in right AAF), while AEPs in response to /a/ showed the largest peak amplitudes (92.12 ㎶ in left AAF and 90.18 ㎶ in right AAF). 

The peak latency, defined as the duration from stimulus onset to the peak amplitude was approximately 0.39 s to 0.5 s, shortest in /i/ (0.39 s in left and right AAFs), and longest in the /o/ sound (0.51 s in left and right AAFs). As shown in Fig 4, similar AEP waveforms were observed from the left and right AAFs.

improve your memory

Time-frequency analysis of the EEG signals

Time-frequency analysis is a powerful method for analyzing nonstationary EEG signals over a time-frequency plane and is used to provide qualitative information for the classification of EEG [79, 80]. Therefore, the TFR of the grand-averaged EEG was calculated for each sound to identify vowel recognition-related changes in the magnitude and phase of EEG oscillations at specific frequencies (Fig 5A). 

From the TFR analysis, high power activation was observed around the delta (1–4 Hz), theta (4–8 Hz), and alpha (8–12 Hz) band at 0.3–0.6 s from the stimulus onset, regardless of the speech sound stimulation.

boost memory

In addition, an ANOVA test with a Bonferroni correction was conducted to analyze the statistically significant TFR components according to each vowel stimulus. 

Subsequently, the power of statistically significant areas (p < 0.01) was represented by the F-value (Fig 5B). In the analysis, most of the EEG frequency bands from 0.2–0.8 s were significantly different according to the vowel stimuli. 

In addition, part of the TFR from 0.8–1 s was also statistically different for each stimulus. Considering the AEP waveforms and the results of the ANOVA tests, it was inferred that the AEPs from 0.2–0.8 s after the vowel stimulus were the most informative neural responses and were related to vowel sound recognition.

Model training and evaluation of the BiLSTM networks

Based on the results of Fig 5B, EEG data that were band-pass filtered between 1–60 Hz with a time window of 0.2–0.8 s were selected. Then, the z-scores of the selected EEG data were used as the input to the BiLSTM network. 

All EEG data were divided into 10 folds within each subject to evaluate the BiLSTM networks. Therefore, the test performance was obtained per fold using the trained model with the remaining folds in a 10-CV scheme. 

The performance of the network was evaluated using metrics of accuracy, f1-score, and Cohen's kappa statistic κ (Fig 6 and Table 1). The average five-class EEG discrimination accuracy of the BiLSTM network was 75.18 ± 7.06% and the f1-score was 0.74 ± 0.08. Cohen's κ was 0.68 ± 0.09, which was interpreted as a moderate agreement [81].

To analyze the performance of the BiLSTM network in more detail, the confusion matrix in Fig 7 was plotted. This indicated that many of the errors were due to the misclassification of the EEG responses to /u/ as /a/ and /e/ as /o/. However, the BiLSTM network classified most of the EEG responses with more than 50% accuracy, a high accuracy in the five-class EEG classification.

10 ways to improve memory

Comparison of the BiLSTM network with other machine learning methods

To validate the effectiveness of the BiLSTM networks in classifying EEG for vowel sound recognition, the results were compared with those of other conventional machine learning methods. Fig 6 and Table 1 show the performance of the machine learning classifiers. 

The RF demonstrated the highest classification accuracy among the conventional machine learning algorithms (accuracy: 63.21 ± 7.41%, f1-score: 0.62 ± 0.09, and Cohen's: 0.52 ± 0.1). In the statistical analysis, the classification performance of RF was not significantly higher than that of SVM_lin and SVM_rbf, while it showed higher performance when compared with those of NB and KNN. 

However, when the performance of conventional machine learning algorithms, including RF, was compared with BiLSTM, it was obvious that the BiLSTM network was superior for all the metrics used in the study (p < 0.01).

In the confusion matrix, conventional machine-learning algorithms cannot discriminate certain EEG responses well. In particular, all the conventional machine learning algorithms had difficulty distinguishing the sound /u/. It was noted that the algorithms showed a tendency to misclassify sound /u/ as /a/ on average 30% of the time (25.96% in NB to 36.97% in KNN), resulting in a decrease in the overall classification performance (Fig 7).

improving brain function

Discussion

In this study, rat epidural EEG responses to five categorical vowel sounds (/a/, /e/, /i/, /o/, and /u/) were discriminated using the BiLSTM network. Five-class classifications of epidural EEG signals were performed on a single-trial basis, which is known to be challenging. To maximize learning performance, this study tried to determine specific EEG components that might be related to the recognition of speech sounds in the rat brain and utilized these EEG components as input features. As a result, a relatively high performance in classifying AEPs into five different vowel sounds was achieved using BiLSTM. A comparison of the classification performance of the BiLSTM network with other machine learning algorithms showed that the BiLSTM network outperformed other classical classifiers. These results indicate that the BiLSTM network trained with speech recognition-related EEG components reliably classifies AEPs to each categorical vowel sound with a high degree of accuracy. To our knowledge, LSTM networks have not been applied to the classification of EEG responses to auditory stimuli, and this is the first study to use a deep learning algorithm to analyze EEG signals from rat AAF.

short term memory how to improve

Currently, only a few studies have used LSTM architecture to achieve state-of-the-art results in EEG-based classification. The LSTM architecture is suitable for EEG-based classification because its chain-like structure can capture the temporal sequence of EEG data [82]. In the beginning, research focused on improving the classification results through various LSTM architectures; however, the input features were still extracted manually, as in conventional machine learning methods [83, 84]. 

Tsiouris et al. evaluated the performance of diverse combinations of LSTM network elements to find the most efficient LSTM architectures for detecting epileptic seizures, thus obtaining near-perfect results in seizure prediction (100% sensitivity and 99.86% specificity) [83]. Because LSTM is a powerful structure for processing sequential data, some studies use raw EEG data as input features with minimal preprocessing. As the LSTM network directly learns features from raw EEG data, the performance in emotion recognition studies improved by at least 12% [85], and the results of motor imagery classification studies also improved [86], when compared with other traditional feature extraction techniques. 

Moreover, the BiLSTM architecture was utilized for EEG-based classification because it can access information from both past and future states. Therefore, in detecting various brain states reflected in EEG data, such as seizure, sleep, etc. [63–67], the BiLSTM network generally outperformed the LSTM network that only captures past information from the sequence in the forward direction. For this reason, high performance has been reported in recent EEG-based classification using BiLSTM networks. Sharma et al. achieved 82.01% classification accuracy for four types of emotions based on the BiLSTM algorithm and higher-order statistics [87]. In addition, the BiLSTM networks successfully classified epilepsy types and sleep stages [88, 89].

Similar to previous studies, this study achieved comparatively good results using BiLSTM networks. The proposed algorithm successfully discriminated the EEG responses to five vowel sounds with high values of accuracy, f1-score, and Cohen's κ of 75.18%, 74.43%, and 0.68, respectively. The value of Cohen's κ for five-class classification is higher than that seen in most current studies [90]. As illustrated in Fig 6, the BiLSTM method produced the highest value for all metrics compared with the other machine learning methods. In addition, to determine the statistical difference in classification performance, repeated-measured ANOVA results were analyzed between BiLSTM and other classical machine learning methods using all the metric values. Through statistical analysis, it was determined that the classification performance of the BiLSTM network was significantly higher than that of other classical machine learning methods (p < 0.01). This result was also consistent with the confusion matrix. As shown in Fig 7, the BiLSTM network predicted the true labels of the five vowel sounds well, whereas classical machine learning methods did not. 

The prediction acquired through the conventional machine learning classifier was especially poor at classifying the /u/ sound; the /u/ sound was mainly misinterpreted as /a/. Even RF, which showed the best performance among the five conventional machine learning classifiers, had a classification rate of 34.48% for the /u/ sound, with a 33.89% misclassification rate of the /u/ sound as an /a/ sound. As can be seen in Fig 4, the /a/ and /u/ sounds had a similar peak latency, which is one of the main characteristics of AEP waveforms (peak latency of sound /a/: 0.448, peak latency of sound /u/: 0.444). When classification was performed based on minimally pre-processed single-trial EEG signals, it seems that such similarities could not be distinguished by conventional machine learning algorithms, whereas the BiLSTM network could distinguish them. 

Given that the BiLSTM network can simultaneously access all past and future contexts, rich information can be learned through this network. In addition, even though the features reflecting the characteristics of EEG responses to each vowel sound were extracted directly from the forward and backward directions of the LSTM layer, the classification performance was improved. In this study, we can derive good classification results using a simple BiLSTM architecture without an additional handcrafted feature extraction process.

Classifying the ERP responses to speech stimuli in a single trial is very challenging owing to the characteristics of the low SNR of EEG. Although one of the key advantages of the deep learning method is its ability to learn high-level features without hard-core feature extraction, we attempted to select the most relevant EEG signals related to speech recognition to achieve better performance. In this study, distinct AEP waveforms corresponding to each speech sound stimulus were observed with the high-power activation of the low-frequency band, including the delta, theta, and alpha bands, in the TFR analyses. Neural oscillations in the alpha band have been widely recognized to play an important role in auditory processing. Mazaheri et al. reported that the attenuation of alpha activity is closely related to the discrimination of auditory targets [91].

Staruß et al. proved that cortical alpha oscillations are a pivotal mechanism for selectively inhibiting the processing of noise to improve the auditory selective attention toward target signals [92]. Previously, we also found that alpha power was highly activated in bilateral temporal areas after specific sound stimuli that were statistically different in terms of the type of sound [48]. In addition, the delta and theta bands are known to be associated with shaping the segmentation and perceptual influence of acoustic information [93]. 

Although this study is based on animal experimental data, similar speech-related components, as compared to the previous studies on human subjects, were observed in the TFR analyses. Moreover, in the statistical analysis, all the EEG bands were found to be significant within 1 s after the stimuli and represented the EEG components related to sound perception. These results were somewhat different from those of previous studies, suggesting that only specific EEG bands, such as the alpha band, were related to sound perception. It is expected that even subtle changes across all the EEG band activities are recorded through the epidural EEG recording because it provides a higher SNR by reducing volume conduction and eliminating the artifacts that are inherent to extracranial EEG recordings.

In this study, the speech sound recognition-related EEG components in rats were determined and the AEP components were successfully classified using the BiLSTM network. However, this study had some limitations. First, the number of subjects included was too small, especially for deep learning. Moreover, this study did not evaluate each classifier's performance with external validation, but instead used 10-CV to overcome the limited sample sizes. Besides, we cannot rule out the possibility that the rat's auditory system responds continuously to sound since only a single utterance of each vowel sound was used in this study. In addition, the acquired EEG responses were affected by the anesthetic effects. Although minimal anesthetic dose was used, frequency slowing with an increase in delta power is a typical finding of EEG changes after isoflurane inhalation [94]. Therefore, the vowel recognition EEG components suggested in this study may be different from EEG signals acquired from rats that are awake. However, we believe that the quality of the EEG signal is good enough since we recorded EEG through epidural electrode implantation, and it was not contaminated by motion artifacts.

Conclusions

In conclusion, this study extracted meaningful neural components related to categorical speech perception. Furthermore, based on the characteristics of the LSTM networks, it was proved that the BiLSTM network was suitable for classifying EEG responses with minimally pre-processed AEPs. Since this study is pioneer research with animal data, it may not be directly transferable to other practical applications such as brain-computer interfaces or alternative communication aids for humans. 

supplements to boost memory

Therefore, future studies with human EEG data are required to verify the effectiveness of the BiLSTM network in classifying auditory EEG-based speech recognition. Additionally, it needs to be re-evaluated for optimal parameter tuning and feature extraction. It is expected that this study will provide a novel approach for analyzing EEG signals as well as valuable information regarding the mechanisms of speech perception and recognition in the brain.

ways to improve memory


References

1. Wernicke C. The symptom complex of aphasia. In: Cohen RS, Wartofsky MW, editors. Proceedings of the Boston Colloquium for the Philosophy of Science 1966/1968. Dordrecht: Springer Netherlands; 1969. pp. 34–97.

2. Shi Z, Yan S, Ding Y, Zhou C, Qian S, Wang Z, et al. The anterior auditory field is needed for sound categorization in the fear conditioning task of the adult rats. Front Neurosci. 2019; 13: 1374.

3. Liberman AM, Harris KS, Hoffman HS, Griffith BC. The discrimination of speech sounds within and across phoneme boundaries. J Exp Psychol. 1957; 54: 358–368.

4. Johnson K. Acoustic and auditory phonetics. Chichester: Wiley-Blackwell; 2012.

5. Green PA, Brandley NC, Nowicki S. Categorical perception in animal communication and decision-making. Behav Ecol. 2020; 31: 859–867.

6. Craik A, He Y, Contreras-Vidal JL. Deep learning for electroencephalogram (EEG) classification tasks: A review. J Neural Eng. 2019; 16: 28.

7. Na¨a¨ta¨nen R, Paavilainen P, Rinne T, Alho K. The mismatch negativity (MMN) in basic research of central auditory processing: A review. Clinical Neurophysiology. Elsevier; 2007. pp. 2544–2590.

8. Garrido MI, Kilner JM, Stephan KE, Friston KJ. The mismatch negativity: A review of underlying mechanisms. Clinical Neurophysiology. Elsevier; 2009. pp. 453–463


For more information:1950477648nn@gmail.com




You Might Also Like