An Insight Of Deep Learning Based Demand Forecasting in Smart Grids(1)

Jun 05, 2023

Abstract: Smart grids are able to forecast customers’ consumption patterns, i.e., their energy demand, and consequently electricity can be transmitted after taking into account the expected demand. To face today’s demand forecasting challenges, where the data generated by smart grids is huge, modern data-driven techniques need to be used. In this scenario, Deep Learning models are a good alternative to learn patterns from customer data and then forecast demand for different forecasting horizons. Among the commonly used Artificial Neural Networks, Long Short-Term Memory Networks—based on Recurrent Neural Networks—are playing a prominent role. This paper provides an insight into the importance of the demand forecasting issue, and other related factors, in the context of smart grids, and collects some experiences of the use of Deep Learning techniques, for demand forecasting purposes. To have an efficient power system, a balance between supply and demand is necessary. Therefore, industry stakeholders and researchers should make a special effort in load forecasting, especially in the short term, which is critical for demand response.

Keywords: demand forecasting; load forecasting; demand response; forecasting horizon; smart grid; smart environment; Deep Learning; Long Short-Term Memory networks; Convolutional Neural Networks;Cistanche deserticola

cistanche tubulosa benefits and side effects

Desert ginseng

1. Introduction

Electricity cannot be easily stored for future supply, unlike other commodities such as oil. This means that electricity must be distributed to the consumers immediately after its production. The distribution of electricity to final users has been done with the help of the traditional electrical grid (see definition in Table 1) which allows the delivery of electricity from producers to consumers. To achieve that goal, it connects the electricity generating stations and the transmission lines that deliver the electricity to the final users. Traditional electrical grids vary in size. When these grids started to expand, controlling them became a very complex and difficult task. Additionally, demand forecasting (see definition in Table 1) has not traditionally been considered.

In this context, the concept of the smart grid (see definition in Table 1) arises and starts to play an important role. This concept has been exhaustively reviewed in the literature (e.g., [1–4]). Smart grids provide two-way communication between consumers and suppliers. Smart grids add hardware and software to the traditional electrical grid to provide it with an autonomous response capacity to different events that can affect the electrical grid. The final objective is to achieve optimal daily operational efficiency for electrical power delivery. In [4], the authors define a “smart grid” as a new form of electricity network that offers self-healing, power-fellow control, energy security, and energy reliability using digital technology. In [2], the authors highlight that the concept of the smart grid is transforming the traditional electrical grid by using different types of advanced technology. According to these authors, this concept integrates all the elements that are necessary to generate, distribute, and consume energy efficiently and effectively. In [5], the authors emphasize that the smart grid concept emerged to make the traditional electrical grid more efficient, secure, reliable, and stable, and to be able to implement demand response (see definition in Table 1).

Table 1. Important keywords for reading this article.

Table 2. Main variables affecting electricity demand.  image

The smart grid paradigm allows consumers to find out their energy usage patterns. Consequently, consumers can control their consumption and use energy more efficiently. In the implementation of the smart grid concept, demand response—for both household and industrial purposes—plays an important role. Another useful tool is load forecasting (see definition in Table 1). In [6] the authors mark the importance of this concept in the context of smart grids, as forecasting the electricity needed to meet demand allows power companies to better balance demand and supply. Power companies are especially interested in achieving accurate forecasts for the next 24 h, which is called load profile (see definition in Table 1).

In addition, in recent years, the increased demand for electricity at certain times of the day has created several problems. Load forecasting is especially important during peak hours. Demand response encourages customers to offload non-essential energy consumption during these peak hours.

To face load forecasting challenges, it is necessary to use modern data-driven techniques. Indeed, the incorporation of new technologies, such as Big Data, Machine Learning, Deep Learning, and the Internet of Things (IoT), has upgraded the smart grid concept to another level, as these technologies allow for improved demand forecasting and automated demand response.

This paper provides insight into the importance of demand forecasting and important related factors in the context of smart grids, as well as the possibility of using data-driven techniques for this purpose. More specifically, the authors focus on Deep Learning techniques, as it has emerged as a good option for the implementation of demand forecasting in the context of smart grids. The paper collects some experiences of using different Deep Learning techniques in the energy domain for forecasting purposes. An effigy-client power system must take demand response into account. Additionally, accurate load forecasting, especially in the short term, is essential, which is why industry stakeholders and researchers are putting special efforts into it.

Cistanche deserticola experiment

Cistanche deserticola experiment

Click here to view Cistanche products

【Ask for more】 Email:cindy.xue@wecistanche.com /  Whats App:  0086 18599088692 /  Wechat:  18599088692

The remainder of this paper is organized as follows. Section 2 presents the reasons why demand forecasting is important in the context of smart grids. Section 3 describes the most important factors in demand forecasting. Section 4 presents the different possible classifications of demand forecasting techniques. Section 5 provides some fundamentals and concepts useful to understand the Deep Learning models commonly used in the energy domain. Section 6 collects different experiences of using these models in the context of smart grids for forecasting purposes. Finally, Section 7 summarizes the main conclusions of this work.

2. The Importance of Demand Forecasting

In [7] the authors summarize the main requirements of smart grids as follows: flexible enough to meet users’ needs, able to manage uncertain events, accessible for all users, reliable enough to guarantee high-quality energy delivery to consumers, and innovative enough to manage energy efficiently.

With these requirements in mind, smart grids should aim to develop low-cost, easy-to-deploy technical solutions with distributed intelligence to operate efficiently in today’s increasingly complex scenarios. To upgrade a traditional electrical grid into a smart grid, intelligent and secure communication infrastructures are necessary [4].

According to the study presented in [8], forecasting can be applied in two main areas: grid control and demand response. In [9], the authors highlight that forecasting models are essential to provide optimal quality of the energy supply at the lowest cost. In addition, real-time information on users’ energy consumption patterns will enable more sophisticated and efficient forecasting models to be applied. Forecasting must also consider the need to manage constantly changing information. In [10], the authors highlight that, with the smart grid, demand response programs can make the grid more cost-efficient and resilient.

The authors in [11] remark that there are important challenges in demand forecasting due to the uncertainties in the generation profile of distributed and renewable energy generation resources. Increasing attention is being paid to load forecasting models, especially dealing with renewable energy sources (solar radiation, wind, etc.) [9].

Cistanche supplement near me—Improve memory2

What does cistanche do- Improve Memory

The distributed generation paradigm facilitates the use of renewable energy sources that can be placed near consumption points. When using this paradigm, smart grids have multiple small plants that supply energy to their surroundings. Consequently, the dependence on the distribution and transmission grid is smaller [9]. However, this paradigm makes grid control even more uncertain, especially when the distributed generation sources are renewable and consequently have a random nature. Despite this difficulty, the share in energy production of variable renewable energy sources is expected to increase in the coming years [12].

Another key element is microgrids (see definition in Table 1) [13,14]. Based on this concept, and taking into consideration the intelligence deployed in buildings, new concepts have emerged including smart homes (see definition in Table 1) and smart buildings (see definition in Table 1). Buildings today are complex combinations of structures, systems, and technology. Technology is a great ally in optimizing resources and improving safety. Advances in building technologies are combining networked sensors and data recording in innovative ways [15]. Modern facilities can adjust heating, cooling, and lighting to maximize energy efficiency, providing also detailed reports of energy consumption. In these new smart environments (see definition in Table 1), sensors and smart devices are deployed to obtain enough information about the users’ energy consumption patterns. Once again, this requires forecasting models that must be applied to the specific variables of the scenario to be controlled.

Forecasting models will allow considering variables (climatic, social, economic, habit-related, etc.) that can influence the accuracy of forecasts [9]. These authors remark that energy demand estimates in disaggregated scenarios, such as residential users in smart buildings, are more complex compared to energy demand estimates for an aggregated scenario, such as a country. Disaggregating the demand also facilitates the implementation of demand response, as different prices can be offered based on the criteria set by the power company.

The gradual integration of intelligence at the transmission, distribution, and end-user levels of the electricity system aims to optimize energy production and distribution to adjust producers’ supply to consumers’ demand. Moreover, smart grids seek to improve fault detection algorithms [16]. Accurate demand forecasts are very useful for energy suppliers and other stakeholders in the energy market [17]. Load forecasting has been one of the main problems faced by the power industry since the introduction of electric power [18].

3. Important Factors in Demand Forecasting

Electricity demand is affected by different variables or determinants. These variables include forecasting horizons, the level of load aggregation, weather conditions (humidity, temperature, wind speed, and cloudiness), socio-economic factors (industrial development, population growth, cost of electricity, etc.), customer type (residential, commercial, and industrial), and customer factors about electricity consumption (characteristics of the consumer’s electrical equipment) (e.g., [19–23]).

To fully understand demand forecasting techniques and objectives, it is necessary to examine these determinants. In this section, the authors will focus on (1) period, (2) economic issues, (3) weather conditions, and (4) customer-related factors.

cistanche—Improve memory

Benefits of cistanche tubulosa-Improve Memory

3.1. Period or Forecasting Horizon

The period commonly referred to as forecasting horizon is probably one of the factors that have the greatest impact.

According to different authors (e.g., [17,24]), demand forecasting can be classified into three categories concerning the forecasting horizon:

• Short-term (typically one hour to one week). 

• Medium-term (typically one week to one year). 

•Long-term (typically more than one year).

Factors affecting short-term demand forecasting usually do not last long, such as sudden changes in weather [22]. The quality of short-term demand forecasting is critical for electricity market players [20]. On the other hand, the influencing factors of medium-term demand forecasting often have a certain time duration, such as seasonal weather changes. Finally, the factors influencing long-term demand forecasting last for a long time, typically several forecast periods, e.g., changes in Gross Domestic Product (GDP) [22]. Indeed, economic factors have an important impact on long-term demand forecasting, but also on medium and short-term forecasting [25].

The authors of [26] identify the following categories about the forecasting horizon:

• Very short-term (typically seconds or minutes to several hours). 

• Short-term (typically hours to weeks). 

• Medium-term and long-term (typically months to years).

According to these authors, very short-term demand forecasting models are generally used to control the flow. Short-term demand forecasting models are commonly used to match supply and demand. And, finally, medium-term and long-term demand forecasting models are typically used to plan asset utilities.

The authors in [27] showed that the load curve of grid stations is periodic, not only in the daily load curve, but also in the weekly, monthly, seasonal, and annual load curves. This periodicity makes it possible to forecast the load quite effectively.

Demand also reflects the daily lifestyle of the consumer [28]. Consumers’ daily demand patterns are based on their daily activities, including work, leisure, and sleep hours. In addition, there are other demand variations patterns over time. For example, during holidays and weekends, demand in industries and offices is significantly lower than during weekdays due to a drastic decrease in activity. Finally, power demand also varies cyclically depending on the time of the year, the day of the week, and the time of day [22].

3.2. Socio-Economic Factors

Socio-economic factors, including industrial development, GDP, and the cost of electricity, also significantly affect the evolution of demand. Indeed, as mentioned in the previous section, economic factors considerably affect long-term demand forecasts and also have an important impact on medium- and short-term forecasts.

For example, industrial development will undoubtedly increase energy consumption. The same will be true for population growth. This means that there is a positive correlation between industrial development, population growth, and energy consumption.

GDP is an indicator that captures a country’s economic output. Countries with a higher GDP generate a greater quantity of goods and services and will consequently have a higher standard of living and lifestyle habits, which will stimulate energy demand.

Another economic factor to consider is cost, as it also affects demand. For example, when the price of electricity decreases, wasteful electricity consumption tends to increase [22].

The cost of electricity depends on different factors and is shaped in different ways. For example, in some countries such as Spain, there are two markets (regulated and free) for electricity. In the free market, the cost of electricity is established in the contract signed by the consumer. In contrast, in the regulated market, the price of electricity depends on supply and demand. The price is updated hourly and fluctuates. From the demand side, the more electricity is demanded, the more expensive it is. When less electricity is demanded, the cheaper it is. Normally, it is cheap to use electricity at dawn and expensive to do it when everyone else is using it (e.g., at dinner time).

But it is not only the demand that influences prices but also the supply of energy. The reason is that variations in the price of electricity on the regulated market are caused by differences between demand and supply. Consequently, supply must consider the different ways of generating electricity, which have different costs. The cheapest is electricity generated by renewable energies such as solar, wind, and hydroelectric. The price of nuclear energy is also low; however, in many countries (e.g., Spain), nuclear energy does not cover all energy needs. Thermal (coal), cogeneration, or combined cycle—whose main fuel is gas—tends to be more expensive. It is also important to remember that the main sources of renewable energy, such as hydroelectric or wind, depend on uncontrollable external factors. For example, sufficient rainfall is essential to produce hydroelectric power. However, there is no way to control the weather to make it favorable for producing electrical energy. Given the above, the price is determined by the price of a mix of different sources of power generation, from cheapest to most expensive, until the entire energy demand is met.

cistanche—Improve memory2

Cistanche supplement near me-Improve Memory

3.3. Weather Condition

There are different weather variables relevant for demand forecasting such as temperature, humidity, and wind speed.

The influence of weather conditions on demand forecasting has attracted the interest of many researchers. As an example, the authors in [29] proposed different models to forecast the next day’s aggregated load using Artificial Neural Networks (ANNs), considering the most relevant weather variables—more specifically, mean temperature, relative humidity, and aggregate solar radiation—to analyze the influence of the weather.

Some authors have studied the relationship between temperature and electricity consumption and claim that the correlation between temperature and the electricity load curve is positive, especially in summer (e.g., [25]).

Currently, heat waves have become more common around the world, as well as the possibility of extreme temperatures. In addition, heat waves are not only more frequent but also more intense and longer lasting. Moreover, the nights are getting warmer, which is an added problem. The main effect of a heat wave is an increase in energy consumption as the consumer turns on the air conditioning more and for longer periods. Additionally, cooling systems must work harder as they must cope with higher temperatures.

During the summer, heat waves force the grid to be at maximum capacity. One of how a heat wave affects consumption is through the increased saturation of the electrical grid. While cold waves are counteracted with electricity, gas, wood, etc., heat waves can only be fought with electricity. In other words, the devices that consumers use for cooling are mainly powered by electricity. For this reason, heat waves generate more stress on power lines, as well as higher consumption.

It should be noted that, in colder countries, the increase in consumption during a heat wave is usually lower. This is because the installation of air conditioning systems is not as common as in warmer countries. However, these colder countries are facing heat waves that did not occur in previous years (before climate change) and this is causing them all types of problems, as they are less prepared. This situation is forcing these countries to make changes such as increasing the use of cooling systems.

On the other side, the experience of the harshness of temperature increases with humidity, especially during the rainy season and summer. For this reason, electricity consumption increases during humid summer days. It is also important to note that in coastal areas, such as the Mediterranean area in Spain, electricity consumption tends to be higher. This is both because houses tend to have more electrical equipment than in other areas, and because of the high degree of humidity due to the proximity of the sea.

Wind speed also affects electricity consumption. When it is windy, the human body feels that the temperature is much lower and more heating is needed, which increases electricity consumption. However, it should also be noted that wind energy is one of the main renewable energies. In other words, when there are wind, electricity consumption increases, but at the same time its price decreases. This is because, as explained in the previous section, the price of electricity is usually determined as a mix of the different energy sources, from the cheapest energies (renewables, including wind, and nuclear) to the most expensive generation sources (thermal, combined cycle).

Temperature, humidity, and wind affect the use of electricity. Humidity and temperature are also the main weather variables used in electricity demand prediction systems to minimize operating costs. However, other factors, such as clouds, also play a role. For example, during the day, when clouds disrupt sunlight there is usually a drop in temperature and, consequently, higher electricity consumption.

3.4. Customer Factors

The type of customer (residential, commercial, and industrial), as well as other customer factors related to electricity consumption (characteristics of the consumer’s electrical equipment), can also affect demand. This is important because most energy companies have different types of customers (residential, commercial, and industrial consumers), who have equipment that varies in type and size. These different types of customers have different load curves, although there are some similarities between commercial and industrial customers.

Table 2 summarizes the main determinants affecting electricity demand described in this section.

Table 2. Main variables affecting electricity demand.

Table 2. Main variables affecting electricity demand.  image

4. Classification of Demand Forecasting Techniques

This section classifies demand forecasting models according to three different criteria: (1) period, (2) forecasting objective, and (3) type of model used.

The first classification focuses on the point of view of the period to be forecasted, i.e., the forecasting horizon. To select this criterion, the electricity demand determinants presented in the previous section have been considered. The second classification focuses on the point of view of the forecasting objective, differentiating between forecasting techniques that produce a single value and those that produce multiple values. Finally, the third classification focuses on the point of view of the model used.

4.1. Classification of Demand Forecasting Techniques according to the Forecasting Horizon

As explained in the previous section, the main forecasting horizons that can be identified are the following:

• Very short-term: typically from seconds or minutes to several hours. 

• Short-Term: typically from hours to weeks. 

• Medium-Term: typically from a week to a year. 

• Long-Term: typically more than a year.

The main difference is the scope of the variables used in each case. Very short-term forecasting models use recent inputs (typically minutes or hours), short-term forecasting models use inputs typically in the range of days, and medium and long-term forecasting models use inputs typically in the range of weeks or even months.

Power companies are particularly interested in producing accurate forecasts for the load profile (e.g., [9,30,31]). This is because it can directly affect the optimal scheduling of power generation units. However, due to the non-linear and stochastic behavior of consumers, the load profile is complex, and although research has been done in this area, accurate forecasting models are still needed [32].

4.2. Classification of Demand Forecasting Techniques by Forecasting Objective

Forecasting models can be also classified according to the number of values to be forecasted. In this case, two main categories can be considered.

The first category refers to forecasting techniques that produce only one value (e.g., the next day’s total load, the next day’s peak load, the next hour’s load, etc.). Examples are found in [33,34].

The second category refers to forecasting techniques that produce multiple values, e.g., the next hour’s peak load plus another parameter (e.g., the aggregate load) or the load profile. Examples are found in [35–37].

Generally speaking, one-value forecasts are useful for optimizing the performance of load fellows. On the other hand, multiple-value forecasts are mainly used for energy generation scheduling [9].

4.3. Classification of Demand Forecasting Techniques according to the Model Used

The model to be used is usually decided by the practitioner. In terms of models, the main groups are linear and non-linear approaches.

Linear models include Spectral Decomposition (SD), Partial Least-Square (PLS), AutoRegressive Integrated Moving Average (ARIMA), Auto-Regressive Conditional Heteroscedasticity (ARCH), Auto-Regressive (AR), Auto-Regressive and Moving Average (ARMA), Moving Average Model (MAM), Linear Regression (LR), and State-Space (SS).

Linear techniques have progressively lost importance and interest in favor of nonlinear techniques based on ANNs. Deep Learning models use ANNs, inspired by the human nervous system. These models can learn patterns from the data generated and forecast peak demand in the context of today’s complex smart scenarios, where a large amount of data is continuously generated from different sources [7].

Table 3 summarizes the criteria commonly used to classify demand forecasting models.

Table 3. Main criteria commonly used to classify demand forecasting models.

Table 3. Main criteria commonly used to classify demand forecasting models.  image

5. Fundamentals and Concepts of Machine Learning and Deep Learning Systems

image

Artificial Intelligence is a complex concept that, in a nutshell, refers to machine intelligence [38]. Unlike humans, Artificial Intelligence can identify patterns within a large amount of data using a quite limited amount of time and resources. Furthermore, the computational capacity of machines does not decrease with time and/or fatigue [39].

Artificial Intelligence systems use different types of learning methods, such as Machine Learning and Deep Learning.

5.1. Machine Learning

Machine Learning algorithms are pre-trained to produce an outcome when confronted with a never-before-seen dataset or situation [40]. However, the computer needs more examples to learn than humans do [41]. Machine Learning allows the introduction of intelligent decision-making in many areas and applications where developing algorithms would be complex and excellent results are needed [42].

There are different categories of Machine Learning algorithms including supervised, semisupervised, unsupervised, and reinforcement learning. These different categories of algorithms are briefly described below.

5.1.1. Supervised Learning

After being trained with a set of labeled data examples, these algorithms can predict label values when the input has unlabeled data. The problems typically associated with this type of learning are (1) regression and (2) classification [43].

In regression, the algorithm focuses on understanding the relationship between dependent and independent variables. In classification, the algorithm is used to predict the class label of the data. Common classification problems include (1) binary classification, between two class labels; (2) multi-class classification, between more than two class labels; and (3) multi-label classification where one piece of data is associated with several classes or labels, as opposed to traditional classification problems with mutually exclusive class labels [44].

Methods used for supervised learning include Linear Discriminant Analysis (LDA), Naive Bayes (NB), K-nearest Neighbors (KNN), Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), Adaptive Boosting (AdaBoost), Extreme Gradient Boosting (XGBoost), Stochastic Gradient Descent (SGD), Rule-based Classification (for classification); and LR, Non-Linear Regression (NLR), and Ordinary Least Squares Regression (OLS) (for regression) [44,45]. The most widely cited and implemented supervised learners in the literature are DT, NB, and SVM algorithms [46].

Some interesting practical applications are text classification, predicting the sentiment of a text (such as a Tweet or other social media), assessing the environmental sustainability of clothing products [47], characterizing, predicting, and treating mental disorders [48], and estimating peak energy demand.

5.1.2. Unsupervised Learning

This type of learning uses unlabeled data. In this case, the system explores the unlabeled data to find hidden structures, rather than predicting the correct output. This type of learning is not directly applicable to regression or classification problems, as the possible values of the output are unknown [49]. Instead, it is often used for (1) clustering, (2) association, and (3) dimensionality reduction [43].

Clustering allows unlabeled data to be grouped based on their similarities or differences [49,50]. Association uses different rules to identify new and relevant insights between the objects of a set. Finally, dimension reduction allows a reduction of the number of features (or dimensions) of a dataset to eliminate irrelevant or less important features and thus reduce the complexity of the model [44]. This reduction in the number of features can be done by keeping a subset of the original features (feature selection) or by creating completely new features (feature extraction).

The most popular clustering algorithm is probably K-means clustering, where the k value represents the size of the cluster [44,45,51]. Association algorithms include Apriori, Equivalence Class Transformation (ECLAT), and Frequent Pattern (F-P) Growth algorithms. Finally, dimensionality reduction typically uses the Chi-squared test, Analysis of Variance (ANOVA) test, Pearson’s correlation coefficient, Recursive Feature Elimination (RFE) for feature selection, and Principal Components Analysis (PCA) for feature extraction.

According to [46], the most commonly used unsupervised learners are K-means, hierarchical clustering, and PCA.

These unsupervised learners can have many practical applications, such as facial recognition, customer classification, patient classification, detecting cyber-attacks or intrusions [52], and data analysis in the astronomical field [53].

5.1.3. Semisupervised Learning

Conceptually situated between supervised and unsupervised learning, this type of learning allows the taking advantage of the large unlabeled datasets that are available in some cases combined with (usually smaller) amounts of labeled data [54,55]. This opens up interesting possibilities as labeled data are often scarce, while unlabeled data are more frequent, and a semisupervised learner can obtain better predictions than those produced using only labeled data [44].

Candidate applications are those where there is only a small set of labeled examples and many more unlabeled ones, or when the labeling effort is too high. An example is medical imaging, where a small amount of training data can provide a large improvement in accuracy [43,56].

Table 4 compares Supervised and Unsupervised learning, focusing on the type of input data used in each case (labeled versus unlabeled data), and the main tasks for which both types of learning are used (classification, regression versus clustering, association, and dimensionality reduction).

Table 4. Supervised vs unsupervised learning.

Table 4. Supervised vs unsupervised learning.  image

5.1.4. Reinforcement Learning

This learning technique depends on the relationship between an agent performing an activity and its environment, which provides positive or negative feedback [57,58]. The agent must choose actions that maximize the reward in that environment. Popular methods include Monte Carlo, Q-learning, and Deep Q-learning [44].

Traditionally common applications include strategy games such as chess, autonomous driving, supply chain logistics and manufacturing, genetic algorithms [57], 5G mobility management [59], and personalized care delivery [60].

5.2. Deep Learning

Machine Learning can be classified into shallow and deep, considering the complexity and structure of the algorithm [41]. Deep Learning uses multiple layers of neurons composed of complex structures to model high-level data abstractions [61]. The type of output and the characteristics of the data determine the algorithm to be used for a particular use case [62].

Deep Learning uses ANNs inspired by the human nervous system [63]. This type of network typically has two layers of input and output nodes respectively, connected to each other by one or more layers of hidden nodes. Possible deep ANN architectures include Multilayer Perceptron (MLP), Long Short-Term Memory Recurrent Neural Networks (LSTMRNN), Generative Adversarial Networks (GAN), and Convolutional Neural Networks (CNN or ConvNet).

According to our literature review, the most widely used models in the energy domain are Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM), Deep Q-Networks (DQNs) and Conditional Restricted Boltzmann Machine (CRBM) and a variation of any of them, a combination of two or more of them, or the combination of any of them with other techniques. These models are briefly described below.

5.2.1. Convolutional Neural Networks

These networks are biologically inspired networks, like ordinary neural networks. However, in this type of network, the inputs are assumed to have a specific structure such as images [64]. Being one of the most widely used and effective models for Deep Learning, these networks usually include two types of layers (i.e., pooling and convolution layers). A typical CNN architecture usually consists of an input layer, a convolutional layer, a Max pooling layer, and the final fully connected layer, as shown in Figure 1 [65].

Figure 1. Standard Convolutional Neural Network Architecture [65]

Figure 1. Standard Convolutional Neural Network Architecture [65]

Different architectural designs explore the effect of multilevel transformations on the learning ability of such networks. One of these possible architectural designs is Pyramid. The Pyramid Architecture of Convolutional Neural Networks is commonly known as Pyramid-CNN [66].

5.2.2. Recurrent Neural Networks

In this type of network, the connections between nodes form a directed or undirected graph along a time sequence. Figure 2 shows a typical RNN structure [65].

Figure 2. The framework of a Recurrent Neural Network: Input layer (Xt); output layer (Ot); hidden layer (St); parameter matrices and vectors (U, V, W); activation function of output layer (σy); and activation function of the hidden layer (σh ) [65].

Figure 2. The framework of a Recurrent Neural Network: Input layer (Xt); output layer (Ot); hidden layer (St); parameter matrices and vectors (U, V, W); activation function of output layer (σy); and activation function of the hidden layer (σh ) [65].

This network can use a gating mechanism called Gated Recurrent Units (GRUs) and introduced in 2014 by the authors [67]. GRU are like LSTM networks but with a forgetting gate and fewer parameters as they lack an output gate.

Another variation of this type of network, proposed by Elman [68], is the Elman RNN which includes modifiable feedforward connections and fixed recurrent connections. It uses a set of context nodes to store internal states, which gives it certain unique dynamic characteristics over static ones [69].

5.2.3. Long Short-Term Memory

These networks are a special kind of RNN. Unlike standard feedforward neural networks, these networks have feedback connections, and can even process entire sequences of data (such as speech or video), in addition to individual data points (such as images). This type of RNN contains an input layer, a recurrent hidden layer, and an output layer, with a memory block structure as shown in Figure 3 [70].

image FIgure 3. LSTM memory block [70].

FIgure 3. LSTM memory block [70].

The LSTM memory block can be described according to the following equations [70]:

image


where tx is the model input at time t; Wi, Wf, Wc, W0, Ui, Uf, Uc, U0, V0 are weight matrices; bi, bf , bc, b0 are bias vectors; it, ft, 0t are respectively the activations of the three gates at time t; it is the state of the memory cell at time t; ht is the output of the memory block at time t;   represents the scalar product of two vectors; σ(x) is the gate activation function; g(x) is the cell input activation function; h(x) is the cell output activation function.

A possible extension of this model is the Bidirectional LSTM (B-LSTM). This type of LSTM network aims to analyze sequences from both front-to-back and back-to-front, i.e., the sequence information fellows in both directions backward and forwards, unlike in a normal LSTM.

5.2.4. Deep Q Network and Dueling Deep-Q Network

Deep Q Networks (DQN) and Dueling Deep-Q Networks (DDQN) are a type of ANN using the Deep Q learning algorithm, which is popular in reinforcement learning. In a dueling network, there are two streams to separately estimate the state value as well as the advantages of each action. The main objective of the Deep-Q Network is to choose the best action in a certain state. Considering π is the policy followed by an agent in a given environment, the function Qπ can be defined as follows [71]:

image

where s is a state; a is an action; ri is the potential reward; γ ∈ [0, 1] is a discount factor for making the immediate reward more important than the future ones. Therefore, the objective of Q-learning is to maximize the optimized value function Q* (s, a) = max πQ π(s, a). Figure 4 shows the scheme of a typical DQN architecture [71]. 5.2.5. Conditional Restricted Boltzmann Machine A Restricted Boltzmann Machine (RBM) is a stochastic RNN with two layers, one with visible units and one with binary hidden units. This type of network can learn a probability distribution over its set of inputs. RBMs are a variant of Boltzmann Machines Deep Q-Network architecture [71]. 

Figure 4. Figure 4 shows the scheme of a typical DQN architecture [71].

Figure 4. Figure 4 shows the scheme of a typical DQN architecture [71]. 

You Might Also Like