Injecting Prior Knowledge for Transfer Learning into Reinforcement Learning Algorithms using Logic Tensor Networks

Samy Badreddine, Michael Spranger

Introduction

Recently, much of AI and ML is concerned with end-to-end learning for solving various complex tasks such as Atari Games and Go Mnih et al. (2015); Silver et al. (2017). Almost all recent progress in this field has been driven by applications of deep neural networks to Reinforcement Learning (RL) tasks – a subfield known as Deep Reinforcement Learning (DRL). While DRL has shown important progress over the past years, state-of-the-art algorithms still struggle to achieve human-like performance within limited training time. Traditional DRL agents’ learning is slow and algorithms require lots of training data. Humans on the other hand are able to quickly understand and solve tasks they have never seen before in complex environments. It is likely that a large part of the explanation for such efficiency lies in the usage of prior knowledge. For example, experiments with human participants Dubey et al. (2018) show that humans are able to solve complex tasks using specific priors whereas humans fail if such priors are not applicable. Such priors are illustrated in Figure 1.

Inspired by human cognition, this paper proposes to exploit prior knowledge for transfer learning in DRL agents. More specifically, we focus on object semantics priors and describe high-level symbolic facts about objects in the environment –e.g. object xx is an enemy, yy a key, zz a door, etc.– using a first-order language. The knowledge is provided by the human (as prior knowledge) and joined to the image describing the environment. A DRL algorithm is then trained on conjoint image and semantic data and can choose to exploit prior information if it helps performance and learning. That is the system proposed in this paper can learn to take advantage of both the symbolic layer and the conventional layer in a single decision selection module. Also because knowledge is provided in a first-order language, the system is easily extended with new facts and relationships about objects and the environment. We test our framework in a simple grid-world environment.

The system presented in this paper relies on Logic Tensor Networks (LTN) Serafini and Garcez (2016) for representing prior knowledge. LTN has been applied to image segmentation and interpretation Donadello et al. (2017) and also hierarchy learning. So far LTN has not been applied to Reinforcement Learning. To the best of our knowledge this paper presents the first attempt at applying a first-order, neural network grounded system (such as LTN) to Reinforcement Learning.

The paper proceeds by discussing related work, we then illustrate the target environments followed by a detailed description of the proposed framework and system. Following results, we discuss some of the implications of our approach and future work.

Related Work

The combination of Deep Neural Networks and Reinforcement Learning has been the most successful approach to Reinforcement Learning in the last ten years. Most famously DRL has solved ATARI games Mnih et al. (2015) and the game of Go Silver et al. (2017). ATARI and Go both are essentially discrete state and action space Markov Decision Problems. But, DRL has also been applied to continuous control problems with multiple proposals existing for continuous state and action spaces Duan et al. (2016); Lillicrap et al. (2015); Mnih et al. (2016). In principle our approach is compatible with all discrete time Reinforcement Learning problems. Our approach combines rich (continuous) input from images or other sensori information with symbolic information and we then apply any DRL learning system. Our overall system differs from pure DRL systems by being able to easily incorporate prior and background knowledge available in a first-order grounded language.

All of the approaches mentioned in the previous paragraph are examples of end-to-end trainable systems (in the case of DRL). DRL takes as input raw images or sensor data and outputs discrete or continuous actions. It is difficult to add prior knowledge to such systems in a systematic manner. Similarly, traditional relational RL approaches do not take into account the grounding in sensorimotor spaces and how to learn in those spaces conjointly with the symbolic information.

How to combine such systems effectively has led recently to new work on how to use traditional symbolic knowledge representations in Reinforcement Learning. For instance, some recent work introduces symbolic front ends on top of neural back ends Garnelo et al. (2016). The neural back-end is responsible of conceptual abstraction from the image and maps the raw input to symbolic representations. A symbolic layer then represents information in separate streams for each symbol, before a decision module aggregates them using heuristics. Others have tried to add common sense priors in the heuristic aggregation Garcez et al. (2018). The system presented in this paper integrates different sources of information in a single representation before a DRL algorithm can learn to make choices using either symbolic information or raw pixel data.

Another recent example of integrating symbolic information in Reinforcement Learning is Bougie et al. Bougie et al. (2018). They propose two streams of processing, one where the symbolic information is processed and one where the pixel data is used. In the end there are two actions and a supervised learning module selects which action to take either the symbolic or DRL output. Our architecture presented in this paper simplifies into a single action selection stream and can learn to take advantage of both the symbolic layer and the image layer in a single decision selection module.

Experimental Setup and Task

We test our system in a simple game environment inspired by previous work Garnelo et al. (2016). The environment consists of different types of objects and an agent (see Figure 2). Objects differ in shape: circle, square and cross.

In each environment objects of all three types are present. The task for the agent is to collect all objects of a particular target type – e.g. all circles – and avoid all objects of avoid type – e.g. squares –. The agent receives a +1 reward upon collecting a target type object, a -1 reward upon collecting an avoid type object, a zero reward in all other cases.

The environment is represented by a 50×5050\times 50 image with objects of size 10×1010\times 10 in 5×55\times 5 cells. The agent can move in the environment with the following actions: move-up, move-down, move-left, move-right. An object is collected when the agent enters a field.

Number and position of objects as well as the position of the agent are randomized for each trial. A trial ends when the agent has collected all objects associated with a positive reward or after 50 steps.

From this basic environment structure with 3 types of objects and 1 agent, we create two experiments: Experiment I - Symbolic abstraction and Experiment II - Fact derivation

The agent is trained on the same game but rendered with different colors.

The different settings represent separate video games where objects and backgrounds look visually different but have the same meaning – e.g. enemies, hero, etc.–. The goal of this experiment is to highlight how a symbolic layer helps to transfer collection/avoidance strategies across object properties (a circle is a circle if drawn in red or in white).

Experiment II - Fact derivation

The types of objects that have to be collected change over time.

target cross (+1), avoid square and circle (-1)

The scenarios represent different video games where objects are similar in type but interaction is different. The goal of this experiment is to highlight how a symbolic layer can help to transfer (collection/avoidance) strategies across object types.

Architecture Overview

Our task setting is a sequential task and it can be principally solved using Reinforcement Learning (RL). The basic idea for RL is having an agent trying to solve a task by observing the environment through its sensors, choosing to act accordingly and occasionally receiving a reward for its actions. The goal of the agent is to learn a policy that maximizes expected reward across trials. We propose the following agent architecture to solve tasks such as the one discussed here (see also Figure 4). Overall our architecture consists of the following parts.

The current image of the environment. Here we use a 50x50 pixel image.

The image is processed with prior knowledge. Derived representations similar to feature maps are computed and augment the raw pixel information.

Both the original input and the prior knowledge derived maps are combined into a single representation.

The concatenation result is the input for a Deep Reinforcement Learning algorithm. For the environment discussed in this paper, algorithms that can deal with image-like input spaces and discrete action spaces are appropriate. Here, we apply a Double Dueling Deep Q Network architecture with experience replay.

The output is one of 4 possible actions (move-up, move-down, move-left, move-right)

The following Sections give more details on the key parts of the architecture and how they interact.

Prior Knowledge Representation with Logic Tensor Networks

We inject prior knowledge into the agent using a three step process (Figure 4).

We discretize the environment in a patch representation similar to feature maps. Being a grid, our game environment is easily described with 5×55\times 5 maps. Each channel of the maps represents an object type predicate and is filled with the detection results of the previous step. Each map is filled with the detection results of the previous step, using interpolation methods if the dimensions mismatch.

Facts about which objects to avoid and which to goto are computed based on background knowledge (axioms) for the currently active scenario.

This leads to the following grounded theory.

We assume a domain of objects O\mathcal{O} that consists in each cell of the object maps – that is, the image patches of size 10×1010\times 10.

For each scenario we provide background knowledge in the form of axioms about which objects to avoid and which to go to. For instance, in Scenario 1 the following axioms hold and are used:

Action Selection

Both the raw original input and the prior knowledge derived maps are fed into the action selection module. Figure 6 shows our architecture applied to input/output streams. The original image input and the prior knowledge input are conjointly fed into the architecture. The action selection in our system is based on the Double Dueling Deep Q Network architecture van Hasselt et al. (2015) augmented by streams for processing symbolic and raw image information. The Double Dueling Deep Q Network is a variation of the Q-learning algorithm DQN Mnih et al. (2015), a value-based RL algorithm that learns to predict the expected discounted reward for state-action pairs Q(s,a)Q(s,a). DQN approximates the Q values using a deep neural network with various convolutional and fully-connected layers. The training uses an experience replay buffer for a more stable learning.

The Double Dueling Deep Q Network architecture Hessel et al. (2017) decouples the selection of the action from its evaluation van Hasselt et al. (2015). Dueling approaches of value-based algorithms feature two stream of computation – advantage depending on state-action and value depending on state only – for a more robust estimation Wang et al. (2015). These improvements are compatible with our approach as we only investigate the input and not the action selection algorithm.

Actions are selected ϵ\epsilon-greedily, i.e. maximum valued action with probability (1−ϵ)(1-\epsilon), random action with probability ϵ\epsilon. The rate of exploration ϵ\epsilon is decreased over time.

Results

We train the framework on the experiments presented in Section 2.

Experiment I tests an agent on different texture rendering settings of the game. We change setting every 50 epochs. It highlights the importance of symbolic abstraction and representing objects out of their visual context into types predicates.

Experiment II tests an agent on different scenarios of target type and avoid type objects. We change scenario every 50 epochs. It highlights the importance of providing facts on the objects to describe a task.

The background knowledge (object type classification and derived facts) is adjusted to the currently active scenario. We systematically vary the symbolic information injected into the agent architecture to show the effect. In our experiments, we test:

The agent only has the raw image as input for the action selection.

The results of the object recognition are joined to the raw image input.

The results of the object recognition and fact derivation are joined to the raw image input.

Hyper-parameters of the RL algorithm such as ϵ\epsilon can be optimized. We investigate two strategies for adjusting ϵ\epsilon.

The decreasing exploration rate of the ϵ\epsilon-greedy action selection is reset at each new environment setting.

The decreasing exploration rate of the ϵ\epsilon-greedy action selection is unchanged at each new environment setting.

The agents’ network is trained every 2\times1022\text{\times}{10}^{2} timesteps. It is evaluated every 2\times1052\text{\times}{10}^{5} timesteps, defined as an epoch. Evaluations measure the collected rewards normalized by the potential maximum reward for a particular environment. Evaluations are averaged on 50 trajectories. The position of the agent and the position and number of objects are randomized at each trajectory.

Figure 7(a) and Figure 7(b) show results for 5 experimental runs per experiment. Also the 95% confidence interval is plotted. We change scenarios and settings every 50 epochs during 400 epochs, for a total of 8 changes. Experiment I proves that the agent successfully leverages the knowledge on object types through time. Experiment II proves that the agent successfully leverages the knowledge on facts through time.

Figure 7(a) shows results for Experiment I. The baseline no prior condition shows that every time the scenario changes, the algorithm is relearning the task. This behavior does not actually depend on whether ϵ\epsilon is reset or not.

If we check the performance of priors on object types, then we can see 2 trends vs the baseline. 1) The learning seems to be much faster. 2) The drop in performance at each scenario change becomes smaller and smaller. In other words the system still has to relearn the task initially but over time becomes more and more immune to scenario/color changes. The system has learned to rely on object type information to become performant even when the color of the object changes. In summary, object priors help the system to learn faster and to be able to transfer task performance between different scenarios. Notice that the effect is more pronounced when ϵ\epsilon is not reset. As opposed to the baseline, ϵ\epsilon reset actually matters as it will favor exploration over exploitation. So because the system has learned to generalize across scenarios, ϵ\epsilon should not be reset.

Lastly, if we check the priors on object types and facts then we can see almost no difference to the performance of priors on object types. This makes sense as only the color of objects change and not their avoidance/collection semantics. Consequently, knowledge about the object semantics is not relevant for learning to generalize over Experiment I scenario changes.

2 Results Experiment II

Figure 7(b) shows results for Experiment II. Similar to Experiment I, the graph shows that in the baseline no prior condition, the algorithm is relearning the task every time the scenario changes. While there is some impact of whether ϵ\epsilon is reset or not, overall the system is mostly relearning the task (unless there is some overlap between succeeding scenarios).

If we check the performance of priors on object types condition, then we can see that it mirrors the baseline condition. Remember that in this experiment a circle can be an object to avoid in one scenario and is an object to collect in another. Consequently, knowledge about object types does not help to learn faster or be able to transfer knowledge from one scenario to another.

If we check the priors on object types and facts condition, then we can see 2 trends vs the baseline (and the priors on object types condition). 1) The learning seems to be much faster. 2) The drop in performance at each scenario change becomes smaller and smaller. In other words the system still has to relearn the task initially but over time becomes more and more immune to scenario changes. The system has learned to rely on derived facts to become performant even when the semantics of an object with respect to the task changes. In summary, object priors help the system to learn faster and to be able to transfer task performance between different scenarios.

In this last condition resetting ϵ\epsilon does have impact on the general trend in learning. Resetting ϵ\epsilon hurts the baseline system and the system using only object type priors, but aids the system using all background knowledge. Overall performance is higher and learning is faster. This makes sense, because if the system has learned to transfer across scenario changes, then resetting ϵ\epsilon is undesirable. On the other hand, resetting ϵ\epsilon does aid the baseline and object type prior systems. For systems that do not learn to transfer, resetting ϵ\epsilon at least allows them to learn each scenario over and over again.

Conclusion

This paper discussed a new approach to injecting prior knowledge conjointly with original raw input into reinforcement learning algorithms. By grounding these priors in predicates, we showed how symbolic semantics on objects are useful for transfer learning. As proof-of-concept, we demonstrated the architecture in a simple grid world. The experiments show that symbolic abstractions can help to solve tasks across various scenarios without relearning the decision module. We demonstrated how agents learn to leverage the appropriate knowledge for a particular task and learn to select and exploit information from prior knowledge sources.

Further work should test this approach in more complex environments with more or less human-provided information. Here, we investigated priors as a state expansion method for transfer learning. We plan to explore other ways to use the priors in a RL architecture, such as a policy transfer method or reward function indicator. We also plan to further investigate the impact of exploration strategies on the system, to prevent misguidance from inaccurate priors. More elaborate exploration frameworks have recently been investigated in the literature Oudeyer and Kaplan (2009); Pathak et al. (2017). Such algorithms might become essential when using prior knowledge in order to tradeoff exploration and exploitation at the correct time in the experiment.

References