Disentangling the independently controllable factors of variation by interacting with the world

Valentin Thomas, Emmanuel Bengio, William Fedus, Jules Pondard, Philippe Beaudoin, Hugo Larochelle, Joelle Pineau, Doina Precup, Yoshua Bengio

Introduction

When solving Reinforcement Learning problems, what separates great results from random policies is often having the right feature representation. Even with function approximation, learning the right features can lead to faster convergence than blindly attempting to solve given problems (Jaderberg et al., 2016).

The idea that learning good representations is vital for solving most kinds of real-world problems is not new, both in the supervised learning literature (Bengio, 2009; Goodfellow et al., 2016), and in the RL literature (Dayan, 1993; Precup, 2000). An alternate idea is that these representations do not need to be learned explicitly, and that learning can be guided through internal mechanisms of reward, usually called intrinsic motivation (Barto et al., ; Oudeyer and Kaplan, 2009; Salge et al., 2013; Gregor et al., 2017).

We build on a previously studied (Thomas et al., 2017) mechanism for representation learning that has close ties to intrinsic motivation mechanisms and causality. This mechanism explicitly links the agent’s control over its environment to the representation of the environment that is learned by the agent. More specifically, this mechanism’s hypothesis is that most of the underlying factors of variation in the environment can be controlled by the agent independently of one another.

We propose a general and easily computable objective for this mechanism, that can be used in any RL algorithm that uses function approximation to learn a latent space. We show that our mechanism can push a model to learn to disentangle its input in a meaningful way, and learn to represent factors which take multiple actions to change and show that these representations make it possible to perform model-based predictions in the learned latent space, rather than in a low-level input space (e.g. pixels).

Learning disentangled representations

Other authors have proposed mechanisms to disentangle underlying factors of variation. Many deep generative models, including variational autoencoders (Kingma and Welling, 2014) , generative adversarial networks (Goodfellow et al., 2014) or non-linear versions of ICA (Dinh et al., 2014; Hyvarinen and Morioka, 2016) attempt to disentangle the underlying factors of variation by assuming that their joint distribution (marginalizing out the observed ss) factorizes, i.e., that they are marginally independent.

Here we explore another direction, trying to exploit the ability of a learning agent to act in the world in order to impose a further constraint on the representation. We hypothesize that interactions can be the key to learning how to disentangle the various causal factors of the stream of observations that an agent is faced with, and that such learning can be done in an unsupervised way.

The selectivity objective

To discover meaningful factors of variation ϕ\phi and their associated policies πϕ\pi_{\phi}, we consider the following general quantity S\mathcal{S} which we refer to as selectivity and that is used as a reward signal for πϕ\pi_{\phi}:

Here h=f(s)h=f(s) is the encoded initial state before executing πϕ\pi_{\phi} and h′=f(s′)h^{\prime}=f(s^{\prime}) is the encoded terminal state. ϕ\phi and φ\varphi represent factors of variation a factor. A(h′,h,ϕ)A(h^{\prime},h,\phi) should be understood as a score describing how close ϕ\phi is to the variation it caused in (h′,h)(h^{\prime},h). For example in the experiments of section 4.1, we choose AA to be a gaussian kernel between h′−hh^{\prime}-h and ϕ\phi, while in the experiments of section 4.2, we choose A(h′,h,ϕ)=max⁡{0,⟨h′−h,ϕ⟩}A(h^{\prime},h,\phi)=\max\{0,\langle h^{\prime}-h,\phi\rangle\}. The intuition behind these objectives is that in expectation, a factor ϕ\phi should be close to the variation it caused (h′,h)(h^{\prime},h) when following πϕ\pi_{\phi} compared to other factors φ\varphi that could have been sampled and followed thus encouraging independence within the factors.

where θ\theta is the set of weights shared by the factor generator, the policy network and the encoder.

Thus, our total objective along entire trajectories is a lower bound on the causal (Ziebart, 2010) or directed (Massey, 1990) information Ip(ϕ↦h)=∑tIp(ϕ1:t,ht∣ht−1)\mathcal{I}_{p}(\phi\mapsto h)=\sum_{t}\mathcal{I}_{p}(\phi_{1:t},h_{t}|h_{t-1}) which is a measure of the causality the process ϕ\phi exercises on the process hh. See Appendix C for details.

Experiments

We use MazeBase (Sukhbaatar et al., 2015) to assess the performance of our approach. We do not aim to solve the game. In this setting, the agent (a red circle) can move in a small environment (64×6464\times 64 pixels) and perform the actions down, left, right, up. The agent can go anywhere except on the orange blocks.

After jointly training the reconstruction and selectivity losses, our algorithm disentangles four directed factors of variations as seen in Figure 1: ±x\pm x-position and ±y\pm y-position of the agent. For visualization purposes we chose the bottleneck of the autoencoder to be of size K=2K=2. To complicate the disentanglement task, we added the redundant action up as well as the action down+left in this experiment.

The disentanglement appears clearly as the latent features corresponding to the xx and yy position are orthogonal in the latent space. Moreover, we notice that our algorithm assigns both actions up (white and pink dots in Figure 1.a) to the same feature. It also does not create a significant mode for the feature corresponding to the action down+left (light blue dots in Figure 1.a) as this feature is already explained by features down and left.

2 Multistep embedding of policies

In this experiment, ϕ\phi are embeddings of 33-steps policies πϕ\pi_{\phi}. We add a model-based loss LMB=∣∣ht+3−Tθ(ht,ϕ)∣∣2\mathcal{L}_{MB}=||h_{t+3}-T_{\theta}(h_{t},\phi)||^{2} defined only in the latent space, and jointly train a decoder alongside with the encoder. Notice that we never train our model-based cost at pixel level. While we currently suffer from mode collapsing of some factors of variations, we show that we are successfully able to do predictions in latent space, reconstruct the latent prediction with the decoder, and that our factor space disentangles several types of variations.

Conclusion, success and limitations

Pushing representations to model independently controllable features currently yields some encouraging success. Visualizing our features clearly shows the different controllable aspects of simple environments, yet, our learning algorithm is unstable. What seems to be the strength of our approach could also be its weakness, as the independence prior forces a very strict separation of concerns in the learned representation, and should maybe be relaxed.

Some sources of instability also seem to slow our progress: learning a conditional distribution on controllable aspects that often collapses to fewer modes than desired, learning stochastic policies that often optimistically converge to a single action, tuning many hyperparameters due to the multiple parts of our model. Nonetheless, we are hopeful in the steps that we are now taking. Disentangling happens, but understanding our optimization process as well as our current objective function will be key to further progress.

References

Appendix A Additional details

Through our research, we experiment with different outputs for our generator Φ(h,z)\Phi(h,z). We explored embedding the ϕ\phi-vectors into a hypercube, a hypersphere, a simplex and also a simplex multiplied by the output of a tanh(⋅)tanh(\cdot) operation on a scalar.

A.2 First experiment

In the first experiment, figure 1, we used a gaussian similarity kernel i.e A(h′,h,ϕ)=exp⁡(−∣∣h′−(h+ϕ)∣∣22σ2)A(h^{\prime},h,\phi)=\exp(-\frac{||h^{\prime}-(h+\phi)||^{2}}{2\sigma^{2}}) with σ=dim(h)\sigma=\sqrt{dim(h)}. In this experiment only, for clarity of the figure, we only allowed permissible actions in the environment (no no-op action).

Appendix B Additional Figures

Here we consider the case where we learn a latent space HH of size KK, with KK factors corresponding to the coordinates of hh (hi,  i∈[k])h_{i},\;i\in[k]), and learn KK separately parameterized policies πi(a∣h),  i∈[k]\pi_{i}(a|h),\;i\in[k]. We train our model with the selectivity objective, but no autoencoder loss, and find that we correctly recover independently controllable features on a simple environment. Albeit slower than when jointly training an autoencoder, this shows that the objective we propose is strong enough to provide a learning signal for discovering a disentangled latent representation.

We train such a model on a gridworld MNIST environment, where there are two MNIST digits . The two digits can be moved on the grid via 4 directional actions (so there are 8 actions total), the first digit is always odd and the second digit always even, so they are distiguishable. In Figure 4 we plot each latent feature hkh_{k} as a curve, as a function of each ground truth. For example we see that the black feature recovers +x1+x_{1}, the horizontal position of the first digit, or that the purple feature recovers −y2-y_{2}, the vertical position of the second digit.

B.2 Planning and policy inference example in 1-step

This disentangled structure could be used to address many challenging issues in reinforcement learning. We give two examples in figure 5:

Model-based predictions: Given an initial state, s0s_{0}, and an action sequence a{0:T−1}a_{\{0:T-1\}}, we want to predict the resulting state sTs_{T}.

A simplified deterministic policy inference problem: Given an initial state sstarts_{start} and a terminal state sgoals_{goal}, we aim to find a suitable action sequence a{0:T−1}a_{\{0:T-1\}} such that sgoals_{goal} can be reached from sstarts_{start} by following it.

Because of the tanhtanh activation on the last layer of Φ(h,z)\Phi(h,z), the different factors of variation dh=h′−hdh=h^{\prime}-h are placed on the vertices of a hypercube of dimension KK, and we can think of the the policy inference problem as finding a path in that simpler space, where the starting point is hstarth_{start} and the goal is hgoalh_{goal}. We believe this could prove to be a much easier problem to solve.

B.3 Multistep Example

We demonstrate an instance of ICF operating in a 4×\times4 Mazebase enviroment over five time steps in Figure 6. We consistently witness a failure of mode collapse in our generator Φ\Phi and therefore the generator only produces a subset of all possible ϕ\phi-variations. In Figure 6, we observe the ϕ\phi governing the agent’s policy πϕ\pi_{\phi} appears to correspond to moving two positions down and then to repeatedly toggle the switch. A random action due to ϵ\epsilon-greedy led to the agent moving up and off the switch at time step-4. This perturbation is corrected by the policy πϕ\pi_{\phi} by moving down in order to return to toggling the relevant switch.

Appendix C Variational bound and the selectivity

Let us call p(ht+1∣ϕt+1,ht)=Ph′,hϕp(h_{t+1}|\phi_{t+1},h_{t})=\mathcal{P}^{\phi}_{h^{\prime},h} the probability distribution over final hidden states starting from hh and using the policy parametrized by the embedding ϕ\phi.

p(ht+1∣ϕt+1,ht)=Πk=1Kπϕt+1(at+k−1K∣ht+k−1K)penv(st+kK∣at+k−1K,st+k−1K)p(h_{t+1}|\phi_{t+1},h_{t})=\Pi_{k=1}^{K}\pi_{\phi_{t+1}}(a_{t+\frac{k-1}{K}}|h_{t+\frac{k-1}{K}})p_{env}(s_{t+\frac{k}{K}}|a_{t+\frac{k-1}{K}},s_{t+\frac{k-1}{K}}). where penvp_{env} is the transition probability of the environment.

For simplicity, let’s refer to hth_{t} as hh, ht+1h_{t+1} as h′h^{\prime} and ϕt+1\phi_{t+1} as ϕ\phi.

can be proven by using Donsker-Varadhan variational representation of the KL divergence [Donsker and Varadhan, 1975, Ruderman et al., 2012]:

As we sample the factors ϕ\phi uniformly, our total objective is then a lower bound on ∑tI(ϕt,ht∣ht−1)\sum_{t}\mathcal{I}(\phi_{t},h_{t}|h_{t-1}) which corresponds here to the directed information [Massey, 1990] Ziebart as ϕt\phi_{t} is sampled independently from ϕ1:t−1\phi_{1:t-1}.

Appendix D Additional information on the training

In our experiments, we use the selectivity objective, an autoencoding loss and an entropy regularization loss H(πϕ)\mathcal{H}(\pi_{\phi}) for each of the policies πϕ\pi_{\phi}. Furthermore, in experiment 4.2 we added the model-based cost ∣∣h′−T(h,ϕ)∣∣2||h^{\prime}-T(h,\phi)||^{2} with TT a learned two layer fully connected neural network.

The selectivity is used to update the parameters of the encoder, factor generator and policy networks. We use the following equation for computing the gradients

We also use a state dependent baseline VV as a control variate to reduce the variance of the REINFORCE estimator.

Furthermore, to be able to train the factor generator efficiently, we train all ϕ\phi sampled in a mini-batch (of size 10241024) by importance sampling on the probability ratio of the trajectory under each ϕ\phi