A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning

Marco Fraccaro, Simon Kamronn, Ulrich Paquet, Ole Winther

Introduction

From the earliest stages of childhood, humans learn to represent high-dimensional sensory input to make temporal predictions. From the visual image of a moving tennis ball, we can imagine its trajectory, and prepare ourselves in advance to catch it. Although the act of recognising the tennis ball is seemingly independent of our intuition of Newtonian dynamics , very little of this assumption has yet been captured in the end-to-end models that presently mark the path towards artificial general intelligence. Instead of basing inference on any abstract grasp of dynamics that is learned from experience, current successes are autoregressive: to imagine the tennis ball’s trajectory, one forward-generates a frame-by-frame rendering of the full sensory input .

To disentangle two latent representations, an object’s, and that of its dynamics, this paper introduces Kalman variational auto-encoders (KVAEs), a model that separates an intuition of dynamics from an object recognition network (section 3). At each time step tt, a variational auto-encoder compresses high-dimensional visual stimuli xt\mathbf{x}_{t} into latent encodings at\mathbf{a}_{t}. The temporal dynamics in the learned at\mathbf{a}_{t}-manifold are modelled with a linear Gaussian state space model that is adapted to handle complex dynamics (despite the linear relations among its states zt\mathbf{z}_{t}). The parameters of the state space model are adapted at each time step, and non-linearly depend on past at\mathbf{a}_{t}’s via a recurrent neural network. Exact posterior inference for the linear Gaussian state space model can be preformed with the Kalman filtering and smoothing algorithms, and is used for imputing missing data, for instance when we imagine the trajectory of a bouncing ball after observing it in initial and final video frames (section 4). The separation between recognition and dynamics model allows for missing data imputation to be done via a combination of the latent states zt\mathbf{z}_{t} of the model and its encodings at\mathbf{a}_{t} only, without having to forward-sample high-dimensional images xt\mathbf{x}_{t} in an autoregressive way. KVAEs are tested on videos of a variety of simulated physical systems in section 5: from raw visual stimuli, it “end-to-end” learns the interplay between the recognition and dynamics components. As KVAEs can do smoothing, they outperform an array of methods in generative and missing data imputation tasks (section 5).

Background

Linear Gaussian state space models (LGSSMs) are widely used to model sequences of vectors a=a1:T=[a1,..,aT]\mathbf{a}=\mathbf{a}_{1:T}=[\mathbf{a}_{1},..,\mathbf{a}_{T}]. LGSSMs model temporal correlations through a first-order Markov process on latent states z=[z1,..,zT]\mathbf{z}=[\mathbf{z}_{1},..,\mathbf{z}_{T}], which are potentially further controlled with external inputs u=[u1,..,uT]\mathbf{u}=[\mathbf{u}_{1},..,\mathbf{u}_{T}], through the Gaussian distributions

Matrices γt=[At,Bt,Ct]\gamma_{t}=[\mathbf{A}_{t},\mathbf{B}_{t},\mathbf{C}_{t}] are the state transition, control and emission matrices at time tt. Q\mathbf{Q} and R\mathbf{R} are the covariance matrices of the process and measurement noise respectively. With a starting state \mathbf{z}_{1}\sim\mathcal{N}(\mathbf{z}_{1};\mathbf{0},\text{\boldmath\Sigma}), the joint probability distribution of the LGSSM is given by

where γ=[γ1,..,γT]\gamma=[\gamma_{1},..,\gamma_{T}]. LGSSMs have very appealing properties that we wish to exploit: the filtered and smoothed posteriors p(zt∣a1:t,u1:t)p(\mathbf{z}_{t}|\mathbf{a}_{1:t},\mathbf{u}_{1:t}) and p(zt∣a,u)p(\mathbf{z}_{t}|\mathbf{a},\mathbf{u}) can be computed exactly with the classical Kalman filter and smoother algorithms, and provide a natural way to handle missing data.

Variational auto-encoders.

A variational auto-encoder (VAE) defines a deep generative model pθ(xt,at)=pθ(xt∣at)p(at)p_{\theta}(\mathbf{x}_{t},\mathbf{a}_{t})=p_{\theta}(\mathbf{x}_{t}|\mathbf{a}_{t})p(\mathbf{a}_{t}) for data xt\mathbf{x}_{t} by introducing a latent encoding at\mathbf{a}_{t}. Given a likelihood pθ(xt∣at)p_{\theta}(\mathbf{x}_{t}|\mathbf{a}_{t}) and a typically Gaussian prior p(at)p(\mathbf{a}_{t}), the posterior pθ(at∣xt)p_{\theta}(\mathbf{a}_{t}|\mathbf{x}_{t}) represents a stochastic map from xt\mathbf{x}_{t} to at\mathbf{a}_{t}’s manifold. As this posterior is commonly analytically intractable, VAEs approximate it with a variational distribution qϕ(at∣xt)q_{\phi}(\mathbf{a}_{t}|\mathbf{x}_{t}) that is parameterized by ϕ\phi. The approximation qϕq_{\phi} is commonly called the recognition, encoding, or inference network.

Kalman Variational Auto-Encoders

The useful information that describes the movement and interplay of objects in a video typically lies in a manifold that has a smaller dimension than the number of pixels in each frame. In a video of a ball bouncing in a box, like Atari’s game Pong, one could define a one-to-one mapping from each of the high-dimensional frames x=[x1,..,xT]\mathbf{x}=[\mathbf{x}_{1},..,\mathbf{x}_{T}] into a two-dimensional latent space that represents the position of the ball on the screen. If the position was known for consecutive time steps, for a set of videos, we could learn the temporal dynamics that govern the environment. From a few new positions one might then infer where the ball will be on the screen in the future, and then imagine the environment with the ball in that position.

The Kalman variational auto-encoder (KVAE) is based on the notion described above. To disentangle recognition and spatial representation, a sensory input xt\mathbf{x}_{t} is mapped to at\mathbf{a}_{t} (VAE), a variable on a low-dimensional manifold that encodes an object’s position and other visual properties. In turn, at\mathbf{a}_{t} is used as a pseudo-observation for the dynamics model (LGSSM). xt\mathbf{x}_{t} represents a frame of a videoWhile our main focus in this paper are videos, the same ideas could be applied more in general to any sequence of high dimensional data. x=[x1,..,xT]\mathbf{x}=[\mathbf{x}_{1},..,\mathbf{x}_{T}] of length TT. Each frame is encoded into a point at\mathbf{a}_{t} on a low-dimensional manifold, so that the KVAE contains TT separate VAEs that share the same decoder pθ(xt∣at)p_{\theta}(\mathbf{x}_{t}|\mathbf{a}_{t}) and encoder qϕ(at∣xt)q_{\phi}(\mathbf{a}_{t}|\mathbf{x}_{t}), and depend on each other through a time-dependent prior over a=[a1,..,aT]\mathbf{a}=[\mathbf{a}_{1},..,\mathbf{a}_{T}]. This is illustrated in figure 1.

We assume that a\mathbf{a} acts as a latent representation of the whole video, so that the generative model of a sequence factorizes as pθ(x∣a)=∏t=1Tpθ(xt∣at)p_{\theta}(\mathbf{x}|\mathbf{a})=\prod_{t=1}^{T}p_{\theta}(\mathbf{x}_{t}|\mathbf{a}_{t}). In this paper pθ(xt∣at)p_{\theta}(\mathbf{x}_{t}|\mathbf{a}_{t}) is a deep neural network parameterized by θ\theta, that emits either a factorized Gaussian or Bernoulli probability vector depending on the data type of xt\mathbf{x}_{t}. We model a\mathbf{a} with a LGSSM, and following (2), its prior distribution is

so that the joint density for the KVAE factorizes as p(x,a,z∣u)=pθ(x∣a) pγ(a∣z) pγ(z∣u)p(\mathbf{x},\mathbf{a},\mathbf{z}|\mathbf{u})=p_{\theta}(\mathbf{x}|\mathbf{a})\,p_{\gamma}(\mathbf{a}|\mathbf{z})\,p_{\gamma}(\mathbf{z}|\mathbf{u}). A LGSSM forms a convenient back-bone to a model, as the filtered and smoothed distributions pγ(zt∣a1:t,u1:t)p_{\gamma}(\mathbf{z}_{t}|\mathbf{a}_{1:t},\mathbf{u}_{1:t}) and pγ(zt∣a,u)p_{\gamma}(\mathbf{z}_{t}|\mathbf{a},\mathbf{u}) can be obtained exactly. Temporal reasoning can be done in the latent space of zt\mathbf{z}_{t}’s and via the latent encodings a\mathbf{a}, and we can do long-term predictions without having to auto-regressively generate high-dimensional images xt\mathbf{x}_{t}. Given a few frames, and hence their encodings, one could “remain in latent space” and use the smoothed distributions to impute missing frames. Another advantage of using a\mathbf{a} to separate the dynamics model from x\mathbf{x} can be seen by considering the emission matrix Ct\mathbf{C}_{t}. Inference in the LGSSM requires matrix inverses, and using it as a model for the prior dynamics of at\mathbf{a}_{t} allows the size of Ct\mathbf{C}_{t} to remain small, and not scale with the number of pixels in xt\mathbf{x}_{t}. While the LGSSM’s process and measurement noise in (1) are typically formulated with full covariance matrices , we will consider them as isotropic in a KVAE, as at\mathbf{a}_{t} act as a prior in a generative model that includes these extra degrees of freedom.

What happens when a ball bounces against a wall, and the dynamics on at\mathbf{a}_{t} are not linear any more? Can we still retain a LGSSM backbone? We will incorporate nonlinearities into the LGSSM by regulating γt\gamma_{t} from outside the exact forward-backward inference chain. We revisit this central idea at length in section 3.3.

2 Learning and inference for the KVAE

We learn θ\theta and γ\gamma from a set of example sequences {x(n)}\{\mathbf{x}^{(n)}\} by maximizing the sum of their respective log likelihoods L=∑nlog⁡pθγ(x(n)∣u(n))\mathcal{L}=\sum_{n}\log p_{\theta\gamma}(\mathbf{x}^{(n)}|\mathbf{u}^{(n)}) as a function of θ\theta and γ\gamma. For simplicity in the exposition we restrict our discussion below to one sequence, and omit the sequence index nn. The log likelihood or evidence is an intractable average over all plausible settings of a\mathbf{a} and z\mathbf{z}, and exists as the denominator in Bayes’ theorem when inferring the posterior p(a,z∣x,u)p(\mathbf{a},\mathbf{z}|\mathbf{x},\mathbf{u}). A more tractable approach to both learning and inference is to introduce a variational distribution q(a,z∣x,u)q(\mathbf{a},\mathbf{z}|\mathbf{x},\mathbf{u}) that approximates the posterior. The evidence lower bound (ELBO) F\mathcal{F} is

and a sum of F\mathcal{F}’s is maximized instead of a sum of log likelihoods. The variational distribution qq depends on ϕ\phi, but for the bound to be tight we should specify qq to be equal to the posterior distribution that only depends on θ\theta and γ\gamma. Towards this aim we structure qq so that it incorporates the exact conditional posterior pγ(z∣a,u)p_{\gamma}(\mathbf{z}|\mathbf{a},\mathbf{u}), that we obtain with Kalman smoothing, as a factor of γ\gamma:

The benefit of the LGSSM backbone is now apparent. We use a “recognition model” to encode each xt\mathbf{x}_{t} using a non-linear function, after which exact smoothing is possible. In this paper qϕ(at∣xt)q_{\phi}(\mathbf{a}_{t}|\mathbf{x}_{t}) is a deep neural network that maps xt\mathbf{x}_{t} to the mean and the diagonal covariance of a Gaussian distribution. As explained in section 4, this factorization allows us to deal with missing data in a principled way. Using (5), the ELBO in (4) becomes

The lower bound in (6) can be estimated using Monte Carlo integration with samples {a~(i),z~(i)}i=1I\{\widetilde{\mathbf{a}}^{(i)},\widetilde{\mathbf{z}}^{(i)}\}_{i=1}^{I} drawn from qq,

Note that the ratio pγ(a~(i),z~(i)∣u)/pγ(z~(i)∣a~(i),u)p_{\gamma}(\widetilde{\mathbf{a}}^{(i)},\widetilde{\mathbf{z}}^{(i)}|\mathbf{u})/p_{\gamma}(\widetilde{\mathbf{z}}^{(i)}|\widetilde{\mathbf{a}}^{(i)},\mathbf{u}) in (7) gives pγ(a~(i)∣u)p_{\gamma}(\widetilde{\mathbf{a}}^{(i)}|\mathbf{u}), but the formulation with {z~(i)}\{\widetilde{\mathbf{z}}^{(i)}\} allows stochastic gradients on γ\gamma to also be computed. A sample from qq can be obtained by first sampling a~∼qϕ(a∣x)\widetilde{\mathbf{a}}\sim q_{\phi}(\mathbf{a}|\mathbf{x}), and using a~\widetilde{\mathbf{a}} as an observation for the LGSSM. The posterior pγ(z∣a~,u)p_{\gamma}(\mathbf{z}|\widetilde{\mathbf{a}},\mathbf{u}) can be tractably obtained with a Kalman smoother, and a sample z~∼pγ(z∣a~,u)\widetilde{\mathbf{z}}\sim p_{\gamma}(\mathbf{z}|\widetilde{\mathbf{a}},\mathbf{u}) obtained from it. Parameter learning is done by jointly updating θ\theta, ϕ\phi, and γ\gamma by maximising the ELBO on L\mathcal{L}, which decomposes as a sum of ELBOs in (6), using stochastic gradient ascent and a single sample to approximate the intractable expectations.

3 Dynamics parameter network

The LGSSM provides a tractable way to structure pγ(z∣a,u)p_{\gamma}(\mathbf{z}|\mathbf{a},\mathbf{u}) into the variational approximation in (5). However, even in the simple case of a ball bouncing against a wall, the dynamics on at\mathbf{a}_{t} are not linear anymore. We can deal with these situations while preserving the linear dependency between consecutive states in the LGSSM, by non-linearly changing the parameters γt\gamma_{t} of the model over time as a function of the latent encodings up to time t−1t-1 (so that we can still define a generative model). Smoothing is still possible as the state transition matrix At\mathbf{A}_{t} and others in γt\gamma_{t} do not have to be constant in order to obtain the exact posterior pγ(zt∣a,u)p_{\gamma}(\mathbf{z}_{t}|\mathbf{a},\mathbf{u}).

Recall that γt\gamma_{t} describes how the latent state zt−1\mathbf{z}_{t-1} changes from time t−1t-1 to time tt. In the more general setting, the changes in dynamics at time tt may depend on the history of the system, encoded in a1:t−1\mathbf{a}_{1:t-1} and possibly a starting code a0\mathbf{a}_{0} that can be learned from data. If, for instance, we see the ball colliding with a wall at time t−1t-1, then we know that it will bounce at time tt and change direction. We then let γt\gamma_{t} be a learnable function of a0:t−1\mathbf{a}_{0:t-1}, so that the prior in (2) becomes

During inference, after all the frames are encoded in a\mathbf{a}, the dynamics parameter network returns γ=γ(a)\gamma=\gamma(\mathbf{a}), the parameters of the LGSSM at all time steps. We can now use the Kalman smoothing algorithm to find the exact conditional posterior over z\mathbf{z}, that will be used when computing the gradients of the ELBO.

We globally learn KK basic state transition, control and emission matrices A(k)\mathbf{A}^{(k)}, B(k)\mathbf{B}^{(k)} and C(k)\mathbf{C}^{(k)}, and interpolate them based on information from the VAE encodings. The weighted sum can be interpreted as a soft mixture of KK different LGSSMs whose time-invariant matrices are combined using the time-varying weights \text{\boldmath\alpha}_{t}. In practice, each of the KK sets {A(k),B(k),C(k)}\{\mathbf{A}^{(k)},\mathbf{B}^{(k)},\mathbf{C}^{(k)}\} models different dynamics, that will dominate when the corresponding αt(k)\alpha_{t}^{(k)} is high. The dynamics parameter network resembles the locally-linear transitions of ; see section 6 for an in depth discussion on the differences.

Missing data imputation

Let xobs\mathbf{x}_{\mathsf{obs}} be an observed subset of frames in a video sequence, for instance depicting the initial movement and final positions of a ball in a scene. From its start and end, can we imagine how the ball reaches its final position? Autoregressive models like recurrent neural networks can only forward-generate xt\mathbf{x}_{t} frame by frame, and cannot make use of the information coming from the final frames in the sequence. To impute the unobserved frames xun\mathbf{x}_{\mathsf{un}} in the middle of the sequence, we need to do inference, not prediction.

The KVAE exploits the smoothing abilities of its LGSSM to use both the information from the past and the future when imputing missing data. In general, if x={xobs,xun}\mathbf{x}=\{\mathbf{x}_{\mathsf{obs}},\mathbf{x}_{\mathsf{un}}\}, the unobserved frames in xun\mathbf{x}_{\mathsf{un}} could also appear at non-contiguous time steps, e.g. missing at random. Data can be imputed by sampling from the joint density p(aun,aobs,z∣xobs,u)p(\mathbf{a}_{\mathsf{un}},\mathbf{a}_{\mathsf{obs}},\mathbf{z}|\mathbf{x}_{\mathsf{obs}},\mathbf{u}), and then generating xun\mathbf{x}_{\mathsf{un}} from aun\mathbf{a}_{\mathsf{un}}. We factorize this distribution as

and we sample from it with ancestral sampling starting from xobs\mathbf{x}_{\mathsf{obs}}. Reading (10) from right to left, a sample from p(aobs∣xobs)p(\mathbf{a}_{\mathsf{obs}}|\mathbf{x}_{\mathsf{obs}}) can be approximated with the variational distribution qϕ(aobs∣xobs)q_{\phi}(\mathbf{a}_{\mathsf{obs}}|\mathbf{x}_{\mathsf{obs}}). Then, if γ\gamma is fully known, pγ(z∣aobs,u)p_{\gamma}(\mathbf{z}|\mathbf{a}_{\mathsf{obs}},\mathbf{u}) is computed with an extension to the Kalman smoothing algorithm to sequences with missing data, after which samples from pγ(aun∣z)p_{\gamma}(\mathbf{a}_{\mathsf{un}}|\mathbf{z}) could be readily drawn.

However, when doing missing data imputation the parameters γ\gamma of the LGSSM are not known at all time steps. In the KVAE, each γt\gamma_{t} depends on all the previous encoded states, including aun\mathbf{a}_{\mathsf{un}}, and these need to be estimated before γ\gamma can be computed. In this paper we recursively estimate γ\gamma in the following way. Assume that x1:t−1\mathbf{x}_{1:t-1} is known, but not xt\mathbf{x}_{t}. We sample a1:t−1\mathbf{a}_{1:t-1} from qϕ(a1:t−1∣x1:t−1)q_{\phi}(\mathbf{a}_{1:t-1}|\mathbf{x}_{1:t-1}) using the VAE, and use it to compute γ1:t\gamma_{1:t}. The computation of γt+1\gamma_{t+1} depends on at\mathbf{a}_{t}, which is missing, and an estimate a^t\widehat{\mathbf{a}}_{t} will be used. Such an estimate can be arrived at in two steps. The filtered posterior distribution pγ(zt−1∣a1:t−1,u1:t−1)p_{\gamma}(\mathbf{z}_{t-1}|\mathbf{a}_{1:t-1},\mathbf{u}_{1:t-1}) can be computed as it depends only on γ1:t−1\gamma_{1:t-1}, and from it, we sample

and sample a^t\widehat{\mathbf{a}}_{t} from the predictive distribution of at\mathbf{a}_{t},

The parameters of the LGSSM at time t+1t+1 are then estimated as γt+1([a0:t−1,a^t])\gamma_{t+1}([\mathbf{a}_{0:t-1},\widehat{\mathbf{a}}_{t}]). The same procedure is repeated at the next time step if xt+1\mathbf{x}_{t+1} is missing, otherwise at+1\mathbf{a}_{t+1} is drawn from the VAE. After the forward pass through the sequence, where we estimate γ\gamma and compute the filtered posterior for z\mathbf{z}, the Kalman smoother’s backwards pass computes the smoothed posterior. While the smoothed posterior distribution is not exact, as it relies on the estimate of γ\gamma obtained during the forward pass, it improves data imputation by using information coming from the whole sequence; see section 5 for an experimental illustration.

Experiments

We motivated the KVAE with an example of a bouncing ball, and use it here to demonstrate the model’s ability to separately learn a recognition and dynamics model from video, and use it to impute missing data. To draw a comparison with deep variational Bayes filters (DVBFs) , we apply the KVAE to ’s pendulum example. We further apply the model to a number of environments with different properties to demonstrate its generalizability. All models are trained end-to-end with stochastic gradient descent. Using the control input ut\mathbf{u}_{t} in (1) we can inform the model of known quantities such as external forces, as will be done in the pendulum experiment. In all the other experiments, we omit such information and train the models fully unsupervised from the videos only. Further implementation details can be found in the supplementary material (appendix A) and in the Tensorflow code released at github.com/simonkamronn/kvae.

We compare the generation and imputation performance of the KVAE with two recurrent neural network (RNN) models that are based on the same auto-encoding (AE) architecture as the KVAE and are modifications of methods from the literature to be better suited to the bouncing ball experiments.We also experimented with the SRNN model from as it can do smoothing. However, the model is probably too complex for the task in hand, and we could not make it learn good dynamics. In the AE-RNN, inspired by the architecture from , a pretrained convolutional auto-encoder, identical to the one used for the KVAE, feeds the encodings to an LSTM network . During training the LSTM predicts the next encoding in the sequence and during generation we use the previous output as input to the current step. For data imputation the LSTM either receives the previous output or, if available, the encoding of the observed frame (similarly to filtering in the KVAE). The VAE-RNN is identical to the AE-RNN except that uses a VAE instead of an AE, similarly to the model from .

Figure 3(a) shows how well missing frames are imputed in terms of the average fraction of incorrectly guessed pixels. In it, the first 4 frames are observed (to initialize the models) after which the next 16 frames are dropped at random with varying probabilities. We then impute the missing frames by doing filtering and smoothing with the KVAE. We see in figure 3(a) that it is beneficial to utilize information from the whole sequence (even the future observed frames), and a KVAE with smoothing outperforms all competing methods. Notice that dropout probability 1 corresponds to pure generation from the models. Figure 3(b) repeats this experiment, but makes it more challenging by removing an increasing number of consecutive frames from the middle of the sequence (T=20T=20). In this case the ability to encode information coming from the future into the posterior distribution is highly beneficial, and smoothing imputes frames much better than the other methods. Figure 3(c) graphically illustrates figure 3(b). We plot three trajectories over at\mathbf{a}_{t}-encodings. The generated trajectories were obtained after initializing the KVAE model with 4 initial frames, while the smoothed trajectories also incorporated encodings from the last 4 frames of the sequence. The encoded trajectories were obtained with no missing data, and are therefore considered as ground truth. In the first three plots in figure 3(c), we see that the backwards recursion of the Kalman smoother corrects the trajectory obtained with generation in the forward pass. However, in the fourth plot, the poor trajectory that is obtained during the forward generation step, makes smoothing unable to follow the ground truth.

The smoothing capabilities of KVAEs make it also possible to train it with up to 40% of missing data with minor losses in performance (appendix C in the supplementary material). Links to videos of the imputation results and long-term generation from the models can be found in appendix B and at sites.google.com/view/kvae.

In our experiments the dynamics parameter network \text{\boldmath\alpha}_{t}=\text{\boldmath\alpha}_{t}(\mathbf{a}_{0:t-1}) is an LSTM network, but we could also parameterize it with any differentiable function of a0:t−1\mathbf{a}_{0:t-1} (see appendix D in the supplementary material for a comparison of various architectures). When using a multi-layer perceptron (MLP) that depends on the previous encoding as mixture network, i.e. \text{\boldmath\alpha}_{t}=\text{\boldmath\alpha}_{t}(\mathbf{a}_{t-1}), figure 4 illustrates how the network chooses the mixture of learned dynamics. We see that the model has correctly learned to choose a transition that maintains a constant velocity in the center (k=1k=1), reverses the horizontal velocity when in proximity of the left and right wall (k=2k=2), the reverses the vertical velocity when close to the top and bottom (k=3k=3).

2 Pendulum experiment

3 Other environments

To test how well the KVAE adapts to different environments, we trained it end-to-end on videos of (i) a ball bouncing between walls that form an irregular polygon, (ii) a ball bouncing in a box and subject to gravity, (iii) a Pong-like environment where the paddles follow the vertical position of the ball to make it stay in the frame at all times. Figure 5 shows that the KVAE learns the dynamics of all three environments, and generates realistic-looking trajectories. We repeat the imputation experiments of figures 3(a) and 3(b) for these environments in the supplementary material (appendix E), where we see that KVAEs outperform alternative models.

Related work

Recent progress in unsupervised learning of high dimensional sequences is found in a plethora of both deterministic and probabilistic generative models. The VAE framework is a common work-horse in the stable of probabilistic inference methods, and it is extended to the temporal setting by . In particular, deep neural networks can parameterize the transition and emission distributions of different variants of deep state-space models . In these extensions, inference networks define a variational approximation to the intractable posterior distribution of the latent states at each time step. For the tasks in section 5, it is preferable to use the KVAE’s simpler temporal model with an exact (conditional) posterior distribution than a highly non-linear model where the posterior needs to be approximated. A different combination of VAEs and probabilistic graphical models has been explored in , which defines a general class of models where inference is performed with message passing algorithms that use deep neural networks to map the observations to conjugate graphical model potentials.

In classical non-linear extensions of the LGSSM like the extended Kalman filter and in the locally-linear dynamics of , the transition matrices at time tt have a non-linear dependence on zt−1\mathbf{z}_{t-1}. The KVAE’s approach is different: by introducing the latent encodings at\mathbf{a}_{t} and making γt\gamma_{t} depend on a1:t−1\mathbf{a}_{1:t-1}, the linear dependency between consecutive states of z\mathbf{z} is preserved, so that the exact smoothed posterior can be computed given a\mathbf{a}, and used to perform missing data imputation. LGSSM with dynamic parameterization have been used for large-scale demand forecasting in . introduces recurrent switching linear dynamical systems, that combine deep learning techniques and switching Kalman filters to model low-dimensional time series. introduces a discriminative approach to estimate the low-dimensional state of a LGSSM from input images. The resulting model is reminiscent of a KVAE with no decoding step, and is therefore not suited for unsupervised learning and video generation. Recent work in the non-sequential setting has focused on disentangling basic visual concepts in an image . models neural activity by finding a non-linear embedding of a neural time series into a LGSSM.

Great strides have been made in the reinforcement learning community to model how environments evolve in response to action . In similar spirit to this paper, extracts a latent representation from a PCA representation of the frames where controls can be applied. introduces action-conditional dynamics parameterized with LSTMs and, as for the KVAE, a computationally efficient procedure to make long term predictions without generating high dimensional images at each time step. As autoregressive models, develops a sequence to sequence model of video representations that uses LSTMs to define both the encoder and the decoder. develops an action-conditioned video prediction model of the motion of a robot arm using convolutional LSTMs that models the change in pixel values between two consecutive frames.

While the focus in this work is to define a generative model for high dimensional videos of simple physical systems, several recent works have combined physical models of the world with deep learning to learn the dynamics of objects in more complex but low-dimensional environments .

Conclusion

The KVAE, a model for unsupervised learning of high-dimensional videos, was introduced in this paper. It disentangles an object’s latent representation at\mathbf{a}_{t} from a latent state zt\mathbf{z}_{t} that describes its dynamics, and can be learned end-to-end from raw video. Because the exact (conditional) smoothed posterior distribution over the states of the LGSSM can be computed, one generally sees a marked improvement in inference and missing data imputation over methods that don’t have this property. A desirable property of disentangling the two latent representations is that temporal reasoning, and possibly planning, could be done in the latent space. As a proof of concept, we have been deliberate in focussing our exposition to videos of static worlds that contain a few moving objects, and leave extensions of the model to real world videos or sequences coming from an agent exploring its environment to future work.

Acknowledgements

We would like to thank Lars Kai Hansen for helpful discussions on the model design. Marco Fraccaro is supported by Microsoft Research through its PhD Scholarship Programme. We thank NVIDIA Corporation for the donation of TITAN X GPUs.

References

Appendix A Experimental details

We will describe here some of the most important experimental details. The rest of the details can be found in the code at github.com/simonkamronn/kvae.

All the videos were generated using the physics engine Pymunk. We generated 5000 videos for training and 1000 for testing.

Encoder/Decoder architecture for the KVAE.

As we only use image-based observations, the encoder is fixed to a three layer convolutional neural network with 32 units in each layer, kernel-size of 3x3, stride of 2, and ReLU activations. The decoder is an equally sized network using the Sub-Pixel procedure for deconvolution. In the pendulum experiment however we also test MLPs.

Optimization.

As optimizer we use ADAM with an initial learning rate of 0.007 and an exponential decay scheme with a rate of 0.85 every 20 epochs. Training one epoch takes 55 seconds on an NVIDIA Titan X and the model converges in roughly 80 epochs.

Training tricks for end-to-end learning.

The biggest challenge of this optimization problem is how to avoid poor local minima, for example where all the focus is given to the reconstruction term, at the expense of the prior dynamics given by the LGSSM. To achieve a quick convergence in all the experiments we found it helpful to

downweight the reconstruction term from of VAEs during training, that is scaled by 0.3. By doing this, we can in fact help the model to focus on learning the temporal dynamics.

learn for the first few epochs only the the VAE parameters θ\theta and ϕ\phi and the globally learned matrices A(k)\mathbf{A}^{(k)}, B(k)\mathbf{B}^{(k)} and C(k)\mathbf{C}^{(k)}, but not the parameters of the dynamics parameter network \text{\boldmath\alpha}_{t}(\mathbf{a}_{0:t-1}). After this phase, all parameters are learned jointly. This allows the model to first learn good VAE embeddings and the scale of the prior, and then learn how to utilize the KK different dynamics.

Choice of hyperparameters for the LGSSM.

Appendix B Videos

Videos are generated from all models by initializing with 4 frames and then sampling. The filtering and smoothing versions are allowed to observe part of the sequence depending on the masking scheme. All the filtering and smoothing videos are generated from sequences applied with a random mask with a masking probability of 80% (as in figure 3(a)) except for the videos with the suffix consecutive in which only the first and last 4 frames are observed (as in figure 3(b)). Only the KVAE models have smoothing videos. For the bouncing ball experiment (named box in the attached folder), we also show the videos from a model trained with 40% missing data.

In most videos the black ball is the ground truth, and the red is the one generated from the model, except for the ones marked long_generation in which the true sequence is not shown.

Videos are available from Google Drive and the website sites.google.com/view/kvae.

Appendix C Training with missing data.

The smoothed posterior described in section 4 can also be used to train the KVAE with missing data. In this case, we only need to modify the ELBO by masking the contribution of the missing data points in the joint probability distribution and variational approximation:

where It\mathcal{I}_{t} is 0 if the data point is missing, 1 otherwise. Figure 6 illustrates a slight degradation in performance when training with respectively 30% and 40% missing data but, remarkably, the accuracy is still better when using smoothing in these conditions than with filtering with all training data available.

Appendix D Dynamics parameter network architecture

As the α\alpha-network governs the non-linear dynamics, it has a significant impact on the modelling capabilities. Here we list the architectural choices considered:

Recurrent Neural Networks with LSTM units.

’First in, first out memory’ (FIFO) MLP with access to 5 time steps.

In all cases, we can also model α\alpha as an (approximate) discrete random variable using the the Concrete distribution . In this case we can recover an approximation to the switching Kalman filter.

In figure 7 the different choices are tested against each other on the bouncing ball data. In this case all the alternative choices result in poorer performances than the LSTM chosen for all the other experiments. We believe that LSTMs are able to better model the discretization errors coming from the collisions and the 32x32 rendering of the trajectories computed by the physics engine.

Appendix E Imputation in all environments