Learning A Physical Long-term Predictor

Sebastien Ehrhardt, Aron Monszpart, Niloy J. Mitra, Andrea Vedaldi

Introduction

Most natural intelligences possess a remarkably accurate understanding of some of the physical properties of the world, as needed to navigate, prey, burrow, or perform any number of other ecologically-motivated activities. In particular, evolutionary pressure has caused most animals to develop the capability to perform fast and accurate predictions of mechanical phenomena. However, the nature of these mental models remains unclear and is being actively investigated (Hamrick et al., 2016).

Humans have developed an excellent formal understanding of physics; for example, at the level of granularity at which animals operate, mechanics is nearly perfectly described by Newton’s laws. However, while these laws are simple, their application to the description of a natural phenomena is anything but trivial. In fact, before such laws can be applied, a physical scenario needs first to be abstracted (e.g., by segmenting the world into rigid objects, describing those by mass volumes, estimating physical parameters). Then, except for the most trivial problems, predictions require the numerical integration of very complex systems of equations. It is therefore unclear whether animals would perform mechanical predictions in this manner.

In this paper, we investigate how an accurate understanding of mechanical phenomena can emerge in artificial systems, mimicking natural intelligence. Inspired by a number of recent works, we look in particular at how deep neural networks can be used to perform mechanical predictions in simple physical scenarios (Fig. 1). Among such prior works, by far the most popular approach is to use neural networks (Wu et al., 2016) to extract from sensory data local predictions of physical parameters, such as mass, velocity, or acceleration, that are then integrated by an external mechanism such as a physical simulator to obtain long term predictions. In other words, these approaches look at how an AI can abstract sensory data in physical parameters, but not how it can integrate such parameters over longer times. Further, such an approach assumes access to a simulator that accurately abstracts the physical world with appropriate Newtonian equations. Other attempts have also tried to replace the physical engine with a neural network (Battaglia et al., 2016) but did not really attempt to observe the physical world and deduce properties from it but rather to integrate the physical equations.

By contrast, in this paper we perform end-to-end prediction of mechanical phenomena with a single neural network, implicitly combining prediction and integration of physical parameters from sensory data. In other words, while most other approaches predict instantaneous parameters such as mass and velocity from a few video frames, our model directly performs long-term predictions of physical parameters such as position well beyond the initial observation interval. Thus, as our main contribution, we propose to do so by learning an internal representation of a physical scenario which is induced by the observation of a few images and then is evolved in time by a recurrent architecture.

One of the challenges of extrapolating physical measurements is that the state of a physical system can be determined only up to a certain accuracy, and such uncertainty rapidly accumulates over time. Since no predictor can be expected to deterministically predict the future, predictors are best formulated as estimating distribution of possible outcomes. In this manner, predictors can properly account for uncertainty in the physical parameters, approximations in the model, or other limitations. Thus, our second contribution is to allow the neural network to explicitly model uncertainty by predicting distributions over physical parameters rather than their value directly.

In our experiments, we let convolutional neural networks choose their own internal representation of physical laws. A soft prior is that convolutional architectures encourage learning structures that are local and spatially homogeneous, similar to the applicable physical laws that are also local and homogeneous. However, the network is never explicitly told what physical laws are. Our third contribution, therefore, is to look at whether such networks can learn physical properties that generalise beyond regimes observed during training.

The relation of our work to the literature is discussed in section 2. The detailed structure of the proposed neural networks is given and motivated in section 3. These networks are extensively evaluated on a large dataset of simulated physical experiments in section 4. A summary of our finding can be found in section 5.

Related Work

In this work we address the problem of long-term prediction of object positions in a physical environment without voluntary perturbation with an implicit learning of physical laws. Our work is closely related to a range of recent works in the machine learning community.

To the best of our knowledge (Battaglia et al., 2013) was the first approach to tackle intuitive physics with the aim to answer a set of intuitive questions (e.g., will it fall?) using physical simulations. Their simulations, however used a sophisticated physics engine that incorporates prior knowledge about Newtonian physical equations. More recently (Mottaghi et al., 2016) also used static images and a graphic rendering engine (Blender) to predict movements and directions of forces from a single RGB image. Motivated by the recent success of deep learning for image processing (e.g., (Krizhevsky et al., 2012; He et al., 2016)) they used a convolutional architecture to understand dynamics and forces acting behind the scenes from a static image and produced a “most likely motion” rendered from a graphics engine. In a different framework (Lerer et al., 2016) and (Li et al., 2016) also used the power of deep learning to extract an abstract representation of the concept of stability of block towers purely from images. These approaches successfully demonstrated that not only was a network able to accurately predict the stability of the block tower but in addition, it could identify the source of the instability. Other approaches such as (Agrawal et al., 2016) or (Denil et al., 2016) also attempted to learn intuitive physics of objects through manipulation. None of these approaches did, however, attempt to precisely model the evolution of the physical world.

Learning dynamics.

Learning the evolution of an object’s position also implies to learn about the object’s dynamics regardless of any physical equations. While most successful techniques used LSTM-s (Hochreiter & Schmidhuber, 1997), recent approaches show that propagation can also be done using a single cross-convolution kernel. The idea was further developed in (Xue et al., 2016) in order to generate a next possible image frame from a single static input image. The concept has been shown to have promising performance regarding longer term predictions on the moving MNIST dataset in (De Brabandere et al., 2016). The work of (Ondruska & Posner, 2016) also shows that an internal hidden state can be propagated through time using a simple deep recurrent architecture. These results motivated us to propagate tensor based state representations instead of a single vector representation using a series of convolutions. In the future we also aim to experiment with approaches inspired by (Xue et al., 2016).

Learning physics.

Works of (Wu et al., 2015) and its extension (Wu et al., 2016) propose methods to learn physical properties of scenes and objects. However in (Wu et al., 2015) the MCMC sampling based approach assumes a complete knowledge of the physical equations to estimate the correct physical parameters. In (Wu et al., 2016) deep learning has been used more extensively to replace the MCMC based sampling but this work also employs an explicit encoding and computation of physical laws to regress the output of their tracker. (Stewart & Ermon, 2016) also used physical laws to predict the movement of a pillow from unlabelled data though their approach was only applied to a fixed number of frames.

In another related approach (Fragkiadaki et al., 2015) attempted to build an internal representation of the physical world. Using a billiard board with an external simulator they built a network which observing four frames and an applied force, was able to predict the 20 next object velocities. Generalization in this work was made using an LSTM in the intermediate representations. The process can be interpreted as iterative since frame generation is made to provide new inputs to the network. This can also be seen as a regularization process to avoid the internal representation of dynamics to decay over time which is different to our approach in which we try to build a stronger internal representation that will attempt to avoid such decay.

Other research attempted to abstract the physics engine enforcing the laws of physics as neural network models. (Battaglia et al., 2016) and (Chang et al., 2016) were able to produce accurate estimations of the next state of the world. Although the results look plausible and promising, long term predictions are still an issue in such frameworks. Note, that their process is an iterative one as opposed to ours, which propagates an internal state of the world through time.

Approximate physics with realistic output.

Other approaches also focused on learning the production of realistic future scenarios ((Tompson et al., 2016) and (Jeong et al., 2015)), or inferring collision parameters from monocular videos (Monszpart et al., 2016). In these approaches the authors used physics based losses to produce visually plausible yet erroneous results. They however show promising results and constructed new losses taking into account additional physical parameters other than velocity.

Mechanics Networks

In this section, we introduce a number of neural network models that can predict the behaviour of a simple mechanical system. We start by describing the physical setup and then we introduce the proposed network architectures.

We set the camera to be located above the plane, at height h>0h>0, centered at point (0,0,h)(0,0,h), and looking downwards along the Z-direction (0,0,−1)(0,0,-1). The camera axes are aligned to the world axes and the camera projection model is orthographic; in this setting, a world point p\mathbf{p} simply projects to pixel (px,py)(p_{x},p_{y}) in the image. Note that hh does not have an influence on the generated image and can be dropped.

The sliding object is a cube with center of mass q(t)=(qx(t),qy(t),qz(t))\mathbf{q}(t)=(q_{x}(t),q_{y}(t),q_{z}(t)) at time tt, which projects to pixel y(t)=α(qx(t),qy(t))+β\mathbf{y}(t)=\alpha(q_{x}(t),q_{y}(t))+\beta in the image. Here we consider Hi×Wi=128×128H_{i}\times W_{i}=128\times 128 images with pixels of size α=1\alpha=1 and offset β=(−64,−64)\beta=(-64,-64). Initially, the cube is placed at rest on top of the plane at a random location (qx0,qy0)(q_{x}^{0},q_{y}^{0}), and then starts to slide under the effect of gravity. The cube motion is also affected by friction.

An experiment instance is a tuple α=(qx0,qy0,n,ρ)\alpha=(q_{x}^{0},q_{y}^{0},\mathbf{n},\rho), consisting of the values of the initial object position, the plane normal, and the friction coefficient or distribution. The inclination n\mathbf{n} is arbitrary (within limits), such that the object can slide in any direction. These parameters, as well as many other constant parameters described in section 4.1, are passed to a physical simulator and rendered to simulate the experiment, resulting in a sequence of color images XTα=(xα(0),xα(1),…,xα(T−1))\mathcal{X}_{T}^{\alpha}=(\mathbf{x}^{\alpha}(0),\mathbf{x}^{\alpha}(1),\dots,\mathbf{x}^{\alpha}(T-1)). The simulator also produces the ground-truth center of mass projections YTα=(yα(0),…,yα(T−1))\mathcal{Y}_{T}^{\alpha}=(\mathbf{y}^{\alpha}(0),\dots,\mathbf{y}^{\alpha}(T-1)). Note that, for the purpose of learning predictors, physics needs not to be specified further; while in fact a complete set of parameters are required to run the physical simulation, the predictor learns automatically to extract the required information from the observed images.

2 Neural network architectures

We focus on long-term predictors Φ:XT0↦YT\Phi:\mathcal{X}_{T_{0}}\mapsto\mathcal{Y}_{T} that take as input the first T0=4T_{0}=4 frames XT0\mathcal{X}_{T_{0}} of a video sequence XT\mathcal{X}_{T} and produce as output a long-term estimate YT\mathcal{Y}_{T} of the location of the object’s center of mass at times t=0,1,…,Tt=0,1,\dots,T, where T≫T0T\gg T_{0}.

Our method comprises three building blocks (Fig. 2): a feature extractor, a propagation network and an estimation network. The core of our model is the internal representation of the physics, initialized by the feature extractor, updated by the propagation module, and decoded by the estimation module. We compare two representation types: a vector representation, in which each frame is encoded as CC-dimensional vector (or 1×1×C1\times 1\times C tensor), and a H×W×CH\times W\times C tensor representation. The importance of this difference is that the vector representation is spatially concentrated, whereas the tensor representation is spatially distributed.

Next, the three modules are discussed in detail.

(ii) Propagation network.

(iii) Estimation network.

As discussed above, however, it is preferable to predict the uncertainty of the estimate as well. While in some cases this cannot improve accuracy directly (i.e in the bivariate gaussian case), it is interesting to see if a network is able to develop an internal sense of prediction errors. Further, probabilistic modelling may help the network discount difficult-to-predict points during training, which may otherwise work as outliers negatively affecting training.

We propose to do so in two ways. In the first approach, we predict the mean and variance Y^T=(μ(t),Σ(t);t=0,…,T−1)\widehat{\mathcal{Y}}_{T}=(\mu(t),\Sigma(t);t=0,\dots,T-1) of a bivariate Gaussian distribution N(⋅;μ,Σ)\mathcal{N}(\cdot;\mu,\Sigma). The loss is the negative log-likelihood of the measured object locations:

In practice, the neural network estimates the two dimensional vector μ(t)\mu(t) as well as a three dimensional vector λ1(t),λ2(t),θ(t)\lambda_{1}(t),\lambda_{2}(t),\theta(t) with the first two being the eigenvalues of Σ(t)\Sigma(t), and the third entry being the angle of the rotation matrix in the decomposition Σ(t)=R(−θ(t))[λ1(t)00λ2(t)]R(θ(t))\Sigma(t)=R(-\theta(t))\begin{bmatrix}\lambda_{1}(t)&0\\ 0&\lambda_{2}(t)\end{bmatrix}R(\theta(t)). In order to ensure numerical stability, eigenvalues are constrained to be in the range [0.01…100][0.01\ldots 100] by setting them as the output of a scaled and translated sigmoid λi(t)=σλ,α(βi(t))\lambda_{i}(t)=\sigma_{\lambda,\alpha}(\beta_{i}(t)), where σλ,α(z)=λ/(1+exp⁡(−z))+α\sigma_{\lambda,\alpha}(z)=\lambda/(1+\exp(-z))+\alpha. For more details regarding the training procedure of this model please see section 4.3.

All the output predictions at time t+T0−1t+T_{0}-1 are extracted from the internal state StS_{t} by a single layer L(St)L(S_{t}). The layer LL is linear and fully-connected layer for L\mathcal{L} and Lnrm\mathcal{L}_{\text{nrm}}, and a deconvolutional layer similar to (Long et al., 2015) in the case of Lheat\mathcal{L}_{\text{heat}}. Outputs at times t=0,1,…,T0t=0,1,\dots,T_{0} are all predicted from S0S_{0} using an analogous but independently-trained fully-connected layer L′L^{\prime}. Overall, the output of the predictor is given by:

Experiments

Experiments consider three variants of the physical setup described in section 3.1, called Scenarios S0, S1, and S2. Different scenarios sample experiments α=(qx0,qy0,n,ρ)\alpha=(q_{x}^{0},q_{y}^{0},\mathbf{n},\rho) of increasing difficulty. The parameters of each scenario are summarised in Table 1 and described next. The plane normal n\mathbf{n} was obtained by rotating the ZZ axis around the XX and YY axis by random angles θx,θy\theta_{x},\theta_{y} (Scenario S0 uses a fixed inclination). For Scenarios S1 and S2, the Coulomb friction coefficient ρ\rho of the plane is homogeneous and sampled uniformly at random. For Scenario S2, the plane is split in 10×1010\times 10 patches, each with a random friction coefficient sampled independently. The friction upper bound was chosen so that the object always slides along the slope.

The plane was rendered as a black object so that no static visual cues allow deducing any of the physical parameters except the initial position of the cube; instead the predictor has to approximate physics as needed by observing the motion of the object during the first T0=4T_{0}=4 frames of each experiment.

Each experiment was run for at most 240 frames, or terminated early if the object left the field of view. In order to observe enough movement in each recorded sequence, the first 30 video frames of each video were removed, and the rest of the video was sub-sampled by a factor of 3. In practice, most experiments consist of 40-50 images.

The dataset contains 12,500 experiments for each Scenario, 70% of which are used for training, 15% for validation, and 15% for test.

The object’s starting position is initialized randomly using rejection sampling in such a way that (qx0,qy0)(q_{x}^{0},q_{y}^{0}) falls in the slope quadrant that contains the largest visible hh coordinate. This procedure generates samples that have most of their trajectory visible to the camera.

Rendering and physical simulation use Blender 2.77’s OpenGL renderer and the Blender Game Engine (relying on Bullet 2 as physics engine) respectively. The object is a cube of side 0.130.1^{3} Blender units with mass = 1. The simulation parameters are: max physics steps = 5, physics substeps = 3, max logic steps = 5, FPS = 120. Rendering used white environment lighting (energy = 0.7) and no other light sources. The object color was set to Lambertian red (RGB: 0.8, 0.04, 0.040.8,~{}0.04,~{}0.04) with no specular component. The slope is completely black, covering the whole field of view. The output images were stored as 128×128128\times 128 color JPG files. See Fig. 1 for an example initial setting from Scenario S2.

2 Baseline predictors

We compare the performance of our methods to two simple least squares baselines: Linear and Quadratic. In both cases we fit two least squares polynomials to the estimated screen-space coordinates of the first T=10T=10 frames. The polynomials are of first and second degree(s), respectively. We estimate the object’s position in this case by using the maximum location of the red channel of the input image. Note, that being able to observe the first 10 frames is a large advantage compared to the networks, which only see the first T0=4T_{0}=4 frames.

Physics simulator.

The SimNet baseline is used to evaluate the long term prediction ability of a neural network that has access to an explicit physics simulator, in a manner analogous to the work of (Wu et al., 2015). Similarly to the other networks, SimNet observes the first T0T_{0} images and aims to regress the physical parameters necessary to predict the trajectory of the object using the physics engine. The simulator is assumed to have access to a perfect model of the underlying physical laws. The regression architecture constitutes of the vector based feature extraction network described in Section 3.2 with an extra fully-connected layer on top to regress physical parameters. The network is trained with an L2L^{2} loss to infer the current slope rotation angles and friction coefficient (θx,θy,ρ\theta_{x},\theta_{y},\rho), the object’s position at the observed frames (t0,…,T0−1t_{0},\ldots,T_{0}-1), and its final velocity at frame T0−1T_{0}-1.

We input the regressed parameters to the same physics simulation system that generated the dataset (Section 4.1) and run the simulation to predict the following object positions T0…TT_{0}\ldots T. Note that, since the simulator used by the network is the same as the one used to generate the data, this network is given a significant advantage over the other models.

3 MechaNets

Experiments consider four different variants of the mechanical prediction networks (MechaNet1 to 4 for short). MechaNet1 and MechaNet2 are trained to optimise the L\mathcal{L}; MechaNet1 uses the LSTM propagation network and the spatially concentrated internal representation, whereas MechaNet2 uses the simpler convolutional propagation network but the distributed representation. MechaNet3 and MechaNet4 are similar to MechaNet2, but they use probabilist predictors, using the Gaussian and probability map outputs, respectively. The four variants are summarized in Table 2.

Network weights are initialized by sampling from a Gaussian distribution. Training uses a batch size of 50 using the first 10 to 40 frames of each video sequence using RMSProp (Tieleman & Hinton, 2012). Training is halted when there is no improvement of the L2L^{2} loss after 40 consecutive epochs; 1,000 epochs were found to be usually sufficient for convergence.

Since during the initial phases of training the network is very uncertain, the model using the Gaussian log-likelihood loss Lnrm\mathcal{L}_{\text{nrm}} was found to get stuck on solutions with very high variance Σ(t)\Sigma(t); to solve this issue, the regularizer λ∑tdet⁡Σ(t)\lambda\sum_{t}\det\Sigma(t) was added to the loss, setting λ=10\lambda=10 for the first few epochs and then lowering it to λ=0\lambda=0 when the value of the determinant stablized under 100 on average (this variance is comparable to the image size).

In all our experiments we used Tensorflow (Abadi et al., 2015) r0.12 on a single NVIDIA Titan X GPU.

4 Results

Table 2 and Fig. 3 compare the baseline predictors and the four MechaNets on the task of long term prediction of the object trajectory. We call this “long term” in the sense that all methods observe only the first T0=4T_{0}=4 frames of a video (except the linear and quadratic extrapolators which observe the first 10 frames instead), to then extrapolate the trajectory to 40 time steps.

All networks can be used to perform arbitrary long predictions; the table, in particular, reports the the average L2L^{2} prediction errors at time Ttest=20T_{\text{test}}=20 and 4040. However, models in this table are only shown the first Ttrain=20T_{\text{train}}=20 frames of each video during training.

Consider first the prediction results for Ttest=Ttrain=20T_{\text{test}}=T_{\text{train}}=20. All networks perform considerably better than the linear and quadratic extrapolators in all scenarios, with error rates 5-40×\times smaller. As expected, Scenario S1 and S2 are harder than Scenario S0, which uses a fixed slope and homogeneous friction, but the network prediction errors are still small, in the order of 1-2 pixels. All networks perform similarly well, particularly in Scenarios S1, with a slight advantage for the LSTM-based propagation networks. SimNet is very competitive, as may be expected given that it uses the ground-truth physics engine for integration. However, in Scenario S2 this method does not work well since the variable friction distribution is not observable from the first T0=4T_{0}=4 frames of a video; MechaNet2 and MechaNet4, which can better learn such effects, can account for such uncertainty and significantly outperform SimNet.

Results are different for predictions at time Ttest=40≫TtrainT_{\text{test}}=40\gg T_{\text{train}}. All networks still outperform the extrapolators, but in Scenarios S0 and S1 SimNet performs better than the other networks: by having access to the physics simulator, generalization is not an issue. On the other hand, this experiment shows that the deep networks have a limited capacity to generalize physics beyond the regimes observed during training. Among such networks, the ones modelling uncertainty (MechaNet3 and MechaNet4) are able to generalize better. Scenario S2 still breaks the assumptions made by SimNet, and the other networks outperform it.

Generalization.

The issue of generalization is explored in detail in Table 3, focusing on MechaNet4 that exhibits the best generalization capabilities. The table reports prediction errors at T=10,20,30,40T=10,20,30,40 for networks trained with video sequence of length T=10,20,30,40T=10,20,30,40 respectively. Recall that predictors always observe only the first T0=4T_{0}=4 frames of each sequence; the only change is to allow the training loss to assess the predictors’ performance on progressively longer videos during training.

As expected, training on longer sequences dramatically improves the accuracy of longer term predictions, but also the shorter term ones. Training on the full sequences, in particular, performs ∼20%\sim 20\% better than SimNet. This confirms that, while deep networks are able to learn physical rules accurately for the range of physical experiences observed during training, they do not necessarily learn rules that generalize as readily as conventional physical laws.

Predicting uncertainty.

MechaNet3 and MechaNet4 predict a posterior distribution of possible object locations, using a Gaussian and a probability map model respectively. Table 2 shows that the latter model has significantly lower perplexity, suggesting that the Gaussian model is somewhat too constrained. Qualitatively, Fig. 3 and 4 show that both models make very reasonable predictions of uncertainty, with the uncertain area growing over time as expected.

Conclusions

In this paper we explored the possibility of using a single neural network for long-term prediction of mechanical phenomena. We considered in particular the problem of predicting the long-term motion of a cuboid sliding down a slope of unknown inclination and heterogeneous friction. Differently from many other approaches, we use the network not to predict some physical quantities to be integrated by a simulator, but to directly predict the complete trajectory of the object end-to-end.

Our results, obtained from extensive synthetic simulation, indicate that deep neural networks can successfully predict long-term trajectories without requiring explicit modeling of the underlying physics. They can also reliably estimate a distribution over such predictions to account for uncertainty in the data. Remarkably, these models are competitive with alternative predictors that have access to the ground-truth physical simulator, and outperform them when some of the physical parameters are not observable or known a-priori. However, neural networks exhibit a limited capability to perform predictions outside the physical regimes observed during training. In other words, the internal representation of physics learned by such model is not as general as standard physical laws.

Several future directions remain to be explored. Given the accuracy of mechanical simulators, synthetic experiments are sufficient to assess the capability of networks to learn mechanical phenomena. However, the obvious next phase will be to test the framework on video footage obtained from real-world data in order to assess the ability to do so from visual data affected by real nuisance factors. The other important generalization is to consider more complex physical phenomena, including multiple sliding objects with possible interactions, rolling motion, and sliding over non-flat surfaces.

References