Learning Multi-Object Dynamics with Compositional Neural Radiance Fields
Danny Driess, Zhiao Huang, Yunzhu Li, Russ Tedrake, Marc Toussaint
Introduction
Learning models from observations that predict the future state of a scene is a fundamental concept for enabling an agent to reason about actions to achieve a desired goal. A major challenge in learning predictive models is that raw observations such as images are usually high-dimensional. Therefore, a common approach is to map the observation space into a lower-dimensional latent representation of the scene via an auto-encoder structure. Based on those latent vectors, a dynamics model can be learned that predicts the next latent state, conditioned on actions an agent takes. An intuition for this is that if a latent vector is sufficient to reconstruct the observations, then it contains enough information about the scene to learn a dynamics model on top of it. While an auto-encoder structure combined with a latent dynamics model is a general approach that is applicable for a large variety of tasks, it raises multiple challenges. First, scenes in our world are composed of multiple objects. Therefore, a fixed-size latent vector has difficulties in generalizing over different and changing numbers of objects in the scene than during training, both due to the limited capacity of fixed-size vectors and lack of diversity in the training distribution. Second, image observations are 2D, but the 3D structure of our world is essential for many tasks to reason about the underlying physical processes governing the dynamics the model should predict. Dealing with occlusions, object permanence, and ambiguities in 2D views is challenging for 2D image representations. Importantly, many forward predictive models in visual observation spaces suffer from instabilities in making long-term predictions, often manifested in blurry image predictions .
One way to address these issues is to incorporate inductive biases and structural priors in the model architectures. Li et al. proposed to use Neural Radiance Fields (NeRFs) as a decoder within an auto-encoder to learn dynamics models in latent spaces. NeRFs exhibit strong structural priors about the 3D world, leading to increased performance over 2D baselines. However, the approach of represents the whole scene as a single latent vector, which we found insufficient for scenes that are composed of multiple, different numbers of objects, both in terms of representation and dynamics prediction.
In the present work, we aim to overcome these challenges by incorporating inductive biases on the compositional nature and underlying 3D structure of our world both in learning the latent representations themselves and the dynamics model. We propose a compositional, object-centric auto-encoder framework whose latent vectors are used to learn a compositional forward dynamics model in that learned latent space based on graph neural networks (GNN). More specifically, we learn an implicit object encoder that maps image observations of the scene from multiple views to a set of latent vectors that each represent an object in the scene separately. These latent object encodings then parameterize individual NeRFs for each object. We apply compositional rendering techniques to synthesize images from multiple viewpoints, which forces the object-centric NeRF functions and the corresponding latent vectors to learn precise 3D configurations of the constituting objects. This 3D inductive bias both in the encoder and the compositional NeRF decoder enables us to incorporate priors from the models’ own predictions about objects interactions via an estimated adjacency matrix into learning the GNN dynamics model, making long-term dynamics predictions more stable. This long term-stability allows us utilize a planning method based on RRTs in the latent space.
In our evaluations, we show through comparisons that non-compositional auto-encoder frameworks and non-compositional dynamics models struggle with tasks containing multiple objects, while our framework generalizes well over different numbers of objects than during training and is capable of generating sharp and stable long-term predictions. Relative to more traditional multibody system identification , these models learn the geometry of unknown objects in addition to (implicitly) learning the inertial and contact parameters. We demonstrate the performance of the approach in terms of image reconstruction error, dynamics prediction error, and planning, generalizing over different numbers of objects than during training. Our experiments include rigid and deformable objects in simulation and with a real robot. To summarize, our main contributions are
A compositional scene encoding framework that uses implicit object encoders and NeRF decoders for each object, forcing the view-invariant latent representation to learn about the 3D structure of the problem in a composable way.
A factored dynamics model in the latent space as a graph neural network (GNN), exploiting the compositional nature of the scene representation and an adaptive adjacency matrix estimated from the model itself to yield stable long-term predictions.
Related Work
Learning Dynamics Models for Compositional Systems. Graph neural networks (GNNs) have shown great promise in introducing relational inductive biases , enabling them to model the dynamics of compositional systems consisting of interactions between multiple objects , large-scale dynamical systems represented using particles and meshes , or from visual observations . Our method differs from prior work by learning compositional scene representations grounded in 3D space directly from visual observations. Our novel combination of implicit object encoders and graph-based neural dynamics models reflects the structure of the underlying scene, which endows our agent with better generalization ability in handling complicated compositional dynamic environments.
NeRF for Compositional and Dynamic Scenes. Recent advances on neural implicit representations have demonstrated widespread success in image synthesis or 3D reconstruction . Notably, Neural Radiance Fields (NeRF) show impressive results on novel-view synthesis . Initial NeRF approaches were trained on a single scene without generalization. Prior work have since proposed to modify neural scene representations to make them compositional for static scenes without considering dynamics of object interactions. People have also extended NeRF to enable view synthesis from a sparse set of views , as well as modeling dynamic scenes by learning implicitly represented flow fields or time-variant latent codes . However, these approaches for dynamic environments typically interpolate over a single time sequence and are not able to handle scenes of different initial configurations or different action sequences, limiting their use in downstream planning and control tasks. Li et al. addressed this issue by combining an NeRF auto-encoding framework with modeling the dynamics in a latent space. Yet, they employed a single latent vector as the whole scene representation, which we will show is insufficient at modeling compositional systems. In contrast, our method considers a graph-based scene representation to capture the structure of the underlying scene and achieves significantly better generalization performance than .
Implicit Models in Robotics. Implicit models in robotics have been explored, e.g., for grasping or more general manipulation planning constraints . Analytic signed distance functions or learned NeRFs are used for trajectory planning. One assumption in and is that signed-distance values are available during training. Our work, in contrast, directly operates on RGB images without requiring explicit 3D shape supervision.
Overview – Compositional Visual Dynamics Learning
Our dynamics learning framework (Fig. 1) consists of three parts, an object encoder turning observations into a set of latent vectors , a compositional NeRF-based decoder that renders the latent vectors back into images of the scene to train the encoder, and a graph neural network dynamics model predicting the evolution of the scene in the latent space. This section gives a high-level overview, while Sec. 4, Sec. 5 as well as the appendix Sec. B, Sec. C provide details.
represents the object separately. is trained end-to-end with a NeRF decoder reconstructing
for arbitrary views specified by the camera matrix from the set of latent object representations . The initial observation of the scene is encoded with into the initial latent vectors . The GNN dynamics model then generates long-term predictions of future latent states that can also be decoded with to yield visual predictions from arbitrary views.
Encoding Scenes with Compositional Image-Conditioned NeRFs
Instead of learning defined in (1) as a direct mapping from images, camera matrices and masks to the latent vectors, we first encode each object in the scene as a feature-valued function over 3D space, conditioned on the image observations. This allows us to incorporate multiple views of the objects in a geometrically consistent way, as well as to apply 3D affine transformations to the objects, which will be important for the dynamics model (Sec. 5). This function is then turned into a latent vector by evaluating it on a workspace set followed by a 3D convolutional network.
Note that the same workspace set is used for all objects. The appendix contains visualizations of the architectures of , and .
In summary, the object encoder maps images from multiple views, object masks and the set to latent vectors. The resulting ’s contain not only the appearance of the objects, but also their spatial configurations in the scene relative to other objects.
2 Decoder as Compositional, Conditional NeRF Model
Compared to this standard NeRF formulation where one single model is used to represent the whole scene, we associate separate NeRFs with each object, meaning that the NeRF for object
is conditioned on for . is the density and the color prediction for object , respectively. To turn those back into a global NeRF model that can be rendered to an image, we sum the individual predicted object densities and obtain the colors as their density weighted combination . These composition formulas have been proposed multiple times in the literature, e.g. . This composition forces the individual NeRFs to learn the 3D configuration of each object individually and therefore ensures that each only predicts the object where it is located in the 3D space.
To summarize, the compositional NeRF-decoder takes the set of latent vectors for objects and the camera matrix for a desired view as input to render Since we only represent the objects and not the background as NeRFs, rendering the composed NeRF will yield an image with the background subtracted. In the experiments, we investigate the importance of the decoder being both compositional and a NeRF.
3 Training
The auto-encoder framework is trained end-to-end on an image reconstruction loss for view
Since solely the objects are represented as NeRFs and not the background, we compute the union of the masks of the individual objects and define the target image as with being a slightly enlarged union mask. Please refer to the appendix Sec. B for more details.
Latent Dynamics Model with Graph Neural Networks
Having trained the auto-encoder framework, we learn a graph neural network dynamics model
in the latent space, where is the adjacency matrix at time . Following , we use multi-step message passing to deal with cases where multiple objects interact within one prediction step. Refer to the appendix Sec. C and Algo. 1 for more details about our GNN dynamics model.
Adjacency Matrix from Learned Model. The adjacency matrix in the GNN dynamics model (7) plays an important role in indicating which objects interact. While a dense adjacency matrix, i.e. a graph where each object interacts with all other objects, would in principle work as the GNN could figure out from the latent vectors which objects interact, we found that the long-horizon prediction performance is greatly increased if is more selective in reflecting which objects actually interact.
We propose to utilize the NeRF decoder density prediction for each object to determine the adjacency matrix from the models’ own predictions during training and planning. In order to do so, we define the entries of the adjacency matrix between objects and based on the collision integral
over the density predictions of the learned NeRF model for a threshold . Estimating this way takes the actual 3D geometry of the objects in the scene into account and thereby informs the GNN dynamics model, leading to more stable predictions. Please refer to the appendix Sec. C for more details about and how it is used in the forward prediction Algo. 1.
Experiments
We demonstrate our framework on pushing tasks in scenes containing multiple objects both in simulation and in the real world. For a quantitative analysis and comparison to multiple baselines, we investigate here the forward prediction error of the model in the image space rendered from the learned model over long-horizons. Please refer to the supplementary video showing the reconstructions of the model, novel scene generation, forward predictions and planning/execution results as well as the appendix for more details and further experiments. Our scenarios are challenging, as they are composed of multiple, interacting objects, sparse rewards, and complex dynamics .
We compare our framework to non-compositional scene representations, non-compositional dynamics models, 2D CNN baselines (visual foresight) without NeRF as decoder, and the importance of estimating the adjacency matrix from the model itself.
Reconstruction and Prediction Performance for Generalization over Numbers of Objects. Fig. 4a shows predictions of the model forward unrolled in time for an action sequence of the red pusher, i.e. applying Algo. 1 (appendix) to an initial scene observation and rendering the predicted latent vectors with the NeRF decoder. Despite the movements in this scene leading to multiple object interactions, even after 38 time steps, the rendered predictions from the model are still sharp and reflect the underlying dynamics. By utilizing the estimated adjacency matrix, there is little drift in the objects, leading to long-term prediction stability. Due to its compositional nature, our model generalizes to scenes that contain more or less objects than in the training set, as shown in Fig. 5 where eight objects plus the pusher are observed and reconstructed with high quality from novel views, although during training the model has seen only and exactly 4 objects.
Comparison to Non-Compositional Scene Representation Baselines. We compare to two non-compositional baselines where the scene is represented globally with one single latent vector per time-step. The dynamics model for these baselines is an MLP that takes the action as an additional input. The first baseline (Global NeRF) is the approach from , i.e. we use their CNN encoder instead of our implicit object encoder producing one latent vector conditioning a global NeRF that reconstructs the whole scene. The second baseline (Global 2D CNN auto encoder) uses both a 2D CNN encoder and 2D CNN decoder as well as a single latent vector representing the whole scene. Such frameworks have been used many times in the literature, e.g. and are known as visual foresight. Fig. 3 shows that both global baselines are significantly inferior in our scenarios to our proposed compositional framework, especially for long horizons.
Comparison to 2D Baselines – Importance of NeRF as Decoder. In this section, we replace the NeRF decoder with a 2D CNN decoder to investigate the importance of NeRFs. This decoder takes as input one single latent vector and the camera matrix, i.e. . In order to make it compositional, we aggregate the set of latent vectors from with a mean operation and then pass the aggregated feature through an MLP to produce the single for . The rest of the architecture, i.e. implicit object encoder and GNN, stays the same. Since there is no clear way to estimate the adjacency matrix from , we use a dense adjacency matrix for the GNN. As one can see in Fig. 3, the long-term prediction performance of the CNN decoder is significantly worse than with a compositional NeRF model as the decoder, especially when asking for numbers of objects that differ from the training distribution. Qualitatively, one can see in Fig. 4c that not only the initial reconstruction is much less sharp compared to the NeRF-based models, but especially also that even after only a few time-steps, the predictions with the CNN decoder are of little use.
Comparison to CNN Encoder. Exchanging the implicit object encoder with a 2D CNN compositional encoder leads to an auto-encoder framework similar to . As seen in Fig. 3, the performance is better compared to the other baselines, but still clearly worse than with the proposed method.
Importance of Estimating the Adjacency Matrix. In Sec. 5, we propose how the adjacency matrix of the GNN can be estimated from the learned NeRFs to increase the long-term stability of the predictions. Here we compare to a dense adjacency matrix, i.e. where the network has to figure out from the latent vectors themselves which objects interact. As one can see in Fig. 3 and Fig. 4b, a dense has significantly worse long-horizon prediction performance compared to our proposed way of estimating through the learned NeRF model. In the 2 and 8 object case (generalization over numbers of objects), the predictions with the dense are useless after only a few time-steps.
Non-Compositional Dynamics. Replacing the GNN with a fully connected MLP leads to worse performance than with a GNN with dense adjacency matrix. This model cannot generalize to different numbers of objects due to its fixed input size.
Summary of Performance Comparisons Our method outperforms all baselines both in terms of pure reconstruction error (as can be seen in Fig. 3 by the error after 0 prediction steps) and its ability to perform long-term predictions forward unrolled on the model’s own predictions. Estimating the adjacency matrix from the model itself is important for long-term stability as it prevents objects from drifting away. Too large drift makes future predictions for a pushing tasks meaningless. Since the reconstruction error of our proposed method without dynamics is better than the baselines, the question arises if the increased performance is an artifact of the lower reconstruction error. We show in Fig. 3d the error in the image space between renderings when having access to the observations at each step and the renderings from the predicted latent vectors into the future after observing the scene only at the beginning. This shows the increase in error relative to the reconstruction process. The results indicate that not solely the reconstruction itself is the reason for the better performance, but that the structural choices of our framework also enable to learn the dynamics more precisely.
2 Planning and Execution Results on Object Sorting Task
To demonstrate the effectiveness of the learned model, we utilize it to solve a box sorting task, where the pusher needs to push colored boxes into their corresponding goal regions as shown in Fig. 6. This task is inspired by and involves multiple challenges: As multiple objects interact, a greedy strategy of pushing objects straight to the goal region fails. Movements, i.e. actions, of the pusher do often not immediately lead to a change in the cost function, since contact with the object from a suitable side has to be established . In the appendix Sec. D we propose a latent space RRT that uses our framework for planning. Refer to the appendix and the video for more details about our proposed planning algorithm and comparisons to baselines.
3 Real World Experiments
Fig. 1a shows the rendered forward predictions of our model for a real world scenario where a robot pushes a shoe and a giraffe-shaped toy. Fig. 1b are renderings from novel view points.
4 Applicability to Deformable Objects
The experiments so far focused on objects that behave mainly like rigid objects when being pushed. In Fig. 7 we show that our method is also applicable to deformable objects simulated with .
Discussion & Limitations
Computational Efficiency. Our framework is computationally more demanding during inference time than 2D CNN decoder baselines, mainly due to NeRF evaluations. Many methods have been developed to increase the speed of NeRF , from which our framework could benefit.
Object Masks. The compositional scene encoding framework requires object masks to achieve compositionality. Many mature methods for instance segmentation have been developed such that we believe having masks is a reasonable assumption. However, one could add an output to our implicit encoder that provides object labels or use mechanisms similar to slot attention .
Latent Representations. We have shown the great benefits of a compositional latent representation as it not only provides generalization over different numbers of objects in the scene, but also leads to increased reconstruction and dynamics prediction performance compared to non-compositional baselines. Furthermore, latent representations compress observations, enabling efficient dynamics prediction. However, as still each object in the scene is represented as a latent vector of finite size, latent models are capable of mainly representing objects with shapes similar to the training distribution. To address this, compositionality could not only be introduced on the scene level, but also by representing objects themselves in a composable way.
Long-Term Prediction Stability. Our dynamics model framework exhibits significantly better long-term prediction stability compared to baselines. Our experiments indicate that this is due the structural biases enabled through (compositional) NeRFs. This stability allowed us to use the model for planning scenarios requiring long-horizons, which none of the baseline methods could support. However, we believe that there is still room for improvement regarding the prediction stability. Especially for deformable objects, we observed that after many prediction steps objects in the scene are predicted to penetrate or move through each other, which could be improved in future work.
Conclusion
Visual dynamics models are of high interest to the computer vision and robotics community, as they avoid explicit shape model assumptions and imply end-to-end perception. However, to support manipulation planning and reasoning, we need models that generalize strongly over objects and provide stable long-term predictions. In this paper we proposed a system that introduces 3D structural and compositional priors at various levels, namely compositional NeRFs, 3D implicit object encoders, and GNNs dynamics with an adaptive adjacency matrix. Together our system exhibits significantly stronger long-term prediction performance compared to multiple baselines without these priors or without compositionality, and supports using a latent space RRT planner. We have shown generalization over different numbers of objects, notably up to two times more than during training.
This research has been supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy – EXC 2002/1 “Science of Intelligence” – project number 390523135. Danny Driess thanks the International Max-Planck Research School for Intelligent Systems (IMPRS-IS) for the support. The authors thank Valentin Hartmann for discussions regarding RRTs.
References
Appendix A Expanded Related Work
Model-based planning algorithms typically build a dynamics model of the environment and then use the model to plan the agent’s behavior in order to minimize some task objectives. We can roughly categorize the methods by whether the model is constructed from first principles (i.e., physical rules) or learned from data (i.e., data-driven models). Physics-based models typically require complete information about the objects’ geometry and the system’s state , which limits their applicability in robotic manipulation tasks involving unknown object models and partially observable states. Data-driven methods, on the other hand, learn a dynamics model directly from the robot’s interaction with the environment and have shown impressive results in manipulation tasks ranging from closed-loop planar pushing to complicated dexterous manipulation . Many of the data-driven planning frameworks learn dynamics models directly from visual observation based on representations defined at different levels of abstraction, such as pixel space , 3D volumetric space , signed-distance fields , keypoint space , and low-dimensional latent space . Approaches commonly employ an image reconstruction loss , an self-supervised time contrastive loss , or jointly train a forward and an inverse dynamics model to make sure that the representation encodes meaningful information about the environment. Our method takes a step forward by learning graph-based latent representations from visual observations. The learned model accurately encodes the underlying 3D contents, allowing our learned model to achieve precise manipulation of compositional environments and generalize outside the training distribution, i.e. to scenes with more (and less) objects than during training.
Appendix B Details – Encoding Scenes with Compositional Image-Conditioned NeRFs
This section provides details – especially visualizations of the network architectures – on our proposed compositional auto-encoder framework.
Fig. 8 visualizes the whole auto-encoder architecture, where the encoder
maps posed images from -many views of -many objects including their object masks as well as an workspace set to a set of latent vectors describing the objects in the scene. The same object encoder and workspace set is used for all objects. In particular, is not a 3D bounding box for an individual object, but covers the whole workspace of the scene. See Fig. 13 for a visualization of the workspace set .
Similar architectures of computing such pixel features from world coordinates have been proposed, e.g., in for the single object case. However, we use quite differently compared to these works, as we compute a latent vector from pixel-aligned features with 3D convolutions. The object feature function , cf. (3), aggregates the features from the individual views by taking into account the masks of object in each view. The architectures of and are visualized in Fig. 9. Fig. 10 shows how an object feature function for object is turned by into its corresponding latent vector by querying it on the workspace set followed by a 3D convolutional network.
B.2 Background on Neural Radiance Fields
B.3 Training
The auto-encoder framework is trained end-to-end on an image reconstruction loss. Since, as mentioned, solely the objects are represented as NeRFs and not the background, we compute the union of the masks of the individual objects
and define the target image in a view as with denoting the element-wise product.
A known issue of NeRF is its computational efficiency , since for every pixel all ’s have to be queried on many points along the camera ray. We make two simple, but important, improvements to reduce the computational demand.
First, the near and far bounds , are determined individually for each camera ray such that only those points along the rays that are within the workspace set are considered. This is a reasonable assumption since we assumed that the objects are in the workspace set in the first place. That way, the computational efficiency is already greatly increased by reducing the number of points where functions ’s have to be queried.
Moreover, as the scenes we consider in our experiments (see for example Fig. 13) are composed of multiple smaller objects, when masking out the background, the majority of pixels in each view is black and therefore does not contain information about the scene, although the model is evaluated on those areas. To further decrease the number of points where the NeRFs have to be queried, we only consider those rays for a view that pass through the mask of at least one object in that view. It turned out, however, that training only on rays that go through leads to blurry reconstructions, since there is no loss indicating that the objects should end outside of the masks. In order to resolve this, we enlarge the combined mask of a view with a convolution operation by a few pixels. We denote this enlarged mask by . Together, these techniques ensure that the model learns sharp object boundaries, while significantly reducing the number of considered rays and required NeRF evaluations. See Fig. 11 for a visualization of this procedure.
These considerations lead to the following training objective of and for a view
During training, we randomly sample a view from the dataset for each mini-batch and update the parameters of and using the ADAM optimizer .
Another side effect of training on the enlarged masks only is that it improved the training stability and reconstruction qualities of the model. Indeed, when we trained the model on the whole image, depending on the weight initialization of the network, the model sometimes very quickly converged to a state where it only predicted a black image, since the majority of pixels are actually black and hence a low loss could be achieved. Training on the enlarged masks prevents this reliably.
B.4 Rigid Transformations and Novel Scene Generation
Please note the slight abuse of notation here, the term has to be understood elementwise for each entry in .
Composing scenes via rigid transformations applied to the input of the individual NeRFs has been considered before, e.g. in . However, transforming a NeRF by applying the rigid transformation to its input only leads to changes in the rendered visual space, i.e. it has, in particular, no influence on the latent vectors of the objects. Since we want the latent vectors to represent not only the appearance of an individual object, but the geometric information of the object within the scene relative to other objects, just transforming the NeRF models is not sufficient. Therefore, using (15) we get the latent vector rigidly transformed by , which is crucial for our downstream dynamics prediction task.
This rigid transformation is important to define actions for the GNN dynamics model, as described in Sec. C.3 and Algo. 1.
Appendix C Details – Graph Neural Network Latent Dyanmics Model
This section provides details about the latent dynamics model. Especially relevant is Sec. C.6 and Algo. 1 where we describe the forward prediction algorithm.
Due to the compositional nature of the scenarios we consider, we require a dynamics model that maintains the capabilities of our auto-encoder to generalize over changing numbers of objects, for which graph neural networks (GNNs) are a natural choice.
The general idea behind learning dynamics models with GNNs is to associate each object in the scene with a node in a graph, which, in our case, means that each node in the graph is a latent vector . Edges between the nodes indicate if objects interact, e.g. by exchanging forces due to contact. As argued in , applying a simple GNN to the problem of dynamics prediction is problematic, since interactions caused at one node can influence not only the neighboring nodes, but higher-order neighbors. For example, if three objects touch, the effects of applying a force at the first object have to propagate. The scenarios we consider in the experiments contain multiple objects such that more than two objects can interact in one time step. To take this into account, we use a message passing architecture inspired by .
Let and be the latent vectors of objects and . An edge encoder network determines a feature
describing the interaction between the objects , . An adjacency matrix has entry if object is influenced by object . Assume the state of all latent vectors at time is known. The node propagator network recursively is queried many times to propagate the state to the next time step as follows:
with , for and the final new predicted state .
C.2 Adjacency Matrix from Learned Model
The adjacency matrix in the GNN dynamics model (17) plays an important role in indicating which objects interact. While a dense adjacency matrix, i.e. a graph where each node is connected to every other node implying that each object interacts with all other objects in the scene, would in principle work as the network could figure out from the latent representations itself which objects interact, we found that the long-horizon prediction performance is greatly increased if is more selective in reflecting which objects actually interact (refer to the experiments in Sec. E.2). This is especially relevant for compositional scenes as considered in this work where there are many objects, but which often do not interact with each other in every timestep.
A central question is how the adjacency matrix can be obtained from the observations of the scene without manually specifying it. Due to our model having strong 3D priors, we can exploit the density prediction as defined in (5) for each object to determine the adjacency matrix from the models’ own predictions during training and planning. In order to do so, for a threshold , the collision integral
over the density predictions of the learned NeRF model for objects and indicates if the two objects overlap or not. A similar integral as in (18) has been proposed in to estimate collisions from signed-distance functions. Based on this integral, we define the entries of the adjacency matrix between objects and as
which implies that only those objects that are or are close to being in contact potentially interact. Estimating this way takes the actual geometry of the objects in the scene into account. In relation to the node propagation network , this means that the adjacency matrix at step of the propagation becomes a function of the node encodings itself, i.e.
For training the GNN, however, changing the adjacency matrix during prediction is not differentiable. Therefore, we compute the adjacency matrices from the model such that they are constant within one time-step as follows. We first compute an occupancy grid
for each object over the discretized workspace set and then apply a 3D convolution operation on with a kernel consisting of only ones to expand the occupancy grid. The now constant within one time-step entries for all are then determined by checking if there is a voxel cell where both enlarged and have value one. The size of the convolution kernel is chosen large enough such that the adjacency matrix is not going to change within one timestep. This allows for a trade-off between sufficient sparsity of while ensuring that all objects that potentially interact have corresponding entries in .
In the experiments in Sec. E.2, we investigate the influence on the prediction performance for multiple different ways of predicting/using the adjacency matrix.
C.3 Actions
So far, the way we have formulated the graph neural network dynamics model in Sec. C.1 does not contain a notion of actions. Instead, we interpret an action as a modification to a node in the graph and train the GNN to predict the state of the nodes at the next time step as a result to this modification. This allows us to not explicitly distinguish between controlled and uncontrolled/passive objects.
C.4 Quasi-Static Dynamics
If we assume quasi-static dynamics, meaning that the next system state only depends on the current latent state without history and immediate actions (refer to the discussion in the last paragraph Sec. C.3 about the notion of actions in this work), we can further increase the long-term stability by utilizing the adjacency matrices estimated by the learned model. When an object is not involved in any interactions with other objects, then, under quasi-static assumptions, it does not change between time steps, i.e. its latent vector stays constant, which means we can set
The condition means that the node associated with in the graph has no incoming edges. As we will show in the experiments (Sec. E.2), this can greatly increase the stability for long-term open-loop model predictions as it will prevent drift in objects that do not take part in any interaction with other objects.
C.5 Training
We first train the compositional NeRF auto-encoder framework on training data, which gives us a dataset of trajectories of latent vectors. The GNN dynamics model is then trained on the one-step mean squared error between and of samples of such trajectories using the ADAM optimizer.
Importantly, training the dynamics model does not require a dataset containing the actions. It is sufficient to have video sequences and knowledge which object was the articulated one. At inference time, one can also choose different objects to apply actions to in terms of rigid transformations.
C.6 Forward Prediction Algorithm
Algorithm 1 summarizes the forward prediction procedure. The algorithm is visualized in Fig. 12.
Starting from a single initial observation of the scene in terms of the images from many views and the objects masks of the many objects, Algorithm 1 predicts the latent vectors at times for all objects in the scene given a desired action sequence of rigid transformations applied to object . At every point in time, the scene can be rendered from arbitrary view points from the predicted . Note that the masks of the objects are required only for the initial observation, i.e. no mask prediction has to be performed as compared to .
In line 2, the initial object encodings from the scene observation are computed. Line 5 applies the action to the object with index . Lines 6-10 then perform the prediction step of the GNN in the latent space using message passing. Crucially, in line 7, the adjacency matrix is estimated from the current predictions during message passing (Sec. C.2). Note that here the original collision integral (18), computed on the grid , can be used without enlarging the intermediate occupancy grids, since, if during the message passing step objects interact that previously did not, it will be captured, as is estimated in every step of the message passing part. This leads to further increased prediction stability, as we will show in the experiments. Finally, in line 12, objects that have not interacted with other objects as predicted by the adjacency matrix are kept at their previous latent state (quasi-static assumption from Sec. C.4)
Appendix D Planning with Latent space RRT
In this section, we propose a planning algorithm to manipulate objects to achieve a desired goal using our scene encoding and dynamics model framework. Note that planning and control is not the main focus of this work, however, the algorithm still contains important insights.
The main part of the planning algorithm is an RRT in the latent space. Such latent space RRTs have been considered, for example, in . One central question here is how one can sample in the latent space effectively, since a uniform random sample in the latent space not necessarily is a valid (and/or uniform) sample in the original space. In , they assume to have access to a set of valid latent vectors from which they can sample. In contrast, we can produce valid samples in the latent space directly by exploiting the properties of our model.
On a high level, our model iteratively perceives multi-view images of the scene, finds a plan using a Latent-Space RRT (LS-RRT) based on the forward predictions of the model over a long horizon, and then executes the found plan with Model-Predictive Control (MPC) for a shorter horizon. We describe the algorithm here with pushing scenarios as considered in the experiments in mind.
Algorithm 2 summarizes the LS-RRT algorithm. We grow a tree in the latent space, starting at the latent vector that represents the current state of the environment, encoded by the implicit object encoder from the current visual observation of the scene. In standard RRTs, a target is uniformly sampled in the configuration space to steer the growth of the tree towards a Voronoi bias. To introduce a particular goal-targeted sampling bias and as we do not have an inverse model or steering function, we modify the standard approach as follows:
In a latent space RRT, sampling uniformly in the latent space does neither guarantee that the samples are from the latent space manifold nor that they explore the original space. Therefore, we sample random targets not in the latent space directly, but only in the space of center of mass configurations of all objects, which is of dimension in the experiments (objects and pusher). In this way, we can design a sampling distribution biased to target configurations that have low costs, i.e., more objects within the goal region, or targets in which the articulated object (the pusher in the experiments) is close to one of the objects, inducing a bias for interaction. This sampling distribution and cost evaluation is possible because we can apply rigid transformations to the objects through our object encoder being an implicit function, since, for a sampled random target, we have to move the objects to this target to check the cost on the transformed configuration. Further, the metric to select the expanded node is the -norm in the full configurations between and the centers-of-masses computed from . Using the predictions of the NeRF model, we can estimate (under homogeneous density assumption) the center of mass of an object with latent vector as
i.e. . Note that the sampling distribution, cost function evaluation and metric calculation are done solely based on predictions of the model. At no point the model has access to ground truth center-of-mass information.
Finally, as we do not have an inverse dynamics model or an other kind of steering function, we expand the tree using a random action , similar to control trees. However, our goal-targeted node selection ensures that the tree expands effectively.
D.2 Cost-Functions
As our decoder is based on NeRFs, our method is able to provide a lot of flexibility in defining cost functions. At any time instance forward predicted by the dynamics model, we can render an image from an arbitrary view or reconstruct the objects in 3D. This enables, for example, to have the following options to define cost functions for planning:
Loss on image rendered from arbitrary views (with known camera matrix).
Loss on density prediction of the NeRF reconstruction.
Loss on point-cloud reconstructed from (including color information on each point).
Loss on center-of-mass predicted by the model via (23).
All these options can be defined for specific objects, the whole scene, or anything in between.
D.3 Model-based Control
Although our model achieves impressive performance over a long horizon, the accumulated prediction errors may still lead to a failure when executing the plans open-loop. We therefore apply an MPC scheme, which in each cycle feeds the current visual observation into the model, samples and select actions that match the plan (in terms of the center-of-mass metric) within a short horizon predicted by the learned dynamics model. If there is a significant mismatch between the plan and current observation, the LS-RRT is used again to find a new long-term plan starting from the current observation.
Appendix E Experiments – Simulation
In the simulated experiments, we focus on a pushing task in scenarios with multiple box-shaped objects on a table, see, e.g. Fig. 13 or 11 for such scenes. For a quantitative analysis and comparison to multiple baselines, we investigate the forward prediction error of the model both in the image space (Fig. 3) and, for baselines that use a compositional NeRF decoder, the error in predicting the center of mass of the objects (Fig. 14) over long-horizons. In all plots of Fig. 3 and Fig. 14, the blue curve corresponds to our proposed framework as summarized in Algo. 1. Sec. E.4 presents planning and execution results for a challenging box sorting task.
Please refer to the supplementary material for videos showing the reconstructions of the model, forward predictions, novel scene generation, and planning/execution results.
We consider a rigid-body scenario with multiple objects on a table, see Fig. 13 for an example. In all cases, the red cylinder is the pusher that is articulated in order to push the other objects around.
This scenario is challenging due to multiple reasons. First, it is composed of many objects, which implies not only a broad scene distribution, but especially also that many objects can interact. The mechanics of such multi-body pushing is non-trivial, since, for instance, contact can be established and broken between the objects at multiple phases of the motion. Contact between multiple objects at the same time can occur. Furthermore, we do not assume that the red pusher starts in contact with an object. Hence, if a task implies that an object should be pushed, long-term predictions inherently have to be made in order to establish contact, before any object movement is registered.
All scenes in the training data contain 4 box-shaped objects of randomly sampled sizes, positions and orientations (5 dimensional parameter space for each of the 4 objects) and one cylinder-shaped object with randomly sampled position.
To generate the training data, we randomly sample one of the 4 objects and then move the red pusher towards the center of this chosen object (with Gaussian noise added to the direction vector in each time step) until either the pusher leaves the workspace, in which case a new target object is chosen, or an object is pushed outside the workspace, in which case the data collection for this scene is terminated and a new scene is sampled. In total, the training dataset contains 5752 scenes with an average sequence length of 17. We generate 3 test datasets for evaluating the reconstruction and prediction performance which contain 2, 4, and 8 objects, respectively, plus the pusher. There are 312 scenes for each test dataset, generated with a different random seed than the training data. As visualized in Fig. 13 by the green coordinate systems, we choose 4 camera views for each scene.
E.2 Importance of Estimating the Adjacency Matrix
This section provides a more detailed investigation of the importance of the adjacency matrix than in the main text, where we only have discussed a dense adjacency matrix.
In Sec. C.2, we have proposed how the adjacency matrix of the GNN can be estimated from the density predictions of the learned NeRFs and that under quasi-static assumptions this estimated adjacency matrix can further be exploited to increase the long-term stability of the predictions, cf. Sec. C.4. Here we investigate the consequences of utilizing the adjacency matrix this way by comparing the full Algorithm 1 to the following three ablations. In all these ablations, the rest of the method remains the same, i.e. same object encoder, same GNN, same compositional NeRF decoder.
In this case, line 12 of Algorithm 1 is not used, i.e. the latent vectors of all objects, even when they do not interact with other objects as estimated through the model, are updated using the model forward predictions.
Here, we estimate the adjacency matrix only at the beginning of the message passing step, i.e. before line 6 in Algorithm 1. In order to ensure that it can still capture all object interactions that might occur during the message passing step, we enlarge the determined occupancy grids exactly the same way as for training, see the discussion in Sec.C.2. The effects of this are that objects that are close to each other but do not interact still have entries in indicating that they interact, which means slight errors in the predictions accumulate and lead to drift, although the object would not move in reality.
We further consider a dense adjacency matrix, i.e. where the network has to figure out from the latent vectors themselves if objects interact. Preventing drift in this case is considerably harder.
In Fig. 14 one can see the mean error of the model predicting the center of mass, computed from its density predictions of the NeRFs for each object according to (23), over the number of steps predicted into the future on the test dataset for different numbers of objects in the scene.
As one can see in Fig. 14a, for the two object case, the choices of how the adjacency matrix is used, as long as it is not a dense one, are not significant. For the 4 (Fig. 14b) and 8 (Fig. 14c) object case, however, our proposed utilization of the adjacency matrix, i.e. estimating it during propagation steps and using it to exploit the quasi-static assumption, leads to a significant increase in performance. This can be explained by the fact that utilizing the adjacency matrix as we propose leads to significantly less drift. Especially with the dense adjacency matrix, the predictions are very unstable for all, the 2, 4, and 8 object case. Fig. 4b shows this qualitatively. In the 8 object case, the predictions with the dense are basically useless after only a few time-steps, showing that it has overfitted to the number of 4 objects as in the training data.
E.3 Advantages of Implicit Object Encoder – Comparison to CNN Encoder
This experiment has already be mentioned in the main text, but here we provide more details. We exchange the implicit object encoder with a 2D CNN object encoder. The resulting auto-encoder framework is very similar to the architecture of . More specifically, we encode each masked image observation with a 2D CNN to produce a feature vector. The feature vectors from the different views are aggregated into the final latent vectors for each object. We use the encoder architecture from , but adjust it to the compositional multi-object case by incorporating object masks. Since this encoder is not an implicit function of , we cannot modify the latent vectors by applying rigid transformations and hence need to encode the actions differently. In order to do so, we train a separate MLP network that predicts the latent vector of the pusher resulting from applying an action to it. This gives the modified for line 5 in algorithm 1 for the CNN encoder baseline. The rest of the architecture, i.e. the GNN, the compositional NeRF decoder, estimating the adjacency matrix during message passing from the model, etc., stays the same.
As can be seen in Fig. 14a and Fig. 3, replacing the proposed implicit object encoder with a CNN encoder, the performance is better compared to the other baselines, but still clearly worse than with the proposed method.
E.4 Planning and Execution Results on Box Sorting Task
To demonstrate the effectiveness of the learned model, we utilize it to solve a challenging box sorting task, where the red pusher needs to push the blue and yellow boxes into their corresponding goal regions as shown in Fig. 6. This task is in part inspired by the object sorting task in . The cost function in Algorithm 2 determines how many objects are outside of their goal region, which is computed for each object from their corresponding density and color predictions of the model itself. The goal is fulfilled if all objects are in their respective goal regions.
This object-sorting task is challenging for multiple reasons. First, the dynamics of pushing is non-trivial . In our particular case, many objects potentially interact, which further complicates the setup. Pushing one object could undo an object that is already at the goal, hence a greedy strategy of just pushing the objects straight to the goal region would fail. In addition, movements, i.e. actions, of the pusher do usually not immediately lead to a change in the cost function, since contact with the object from a suitable side has to be established, for which it is often necessary for the pusher to move around objects . Therefore, applying planning methods that are too local like a cross-entropy method would fail for this scenario. Our prediction model combined with the LS-RRT algorithm solves these tasks efficiently, just from image observations of the scene, see Fig. 1 and the video. In the first row of Tab. 1, we show the total size of the exploration trees for solving the tasks for scenes that contain to objects.
As a baseline comparison, we consider planning with a GNN that uses a fully connected adjacency matrix (Sec. E.2) to understand the importance of a precise dynamics model. From the second row of Tab. 1, one can see that planning with a dense either fails to find a solution as the pusher may not be able to move the object correctly, or generates a plan that is tough to follow in the simulation environment. The reason for this is that with a dense , as shown in Fig. 4b, the model induces too much drift of objects that do not interact, which makes planning for pushing scenarios extremely difficult. Objects can also drift closer or away again to/from the goal by model errors.
Finally, we compare the LS-RRT with a naive control tree algorithm, where we still use our full model, but only sample random nodes in the tree to extend. As shown in the last row of Tab. 1, though the planner can find a path to push a single object, it fails to solve tasks containing more objects within the timeout. This demonstrates the benefits of our object encoder being an implicit function and being able to relate information in the 3D world in terms of center of mass predictions via the learned NeRFs to the latent vectors, both of which make planning much more efficient compared to naive control trees.
Appendix F Experiments – Real World
In the real world experiments, we consider a pushing scenario with different objects on a table. The pusher (blue) is articulated by a Franka Emika Panda robot. All real world experiments use the same 4 camera views (see Fig. 16). As one can see, these 4 cameras are side-views, i.e. in particular, no top-down view is available. Obtaining top-down views in such a setup is challenging, as the robot obstructs the objects in a top-down view. We show in the video renderings from the learned NeRF model from top-down views. We use Intel Realsense cameras. To obtain the object masks in each view, we employ a simple color thresholding method.
The training data consists of image sequences from the different views, only. In particular, the movements of the robot, i.e. the rigid transformations applied to the blue pusher by the robot, are not required to be available.
In this experiment, we consider 4 objects, a shoe (red), a giraffe-shaped toy (yellow), a sand mold (green), and a ball of wool (violet). Although these objects are not rigid, they mainly behave as rigid objects when being pushed. See Fig. 16 for these objects.
We train the framework on a dataset of 2000 random pushes in total, which takes about 8 hours to collect. We randomly apply push directions with a simple heuristics to try to prevent that the objects are being pushed outside of the workspace. Further, this heuristic biases the random push direction sampling to a randomly chosen object, for which we use the depth information of one of the Realsense cameras. Note that except for this heuristic to collect the data more efficiently (otherwise pure random pushes would interact with the objects much more rarely), the depth information from the Realsense cameras is not used.
Half of the training scenes contain the giraffe and the shoe, the other half the shoe, the sand mold, and the ball of wool. The giraffe is never in one scene together with the ball of wool. In the video, we show that the model is still able to reasonably predict the dynamics of a scene containing a giraffe and a wool, even when they interact.
Fig. 17 and Fig. 18 show the performance in terms of the forward predicted image reconstruction error for test scenes from the same distribution as during training. As one can see, our method outperforms the dense adjacency and CNN decoder baselines.
Fig. 19 and Fig. 20 show the performance for scenes that contain a different number (in this case less) of objects than during training. Our method achieves a very low error here, even lower as in Fig. 17 and Fig. 18 where the same number of objects are in the test scenes. In contrast, the dense adjacency matrix baseline has overfit to the number of objects in the scene and performs poorly, although the scenes might seem easier, as they contain less objects.
F.2 Deformable Object
Finally, we consider a real world experiment where a rope is pushed on a table. This rope, visualized in Fig. 21, is deformable in the sense that after the interaction with the pusher it remains in its last, deformed state.
We train the framework on a dataset of 500 random pushes with the same method as in Sec. F.1, which takes about 2 hours to collect. No changes were required for applying our method to such a deformable object.
Please refer to the video for a visualization of the forward predictions of this experiment.
Fig. 22 compares on a test dataset containing 50 scenes the performance of our method with a version where the adjacency matrix is dense and where the decoder is a convolutional neural network instead of a compositional NeRF. This shows that compositionality in our framework has a benefit even if the scene only contains one object and one pusher, as representing objects with individual latent vectors that parameterize individual NeRFs allows us to compute the adjacency matrix from the model’s predictions adaptively, leading to better performance.
Appendix G Network Architectures
The dimension of the latent vectors is for each object. All hidden activation functions are ReLUs.
The MLP that encodes the projected coordinate (see Fig.9a) of the implicit object feature encoder has one layer with output dimension 32. The other MLP in has 2 hidden layers with 128 units each and an output dimension of .
The volumetric feature encoder consists of three 3D convolutional layers with kernel size 3 and channel size 128, each. Layers 2 and 3 have strides of 2. After the convolutional layers, the output is flattened and processed with 3 dense layers with 300 hidden units each.
The NeRF network first lifts the 3D input to 64 dimensions with an MLP, where it is concatinated with the latent vector . This is followed by 3 hidden layers with 300 units each. For the density output , we use a softplus activation and a sigmoid for the color outputs .
Both the edge encoder and the node propagator network have 3 hidden layers with 256 units each.