Learning Multi-Object Dynamics with Compositional Neural Radiance Fields

Danny Driess, Zhiao Huang, Yunzhu Li, Russ Tedrake, Marc Toussaint

Introduction

Learning models from observations that predict the future state of a scene is a fundamental concept for enabling an agent to reason about actions to achieve a desired goal. A major challenge in learning predictive models is that raw observations such as images are usually high-dimensional. Therefore, a common approach is to map the observation space into a lower-dimensional latent representation of the scene via an auto-encoder structure. Based on those latent vectors, a dynamics model can be learned that predicts the next latent state, conditioned on actions an agent takes. An intuition for this is that if a latent vector is sufficient to reconstruct the observations, then it contains enough information about the scene to learn a dynamics model on top of it. While an auto-encoder structure combined with a latent dynamics model is a general approach that is applicable for a large variety of tasks, it raises multiple challenges. First, scenes in our world are composed of multiple objects. Therefore, a fixed-size latent vector has difficulties in generalizing over different and changing numbers of objects in the scene than during training, both due to the limited capacity of fixed-size vectors and lack of diversity in the training distribution. Second, image observations are 2D, but the 3D structure of our world is essential for many tasks to reason about the underlying physical processes governing the dynamics the model should predict. Dealing with occlusions, object permanence, and ambiguities in 2D views is challenging for 2D image representations. Importantly, many forward predictive models in visual observation spaces suffer from instabilities in making long-term predictions, often manifested in blurry image predictions .

One way to address these issues is to incorporate inductive biases and structural priors in the model architectures. Li et al. proposed to use Neural Radiance Fields (NeRFs) as a decoder within an auto-encoder to learn dynamics models in latent spaces. NeRFs exhibit strong structural priors about the 3D world, leading to increased performance over 2D baselines. However, the approach of represents the whole scene as a single latent vector, which we found insufficient for scenes that are composed of multiple, different numbers of objects, both in terms of representation and dynamics prediction.

In the present work, we aim to overcome these challenges by incorporating inductive biases on the compositional nature and underlying 3D structure of our world both in learning the latent representations themselves and the dynamics model. We propose a compositional, object-centric auto-encoder framework whose latent vectors are used to learn a compositional forward dynamics model in that learned latent space based on graph neural networks (GNN). More specifically, we learn an implicit object encoder that maps image observations of the scene from multiple views to a set of latent vectors that each represent an object in the scene separately. These latent object encodings then parameterize individual NeRFs for each object. We apply compositional rendering techniques to synthesize images from multiple viewpoints, which forces the object-centric NeRF functions and the corresponding latent vectors to learn precise 3D configurations of the constituting objects. This 3D inductive bias both in the encoder and the compositional NeRF decoder enables us to incorporate priors from the models’ own predictions about objects interactions via an estimated adjacency matrix into learning the GNN dynamics model, making long-term dynamics predictions more stable. This long term-stability allows us utilize a planning method based on RRTs in the latent space.

In our evaluations, we show through comparisons that non-compositional auto-encoder frameworks and non-compositional dynamics models struggle with tasks containing multiple objects, while our framework generalizes well over different numbers of objects than during training and is capable of generating sharp and stable long-term predictions. Relative to more traditional multibody system identification , these models learn the geometry of unknown objects in addition to (implicitly) learning the inertial and contact parameters. We demonstrate the performance of the approach in terms of image reconstruction error, dynamics prediction error, and planning, generalizing over different numbers of objects than during training. Our experiments include rigid and deformable objects in simulation and with a real robot. To summarize, our main contributions are

A compositional scene encoding framework that uses implicit object encoders and NeRF decoders for each object, forcing the view-invariant latent representation to learn about the 3D structure of the problem in a composable way.

A factored dynamics model in the latent space as a graph neural network (GNN), exploiting the compositional nature of the scene representation and an adaptive adjacency matrix estimated from the model itself to yield stable long-term predictions.

Related Work

Learning Dynamics Models for Compositional Systems. Graph neural networks (GNNs) have shown great promise in introducing relational inductive biases , enabling them to model the dynamics of compositional systems consisting of interactions between multiple objects , large-scale dynamical systems represented using particles and meshes , or from visual observations . Our method differs from prior work by learning compositional scene representations grounded in 3D space directly from visual observations. Our novel combination of implicit object encoders and graph-based neural dynamics models reflects the structure of the underlying scene, which endows our agent with better generalization ability in handling complicated compositional dynamic environments.

NeRF for Compositional and Dynamic Scenes. Recent advances on neural implicit representations have demonstrated widespread success in image synthesis or 3D reconstruction . Notably, Neural Radiance Fields (NeRF) show impressive results on novel-view synthesis . Initial NeRF approaches were trained on a single scene without generalization. Prior work have since proposed to modify neural scene representations to make them compositional for static scenes without considering dynamics of object interactions. People have also extended NeRF to enable view synthesis from a sparse set of views , as well as modeling dynamic scenes by learning implicitly represented flow fields or time-variant latent codes . However, these approaches for dynamic environments typically interpolate over a single time sequence and are not able to handle scenes of different initial configurations or different action sequences, limiting their use in downstream planning and control tasks. Li et al. addressed this issue by combining an NeRF auto-encoding framework with modeling the dynamics in a latent space. Yet, they employed a single latent vector as the whole scene representation, which we will show is insufficient at modeling compositional systems. In contrast, our method considers a graph-based scene representation to capture the structure of the underlying scene and achieves significantly better generalization performance than .

Implicit Models in Robotics. Implicit models in robotics have been explored, e.g., for grasping or more general manipulation planning constraints . Analytic signed distance functions or learned NeRFs are used for trajectory planning. One assumption in and is that signed-distance values are available during training. Our work, in contrast, directly operates on RGB images without requiring explicit 3D shape supervision.

Overview – Compositional Visual Dynamics Learning

Our dynamics learning framework (Fig. 1) consists of three parts, an object encoder Ω\Omega turning observations into a set of latent vectors z1:mz_{1:m}, a compositional NeRF-based decoder DNeRFD_{\text{NeRF}} that renders the latent vectors back into images of the scene to train the encoder, and a graph neural network dynamics model FGNNF_{\text{GNN}} predicting the evolution of the scene in the latent space. This section gives a high-level overview, while Sec. 4, Sec. 5 as well as the appendix Sec. B, Sec. C provide details.

represents the object jj separately. Ω\Omega is trained end-to-end with a NeRF decoder DNeRFD_{\text{NeRF}} reconstructing

for arbitrary views specified by the camera matrix KK from the set of latent object representations z1:mz_{1:m}. The initial observation of the scene is encoded with Ω\Omega into the initial latent vectors z1:m0z_{1:m}^{0}. The GNN dynamics model z1:mt+1=FGNN(z1:mt)z_{1:m}^{t+1}=F_{\text{GNN}}\left(z_{1:m}^{t}\right) then generates long-term predictions of future latent states z1:mtz^{t}_{1:m} that can also be decoded with DNeRFD_{\text{NeRF}} to yield visual predictions from arbitrary views.

Encoding Scenes with Compositional Image-Conditioned NeRFs

Instead of learning Ω\Omega defined in (1) as a direct mapping from images, camera matrices and masks to the latent vectors, we first encode each object in the scene as a feature-valued function over 3D space, conditioned on the image observations. This allows us to incorporate multiple views of the objects in a geometrically consistent way, as well as to apply 3D affine transformations to the objects, which will be important for the dynamics model (Sec. 5). This function is then turned into a latent vector by evaluating it on a workspace set followed by a 3D convolutional network.

Note that the same workspace set Xh\mathcal{X}_{h} is used for all objects. The appendix contains visualizations of the architectures of EE, yy and Φ\Phi.

In summary, the object encoder z1:m=Ω(I1:V,K1:V,M1:m1:V,Xh)z_{1:m}=\Omega\left(I^{1:V},K^{1:V},M^{1:V}_{1:m},\mathcal{X}_{h}\right) maps images from multiple views, object masks and the set Xh\mathcal{X}_{h} to latent vectors. The resulting zjz_{j}’s contain not only the appearance of the objects, but also their spatial configurations in the scene relative to other objects.

2 Decoder as Compositional, Conditional NeRF Model

Compared to this standard NeRF formulation where one single model is used to represent the whole scene, we associate separate NeRFs with each object, meaning that the NeRF for object jj

is conditioned on zjz_{j} for j=1,…,mj=1,\ldots,m. σj\sigma_{j} is the density and cjc_{j} the color prediction for object jj, respectively. To turn those f1:mf_{1:m} back into a global NeRF model that can be rendered to an image, we sum the individual predicted object densities σ(x)=∑j=1mσj(x)\sigma(x)=\sum_{j=1}^{m}\sigma_{j}(x) and obtain the colors as their density weighted combination c(x)=1σ(x)∑j=1mσj(x)cj(x)c(x)=\frac{1}{\sigma(x)}\sum_{j=1}^{m}\sigma_{j}(x)c_{j}(x). These composition formulas have been proposed multiple times in the literature, e.g. . This composition forces the individual NeRFs to learn the 3D configuration of each object individually and therefore ensures that each fjf_{j} only predicts the object where it is located in the 3D space.

To summarize, the compositional NeRF-decoder DNeRFD_{\text{NeRF}} takes the set of latent vectors z1:mz_{1:m} for objects j=1,…,mj=1,\ldots,m and the camera matrix KK for a desired view as input to render I=DNeRF(z1:m,K).I=D_{\text{NeRF}}(z_{1:m},K). Since we only represent the objects and not the background as NeRFs, rendering the composed NeRF will yield an image with the background subtracted. In the experiments, we investigate the importance of the decoder being both compositional and a NeRF.

3 Training

The auto-encoder framework is trained end-to-end on an L2\text{L}_{2} image reconstruction loss for view ii

Since solely the objects are represented as NeRFs and not the background, we compute the union of the masks of the individual objects Mtoti=⋁j=1mMjiM_{\text{tot}}^{i}=\bigvee_{j=1}^{m}M_{j}^{i} and define the target image as Ii∘M^totiI^{i}\circ\hat{M}_{\text{tot}}^{i} with M^toti\hat{M}_{\text{tot}}^{i} being a slightly enlarged union mask. Please refer to the appendix Sec. B for more details.

Latent Dynamics Model with Graph Neural Networks

Having trained the auto-encoder framework, we learn a graph neural network dynamics model

in the latent space, where At∈{0,1}m×mA^{t}\in\left\{0,1\right\}^{m\times m} is the adjacency matrix at time tt. Following , we use multi-step message passing to deal with cases where multiple objects interact within one prediction step. Refer to the appendix Sec. C and Algo. 1 for more details about our GNN dynamics model.

Adjacency Matrix from Learned Model. The adjacency matrix AA in the GNN dynamics model (7) plays an important role in indicating which objects interact. While a dense adjacency matrix, i.e. a graph where each object interacts with all other objects, would in principle work as the GNN could figure out from the latent vectors which objects interact, we found that the long-horizon prediction performance is greatly increased if AA is more selective in reflecting which objects actually interact.

We propose to utilize the NeRF decoder density prediction σj\sigma_{j} for each object to determine the adjacency matrix from the models’ own predictions during training and planning. In order to do so, we define the entries of the adjacency matrix between objects ii and jj based on the collision integral

over the density predictions of the learned NeRF model for a threshold κ≥0\kappa\geq 0. Estimating AA this way takes the actual 3D geometry of the objects in the scene into account and thereby informs the GNN dynamics model, leading to more stable predictions. Please refer to the appendix Sec. C for more details about AA and how it is used in the forward prediction Algo. 1.

Experiments

We demonstrate our framework on pushing tasks in scenes containing multiple objects both in simulation and in the real world. For a quantitative analysis and comparison to multiple baselines, we investigate here the forward prediction error of the model in the image space rendered from the learned model over long-horizons. Please refer to the supplementary video showing the reconstructions of the model, novel scene generation, forward predictions and planning/execution results as well as the appendix for more details and further experiments. Our scenarios are challenging, as they are composed of multiple, interacting objects, sparse rewards, and complex dynamics .

We compare our framework to non-compositional scene representations, non-compositional dynamics models, 2D CNN baselines (visual foresight) without NeRF as decoder, and the importance of estimating the adjacency matrix from the model itself.

Reconstruction and Prediction Performance for Generalization over Numbers of Objects. Fig. 4a shows predictions of the model forward unrolled in time for an action sequence of the red pusher, i.e. applying Algo. 1 (appendix) to an initial scene observation and rendering the predicted latent vectors with the NeRF decoder. Despite the movements in this scene leading to multiple object interactions, even after 38 time steps, the rendered predictions from the model are still sharp and reflect the underlying dynamics. By utilizing the estimated adjacency matrix, there is little drift in the objects, leading to long-term prediction stability. Due to its compositional nature, our model generalizes to scenes that contain more or less objects than in the training set, as shown in Fig. 5 where eight objects plus the pusher are observed and reconstructed with high quality from novel views, although during training the model has seen only and exactly 4 objects.

Comparison to Non-Compositional Scene Representation Baselines. We compare to two non-compositional baselines where the scene is represented globally with one single latent vector per time-step. The dynamics model for these baselines is an MLP zt+1=FMLP(zt,q)z^{t+1}=F_{\text{MLP}}(z^{t},q) that takes the action qq as an additional input. The first baseline (Global NeRF) is the approach from , i.e. we use their CNN encoder instead of our implicit object encoder producing one latent vector conditioning a global NeRF that reconstructs the whole scene. The second baseline (Global 2D CNN auto encoder) uses both a 2D CNN encoder and 2D CNN decoder as well as a single latent vector representing the whole scene. Such frameworks have been used many times in the literature, e.g. and are known as visual foresight. Fig. 3 shows that both global baselines are significantly inferior in our scenarios to our proposed compositional framework, especially for long horizons.

Comparison to 2D Baselines – Importance of NeRF as Decoder. In this section, we replace the NeRF decoder with a 2D CNN decoder to investigate the importance of NeRFs. This decoder takes as input one single latent vector and the camera matrix, i.e. I=DCNN(z,K)I=D_{\text{CNN}}(z,K). In order to make it compositional, we aggregate the set of latent vectors z1:mz_{1:m} from Ω\Omega with a mean operation and then pass the aggregated feature through an MLP to produce the single zz for DCNND_{\text{CNN}}. The rest of the architecture, i.e. implicit object encoder and GNN, stays the same. Since there is no clear way to estimate the adjacency matrix from DCNND_{\text{CNN}}, we use a dense adjacency matrix for the GNN. As one can see in Fig. 3, the long-term prediction performance of the CNN decoder is significantly worse than with a compositional NeRF model as the decoder, especially when asking for numbers of objects that differ from the training distribution. Qualitatively, one can see in Fig. 4c that not only the initial reconstruction is much less sharp compared to the NeRF-based models, but especially also that even after only a few time-steps, the predictions with the CNN decoder are of little use.

Comparison to CNN Encoder. Exchanging the implicit object encoder with a 2D CNN compositional encoder leads to an auto-encoder framework similar to . As seen in Fig. 3, the performance is better compared to the other baselines, but still clearly worse than with the proposed method.

Importance of Estimating the Adjacency Matrix. In Sec. 5, we propose how the adjacency matrix of the GNN can be estimated from the learned NeRFs to increase the long-term stability of the predictions. Here we compare to a dense adjacency matrix, i.e. where the network has to figure out from the latent vectors themselves which objects interact. As one can see in Fig. 3 and Fig. 4b, a dense AA has significantly worse long-horizon prediction performance compared to our proposed way of estimating AA through the learned NeRF model. In the 2 and 8 object case (generalization over numbers of objects), the predictions with the dense AA are useless after only a few time-steps.

Non-Compositional Dynamics. Replacing the GNN with a fully connected MLP z1:mt+1=FMLP(z1:mt)z_{1:m}^{t+1}=F_{\text{MLP}}(z_{1:m}^{t}) leads to worse performance than with a GNN with dense adjacency matrix. This model cannot generalize to different numbers of objects due to its fixed input size.

Summary of Performance Comparisons Our method outperforms all baselines both in terms of pure reconstruction error (as can be seen in Fig. 3 by the error after 0 prediction steps) and its ability to perform long-term predictions forward unrolled on the model’s own predictions. Estimating the adjacency matrix from the model itself is important for long-term stability as it prevents objects from drifting away. Too large drift makes future predictions for a pushing tasks meaningless. Since the reconstruction error of our proposed method without dynamics is better than the baselines, the question arises if the increased performance is an artifact of the lower reconstruction error. We show in Fig. 3d the error in the image space between renderings when having access to the observations at each step and the renderings from the predicted latent vectors into the future after observing the scene only at the beginning. This shows the increase in error relative to the reconstruction process. The results indicate that not solely the reconstruction itself is the reason for the better performance, but that the structural choices of our framework also enable to learn the dynamics more precisely.

2 Planning and Execution Results on Object Sorting Task

To demonstrate the effectiveness of the learned model, we utilize it to solve a box sorting task, where the pusher needs to push colored boxes into their corresponding goal regions as shown in Fig. 6. This task is inspired by and involves multiple challenges: As multiple objects interact, a greedy strategy of pushing objects straight to the goal region fails. Movements, i.e. actions, of the pusher do often not immediately lead to a change in the cost function, since contact with the object from a suitable side has to be established . In the appendix Sec. D we propose a latent space RRT that uses our framework for planning. Refer to the appendix and the video for more details about our proposed planning algorithm and comparisons to baselines.

3 Real World Experiments

Fig. 1a shows the rendered forward predictions of our model for a real world scenario where a robot pushes a shoe and a giraffe-shaped toy. Fig. 1b are renderings from novel view points.

4 Applicability to Deformable Objects

The experiments so far focused on objects that behave mainly like rigid objects when being pushed. In Fig. 7 we show that our method is also applicable to deformable objects simulated with .

Discussion & Limitations

Computational Efficiency. Our framework is computationally more demanding during inference time than 2D CNN decoder baselines, mainly due to NeRF evaluations. Many methods have been developed to increase the speed of NeRF , from which our framework could benefit.

Object Masks. The compositional scene encoding framework requires object masks to achieve compositionality. Many mature methods for instance segmentation have been developed such that we believe having masks is a reasonable assumption. However, one could add an output to our implicit encoder that provides object labels or use mechanisms similar to slot attention .

Latent Representations. We have shown the great benefits of a compositional latent representation as it not only provides generalization over different numbers of objects in the scene, but also leads to increased reconstruction and dynamics prediction performance compared to non-compositional baselines. Furthermore, latent representations compress observations, enabling efficient dynamics prediction. However, as still each object in the scene is represented as a latent vector of finite size, latent models are capable of mainly representing objects with shapes similar to the training distribution. To address this, compositionality could not only be introduced on the scene level, but also by representing objects themselves in a composable way.

Long-Term Prediction Stability. Our dynamics model framework exhibits significantly better long-term prediction stability compared to baselines. Our experiments indicate that this is due the structural biases enabled through (compositional) NeRFs. This stability allowed us to use the model for planning scenarios requiring long-horizons, which none of the baseline methods could support. However, we believe that there is still room for improvement regarding the prediction stability. Especially for deformable objects, we observed that after many prediction steps objects in the scene are predicted to penetrate or move through each other, which could be improved in future work.

Conclusion

Visual dynamics models are of high interest to the computer vision and robotics community, as they avoid explicit shape model assumptions and imply end-to-end perception. However, to support manipulation planning and reasoning, we need models that generalize strongly over objects and provide stable long-term predictions. In this paper we proposed a system that introduces 3D structural and compositional priors at various levels, namely compositional NeRFs, 3D implicit object encoders, and GNNs dynamics with an adaptive adjacency matrix. Together our system exhibits significantly stronger long-term prediction performance compared to multiple baselines without these priors or without compositionality, and supports using a latent space RRT planner. We have shown generalization over different numbers of objects, notably up to two times more than during training.

This research has been supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy – EXC 2002/1 “Science of Intelligence” – project number 390523135. Danny Driess thanks the International Max-Planck Research School for Intelligent Systems (IMPRS-IS) for the support. The authors thank Valentin Hartmann for discussions regarding RRTs.

References

Appendix A Expanded Related Work

Model-based planning algorithms typically build a dynamics model of the environment and then use the model to plan the agent’s behavior in order to minimize some task objectives. We can roughly categorize the methods by whether the model is constructed from first principles (i.e., physical rules) or learned from data (i.e., data-driven models). Physics-based models typically require complete information about the objects’ geometry and the system’s state , which limits their applicability in robotic manipulation tasks involving unknown object models and partially observable states. Data-driven methods, on the other hand, learn a dynamics model directly from the robot’s interaction with the environment and have shown impressive results in manipulation tasks ranging from closed-loop planar pushing to complicated dexterous manipulation . Many of the data-driven planning frameworks learn dynamics models directly from visual observation based on representations defined at different levels of abstraction, such as pixel space , 3D volumetric space , signed-distance fields , keypoint space , and low-dimensional latent space . Approaches commonly employ an image reconstruction loss , an self-supervised time contrastive loss , or jointly train a forward and an inverse dynamics model to make sure that the representation encodes meaningful information about the environment. Our method takes a step forward by learning graph-based latent representations from visual observations. The learned model accurately encodes the underlying 3D contents, allowing our learned model to achieve precise manipulation of compositional environments and generalize outside the training distribution, i.e. to scenes with more (and less) objects than during training.

Appendix B Details – Encoding Scenes with Compositional Image-Conditioned NeRFs

This section provides details – especially visualizations of the network architectures – on our proposed compositional auto-encoder framework.

Fig. 8 visualizes the whole auto-encoder architecture, where the encoder

maps posed images (I,K)1:V(I,K)^{1:V} from VV-many views of mm-many objects including their object masks M1:m1:VM^{1:V}_{1:m} as well as an workspace set Xh\mathcal{X}_{h} to a set of latent vectors z1:mz_{1:m} describing the objects in the scene. The same object encoder Ω\Omega and workspace set Xh\mathcal{X}_{h} is used for all objects. In particular, Xh\mathcal{X}_{h} is not a 3D bounding box for an individual object, but covers the whole workspace of the scene. See Fig. 13 for a visualization of the workspace set Xh\mathcal{X}_{h}.

Similar architectures of computing such pixel features from world coordinates have been proposed, e.g., in for the single object case. However, we use EE quite differently compared to these works, as we compute a latent vector from pixel-aligned features with 3D convolutions. The object feature function yjy_{j}, cf. (3), aggregates the features from the individual views by taking into account the masks of object jj in each view. The architectures of EE and yy are visualized in Fig. 9. Fig. 10 shows how an object feature function yj(⋅)y_{j}(\cdot) for object jj is turned by Φ\Phi into its corresponding latent vector zjz_{j} by querying it on the workspace set Xh\mathcal{X}_{h} followed by a 3D convolutional network.

B.2 Background on Neural Radiance Fields

B.3 Training

The auto-encoder framework is trained end-to-end on an L2\text{L}_{2} image reconstruction loss. Since, as mentioned, solely the objects are represented as NeRFs and not the background, we compute the union of the masks of the individual objects

and define the target image in a view as Ii∘MtotiI^{i}\circ M_{\text{tot}}^{i} with ∘\circ denoting the element-wise product.

A known issue of NeRF is its computational efficiency , since for every pixel all fjf_{j}’s have to be queried on many points along the camera ray. We make two simple, but important, improvements to reduce the computational demand.

First, the near and far bounds αn\alpha_{n}, αf\alpha_{f} are determined individually for each camera ray such that only those points along the rays that are within the workspace set X\mathcal{X} are considered. This is a reasonable assumption since we assumed that the objects are in the workspace set in the first place. That way, the computational efficiency is already greatly increased by reducing the number of points where functions fjf_{j}’s have to be queried.

Moreover, as the scenes we consider in our experiments (see for example Fig. 13) are composed of multiple smaller objects, when masking out the background, the majority of pixels in each view is black and therefore does not contain information about the scene, although the model is evaluated on those areas. To further decrease the number of points where the NeRFs have to be queried, we only consider those rays for a view that pass through the mask of at least one object in that view. It turned out, however, that training only on rays that go through MtotM_{\text{tot}} leads to blurry reconstructions, since there is no loss indicating that the objects should end outside of the masks. In order to resolve this, we enlarge the combined mask MtotiM_{\text{tot}}^{i} of a view with a convolution operation by a few pixels. We denote this enlarged mask by M^toti\hat{M}_{\text{tot}}^{i}. Together, these techniques ensure that the model learns sharp object boundaries, while significantly reducing the number of considered rays and required NeRF evaluations. See Fig. 11 for a visualization of this procedure.

These considerations lead to the following training objective of DNeRFD_{\text{NeRF}} and Ω\Omega for a view ii

During training, we randomly sample a view from the dataset for each mini-batch and update the parameters of DNeRFD_{\text{NeRF}} and Ω\Omega using the ADAM optimizer .

Another side effect of training on the enlarged masks M^toti\hat{M}_{\text{tot}}^{i} only is that it improved the training stability and reconstruction qualities of the model. Indeed, when we trained the model on the whole image, depending on the weight initialization of the network, the model sometimes very quickly converged to a state where it only predicted a black image, since the majority of pixels are actually black and hence a low loss could be achieved. Training on the enlarged masks prevents this reliably.

B.4 Rigid Transformations and Novel Scene Generation

Please note the slight abuse of notation here, the term R(q)T(Xh−s(q))R(q)^{T}\left(\mathcal{X}_{h}-s(q)\right) has to be understood elementwise for each entry in Xh\mathcal{X}_{h}.

Composing scenes via rigid transformations applied to the input of the individual NeRFs ff has been considered before, e.g. in . However, transforming a NeRF by applying the rigid transformation to its xx input only leads to changes in the rendered visual space, i.e. it has, in particular, no influence on the latent vectors of the objects. Since we want the latent vectors to represent not only the appearance of an individual object, but the geometric information of the object within the scene relative to other objects, just transforming the NeRF models is not sufficient. Therefore, using (15) we get the latent vector rigidly transformed by qq, which is crucial for our downstream dynamics prediction task.

This rigid transformation is important to define actions for the GNN dynamics model, as described in Sec. C.3 and Algo. 1.

Appendix C Details – Graph Neural Network Latent Dyanmics Model

This section provides details about the latent dynamics model. Especially relevant is Sec. C.6 and Algo. 1 where we describe the forward prediction algorithm.

Due to the compositional nature of the scenarios we consider, we require a dynamics model that maintains the capabilities of our auto-encoder to generalize over changing numbers of objects, for which graph neural networks (GNNs) are a natural choice.

The general idea behind learning dynamics models with GNNs is to associate each object in the scene with a node in a graph, which, in our case, means that each node in the graph is a latent vector zjz_{j}. Edges between the nodes indicate if objects interact, e.g. by exchanging forces due to contact. As argued in , applying a simple GNN to the problem of dynamics prediction is problematic, since interactions caused at one node can influence not only the neighboring nodes, but higher-order neighbors. For example, if three objects touch, the effects of applying a force at the first object have to propagate. The scenarios we consider in the experiments contain multiple objects such that more than two objects can interact in one time step. To take this into account, we use a message passing architecture inspired by .

Let ziz_{i} and zjz_{j} be the latent vectors of objects ii and jj. An edge encoder network FeF_{e} determines a feature

describing the interaction between the objects ii, jj. An adjacency matrix A∈{0,1}m×mA\in\left\{0,1\right\}^{m\times m} has entry Aij=1A_{ij}=1 if object ii is influenced by object jj. Assume the state of all latent vectors z1:mtz_{1:m}^{t} at time tt is known. The node propagator network FzF_{z} recursively is queried LL many times to propagate the state z1:jtz_{1:j}^{t} to the next time step t+1t+1 as follows:

with \prescript0zit+1=zit\prescript{0}{}{z_{i}}^{t+1}=z_{i}^{t}, \prescriptleijt=Fe(\prescriptlzit+1,\prescriptlzjt+1)\prescript{l}{}{e}_{ij}^{t}=F_{e}\left(\prescript{l}{}{z_{i}}^{t+1},\prescript{l}{}{z_{j}}^{t+1}\right) for l=0,…,L−1l=0,\ldots,L-1 and the final new predicted state zit+1=\prescriptLzit+1z^{t+1}_{i}=\prescript{L}{}{z}_{i}^{t+1}.

C.2 Adjacency Matrix from Learned Model

The adjacency matrix AA in the GNN dynamics model (17) plays an important role in indicating which objects interact. While a dense adjacency matrix, i.e. a graph where each node is connected to every other node implying that each object interacts with all other objects in the scene, would in principle work as the network could figure out from the latent representations itself which objects interact, we found that the long-horizon prediction performance is greatly increased if AA is more selective in reflecting which objects actually interact (refer to the experiments in Sec. E.2). This is especially relevant for compositional scenes as considered in this work where there are many objects, but which often do not interact with each other in every timestep.

A central question is how the adjacency matrix can be obtained from the observations of the scene without manually specifying it. Due to our model having strong 3D priors, we can exploit the density prediction σj\sigma_{j} as defined in (5) for each object to determine the adjacency matrix from the models’ own predictions during training and planning. In order to do so, for a threshold κ≥0\kappa\geq 0, the collision integral

over the density predictions of the learned NeRF model for objects ii and jj indicates if the two objects overlap or not. A similar integral as in (18) has been proposed in to estimate collisions from signed-distance functions. Based on this integral, we define the entries of the adjacency matrix between objects ii and jj as

which implies that only those objects that are or are close to being in contact potentially interact. Estimating AA this way takes the actual geometry of the objects in the scene into account. In relation to the node propagation network \eqrefeq:nodePropNetwork\eqref{eq:nodePropNetwork}, this means that the adjacency matrix at step ll of the propagation becomes a function of the node encodings itself, i.e.

For training the GNN, however, changing the adjacency matrix during prediction is not differentiable. Therefore, we compute the adjacency matrices from the model such that they are constant within one time-step as follows. We first compute an occupancy grid

for each object jj over the discretized workspace set Xh\mathcal{X}_{h} and then apply a 3D convolution operation on SjS_{j} with a kernel consisting of only ones to expand the occupancy grid. The now constant within one time-step tt entries \prescriptlAijt=Aijt\prescript{l}{}{A}_{ij}^{t}=A_{ij}^{t} for all l=0…,L−1l=0\ldots,L-1 are then determined by checking if there is a voxel cell where both enlarged SiS_{i} and SjS_{j} have value one. The size of the convolution kernel is chosen large enough such that the adjacency matrix is not going to change within one timestep. This allows for a trade-off between sufficient sparsity of AA while ensuring that all objects that potentially interact have corresponding entries in AA.

In the experiments in Sec. E.2, we investigate the influence on the prediction performance for multiple different ways of predicting/using the adjacency matrix.

C.3 Actions

So far, the way we have formulated the graph neural network dynamics model in Sec. C.1 does not contain a notion of actions. Instead, we interpret an action as a modification to a node in the graph and train the GNN to predict the state of the nodes at the next time step as a result to this modification. This allows us to not explicitly distinguish between controlled and uncontrolled/passive objects.

C.4 Quasi-Static Dynamics

If we assume quasi-static dynamics, meaning that the next system state only depends on the current latent state z1:mz_{1:m} without history and immediate actions (refer to the discussion in the last paragraph Sec. C.3 about the notion of actions in this work), we can further increase the long-term stability by utilizing the adjacency matrices estimated by the learned model. When an object is not involved in any interactions with other objects, then, under quasi-static assumptions, it does not change between time steps, i.e. its latent vector stays constant, which means we can set

The condition ∀j≠i : Aij=0\forall_{j\neq i}~{}:~{}A_{ij}=0 means that the node associated with ziz_{i} in the graph has no incoming edges. As we will show in the experiments (Sec. E.2), this can greatly increase the stability for long-term open-loop model predictions as it will prevent drift in objects that do not take part in any interaction with other objects.

C.5 Training

We first train the compositional NeRF auto-encoder framework on training data, which gives us a dataset of trajectories of latent vectors. The GNN dynamics model is then trained on the one-step mean squared error between z1:mt+1z^{t+1}_{1:m} and z1:mtz^{t}_{1:m} of samples of such trajectories using the ADAM optimizer.

Importantly, training the dynamics model does not require a dataset containing the actions. It is sufficient to have video sequences and knowledge which object was the articulated one. At inference time, one can also choose different objects to apply actions to in terms of rigid transformations.

C.6 Forward Prediction Algorithm

Algorithm 1 summarizes the forward prediction procedure. The algorithm is visualized in Fig. 12.

Starting from a single initial observation of the scene in terms of the images I1:VI^{1:V} from VV many views and the objects masks M1:m1:VM^{1:V}_{1:m} of the mm many objects, Algorithm 1 predicts the latent vectors z1:mtz_{1:m}^{t} at times t=1,…,Tt=1,\ldots,T for all objects in the scene given a desired action sequence q1:Tq_{1:T} of rigid transformations applied to object aa. At every point tt in time, the scene can be rendered from arbitrary view points from the predicted z1:mtz_{1:m}^{t}. Note that the masks of the objects are required only for the initial observation, i.e. no mask prediction has to be performed as compared to .

In line 2, the initial object encodings from the scene observation are computed. Line 5 applies the action to the object with index aa. Lines 6-10 then perform the prediction step of the GNN in the latent space using message passing. Crucially, in line 7, the adjacency matrix is estimated from the current predictions during message passing (Sec. C.2). Note that here the original collision integral (18), computed on the grid Xh\mathcal{X}_{h}, can be used without enlarging the intermediate occupancy grids, since, if during the message passing step objects interact that previously did not, it will be captured, as AA is estimated in every step of the message passing part. This leads to further increased prediction stability, as we will show in the experiments. Finally, in line 12, objects that have not interacted with other objects as predicted by the adjacency matrix are kept at their previous latent state (quasi-static assumption from Sec. C.4)

Appendix D Planning with Latent space RRT

In this section, we propose a planning algorithm to manipulate objects to achieve a desired goal using our scene encoding and dynamics model framework. Note that planning and control is not the main focus of this work, however, the algorithm still contains important insights.

The main part of the planning algorithm is an RRT in the latent space. Such latent space RRTs have been considered, for example, in . One central question here is how one can sample in the latent space effectively, since a uniform random sample in the latent space not necessarily is a valid (and/or uniform) sample in the original space. In , they assume to have access to a set of valid latent vectors from which they can sample. In contrast, we can produce valid samples in the latent space directly by exploiting the properties of our model.

On a high level, our model iteratively perceives multi-view images of the scene, finds a plan using a Latent-Space RRT (LS-RRT) based on the forward predictions of the model over a long horizon, and then executes the found plan with Model-Predictive Control (MPC) for a shorter horizon. We describe the algorithm here with pushing scenarios as considered in the experiments in mind.

Algorithm 2 summarizes the LS-RRT algorithm. We grow a tree in the latent space, starting at the latent vector z0=z1:m0z^{0}=z_{1:m}^{0} that represents the current state of the environment, encoded by the implicit object encoder Ω\Omega from the current visual observation of the scene. In standard RRTs, a target is uniformly sampled in the configuration space to steer the growth of the tree towards a Voronoi bias. To introduce a particular goal-targeted sampling bias and as we do not have an inverse model or steering function, we modify the standard approach as follows:

In a latent space RRT, sampling uniformly in the latent space does neither guarantee that the samples are from the latent space manifold nor that they explore the original space. Therefore, we sample random targets g∼Gsamplerg\sim\mathcal{G}_{\text{sampler}} not in the latent space directly, but only in the space of center of mass configurations of all objects, which is of dimension 2m2m in the experiments (objects and pusher). In this way, we can design a sampling distribution Gsampler\mathcal{G}_{\text{sampler}} biased to target configurations that have low costs, i.e., more objects within the goal region, or targets in which the articulated object (the pusher in the experiments) is close to one of the objects, inducing a bias for interaction. This sampling distribution and cost evaluation Cg\mathcal{C}_{g} is possible because we can apply rigid transformations to the objects through our object encoder being an implicit function, since, for a sampled random target, we have to move the objects to this target to check the cost on the transformed configuration. Further, the metric dd to select the expanded node is the L2L_{2}-norm in the full configurations between gg and the centers-of-masses computed from zz. Using the predictions of the NeRF model, we can estimate (under homogeneous density assumption) the center of mass of an object with latent vector zjz_{j} as

i.e. d(z,g)=∥x1:mcom(z)−g∥2d(z,g)=\left\|x_{1:m}^{\text{com}}(z)-g\right\|_{2}. Note that the sampling distribution, cost function evaluation and metric calculation are done solely based on predictions of the model. At no point the model has access to ground truth center-of-mass information.

Finally, as we do not have an inverse dynamics model or an other kind of steering function, we expand the tree using a random action qq, similar to control trees. However, our goal-targeted node selection ensures that the tree expands effectively.

D.2 Cost-Functions

As our decoder is based on NeRFs, our method is able to provide a lot of flexibility in defining cost functions. At any time instance forward predicted by the dynamics model, we can render an image from an arbitrary view or reconstruct the objects in 3D. This enables, for example, to have the following options to define cost functions for planning:

Loss on image rendered from arbitrary views (with known camera matrix).

Loss on density prediction of the NeRF reconstruction.

Loss on point-cloud reconstructed from σj\sigma_{j} (including color information on each point).

Loss on center-of-mass predicted by the model via (23).

All these options can be defined for specific objects, the whole scene, or anything in between.

D.3 Model-based Control

Although our model achieves impressive performance over a long horizon, the accumulated prediction errors may still lead to a failure when executing the plans open-loop. We therefore apply an MPC scheme, which in each cycle feeds the current visual observation into the model, samples and select actions that match the plan (in terms of the center-of-mass metric) within a short horizon predicted by the learned dynamics model. If there is a significant mismatch between the plan and current observation, the LS-RRT is used again to find a new long-term plan starting from the current observation.

Appendix E Experiments – Simulation

In the simulated experiments, we focus on a pushing task in scenarios with multiple box-shaped objects on a table, see, e.g. Fig. 13 or 11 for such scenes. For a quantitative analysis and comparison to multiple baselines, we investigate the forward prediction error of the model both in the image space (Fig. 3) and, for baselines that use a compositional NeRF decoder, the error in predicting the center of mass of the objects (Fig. 14) over long-horizons. In all plots of Fig. 3 and Fig. 14, the blue curve corresponds to our proposed framework as summarized in Algo. 1. Sec. E.4 presents planning and execution results for a challenging box sorting task.

Please refer to the supplementary material for videos showing the reconstructions of the model, forward predictions, novel scene generation, and planning/execution results.

We consider a rigid-body scenario with multiple objects on a table, see Fig. 13 for an example. In all cases, the red cylinder is the pusher that is articulated in order to push the other objects around.

This scenario is challenging due to multiple reasons. First, it is composed of many objects, which implies not only a broad scene distribution, but especially also that many objects can interact. The mechanics of such multi-body pushing is non-trivial, since, for instance, contact can be established and broken between the objects at multiple phases of the motion. Contact between multiple objects at the same time can occur. Furthermore, we do not assume that the red pusher starts in contact with an object. Hence, if a task implies that an object should be pushed, long-term predictions inherently have to be made in order to establish contact, before any object movement is registered.

All scenes in the training data contain 4 box-shaped objects of randomly sampled sizes, positions and orientations (5 dimensional parameter space for each of the 4 objects) and one cylinder-shaped object with randomly sampled position.

To generate the training data, we randomly sample one of the 4 objects and then move the red pusher towards the center of this chosen object (with Gaussian noise added to the direction vector in each time step) until either the pusher leaves the workspace, in which case a new target object is chosen, or an object is pushed outside the workspace, in which case the data collection for this scene is terminated and a new scene is sampled. In total, the training dataset contains 5752 scenes with an average sequence length of 17. We generate 3 test datasets for evaluating the reconstruction and prediction performance which contain 2, 4, and 8 objects, respectively, plus the pusher. There are 312 scenes for each test dataset, generated with a different random seed than the training data. As visualized in Fig. 13 by the green coordinate systems, we choose 4 camera views for each scene.

E.2 Importance of Estimating the Adjacency Matrix

This section provides a more detailed investigation of the importance of the adjacency matrix than in the main text, where we only have discussed a dense adjacency matrix.

In Sec. C.2, we have proposed how the adjacency matrix of the GNN can be estimated from the density predictions of the learned NeRFs and that under quasi-static assumptions this estimated adjacency matrix can further be exploited to increase the long-term stability of the predictions, cf. Sec. C.4. Here we investigate the consequences of utilizing the adjacency matrix this way by comparing the full Algorithm 1 to the following three ablations. In all these ablations, the rest of the method remains the same, i.e. same object encoder, same GNN, same compositional NeRF decoder.

In this case, line 12 of Algorithm 1 is not used, i.e. the latent vectors of all objects, even when they do not interact with other objects as estimated through the model, are updated using the model forward predictions.

Here, we estimate the adjacency matrix only at the beginning of the message passing step, i.e. before line 6 in Algorithm 1. In order to ensure that it can still capture all object interactions that might occur during the message passing step, we enlarge the determined occupancy grids exactly the same way as for training, see the discussion in Sec.C.2. The effects of this are that objects that are close to each other but do not interact still have entries in AA indicating that they interact, which means slight errors in the predictions accumulate and lead to drift, although the object would not move in reality.

We further consider a dense adjacency matrix, i.e. where the network has to figure out from the latent vectors themselves if objects interact. Preventing drift in this case is considerably harder.

In Fig. 14 one can see the mean error of the model predicting the center of mass, computed from its density predictions of the NeRFs for each object according to (23), over the number of steps predicted into the future on the test dataset for different numbers of objects in the scene.

As one can see in Fig. 14a, for the two object case, the choices of how the adjacency matrix is used, as long as it is not a dense one, are not significant. For the 4 (Fig. 14b) and 8 (Fig. 14c) object case, however, our proposed utilization of the adjacency matrix, i.e. estimating it during propagation steps and using it to exploit the quasi-static assumption, leads to a significant increase in performance. This can be explained by the fact that utilizing the adjacency matrix as we propose leads to significantly less drift. Especially with the dense adjacency matrix, the predictions are very unstable for all, the 2, 4, and 8 object case. Fig. 4b shows this qualitatively. In the 8 object case, the predictions with the dense AA are basically useless after only a few time-steps, showing that it has overfitted to the number of 4 objects as in the training data.

E.3 Advantages of Implicit Object Encoder – Comparison to CNN Encoder

This experiment has already be mentioned in the main text, but here we provide more details. We exchange the implicit object encoder with a 2D CNN object encoder. The resulting auto-encoder framework is very similar to the architecture of . More specifically, we encode each masked image observation with a 2D CNN to produce a feature vector. The feature vectors from the different views are aggregated into the final latent vectors for each object. We use the encoder architecture from , but adjust it to the compositional multi-object case by incorporating object masks. Since this encoder is not an implicit function of X\mathcal{X}, we cannot modify the latent vectors by applying rigid transformations and hence need to encode the actions differently. In order to do so, we train a separate MLP network that predicts the latent vector of the pusher resulting from applying an action to it. This gives the modified zaz_{a} for line 5 in algorithm 1 for the CNN encoder baseline. The rest of the architecture, i.e. the GNN, the compositional NeRF decoder, estimating the adjacency matrix during message passing from the model, etc., stays the same.

As can be seen in Fig. 14a and Fig. 3, replacing the proposed implicit object encoder with a CNN encoder, the performance is better compared to the other baselines, but still clearly worse than with the proposed method.

E.4 Planning and Execution Results on Box Sorting Task

To demonstrate the effectiveness of the learned model, we utilize it to solve a challenging box sorting task, where the red pusher needs to push the blue and yellow boxes into their corresponding goal regions as shown in Fig. 6. This task is in part inspired by the object sorting task in . The cost function Cg\mathcal{C}_{g} in Algorithm 2 determines how many objects are outside of their goal region, which is computed for each object jj from their corresponding density σj\sigma_{j} and color cjc_{j} predictions of the model itself. The goal is fulfilled if all objects are in their respective goal regions.

This object-sorting task is challenging for multiple reasons. First, the dynamics of pushing is non-trivial . In our particular case, many objects potentially interact, which further complicates the setup. Pushing one object could undo an object that is already at the goal, hence a greedy strategy of just pushing the objects straight to the goal region would fail. In addition, movements, i.e. actions, of the pusher do usually not immediately lead to a change in the cost function, since contact with the object from a suitable side has to be established, for which it is often necessary for the pusher to move around objects . Therefore, applying planning methods that are too local like a cross-entropy method would fail for this scenario. Our prediction model combined with the LS-RRT algorithm solves these tasks efficiently, just from image observations of the scene, see Fig. 1 and the video. In the first row of Tab. 1, we show the total size of the exploration trees for solving the tasks for scenes that contain 11 to 66 objects.

As a baseline comparison, we consider planning with a GNN that uses a fully connected adjacency matrix (Sec. E.2) to understand the importance of a precise dynamics model. From the second row of Tab. 1, one can see that planning with a dense AA either fails to find a solution as the pusher may not be able to move the object correctly, or generates a plan that is tough to follow in the simulation environment. The reason for this is that with a dense AA, as shown in Fig. 4b, the model induces too much drift of objects that do not interact, which makes planning for pushing scenarios extremely difficult. Objects can also drift closer or away again to/from the goal by model errors.

Finally, we compare the LS-RRT with a naive control tree algorithm, where we still use our full model, but only sample random nodes in the tree to extend. As shown in the last row of Tab. 1, though the planner can find a path to push a single object, it fails to solve tasks containing more objects within the timeout. This demonstrates the benefits of our object encoder being an implicit function and being able to relate information in the 3D world in terms of center of mass predictions via the learned NeRFs to the latent vectors, both of which make planning much more efficient compared to naive control trees.

Appendix F Experiments – Real World

In the real world experiments, we consider a pushing scenario with different objects on a table. The pusher (blue) is articulated by a Franka Emika Panda robot. All real world experiments use the same 4 camera views (see Fig. 16). As one can see, these 4 cameras are side-views, i.e. in particular, no top-down view is available. Obtaining top-down views in such a setup is challenging, as the robot obstructs the objects in a top-down view. We show in the video renderings from the learned NeRF model from top-down views. We use Intel Realsense cameras. To obtain the object masks in each view, we employ a simple color thresholding method.

The training data consists of image sequences from the different views, only. In particular, the movements of the robot, i.e. the rigid transformations applied to the blue pusher by the robot, are not required to be available.

In this experiment, we consider 4 objects, a shoe (red), a giraffe-shaped toy (yellow), a sand mold (green), and a ball of wool (violet). Although these objects are not rigid, they mainly behave as rigid objects when being pushed. See Fig. 16 for these objects.

We train the framework on a dataset of 2000 random pushes in total, which takes about 8 hours to collect. We randomly apply push directions with a simple heuristics to try to prevent that the objects are being pushed outside of the workspace. Further, this heuristic biases the random push direction sampling to a randomly chosen object, for which we use the depth information of one of the Realsense cameras. Note that except for this heuristic to collect the data more efficiently (otherwise pure random pushes would interact with the objects much more rarely), the depth information from the Realsense cameras is not used.

Half of the training scenes contain the giraffe and the shoe, the other half the shoe, the sand mold, and the ball of wool. The giraffe is never in one scene together with the ball of wool. In the video, we show that the model is still able to reasonably predict the dynamics of a scene containing a giraffe and a wool, even when they interact.

Fig. 17 and Fig. 18 show the performance in terms of the forward predicted image reconstruction error for test scenes from the same distribution as during training. As one can see, our method outperforms the dense adjacency and CNN decoder baselines.

Fig. 19 and Fig. 20 show the performance for scenes that contain a different number (in this case less) of objects than during training. Our method achieves a very low error here, even lower as in Fig. 17 and Fig. 18 where the same number of objects are in the test scenes. In contrast, the dense adjacency matrix baseline has overfit to the number of objects in the scene and performs poorly, although the scenes might seem easier, as they contain less objects.

F.2 Deformable Object

Finally, we consider a real world experiment where a rope is pushed on a table. This rope, visualized in Fig. 21, is deformable in the sense that after the interaction with the pusher it remains in its last, deformed state.

We train the framework on a dataset of 500 random pushes with the same method as in Sec. F.1, which takes about 2 hours to collect. No changes were required for applying our method to such a deformable object.

Please refer to the video for a visualization of the forward predictions of this experiment.

Fig. 22 compares on a test dataset containing 50 scenes the performance of our method with a version where the adjacency matrix is dense and where the decoder is a convolutional neural network instead of a compositional NeRF. This shows that compositionality in our framework has a benefit even if the scene only contains one object and one pusher, as representing objects with individual latent vectors that parameterize individual NeRFs allows us to compute the adjacency matrix from the model’s predictions adaptively, leading to better performance.

Appendix G Network Architectures

The dimension of the latent vectors zjz_{j} is k=64k=64 for each object. All hidden activation functions are ReLUs.

The MLP that encodes the projected coordinate (see Fig.9a) of the implicit object feature encoder EE has one layer with output dimension 32. The other MLP in EE has 2 hidden layers with 128 units each and an output dimension of no=64n_{o}=64.

The volumetric feature encoder Φ\Phi consists of three 3D convolutional layers with kernel size 3 and channel size 128, each. Layers 2 and 3 have strides of 2. After the convolutional layers, the output is flattened and processed with 3 dense layers with 300 hidden units each.

The NeRF network ff first lifts the 3D input to 64 dimensions with an MLP, where it is concatinated with the latent vector zz. This is followed by 3 hidden layers with 300 units each. For the density output σ\sigma, we use a softplus activation and a sigmoid for the color outputs cc.

Both the edge encoder FeF_{e} and the node propagator network FzF_{z} have 3 hidden layers with 256 units each.