Entity Abstraction in Visual Model-Based Reinforcement Learning

Rishi Veerapaneni, John D. Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua B. Tenenbaum, Sergey Levine

Introduction

A powerful tool for modeling the complexity of the physical world is to frame this complexity as the composition of simpler entities and processes. For example, the study of classical mechanics in terms of macroscopic objects and a small set of laws governing their motion has enabled not only an explanation of natural phenomena like apples falling from trees but the invention of structures that never before existed in human history, such as skyscrapers. Paradoxically, the creative variation of such physical constructions in human society is due in part to the uniformity with which human models of physical laws apply to the literal building blocks that comprise such structures – the reuse of the same simpler models that apply to primitive entities and their relations in different ways obviates the need, and cost, of designing custom solutions from scratch for each construction instance.

The challenge of scaling the generalization abilities of learning robots follows a similar characteristic to the challenges of modeling physical phenomena: the complexity of the task space may scale combinatorially with the configurations and number of objects, but if all scene instances share the same set of objects that follow the same physical laws, then transforming the problem of modeling scenes into a problem of modeling objects and the local physical processes that govern their interactions may provide a significant benefit in generalizing to solving novel physical tasks the learner has not encountered before. This is the central hypothesis of this paper.

We test this hypothesis by defining models for perceiving and predicting raw observations that are themselves compositions of simpler functions that operate locally on entities rather than globally on scenes. Importantly, the symmetry that all objects follow the same physical laws enables us to define these learnable entity-centric functions to take as input argument a variable that represents a generic entity, the specific instantiations of which are all processed by the same function. We use the term entity abstraction to refer to the abstraction barrier that isolates the abstract variable, which the entity-centric function is defined with respect to, from its concrete instantiation, which contains information about the appearance and dynamics of an object that modulates the function’s behavior.

Defining the observation and dynamic models of a model-based reinforcement learner as neural network functions of abstract entity variables allows for symbolic computation in the space of entities, but the key challenge for realizing this is to ground the values of these variables in the world from raw visual observations. Fortunately, the language of partially observable Markov decision processes (POMDP) enables us to represent these entity variables as latent random state variables in a state-factorized POMDP, thereby transforming the variable binding problem into an inference problem with which we can build upon state-of-the-art techniques in amortized iterative variational inference to use temporal continuity and interactive feedback to infer the posterior distribution of the entity variables given a sequence of observations and actions.

We present a framework for object-centric perception, prediction, and planning (OP3), a model-based reinforcement learner that predicts and plans over entity variables inferred via an interactive inference algorithm from raw visual observations. Empirically OP3 learns to discover and bind information about actual objects in the environment to these entity variables without any supervision on what these variables should correspond to. As all computation within the entity-centric function is local in scope with respect to its input entity, the process of modeling the dynamics or appearance of each object is protected from the computations involved in modeling other objects, which allows OP3 to generalize to modeling a variable number of objects in a variety of contexts with no re-training.

Contributions: Our conceptual contribution is the use of entity abstraction to integrate graphical models, symbolic computation, and neural networks in a model-based reinforcement learning (RL) agent. This is enabled by our technical contribution: defining models as the composition of locally-scoped entity-centric functions and the interactive inference algorithm for grounding the abstract entity variables in raw visual observations without any supervision on object identity. Empirically, we find that OP3 achieves two to three times greater accuracy than state of the art video prediction models in solving novel single and multi-step block stacking tasks.

Related Work

Representation learning for visual model-based reinforcement learning: Prior works have proposed learning video prediction models to improve exploration and planning in RL. However, such works and others that represent the scene with a single representation vector may be susceptible to the binding problem and must rely on data to learn that the same object in two different contexts can be modeled similarly. But processing a disentangled latent state with a single function or processing each disentangled factor in a permutation-sensitive manner (1) assumes a fixed number of entities that cannot be dynamically adjusted for generalizing to more objects than in training and (2) has no constraints to enforce that multiple instances of the same entity in the scene be modeled in the same way. For generalization, often the particular arrangement of objects in a scene does not matter so much as what is constant across scenes – properties of individual objects and inter-object relationships – which the inductive biases of these prior works do not capture. The entity abstraction in OP3 enforces symmetric processing of entity representations, thereby overcoming the limitations of these prior works.

Unsupervised grounding of abstract entity variables in concrete objects: Prior works that model entities and their interactions often pre-specify the identity of the entities , provide additional supervision , or provide additional specification such as segmentations , crops , or a simulator . Those that do not assume such additional information often factorize the entire scene into pixel-level entities , which do not model objects as coherent wholes. None of these works solve the problem of grounding the entities in raw observation, which is crucial for autonomous learning and interaction. OP3 builds upon recently proposed ideas in grounding entity representations via inference on a symmetrically factorized generative model of static and dynamic scenes, whose advantage over other methods for grounding is the ability to refine the grounding with new information. In contrast to other methods for binding in neural networks , formulating inference as a mechanism for variable binding allows us to model uncertainty in the values of the variables.

Comparison with similar work: The closest three works to OP3 are the Transporter , COBRA , and C-SWMs . The Transporter enforces a sparsity bias to learn object keypoints, each represented as a feature vector at a pixel location, and the method’s focus on keypoints has the advantage of enabling long-term object tracking, modeling articulated composite bodies such as joints, and scaling to dozens of objects. C-SWMs learn entity representations using a contrastive loss, which has the advantage of overcoming the difficulty in attending to small but relevant features as well as the large model capacity requirements usually characteristic of the pixel reconstructive loss. COBRA uses the autoregressive attention-based MONet architecture to obtain entity representations, which has the advantage of being more computationally efficient and stable to train. Unlike works such as that infer entity representations from static scenes, these works represent complementary approaches to OP3 (Figure 2) for representing dynamic scenes.

Symmetric processing of entities – processing each entity representation with the same function, as OP3 does with its observation, dynamics, and refinement networks – enforces the invariance that local properties are invariant to changes in global structure because it prevents the processing of one entity from being affected by other entities. How symmetric the process is for obtaining these entity representations from visual observation affects how straightforward it is to directly transfer models of a single entity across different global contexts, such as in modeling multiple instances of the same entity in the scene in a similar way or in generalizing to modeling different numbers of objects than in training. OP3 can exhibit this type of zero-shot transfer because the learnable components of its refinement process are fully symmetric across entities, which prevents OP3 from overfitting to the global structure of the scene. In contrast, the KeyNet encoder of the Transporter and the CNN-encoder of C-SWMs associate the content of the entity representation with the index of that entity in a global representation vector (Figure 1d), and this permutation-sensitive mapping entangles the encoding of an entity with the global structure of the scene. COBRA lies in between: it uses a learnable autoregressive attention network to infer segmentation masks, which entangles local object segmentations with global structure but may provide a useful bias for attending to salient objects, and symmetrically encodes entity representations given these masksA discussion of the advantages and disadvantages of using an attention-based entity disambiguation method, which MONet and COBRA use, versus an iterative refinement method, which IODINE and OP3 use, is discussed in Greff et al. ..

As a recurrent probabilistic dynamic latent variable model, OP3 can refine the grounding of its entity representations with new information from raw observations by simply applying a belief update similar to that used in filtering for hidden Markov models. The Transporter, COBRA, and C-SWMs all do not have mechanisms for updating the belief of the entity representations with information from subsequent image frames. Without recurrent structure, such methods rely on the assumption that a single forward pass of the encoder on a static image is sufficient to disambiguate objects, but this assumption is not generally true: objects can pop in and out of occlusion and what constitutes an object depends temporal cues, especially in real world settings. Recurrent structure is built into the OP3 inference update (Appendix 2), enabling OP3 to model object permanence under occlusion and refine its object representations with new information in modeling real world videos (Figure 7).

Problem Formulation

Let x∗x^{*} denote a physical scene and h1:K∗h^{*}_{1:K} denote the objects in the scene. Let XX and AA be random variables for the image observation of the scene x∗x^{*} and the agent’s actions respectively. In contrast to prior works that use a single latent variable to represent the state of the scene, we use a set of latent random variables H1:KH_{1:K} to represent the state of the objects h1:K∗h^{*}_{1:K}. We use the term object to refer to hk∗h^{*}_{k}, which is part of the physical world, and the term entity to refer to HkH_{k}, which is part of our model of the physical world. The generative distribution of observations X(0:T)X^{(0:T)} and latent entities H1:K(0:T)H^{(0:T)}_{1:K} from taking TT actions a(0:T−1)a^{(0:T-1)} is modeled as:

where p(X(t) ∣ H1:K(t))p(X^{(t)}\,|\,H^{(t)}_{1:K}) and p(H1:K(t) ∣ H1:K(t−1),A(t−1))p(H^{(t)}_{1:K}\,|\,H^{(t-1)}_{1:K},A^{(t-1)}) are the observation and dynamics distribution respectively shared across all timesteps tt. Our goal is to build a model that, from simply observing raw observations of random interactions, can generalize to solve novel compositional object manipulation problems that the learner was never trained to do, such as building various block towers during test time from only training to predict how blocks fall during training time.

When all tasks follow the same dynamics we can achieve such generalization with a planning algorithm if given a sequence of actions we could compute p(X(T+1:T+d) ∣ X(0:T),A(0:T+d−1))p(X^{(T+1:T+d)}\,|\,X^{(0:T)},A^{(0:T+d-1)}), the posterior predictive distribution of observations dd steps into the future. Approximating this predictive distribution can be cast as a variational inference problem (Appdx. B) for learning the parameters of an approximate observation distribution \mathdutchcalG(X(t) ∣ H1:K(t))\mathdutchcal{G}(X^{(t)}\,|\,H^{(t)}_{1:K}), dynamics distribution \mathdutchcalD(H1:K(t) ∣ H1:K(t−1),A(t−1))\mathdutchcal{D}(H^{(t)}_{1:K}\,|\,H^{(t-1)}_{1:K},A^{(t-1)}), and a time-factorized recognition distribution \mathdutchcalQ(H1:K(t) ∣ H1:K(t−1),X(t),A(t−1))\mathdutchcal{Q}(H^{(t)}_{1:K}\,|\,H^{(t-1)}_{1:K},X^{(t)},A^{(t-1)}) that maximize the evidence lower bound (ELBO), given by L=∑t=0TLr(t)−Lc(t)\mathcal{L}=\sum_{t=0}^{T}\mathcal{L}^{(t)}_{\text{r}}-\mathcal{L}^{(t)}_{\text{c}}, where

The ELBO pushes \mathdutchcalQ\mathdutchcal{Q} to produce states of the entities H1:KH_{1:K} that contain information useful for not only reconstructing the observations via \mathdutchcalG\mathdutchcal{G} in Lr(t)\mathcal{L}^{(t)}_{\text{r}} but also for predicting the entities’ future states via \mathdutchcalD\mathdutchcal{D} in Lc(t)\mathcal{L}^{(t)}_{\text{c}}. Sec. 4 will next offer our method for incorporating entity abstraction into modeling the generative distribution and optimizing the ELBO.

Object-Centric Perception, Prediction, and Planning (OP3)

The entity abstraction is derived from an assumption about symmetry: that the problem of modeling a dynamic scene of multiple entities can be reduced to the problem of (1) modeling a single entity and its interactions with an entity-centric function and (2) applying this function to every entity in the scene. Our choice to represent a scene as a set of entities exposes an avenue for directly encoding such a prior about symmetry that would otherwise not be straightforward with a global state representation.

As shown in Fig. 1, a function FF that respects the entity abstraction requires two ingredients. The first ingredient (Sec. 4.1) is that F(H1:K)F(H_{1:K}) is expressed in part as the higher-order operation map(f,H1:K)\texttt{map}(f,H_{1:K}) that broadcasts the same entity-centric function f(Hk)f(H_{k}) to every entity variable HkH_{k}. This yields the benefit of automatically transferring learned knowledge for modeling an individual entity to all entities in the scene rather than learn such symmetry from data. As ff is a function that takes in a single generic entity variable HkH_{k} as argument, the second ingredient (Sec. 4.2) should be a mechanism that binds information from the raw observation XX about a particular object hk∗h^{*}_{k} to the variable HkH_{k}.

The functions of interest in model-based RL are the observation and dynamics models \mathdutchcalG\mathdutchcal{G} and \mathdutchcalD\mathdutchcal{D} with which we seek to approximate the data-generating distribution in equation 1.

Observation Model: The observation model \mathdutchcalG(X ∣ H1:K)\mathdutchcal{G}(X\,|\,H_{1:K}) approximates the distribution p(X ∣ H1:K)p(X\,|\,H_{1:K}), which models how the observation XX is caused by the combination of entities H1:KH_{1:K}. We enforce the entity abstraction in \mathdutchcalG\mathdutchcal{G} (in Fig. 1g) by applying the same entity-centric function \mathdutchcalg(X ∣ Hk)\mathdutchcal{g}(X\,|\,H_{k}) to each entity HkH_{k}, which we can implement using a mixture model at each pixel (i,j)(i,j):

where \mathdutchcalg\mathdutchcal{g} computes the mixture components that model how each individual entity HkH_{k} is independently generated, combined via mixture weights mm that model the entities’ relative depth from the camera, the derivation of which is in Appdx. A.

Dynamics Model: The dynamics model \mathdutchcalD(H1:K′ ∣ H1:K,A)\mathdutchcal{D}(H^{\prime}_{1:K}\,|\,H_{1:K},A) approximates the distribution p(H1:K′ ∣ H1:K,A)p(H^{\prime}_{1:K}\,|\,H_{1:K},A), which models how an action AA intervenes on the entities H1:KH_{1:K} to produce their future values H1:K′H^{\prime}_{1:K}. We enforce the entity abstraction in \mathdutchcalD\mathdutchcal{D} (in Fig. 1f) by applying the same entity-centric function \mathdutchcald(Hk′ ∣ Hk,H[≠k],A)\mathdutchcal{d}(H^{\prime}_{k}\,|\,H_{k},H_{[\neq k]},A) to each entity HkH_{k}, which reduces the problem of modeling how an action affects a scene with a combinatorially large space of object configurations to the problem of simply modeling how an action affects a single generic entity HkH_{k} and its interactions with the list of other entities H[≠k]H_{[\neq k]}. Modeling the action as an finer-grained intervention on a single entity rather than the entire scene is a benefit of using local representations of entities rather than global representations of scenes.

However, at this point we still have to model the combinatorially large space of interactions that a single entity could participate in. Therefore, we can further enforce a pairwise entity abstraction on \mathdutchcald\mathdutchcal{d} by applying the same pairwise function \mathdutchcaldoo(Hk,Hi)\mathdutchcal{d}_{oo}(H_{k},H_{i}) to each entity pair (Hk,Hi)(H_{k},H_{i}), for i∈[≠k]i\in[\neq k]. Omitting the action to reduce clutter (the full form is written in Appdx. F.2), the structure of the \mathdutchcalD\mathdutchcal{D} therefore follows this form:

The entity abstraction therefore provides the flexibility to scale to modeling a variable number of objects by solely learning a function \mathdutchcald\mathdutchcal{d} that operates on a single generic entity and a function \mathdutchcaldoo\mathdutchcal{d}_{oo} that operates on a single generic entity pair, both of which can be re-used for across all entity instances.

2 Interactive Inference for Binding Object Properties to Latent Variables

For the observation and dynamics models to operate from raw pixels hinges on the ability to bind the properties of specific physical objects h1:K∗h^{*}_{1:K} to the entity variables H1:KH_{1:K}. For latent variable models, we frame this variable binding problem as an inference problem: binding information about h1:K∗h^{*}_{1:K} to H1:KH_{1:K} can be cast as a problem of inferring the parameters of p(H(0:T) ∣ x(0:T),a(0:T−1))p(H^{(0:T)}\,|\,x^{(0:T)},a^{(0:T-1)}), the posterior distribution of H1:KH_{1:K} given a sequence of interactions. Maximizing the ELBO in Sec. 3 offers a method for learning the parameters of the observation and dynamics models while simultaneously learning an approximation to the posterior q(H(0:T) ∣ x(0:T),a(0:T−1))=∏t=0T\mathdutchcalQ(H1:K(t) ∣ H1:K(t−1),x(t),a(t))q(H^{(0:T)}\,|\,x^{(0:T)},a^{(0:T-1)})=\prod_{t=0}^{T}\mathdutchcal{Q}(H^{(t)}_{1:K}\,|\,H^{(t-1)}_{1:K},x^{(t)},a^{(t)}), which we have chosen to factorize into a per-timestep recognition distribution \mathdutchcalQ\mathdutchcal{Q} shared across timesteps. We also choose to enforce the entity abstraction on the process that computes the recognition distribution \mathdutchcalQ\mathdutchcal{Q} (in Fig. 1e) by decomposing it into a recognition distribution \mathdutchcalq\mathdutchcal{q} applied to each entity:

Whereas a neural network encoder is often used to approximate the posterior , a single forward pass that computes \mathdutchcalq\mathdutchcal{q} in parallel for each entity is insufficient to break the symmetry for dividing responsibility of modeling different objects among the entity variables because the entities do not have the opportunity to communicate about which part of the scene they are representing.

We therefore adopt an iterative inference approach to compute the recognition distribution \mathdutchcalQ\mathdutchcal{Q}, which has been shown to break symmetry among modeling objects in static scenes . Iterative inference computes the recognition distribution via a procedure, rather than a single forward pass of an encoder, that iteratively refines an initial guess for the posterior parameters λ1:K\lambda_{1:K} by using gradients from how well the generative model is able to predict the observation based on the current posterior estimate. The initial guess provides the noise to break the symmetry.

For scenes where position and color are enough for disambiguating objects, a static image may be sufficient for inferring \mathdutchcalq\mathdutchcal{q}. However, in interactive environments disambiguating objects is more underconstrained because what constitutes an object depends on the goals of the agent. We therefore incorporate actions into the amortized varitional filtering framework to develop an interactive inference algorithm (Appdx. D and Fig. 5) that uses temporal continuity and interactive feedback to disambiguate objects. Another benefit of enforcing entity abstraction is that preserving temporal consistency on entities comes for free: information about each object remains bound to its respective HkH_{k} through time, mixing with information about other entities only through explicitly defined avenues, such as in the dynamics model.

3 Training at Different Timescales

The variational parameters λ1:K\lambda_{1:K} are the interface through which the neural networks f\mathdutchcalgf_{\mathdutchcal{g}}, f\mathdutchcaldf_{\mathdutchcal{d}}, f\mathdutchcalqf_{\mathdutchcal{q}} that respectively output the distribution parameters of \mathdutchcalG\mathdutchcal{G}, \mathdutchcalD\mathdutchcal{D}, and \mathdutchcalQ\mathdutchcal{Q} communicate. For a particular dynamic scene, the execution of interactive inference optimizes the variational parameters λ1:K\lambda_{1:K}. Across scene instances, we train the weights of f\mathdutchcalgf_{\mathdutchcal{g}}, f\mathdutchcaldf_{\mathdutchcal{d}}, f\mathdutchcalqf_{\mathdutchcal{q}} by backpropagating the ELBO through the entire inference procedure, spanning multiple timesteps. OP3 thus learns at three different timescales: the variational parameters learn (1) across MM steps of inference within a single timestep and (2) across TT timesteps within a scene instance, and the network weights learn (3) across different scene instances.

Beyond next-step prediction, we can directly train to compute the posterior predictive distribution p(X(T+1:T+d) ∣ x(0:T),a(0:T+d))p(X^{(T+1:T+d)}\,|\,x^{(0:T)},a^{(0:T+d)}) by sampling from the approximate posterior of H1:K(T)H_{1:K}^{(T)} with \mathdutchcalQ\mathdutchcal{Q}, rolling out the dynamics model \mathdutchcalD\mathdutchcal{D} in latent space from these samples with a sequence of dd actions, and predicting the observation X(T+d)X^{(T+d)} with the observation model \mathdutchcalG\mathdutchcal{G}. This approach to action-conditioned video prediction predicts future observations directly from observations and actions, but with a bottleneck of KK time-persistent entity-variables with which the dynamics model \mathdutchcalD\mathdutchcal{D} performs symbolic relational computation.

4 Object-Centric Planning

OP3 rollouts, computed as the posterior predictive distribution, can be integrated into the standard visual model-predictive control framework. Since interactive inference grounds the entities H1:KH_{1:K} in the actual objects h1:K∗h_{1:K}^{*} depicted in the raw observation, this grounding essentially gives OP3 access to a pointer to each object, enabling the rollouts to be in the space of entities and their relations. These pointers enable OP3 to not merely predict in the space of entities, but give OP3 access to an object-centric action space: for example, instead of being restricted to the standard (pick_xy, place_xy) action space common to many manipulation tasks, which often requires biased picking with a scripted policy , these pointers enable us to compute a mapping (Appdx. G.2) between entity_id and pick_xy, allowing OP3 to automatically use a (entity_id, place_xy) action space without needing a scripted policy.

5 Generalization to Various Tasks

We consider tasks defined in the same environment with the same physical laws that govern appearance and dynamics. Tasks are differentiated by goals, in particular goal configurations of objects. Building good cost functions for real world tasks is generally difficult because the underlying state of the environment is always unobserved and can only be modeled through modeling observations. However, by representing the environment state as the state of its entities, we may obtain finer-grained goal-specification without the need for manual annotations . Having rolled out OP3 to a particular timestep, we construct a cost function to compare the predicted entity states H1:K(P)H_{1:K}^{(P)} with the entity states H1:K(G)H_{1:K}^{(G)} inferred from a goal image by considering pairwise distances between the entities, another example of enforcing the pairwise entity abstraction. Letting S′S^{\prime} and SS denote the set of goal and predicted entities respectively, we define the form of the cost function via a composition of the task specific distance function \mathdutchcalc\mathdutchcal{c} operating on entity-pairs:

in which we pair each goal entity with the closest predicted entity and sum over the costs of these pairs. Assuming a single action suffices to move an object to its desired goal position, we can greedily plan each timestep by defining the cost to be min⁡a∈S′,b∈S  \mathdutchcalc(Ha(G),Hb(P))\min_{a\in S^{\prime},b\in S}\;\mathdutchcal{c}(H_{a}^{(G)},H_{b}^{(P)}), the pair with minimum distance, and removing the corresponding goal entity from further consideration for future planning.

Experiments

Our experiments aim to study to what degree entity abstraction improves generalization, planning, and modeling. Sec. 5.1 shows that from only training to predict how objects fall, OP3 generalizes to solve various novel block stacking tasks with two to three times better accuracy than a state-of-the-art video prediction model. Sec. 5.2 shows that OP3 can plan for multiple steps in a difficult multi-object environment. Sec. 5.3 shows that OP3 learns to ground its abstract entities in objects from real world videos.

We first investigate how well OP3 can learn object-based representations without additional object supervision, as well as how well OP3’s factorized representation can enable combinatorial generalization for scenes with many objects.

Domain: In the MuJoCo block stacking task introduced by Janner et al. for the O2P2 model, a block is raised in the air and the model must predict the steady-state effects of dropping the block on a surface with multiple objects, which implicitly requires modeling the effects of gravity and collisions. The agent is never trained to stack blocks, but is tested on a suite of tasks where it must construct block tower specified by a goal image. Janner et al. showed that an object-centric model with access to ground truth object segmentations can solve these tasks with about 76% accuracy. We now consider whether OP3 can do better, but without any supervision on object identity.

Setup: We train OP3 on the same dataset and evaluate on the same goal images as Janner et al. . While the training set contains up to five objects, the test set contains up to nine objects, which are placed in specific structures (bridge, pyramid, etc.) not seen during training. The actions are optimized using the cross-entropy method (CEM) , with each sampled action evaluated by the greedy cost function described in Sec. 4.5. Accuracy is evaluated using the metric defined by Janner et al. , which checks that all blocks are within some threshold error of the goal.

Results: The two baselines, SAVP and O2P2, represent the state-of-the-art in video prediction and symmetric object-centric planning methods, respectively. SAVP models objects with a fixed number of convolutional filters and does not process entities symmetrically. O2P2 does process entities symmetrically, but requires access to ground truth object segmentations. As shown in Table 1, OP3 achieves better accuracy than O2P2, even without any ground truth supervision on object identity, possibly because grounding the entities in the raw image may provide a richer contextual representation than encoding each entity separately without such global context as O2P2 does. OP3 achieves three times the accuracy of SAVP, which suggests that symmetric modeling of entities is enables the flexibility to transfer knowledge of dynamics of a single object to novel scenes with different configurations heights, color combinations, and numbers of objects than those from the training distribution. Fig. 8 and Fig. 9 in the Appendix show that, by grounding its entities in objects of the scene through inference, OP3’s predictions isolates only one object at a time without affecting the predictions of other objects.

2 Multi-Step Planning

The goal of our second experiment is to understand how well OP3 can perform multi-step planning by manipulating objects already present in the scene. We modify the block stacking task by changing the action space to represent a picking and dropping location. This requires reasoning over extended action sequences since moving objects out of place may be necessary.

Goals are specified with a goal image, and the initial scene contains all of the blocks needed to build the desired structure. This task is more difficult because the agent may have to move blocks out of the way before placing other ones which would require multi-step planning. Furthermore, an action only successfully picks up a block if it intersects with the block’s outline, which makes searching through the combinatorial space of plans a challenge. As stated in Sec. 4.4, having a pointer to each object enables OP3 to plan in the space of entities. We compare two different action spaces (pick_xy, place_xy) and (entity_id, place_xy) to understand how automatically filtering for pick locations at actual locations of objects enables better efficiency and performance in planning. Details for determining the pick_xy from entity_id are in appendix G.2.

Results: We compare with SAVP, which uses the (pick_xy, place_xy) action space. With this standard action space (Table 2) OP3 achieves between 1.5-2 times the accuracy of SAVP. This performance gap increases to 2-3 times the accuracy when OP3 uses the (entity_id, place_xy) action space. The low performance of SAVP with only two blocks highlights the difficulty of such combinatorial tasks for model-based RL methods, and highlights the both the generalization and localization benefits of a model with entity abstraction. Fig. 6b shows that OP3 is able to plan more efficiently, suggesting that OP3 may be a more effective model than SAVP in modeling combinatorial scenes. Fig. 7a shows the execution of interactive inference during training, where OP3 alternates between four refinement steps and one prediction step. Notice that OP3 infers entity representations that decompose the scene into coherent objects and that entities that do not model objects model the background. We also observe in the last column (t=2t=2) that OP3 predicts the appearance of the green block even though the green block was partially occluded in the previous timestep, which shows its ability to retain information across time.

3 Real World Evaluation

The previous tasks used simulated environments with monochromatic objects. Now we study how well OP3 scales to real world data with cluttered scenes, object ambiguity, and occlusions. We evaluate OP3 on the dataset from Ebert et al. which contains videos of a robotic arm moving cloths and other deformable and multipart objects with varying textures.

We evaluate qualitative performance by visualizing the object segmentations and compare against vanilla IODINE, which does not incorporate an interaction-based dynamics model into the inference process. Fig. 7b highlights the strength of OP3 in preserving temporal continuity and disambiguating objects in real world scenes. While IODINE can disambiguate monochromatic objects in static images, we observe that it struggles to do more than just color segmentation on more complicated images where movement is required to disambiguate objects. In contrast, OP3 is able to use temporal information to obtain more accurate segmentations, as seen in Fig. 7b where it initially performs color segmentation by grouping the towel, arm, and dark container edges together, and then by observing the effects of moving the arm, separates these entities into different groups.

Discussion

We have shown that enforcing the entity abstraction in a model-based reinforcement learner improves generalization, planning, and modeling across various compositional multi-object tasks. In particular, enforcing the entity abstraction provides the learner with a pointer to each entity variable, enabling us to define functions that are local in scope with respect to a particular entity, allowing knowledge about an entity in one context to directly transfer to modeling the same entity in different contexts. In the physical world, entities are often manifested as objects, and generalization in physical tasks such as robotic manipulation often may require symbolic reasoning about objects and their interactions. However, the general difficulty with using purely symbolic, abstract representations is that it is unclear how to continuously update these representations with more raw data. OP3 frames such symbolic entities as random variables in a dynamic latent variable model and infers and refines the posterior of these entities over time with neural networks. This suggests a potential bridge to connect abstract symbolic variables with the noisy, continuous, high-dimensional physical world, opening a path to scaling robotic learning to more combinatorially complex tasks.

The authors would like to thank the anonymous reviewers for their helpful feedback and comments. The authors would also like to thank Sjoerd van Steenkiste, Nalini Singh and Marvin Zhang for helpful discussions on the graphical model, Klaus Greff for help in implementing IODINE, Alex Lee for help in running SAVP, Tom Griffiths, Karl Persch, and Oleg Rybkin for feedback on earlier drafts, Joe Marino for discussions on iterative inference, and Sam Toyer, Anirudh Goyal, Jessica Hamrick, Peter Battaglia, Loic Matthey, Yash Sharma, and Gary Marcus for insightful discussions. This research was supported in part by the National Science Foundation under IIS-1651843, IIS-1700697, and IIS-1700696, the Office of Naval Research, ARL DCIST CRA W911NF-17-2-0181, DARPA, Berkeley DeepDrive, Google, Amazon, and NVIDIA.

References

Appendix A Observation Model

Deterministic rendering engine: Each object HkH_{k} is rendered independently as the sub-image IkI_{k} and the resulting KK sub-images are combined to form the final image observation XX. To combine the sub-images, each pixel Ik(ij)I_{k(ij)} in each sub-image is assigned a depth δk(ij)\delta_{k(ij)} that specifies the distance of object kk from the camera at coordinate (ij)(ij). of the image plane. Thus the pixel X(ij)X_{(ij)} takes on the value of its corresponding pixel Ik(ij)I_{k(ij)} in the sub-image IkI_{k} if object kk is closest to the camera than the other objects, such that

Modeling uncertainty with the observation model: In reality we do not directly observe the depth values, so we must construct a probabilistic model to model our uncertainty:

where every pixel (ij)(ij) is modeled through a set of mixture components \mathdutchcalg(X(ij) ∣ Hk):=p(Xij∣Zk(ij)=1,Hk)\mathdutchcal{g}\left(X_{(ij)}\,|\,H_{k}\right):=p\left(X_{ij}|Z_{k(ij)}=1,H_{k}\right) that model how pixels of the individual sub-images IkI_{k} are generated, as well as through the mixture weights mij(Hk):=p(Zk(ij)=1∣Hk)m_{ij}(H_{k}):=p\left(Z_{k(ij)}=1|H_{k}\right) that model which point of each object is closest to the camera.

Appendix B Evidence Lower Bound

Here we provide a derivation of the evidence lower bound. We begin with the log probability of the observations X(1:T)X^{(1:T)} conditioned on a sequence of actions a(0:T−1)a^{(0:T-1)}:

We have freedom to choose the approximating distribution q(H1:K(0:T) | ⋅)q\left(H^{(0:T)}_{1:K}\,\middle|\,\cdot\right) so we choose it to be conditioned on the past states and actions, factorized across time:

With this factorization, we can use linearity of expectation to decouple Equation 8 across timesteps:

By the Markov property, the marginal q(H1:K(t) ∣ h1:K(0:t−1),X(0:t),a(0:t−1))q(H^{(t)}_{1:K}\,|\,h^{(0:t-1)}_{1:K},X^{(0:t)},a^{(0:t-1)}) is computed recursively as

whose base case is q(H(0) ∣ X(0))q\left(H^{(0)}\,|\,X^{(0)}\right) when t=0t=0.

We approximate observation distribution p(X ∣ H1:K)p(X\,|\,H_{1:K}) and the dynamics distribution p(H1:K′ ∣ H1:K,a)p(H^{\prime}_{1:K}\,|\,H_{1:K},a) by learning the parameters of the observation model \mathdutchcalG\mathdutchcal{G} and dynamics model \mathdutchcalD\mathdutchcal{D} respectively as outputs of neural networks. We approximate the recognition distribution q(H1:K(t) ∣ h1:K(t−1),X(t),a(t−1))q(H^{(t)}_{1:K}\,|\,h^{(t-1)}_{1:K},X^{(t)},a^{(t-1)}) via an inference procedure that refines better estimates of the posterior parameters, computed as an output of a neural network. To compute the expectation in the marginal q(H1:K(t) ∣ h1:K(0:t−1),X(0:t),a(0:t−1))q(H^{(t)}_{1:K}\,|\,h^{(0:t-1)}_{1:K},X^{(0:t)},a^{(0:t-1)}), we follow standard practice in amortized variational inference by approximating the expectation with a single sample of the sequence h1:K(0:t−1)h^{(0:t-1)}_{1:K} by sequentially sampling the latents for one timestep given latents from the previous timestep, and optimizing the ELBO via stochastic gradient ascent .

Appendix C Posterior Predictive Distribution

Here we provide a derivation of the posterior predictive distribution for the dynamic latent variable model with multiple latent states. Section B described how we compute the distributions p(X ∣ H1:K)p(X\,|\,H_{1:K}), p(H1:K′ ∣ H1:K,a)p(H^{\prime}_{1:K}\,|\,H_{1:K},a), q(H1:K(t) ∣ h1:K(t−1),X(t),a(t−1))q(H^{(t)}_{1:K}\,|\,h^{(t-1)}_{1:K},X^{(t)},a^{(t-1)}), and q(H1:K(0:T) ∣ x(1:T),a(1:T))q(H_{1:K}^{(0:T)}\,|\,x^{(1:T)},a^{(1:T)}). Here we show that these distributions can be used to approximate the predictive posterior distribution p(X(T+1:T+d) ∣ x(0:T),a(0:T+d))p(X^{(T+1:T+d)}\,|\,x^{(0:T)},a^{(0:T+d)}) by maximizing the following lower bound:

The numerator p(X(T+1:T+d),h1:K(0:T+d) ∣ x(0:T),a(0:T+d))p(X^{(T+1:T+d)},h_{1:K}^{(0:T+d)}\,|\,x^{(0:T)},a^{(0:T+d)}) can be decomposed into two terms, one of which involving the posterior p(h1:K(0:T+d) ∣ x(0:T),a(0:T+d))p(h_{1:K}^{(0:T+d)}\,|\,x^{(0:T)},a^{(0:T+d)}):

This allows Equation 9 to be broken up into two terms:

Maximizing the second term, the negative KL-divergence between the variational distribution q(H1:K(0:T+d) ∣ ⋅)q(H_{1:K}^{(0:T+d)}\,|\,\cdot) and the posterior p(H1:K(0:T+d) ∣ x(0:T),a(0:T+d))p(H_{1:K}^{(0:T+d)}\,|\,x^{(0:T)},a^{(0:T+d)}) is the same as maximizing the following lower bound:

where the first term is due to the conditional independence between X(0:T)X^{(0:T)} and the future states H1:K(T+1:T+d)H_{1:K}^{(T+1:T+d)} and actions A(T+1:T+d)A^{(T+1:T+d)}. We choose to express q(H1:K(0:T+d) | ⋅)q\left(H_{1:K}^{(0:T+d)}\,\middle|\,\cdot\right) as conditioned on past states and actions, factorized across time:

In summary, Equation 9 can be expressed as

which can be interpreted as the standard ELBO objective for timesteps 0:T0:T, plus an addition reconstruction term for timesteps T+1:T+dT+1:T+d, a reconstruction term for timesteps 0:T0:T. We can maximize this using the same techniques as maximizing Equation 8.

Whereas approximating the ELBO in Equation 9 can be implemented by rolling out OP3 to predict the next observation via teacher forcing , approximating the posterior predictive distribution in Equation 9 can be implemented by rolling out the dynamics model dd steps beyond the last observation and using the observation model to predict the future observations.

Appendix D Interactive Inference

Algorithms 1 and 2 detail MM steps of the interactive inference algorithm at timestep and t∈[1,T]t\in[1,T] respectively. Algorithm 1 is equivalent to the IODINE algorithm described in . Recalling that λ1:K\lambda_{1:K} are the parameters for the distribution of the random variables H1:KH_{1:K}, we consider in this paper the case where this distribution is an isotropic Gaussian (e.g. N(λk)\mathcal{N}(\lambda_{k}) where λk=(μk,σk)\lambda_{k}=(\mu_{k},\sigma_{k})), although OP3 need not be restricted to the Gaussian distribution. The refinement network f\mathdutchcalqf_{\mathdutchcal{q}} produces the parameters for the distribution \mathdutchcalq(Hk(t) ∣ hk(t−1),x(t),a(t))\mathdutchcal{q}(H^{(t)}_{k}\,|\,h^{(t-1)}_{k},x^{(t)},a^{(t)}). The dynamics network f\mathdutchcaldf_{\mathdutchcal{d}} produces the parameters for the distribution \mathdutchcald(Hk(t) ∣ hk(t−1),h[≠k](t−1),a(t))\mathdutchcal{d}(H^{(t)}_{k}\,|\,h^{(t-1)}_{k},h^{(t-1)}_{[\neq k]},a^{(t)}). To implement \mathdutchcalq\mathdutchcal{q}, we repurpose the dynamics model to transform hk(t−1)h_{k}^{(t-1)} into the initial posterior estimate λk(0)\lambda_{k}^{(0)} and then use f\mathdutchcalqf_{\mathdutchcal{q}} to iteratively update this parameter estimate. βk\beta_{k} indicates the auxiliary inputs into the refinement network used in . We mark the major areas where the algorithm at timestep tt differs from the algorithm at timestep in blue.

We can train the entire OP3 system end-to-end by backpropagating through the entire inference procedure, using the ELBO at every timestep as a training signal for the parameters of G\mathcal{G}, D\mathcal{D}, Q\mathcal{Q} in a similar manner as . However, the interactive inference algorithm can also be naturally be adapted to predict rollouts by using the dynamics model to propagate the λ1:K\lambda_{1:K} for multiple steps, rather than just the one step for predicting λ1:K(t,0)\lambda_{1:K}^{(t,0)} in line 2 of Algorithm 2. To train OP3 to rollout the dynamics model for longer timescales, we use a curriculum that increases the prediction horizon throughout training.

Appendix E Cost Function

Let I^(Hk):=m(Hk)⋅\mathdutchcalg(X ∣ Hk)\hat{I}(H_{k}):=m(H_{k})\cdot\mathdutchcal{g}\left(X\,|\,H_{k}\right) be a masked sub-image (see Appdx: A). We decompose the cost of a particular configuration of objects into a distance function between entity states, \mathdutchcalc(Ha,Hb)\mathdutchcal{c}(H_{a},H_{b}). For the first environment with single-step planning we use L2L_{2} distance of the corresponding masked subimages: \mathdutchcalc(Ha,Hb)=L2(I^(Ha),I^(Hb))\mathdutchcal{c}(H_{a},H_{b})=L_{2}(\hat{I}(H_{a}),\hat{I}(H_{b})). For the second environment with multi-step planning we a different distance function since the previous one may care more about if a shape matches than if the color matches. We instead use a form of intersection over union but that counts intersection if the mask aligns and pixel color values are close \mathdutchcalc(Ha,Hb)=1−∑i,jmij(Ha)>0.01 and mij(Hb)>0.01 and L2(\mathdutchcalg(Ha)(ij),\mathdutchcalg(Hb)(ij))<0.1∑i,jmij(Ha)>0.01 or mij(Hb)>0.01\mathdutchcal{c}(H_{a},H_{b})=1-\frac{{\sum_{i,j}m_{ij}(H_{a})>0.01\text{ and }m_{ij}(H_{b})>0.01}\text{ and }L_{2}(\mathdutchcal{g}(H_{a})_{(ij)},\mathdutchcal{g}(H_{b})_{(ij)})<0.1}{\sum_{i,j}m_{ij}(H_{a})>0.01\text{ or }m_{ij}(H_{b})>0.01}. We found this version to work better since it will not give low cost to moving a wrong color block to the position of a different color goal block.

Appendix F Architecture and Hyperparameter Details

We use similar model architectures as in and so have rewritten some details from their appendix here. Differences include the dynamics model, inclusion of actions, and training procedure over sequences of data. Like , we define our latent distribution of size RR to be divided into a deterministic component of size RdR_{d} and stochastic component of size RsR_{s}. We found that splitting the latent state into a deterministic and stochastic component (as opposed to having a fully stocahstic representation) was helpful for convergence. We parameterize the distribution of each HkH_{k} as a diagonal Gaussian, so the output of the refinement and dynamics networks are the parameteres of a diagonal Gaussian. We parameterize the output of the observation model also as a diagonal Gaussian with means μ\mu and global scale σ=0.1\sigma=0.1. The observation network outputs the μ\mu and mask mkm_{k}.

Training: All models are trained with the ADAM optimizer with default parameters and a learning rate of 0.0003. We use gradient clipping as in where if the norm of global gradient exceeds 5.0 then the gradient is scaled down to that norm.

Inputs: For all models, we use the following inputs to the refinement network, where LN means Layernorm and SG means stop gradients. The following image-sized inputs are concatenated and fed to the corresponding convolutional network:

The posterior parameters λ1:K\lambda_{1:K} and their gradients are flat vectors, and we concatenate them with the output of the convolutional part of the refinement network and use the result as input to the refinement LSTM:

All models use the ELU activation function and the convolutional layers use a stride equal to 1 and padding equal to 2 unless otherwise noted. For the table below Rs=64R_{s}=64 and R=128R=128.

F.2 Dynamics Model

The dynamics model D\mathcal{D} models how each entity HkH_{k} is affected by action AA and the other entity H[≠k]H_{[\neq k]}. It applies the same function \mathdutchcald(Hk′ ∣ Hk,H[≠k],A)\mathdutchcal{d}(H_{k}^{\prime}\,|\,H_{k},H_{[\neq k]},A) to each state, composed of several functions illustrated and described in Fig. 4:

This architectural choice for the dynamics model is an action-conditioned modification of the interaction function used in Relational Neural Expectation Maximization (RNEM) , which is a latent-space attention-based modification of the Neural Physics Engine (NPE) , which is one of a broader class of architectures known as graph networks .

Appendix G Experiment Details

The training dataset has 60,000 trajectories each containing before and after images of size 64x64 from . Before images are constructed with actions which consist of choosing a shape (cube, rectangle, pyramid), color, and an (x,y,z)(x,y,z) position and orientation for the block to be dropped. At each time step, a block is dropped and the simulation runs until the block settles into a stable position. The model takes in an image containing the block to be dropped and must predict the steady-state effect. Models were trained on scenes with 1 to 5 blocks with K=7K=7 entity variables. The cross entorpy method (CEM) begins from a uniform distribution on the first iteration, uses a population size of 1000 samples per iteration, and uses 10% of the best samples to fit a Gaussian distribution for each successive iteration.

G.2 Multi-Step Block-Stacking

The training dataset has 10,000 trajectories each from a separate environment with two different colored blocks. Each trajectory contains five frames (64x64) of randomly picking and placing blocks. We bias the dataset such that 30% of actions will pick up a block and place it somewhere randomly, 40% of actions will pick up a block and place it on top of a another random block, and 30% of actions contain random pick and place locations. Models were trained with K = 4 slots. We optimize actions using CEM but we optimize over multiple consecutive actions into the future executing the sequence with lowest cost. For a goal with nn blocks we plan nn steps into the future, executing nn actions. We repeat this procedure 2n2n times or until the structure is complete. Accuracy is computed as # blocks in correct position# goal blocks\frac{\text{\# blocks in correct position}}{\text{\# goal blocks}}, where a correct position is based on a threshold of the distance error.

For MPC we use two difference action spaces:

Coordinate Pick Place: The normal action space involves choosing a pick (x,y) and place (x,y) location.

Entity Pick Place: A concern with the normal action space is that successful pick locations are sparse ( 2%) given the current block size. Therefore, the probability of picking nn blocks consecutively becomes 0.02n0.02^{n} which becomes improbable very fast if we just sample pick locations uniformly. We address this by using the pointers to the entity variables to create an action space that involves directly choosing one of the latent entities to move and then a place (x,y)(x,y) location. This allows us to easily pick blocks consecutively if we can successfully map a latent entity_id of a block to a corresponding successful pick location. In order to determine the pick (x,y)(x,y) from an entity_id kk, we sample coordinates uniformly over the pick (x,y)(x,y) space and then average these coordinates weighted by their attention coefficient on that latent:

Appendix H Ablations

We perform ablations on the block stacking task from examining components of our model. Table 3 shows the effect of non-symmetrical models or cost functions. The “Unfactorized Model” and “No Weight Sharing” follow (c) and (d) from Figure 1 and are unable to sufficiently generalize. “Unfactorized Cost” refers to simply taking the mean-squared error of the compositie prediction image and the goal image, rather than decomposing the cost per entity masked subimage. We see that with the same OP3 model trained on the same data, not using an entity-centric factorization of the cost significantly underperforms a cost function that does decompose the cost per entity (c.f. Table 1).

Appendix I Interpretability

We do not explicitly explore interpretability in this work, but we see that an entity-factorized model readily lends itself to be interpretable by construction. The ability to decompose a scene into specific latents, view latents invididually, and explicitly see how these latents interact with each other could lead to significantly more interpretable models than current unfactorized models. Our use of attention values to determine the pick locations of blocks scratches the surface of this potential. Additionally, the ability to construct cost functions based off individual latents allows for more interpretable and customizable cost functions.