Compositional Foundation Models for Hierarchical Planning

Anurag Ajay, Seungwook Han, Yilun Du, Shuang Li, Abhi Gupta, Tommi Jaakkola, Josh Tenenbaum, Leslie Kaelbling, Akash Srivastava, Pulkit Agrawal

Introduction

Consider the task of making a cup of tea in an unfamiliar house. To successfully execute this task, an effective approach is to reason hierarchically at multiple levels: an abstract level, e.g. the high level steps needed to heat up the tea, a concrete geometric level e.g., how we should physically navigate to and in the kitchen, and at a control level e.g. how we should actuate our joints to lift a cup. It is further important that reasoning at each level is self-consistent with each other – an abstract plan to look in cabinets for tea kettles must also be physically plausible at the geometric level and executable given the actuations we are capable of. In this paper, we explore how we can create agents capable of solving novel long-horizon tasks which require hierarchical reasoning.

Large “foundation models" have become a dominant paradigm in solving tasks in natural language processing , computer vision , and mathematical reasoning . In line with this paradigm, a question of broad interest is to develop a “foundation model” that can solve novel and long-horizon decision-making tasks. Some prior works collected paired visual, language and action data and trained a monolithic neural network for solving long-horizon tasks. However, collecting paired visual, language and action data is expensive and hard to scale up. Another line of prior works finetune large language models (LLM) on both visual and language inputs using task-specific robot demonstrations. This is problematic because, unlike the abundance of text on the Internet, paired vision and language robotics demonstrations are not readily available and are expensive to collect. Furthermore, finetuning high-performing language models, such as GPT3.5/4 and PaLM , is currently impossible, as the model weights are not open-sourced.

The key characteristic of the foundation model is that solving a new task or adapting to a new environment is possible with much less data compared to training from scratch for that task or domain. Instead of building a foundation model for long-term planning by collecting paired language-vision-action data, in this work we seek a scalable alternative – can we reduce the need for a costly and tedious process of collecting paired data across three modalities and yet be relatively efficient at solving new planning tasks? We propose Compositional Foundation Models for Hierarchical Planning (HiP), a foundation model that is a composition of different expert models trained on language, vision, and action data individually. Because these models are trained individually, the data requirements for constructing the foundation models are substantially reduced (Figure 1). Given an abstract language instruction describing the desired task, HiP uses a large language model to find a sequence of sub-tasks (i.e., planning). HiP then uses a large video diffusion model to capture geometric and physical information about the world and generates a more detailed plan in form of an observation-only trajectory. Finally, HiP uses a large pre-trained inverse model that maps a sequence of ego-centric images into actions. The compositional design choice for decision-making allows separate models to reason at different levels of the hierarchy, and jointly make expert decisions without the need for ever collecting expensive paired decision-making data across modalities.

Given three models trained independently, they can produce inconsistent outputs that can lead to overall planning failure. For instance, a naïve approach for composing models is to take the maximum-likelihood output at each stage. However, a step of plan which is high likelihood under one model, i.e. looking for a tea kettle in a cabinet may have zero likelihood under a seperate model, i.e. if there is no cabinet in the house. It is instead important to sample a plan that jointly maximizes likelihood across every expert model. To create consistent plans across our disparate models, we propose an iterative refinement mechanism to ensure consistency using feedback from the downstream models . At each step of the language model’s generative process, intermediate feedback from a likelihood estimator conditioned on an image of the current state is incorporated into the output distribution. Similarly, at each step of the video model generation, intermediate feedback from the action model refines video generation. This iterative refinement procedure promotes consensus among the different models and thereby enables hierarchically consistent plans that are both responsive to the goal and executable given the current state and agent. Our proposed iterative refinement approach is computationally efficient to train, as it does not require any large model finetuning. Furthermore, we do not require access to the model’s weights and our approach works with any models that offer only input and output API access.

In summary, we propose a compositional foundation model for hierarchical planning that leverages a composition of foundation models, learned separately on different modalities of Internet and ego-centric robotics data, to construct long-horizon plans. We demonstrate promising results on three long-horizon tabletop manipulation environments.

Compostional Foundation Models for Hierarchical Planning

We propose HiP, a foundation model that decomposes the problem of generating action trajectories for long-horizon tasks specified by a language goal gg into three levels of hierarchy: (1) Task planning – inferring a language subgoal wiw_{i} conditioned on observation xi,1x_{i,1} and language goal gg; (2) Visual planning – generating a physically plausible plan as a sequence of image trajectories τxi={xi,1:T}\tau_{x}^{i}=\{x_{i,1:T}\}, one for each given language subgoal wiw_{i} and observation at first timestep xi,1x_{i,1}; (3) Action planning – inferring a sequence of action trajectories τai={ai,1:T−1}\tau_{a}^{i}=\{a_{i,1:T-1}\} from the image trajectories τxi\tau_{x}^{i} executing the plan. Figure 2 illustrates the model architecture and a pseudocode is provided in Algorithm 1.

Let pΘp_{\Theta} model this hierarchical decision-making process. Given our three levels of hierarchy, pΘp_{\Theta} can be factorized into the following: task distribution pθp_{\theta}, visual distribution pϕp_{\phi}, and action distribution pψp_{\psi}. The distribution over plans, conditioned on the goal and an image of the initial state, can be written under the Markov assumption as:

We seek to find action trajectories τai\tau_{a}^{i}, image trajectories τxi\tau_{x}^{i} and subgoals W={wi}W=\{w_{i}\} which maximize the above likelihood. Please see Appendix A for a derivation of this factorization. In the following sub-sections, we describe the form of each of these components, how they are trained, and how they are used to infer a final plan for completing the long-horizon task.

Given a task specified in language gg and the current observation xi,1x_{i,1}, we use a pretrained LLM as the task planner, which decomposes the goal into a sequence of subgoals. The LLM aims to infer the next subgoal wiw_{i} given a goal gg and models the distribution pLLM(wi∣g)p_{\text{LLM}}(w_{i}|g). As the language model is trained on a vast amount of data on the Internet, it captures powerful semantic priors on what steps should be taken to accomplish a particular task. To adapt the LLM to obtain a subgoal sequence relevant to our task, we prompt it with some examples of domain specific data consisting of high-level goals paired with desirable subgoal sequences.

However, directly sampling subgoals using a LLM can lead to samples that are inconsistent with the overall joint distribution in Eqn (1), as the subgoal wiw_{i} not only affects the marginal likelihood of task planning but also the downstream likelihoods of the visual planning model. Consider the example in Figure 2 where the agent is tasked with packing computer mouse, black and blue sneakers, pepsi box and toy train in brown box. Let’s say the computer mouse is already in the brown box. While the subgoal of placing computer mouse in brown box has high-likelihood under task model pθ(wi∣g)p_{\theta}(w_{i}|g), the resulting observation trajectory generated by visual model pϕ(τxi∣wi,xi,1)p_{\phi}(\tau_{x}^{i}|w_{i},x_{i,1}) will have a low-likelihood under pϕp_{\phi} given the subgoal is already completed. Next, we describe how we use iterative refinement to capture this dependency between language decoding and visual planning to properly sample from Eqn (1).

To ensure that we sample subgoal wiw_{i} that maximizes the joint distribution in Eqn (1), we should sample a subgoal that maximizes the following joint likelihood

i.e. a likelihood that maximizes both conditional subgoal generation likelihood from a LLM and the likelihood of sampled videos τxi\tau_{x}^{i} given the language instruction and current image xi,1x_{i,1}. One way to determine the optimal subgoal wi∗w_{i}^{*} is to generate multiple wiw_{i} from LLM and score them using the likelihood of videos sampled from our video model pϕ(τxi∣wi,xi,1)p_{\phi}(\tau_{x}^{i}|w_{i},x_{i,1}). However, the video generation process is computationally expensive, so we take a different approach.

The likelihood of video generation pϕ(τxi∣wi,xi,1)p_{\phi}(\tau_{x}^{i}|w_{i},x_{i,1}) primarily corresponds to the feasibility of a language subgoal wiw_{i} with respect to the initial image xi,1x_{i,1}. Thus an approximation of Eqn (2) is to directly optimize the conditional density

We estimate the density ratio p(xi,1∣wi,g)p(xi,1∣g)\frac{p(x_{i,1}|w_{i},g)}{p(x_{i,1}|g)} with a multi-class classifier fϕ(xi,1,{wj}i=1M,g)f_{\phi}(x_{i,1},\{w_{j}\}_{i=1}^{M},g) that chooses the appropriate subgoal wi∗w_{i}^{*} from candidate subgoals {wj}j=1M\{w_{j}\}_{j=1}^{M} generated by the LLM. The classifier implicitly estimates the relative log likelihood estimate of p(xi,1∣wi,g)p(x_{i,1}|w_{i},g) and use these logits to estimate the log density ratio with respect to each of the MM subgoals and find wi∗w_{i}^{*} that maximizes the estimate . We use a dataset Dclassify≔{xi,1,g,{wj}j=1M,i}{\mathcal{D}}_{\text{classify}}\coloneqq\{x_{i,1},g,\{w_{j}\}_{j=1}^{M},i\} consisting of observation xi,1x_{i,1}, goal gg, candidate subgoals {wj}j=1M\{w_{j}\}_{j=1}^{M} and the correct subgoal label ii to train fϕf_{\phi}. For further architectural details, please refer to Appendix B.1.

2 Visual Planning with Video Generation

Upon obtaining a language subgoal wiw_{i} from task planning, our visual planner generates a plausible observation trajectory τxi\tau_{x}^{i} conditioned on current observation xi,1x_{i,1} and subgoal wiw_{i}. We use a video diffusion model for visual planning given its success in generating text-conditioned videos . To provide our video diffusion model with a rich prior for physically plausible motions, we pretrain it pϕ(τxi∣wi,xi,1)p_{\phi}(\tau_{x}^{i}|w_{i},x_{i,1}) on a large-scale text-to-video dataset Ego4D . We then finetune it on the task-specific video dataset Dvideo≔{τxi,wi}{\mathcal{D}}_{\text{video}}\coloneqq\{\tau_{x}^{i},w_{i}\} consisting of observation trajectories τxi\tau_{x}^{i} satisfying subgoal wiw_{i}. For further architectural details, please refer to Appendix B.2.

However, analogous to the consistent sampling problem in task planning, directly sampling observation trajectories with video diffusion can lead to samples that are inconsistent with the overall joint distribution in Eqn (1). The observation trajectory τxi\tau_{x}^{i} not only affects the marginal likelihood of visual planning, but also the downstream likelihood of the action planning model.

To ensure observation trajectories τxi\tau_{x}^{i} that correctly maximize the joint distribution in Eqn (1), we optimize an observation trajectory that maximizes the following joint likelihood

i.e. an image sequence that maximizes both conditional observation trajectory likelihood from video diffusion and the likelihood of sampled actions τai\tau_{a}^{i} given the observation trajectory τxi\tau_{x}^{i}.

To sample such an observation trajectory, we could iteratively bias the denoising of video diffusion using the log-likelihood of the sampled actions ∏t=1T−1log⁡pψ(ai,t∣xi,t,xi,t+1)\prod_{t=1}^{T-1}\log p_{\psi}(a_{i,t}|x_{i,t},x_{i,t+1}). While this solution is principled, it is slow as it requires sampling of entire action trajectories and calculating the corresponding likelihoods during every step of the denoising process. Thus, we approximate the sampling and the likelihood calculation of action trajectory ∏t=1T−1pψ(ai,t∣xi,t,xi,t+1)\prod_{t=1}^{T-1}p_{\psi}(a_{i,t}|x_{i,t},x_{i,t+1}) with a binary classifier gψ(τxi)g_{\psi}(\tau_{x}^{i}) that models if the observation trajectory τxi\tau_{x}^{i} leads to a high-likelihood action trajectory.

We learn a binary classifier gψg_{\psi} to assign high likelihood to feasible trajectories sampled from our video dataset τxi∼Dvideo\tau_{x}^{i}\sim{\mathcal{D}}_{\text{video}} and low likelihood to infeasible trajectories generated by randomly shuffling the order of consecutive frames in feasible trajectories. Once trained, we can use the likelihood log⁡gψ(1∣τxi)\log g_{\psi}(1|\tau_{x}^{i}) to bias the denoising of the video diffusion and maximize the likelihood of the ensuing action trajectory. For further details on binary classifier, please refer to Appendix C.2.

3 Action Planning with Inverse Dynamics

After generating an observation trajectory τxi\tau_{x}^{i} from visual planning, our action planner generates an action trajectory τai\tau_{a}^{i} from the observation trajectory. We leverage egocentric internet images for providing our action planner with useful visual priors. Our action planner is parameterized as an inverse dynamics model that infers the action ai,ta_{i,t} given the observation pair (xi,t,xi,t+1)(x_{i,t},x_{i,t+1}):

To imbue the inverse dynamics pψp_{\psi} with useful visual priors, we initialize it with VC-1 weights, pretrained on ego-centric images and ImageNet. We then finetune it on dataset Dinv≔{τxi,τai}{\mathcal{D}}_{\text{inv}}\coloneqq\{\tau_{x}^{i},\tau_{a}^{i}\} consisting of paired observation and action trajectories by optimizing:

For further architectural details, please refer to Appendix B.3.

Experimental Evaluations

We evaluate the ability of HiP to solve long-horizon planning tasks that are drawn from distributions with substantial variation, including the number and types of objects and their arrangements. We then study the effects of iterative refinement and of pretraining on overall performance of HiP. We also compare against an alternative strategy of visually grounding the LLM without any task-specific data. In addition, we study how granularity of subgoals affects HiP’s performance, ablate over choice of visual planning model and analyze sensitivity of iterative refinement to hyperparameters (Appendix E).

We evaluate HiP on three environments, paint-block, object-arrange, and kitchen-tasks which are inspired by combinatorial planning tasks in Mao et al. , Shridhar et al. and Xing et al. respectively.

paint-block: A robot has to manipulate blocks in the environment to satisfy language goal instructions, such as stack pink block on yellow block and place green block right of them. However, objects of correct colors may not be present in the environment, in which case, the robot needs to first pick up white blocks and put them in the appropriate color bowls to paint them. After that, it should perform appropriate pick-and-place operations to stack a pink block on the yellow block and place the green block right of them. A new task T{\mathcal{T}} is generated by randomly selecting 33 final colors (out of 1010 possible colors) for the blocks and then sampling a relation (out of 33 possible relations) for each pair of blocks. The precise locations of individual blocks, bowls, and boxes are fully randomized across different tasks. Tasks have 4∼64\sim 6 subgoals.

object-arrange: A robot has to place appropriate objects in the brown box to satisfy language goal instructions such as place shoe, tablet, alarm clock, and scissor in brown box. However, the environment may have distractor objects that the robot must ignore. Furthermore, some objects can be dirty, indicated by a lack of texture and yellow color. For these objects, the robot must first place them in a blue cleaning box and only afterwards place those objects in the brown box. A new task T{\mathcal{T}} is generated by randomly selecting 77 objects (out of 5555 possible objects), out of which 33 are distractors, and then randomly making one non-distractor object dirty. The precise locations of individual objects and boxes are fully randomized across different tasks. Tasks usually have 3∼53\sim 5 subgoals.

kitchen-tasks: A robot has to complete kitchen subtasks to satisfy language goal instructions such as open microwave, move kettle out of the way, light the kitchen area, and open upper right drawer. However, the environment may have objects irrelevant to the subtasks that the robot must ignore. Furthermore, some kitchen subtasks specified in the language goal might already be completed, and the robot should ignore those tasks. There are 77 possible kitchen subtasks: opening the microwave, moving the kettle, switching on lights, turning on the bottom knob, turning on the top knob, opening the left drawer, and opening the right drawer. A new task T{\mathcal{T}} is generated by randomly selecting 44 out of 77 possible kitchen subtasks, randomly selecting an instance of microwave out of 33 possible instances, randomly selecting an instance of kettle out of 44 possible instances, randomly and independently selecting texture of counter, floor and drawer out of 33 possible textures and randomizing initial pose of kettle and microwave. With 50%50\% probability, one of 44 selected kitchen subtask is completed before the start of the task. Hence, tasks usually have 3∼43\sim 4 subtasks (i.e. subgoals).

For all environments, we sample two sets of tasks Ttrain,Ttest∼p(T){\mathcal{T}}_{train},{\mathcal{T}}_{test}\sim p({\mathcal{T}}). We use the train set of tasks Ttrain{\mathcal{T}}_{train} to create datasets Dclassify{\mathcal{D}}_{\text{classify}}, Dvideo{\mathcal{D}}_{\text{video}}, Dinv{\mathcal{D}}_{\text{inv}} and other datasets required for training baselines. We ensure the test set of tasks Ttest{\mathcal{T}}_{test} contains novel combinations of object colors in paint-block, novel combinations of object categories in object-arrange, and novel combinations of kitchen subtasks in kitchen-tasks.

Evaluation Metrics

We quantitatively evaluate a model by measuring its task completion rate for paint-block and object-arrange, and subtask completion rate for kitchen tasks. We use the simulator to determine if the goal, corresponding to a task, has been achieved. We evaluate a model on Ttrain{\mathcal{T}}_{train} (seen) to test its ability to solve long-horizon tasks and on Ttest{\mathcal{T}}_{test} (unseen) to test its ability to generalize to long-horizon tasks consisting of novel combinations of object colors in paint-block, object categories in object-arrange, and kitchen subtasks in kitchen-tasks. We sample 10001000 tasks from Ttrain\mathcal{T}_{train} and Ttest\mathcal{T}_{test} respectively, and obtain average task completion rate on paint-block and object-arrange domains and average subtask completion rate on kitchen tasks domain. We repeat this procedure over 44 different seeds and report the mean and the standard error over those seeds in Table 1.

2 Baselines

There are several existing strategies for constructing robot manipulation policies conditioned on language goals, which we use as baselines in our experiments:

Goal-Conditioned Policy A goal-conditioned transformer model ai,t∼p(ai,t∣xi,t,wi)a_{i,t}\sim p(a_{i,t}|x_{i,t},w_{i}) that outputs action ai,ta_{i,t} given a language subgoal wiw_{i} and current observation xi,tx_{i,t} (Transformer BC) . We provide the model with oracle subgoals and encode these subgoals with a pretrained language encoder (Flan-T5-Base). We also compare against goal-conditioned policy with Gato transformer architecture.

Video Planner A video diffusion model (UniPi) {τxi}∼p({τxi}∣g,xi,1),ai,t∼p(ai,t∣xi,t,xi,t+1)\{\tau_{x}^{i}\}\sim p(\{\tau_{x}^{i}\}|g,x_{i,1}),a_{i,t}\sim p(a_{i,t}|x_{i,t},x_{i,t+1}) that bypasses task planning, generates video plans for the entire task {τxi}\{\tau_{x}^{i}\}, and infers actions ai,ta_{i,t} using an inverse model.

Action Planners Transformer models (Trajectory Transformer) and diffusion models (Diffuser) {ai,t:T−1}∼p({ai,t:T−1}∣xi,t,wi)\{a_{i,t:T-1}\}\sim p(\{a_{i,t:T-1}\}|x_{i,t},w_{i}) that produce an action sequence {ai,t:T−1}\{a_{i,t:T-1}\} given a language subgoal wiw_{i} and current visual observation xi,tx_{i,t}. We again provide the agents with oracle subgoals and encode these subgoals with a pretrained language encoder (Flan-T5-Base).

LLM as Skill Manager A hierarchical system (SayCan) with LLM as high level policy that sequences skills sampled from a repetoire of skills to accomplish a long-horizon task. We use CLIPort policies as skills and the unnormalized logits over the pixel space it produces as affordances. These affordances grounds the LLM to current observation for producing next subgoal.

3 Results

We begin by comparing the performance of HiP and baselines to solve long-horizon tasks in paint-block, object-arrange, kitchen-tasks environments. Table 1 shows that HiP significantly outperforms the baselines, although the baselines have an advantage and have access to oracle subgoals. HiP’s superior performance shows the importance of (i) hierarchy given it outperforms goal-conditioned policy (Transformer BC and Gato), (ii) task planning since it outperforms video planners (UniPi), and (iii) visual planning given it outperforms action planners (Trajectory Transformer, Action Diffuser). It also shows the importance of representing skills with video-based planners which can be pre-trained on Internet videos and can be applied to tasks (such as kitchen-tasks). SayCan, in contrast, requires tasks to be decomposed into primitives paired with an affordance function, which can be difficult to define for many tasks like the kitchen task. Thus, we couldn’t run SayCan on kitchen-tasks environment. Finally, to quantitatively show how the errors in fθ(xi,1,wi,g)f_{\theta}(x_{i,1},w_{i},g) affect the performance of HiP, we compare it to HiP with oracle subgoals. For further details on the training and evaluation of HiP, please refer to Appendix C. For implementation details on Gato and SayCan, please refer to Appendix D. For runtime analysis of different components of HiP, please refer to Appendix F.

We also quantitatively test the ability of HiP to generalize to unseen long-horizon tasks, consisting of novel combinations of object colors in paint-block, object categories in object-arrange, and subtasks in kitchen-tasks. Table 1 shows that HiP’s performance remains intact when solving unseen long-horizon tasks, and still significantly outperforms the baselines. Figure 4 visualizes the execution of HiP in unseen long-horizon tasks in paint-block.

Pre-training Video Diffusion Model

We investigate how much our video diffusion model benefits from pre-training on the Internet-scale data. We report both the success rate of HiP and Fréchet Video Distance (FVD) score that quantifies the similarity between generated videos and ground truth videos, where lower scores indicate greater similarity in Figure 5. We see that pretraining video diffusion leads to a higher success rate and lower FVD score. If we reduce the training dataset to 75%75\% and 50%50\% of the original dataset, the FVD score for video diffusion models (both, with and without Ego4D dataset pretraining) increases and their success rate falls. However, the video diffusion model with Ego4D dataset pretraining consistently gets higher success rate and lower FVD scores across different dataset sizes. As we decrease the domain-specific training data, it is evident that the gap in performance between the model with and without the Ego4D pre-training widens. For details on how we process the Ego4D dataset, please refer to Appendix C.2.

Pre-training Inverse Dynamics Model

We also analyze the benefit of pre-training our inverse dynamics model and report the mean squared error between the predicted and ground truth actions in Figure 6. The pre-training comes in the form of initializing the inverse dynamics model with weights from VC-1 , a vision-transformer (ViT-B) trained on ego-centric images with masked-autoencoding objective . In paint-block and object-arrange, we see that the inverse dynamics, when initialized with weights from VC-1, only requires 1K labeled robotic trajectories to achieve the same performance as the inverse dynamics model trained on 10K labeled robotic trajectories but without VC-1 initialization. We also compare against an inverse dynamics model parameterized with a smaller network (ResNet-18). However, the resulting inverse dynamics model still requires 2.5K robotic trajectories to get close to the performance of the inverse dynamics model with VC-1 initialization in paint-block and object-arrange. In kitchen-tasks, inverse dynamics, when initialized with weights from VC-1, only requires 3.5k labeled robotic trajectories to achieve the same performance as the inverse dynamics model trained on 10K labeled robotic trajectories but without VC-1 initialization. When parameterized with ResNet-18, the inverse dynamics model still requires 6k robotic trajectories to get close to the performance of the inverse dynamics model with VC-1 initialization.

Importance of Task Plan and Visual Plan Refinements

We study the importance of refinement in task and visual planning in Figure 7. We compare to HiP without visual plan refinement and HiP without visual and task plan refinement in paint block environment. We see that task plan refinement for visual grounding of LLM is critical to the performance of HiP. Without it, the task plan is agnostic to the robot’s observation and predicts subgoals that lead to erroneous visual and action planning. Furthermore, visual plan refinement improves the performance of HiP as well, albeit by a small margin. For a detailed description of the hyperparameters used, please refer to Appendix C.4.

Exploring Alternate Strategies for Visually Grounding LLM

We use a learned classifier fθ(xi,1,wi,g)f_{\theta}(x_{i,1},w_{i},g) to visually ground the LLM. We explore if we can use a frozen pretrained Vision-Language Model (MiniGPT-4 ) as a classifier in place of the learned classifier. Although we didn’t use any training data, we found the prompt engineering using the domain knowledge of the task to be essential in using the Vision-Language Model (VLM) as a classifier (see Appendix C.1 for details). We use subgoal prediction accuracy to quantify the performance of the learned classifier and the frozen VLM. Figure 7 illustrates that while both our learned multi-class classifier and frozen VLM perform comparably in the paint-block environment, the classifier significantly outperforms the VLM in the more visually complex object-arrange environment. We detail the two common failure modes of the VLM approach in object-arrange environment in Appendix C.1. As VLMs continue to improve, it is possible that their future versions match the performance of learned classifiers and thus replace them in visually complex domains as well. For further details on the VLM parameterization, please refer to Appendix C.1.

Related Work

The field of foundation models for decision-making has seen significant progress in recent years. A large body of work explored using large language models as zero-shot planners , but it is often difficult to directly ground the language model on vision. To address this problem of visually grounding the language model, other works have proposed to directly fine-tune large language models for embodied tasks . However, such an approach requires large paired vision and language datasets that are difficult to acquire. Most similar to our work, SayCan uses an LLM to hierarchically execute different tasks by breaking language goals into a sequence of instructions, which are then inputted to skill-based value functions. While SayCan assumes this fixed set of skill-based value functions, our skills are represented as video-based planners , enabling generalization to new skills.

Another set of work has explored how to construct continuous space planners with diffusion models . Existing works typically assume task-specific datasets from which the continuous-space planner is derived . Most similar to our work, UniPi proposes to use videos to plan in image space and similarily relies on internet videos to train image space planners. We build on top of UniPi to construct our foundation model for hierarchical planning, and illustrate how UniPi may be combined with LLMs to construct longer horizon continuous video plans.

Moreover, some works explored how different foundation models may be integrated with each other. In Flamingo , models are combined through joint finetuning with paired datasets, which are difficult to collect. In contrast both Zeng et al. and Li et al. combine different models zero-shot using either language or iterative consensus. Our work proposes to combine language, video, and ego-centric action models together by taking the product of their learned distributions . We use a similar iterative consensus procedure as in Li et al. to sample from the entire joint distribution and use this combined distribution to construct a hierarchical planning system.

Limitations and Conclusion

Our approach has several limitations. As high-quality foundation models for visual sequence prediction and robot action generation do not exist yet, our approach relies on smaller-scale models that we directly train. Once high-quality video foundation models are available, we can use them to guide our smaller-scale video models which would reduce the data requirements of our smaller-scale video models. Furthermore, our method uses approximations to sample from the joint distribution between all the model. An interesting avenue for future work is to explore more efficient and accurate methods to ensure consistent samples from the joint distribution.

Conclusion

In this paper, we have presented an approach to combine many different foundation models into a consistent hierarchical system for solving long-horizon robotics problems. Currently, large pretrained models are readily available in the language domain only. Ideally, one would train a foundation model for videos and ego-centric actions, which we believe will be available in the near future. However, our paper focuses on leveraging separate foundation models trained on different modalities of internet data, instead of training a single big foundation model for decision making. Hence, for the purposes of this paper, given our computational resource limitations, we demonstrate our general strategy with smaller-scale video and ego-centric action models trained in simulation, which serve as proxies for larger pretrained models. We show the potential of this approach in solving three long-horizon robot manipulation problem domains. Across environments with novel compositions of states and goals, our method significantly outperforms the state-of-the-art approaches towards solving these tasks.

In addition to building larger, more general-purposed visual sequence and robot control models, our work suggests the possibility of further using other pretrained models in other modalities, such as touch and sound, which may be jointly combined and used by our sampling approach. Overall, our work paints a direction towards decision making by leveraging many different powerful pretrained models in combination with a tiny bit of training data. We believe that such a system will be substantially cheaper to train and will ultimately result in more capable and general-purpose decision making systems.

Acknowledgements

The authors would like to thank the members of Improbable AI Lab for discussions and helpful feedback. We thank MIT Supercloud and the Lincoln Laboratory Supercomputing Center for providing compute resources. This research was supported by an NSF graduate fellowship, a DARPA Machine Common Sense grant, ARO MURI Grant Number W911NF-21-1-0328, ARO MURI W911NF2310277, ONR MURI Grant Number N00014-22-1-2740, and an MIT-IBM Watson AI Lab grant. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the United States Army Research Office, the Office of Naval Research, or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes, notwithstanding any copyright notation herein.

Author Contributions

Anurag Ajay co-conceived the framework of leveraging pretrained foundation models for decision making, implemented visual planning and action planning in HiP, evaluated HiP on long-horizon tasks, performed ablation studies and helped in paper writing.

Seungwook Han co-conceived in conceiving the framework of leveraging pretrained foundation models for decision making, implemented task planning in HiP and helped in paper writing.

Yilun Du co-conceived the framework of leveraging pretrained foundation models for decision making, implemented trajectory transformer and transformer BC and evaluated them on long-horizon tasks, implemented data generation scripts for object arrange and paint block environment, and lead paper writing.

Shuang Li helped in conceiving the idea of iterative refinement for consistency between pretrained foundation models, participated in research discussions and helped in writing paper.

Abhi Gupta participated in research discussions and helped in making figures.

Tommi Jaakkola participated in research discussions.

Joshua Tenenbaum participated in research discussions.

Leslie Kaelbling participated in research discussions, suggested baselines and ablation studies, conceived the structure of the paper and helped in paper writing.

Akash Srivastava participated in research discussions, suggested the idea of using classifier for consistency between observations and the large language model and provided feedback on paper writing.

Pulkit Agrawal was involved in research discussions, suggested ablation studies related to iterative refinement, provided feedback on writing, positioning of the work and overall advising.

References

Appendix A Factorizing Hierarchical Decision-Making Process

We model the hierarchical decision-making process described in Section 2 with pΘp_{\Theta} which can be factorized into the task distribution pθp_{\theta}, visual distribution pϕp_{\phi}, and action distribution pψp_{\psi}.

Here, given random variables YiY^{i}, Y<iY^{<i} and Y≤iY^{\leq i} represents {Y1,…,Yi−1}\{Y^{1},\ldots,Y^{i-1}\} and {Y1,…,Yi}\{Y^{1},\ldots,Y^{i}\} respectively. Now, we apply Markov assumption: Given current observation xi,1x_{i,1}, future variables (wi,τxi,τaiw_{i},\tau_{x}^{i},\tau_{a}^{i}) and past variables (wj,τxj,τaj    ∀j<iw_{j},\tau_{x}^{j},\tau_{a}^{j}\;\;\forall j<i) are conditionally independent.

We model task distribution pθp_{\theta} with a large language model (LLM) which is independent of observation xi,1x_{i,1}. Since the image trajectory τxi={xi,1:T}\tau_{x}^{i}=\{x_{i,1:T}\} describes a physically plausible plan for achieving subgoal wiw_{i} from observation xi,1x_{i,1}, it is conditionally independent of goal gg given subgoal wiw_{i} and observation xi,1x_{i,1}. Furthermore, we assume that an action ai,ta_{i,t} can be recovered from observation at the same timestep xi,tx_{i,t} and the next timestep xi,t+1x_{i,t+1}. Thus, we can write the factorization as

Appendix B Background and Architecture

Let pp and qq be two densities, such that qq is absolutely continuous with respect to pp, denoted as q<<pq<<p i.e. q(x)>0q(\bm{x})>0 wherever p(x)>0p(\bm{x})>0. Then, their ratio is defined as r(x)=p(x)/q(x)r(\bm{x})=p(\bm{x})/q(\bm{x}) over the support of pp. We can estimate this density ratio r(x)r(\bm{x}) by training a binary classifier to distinguish between samples from pp and qq . More recent work has shown one can introduce auxiliary densities {mi}i=1M\{m_{i}\}_{i=1}^{M} and train a multi-class classifier to distinguish samples between MM classes to learn a better-calibrated and more accurate density ratio estimator. Once trained, the log density ratio can be estimated by log⁡r(x)=hp^(x)−hq^(x)\log{r(\bm{x})}=\hat{h_{p}}(\bm{x})-\hat{h_{q}}(\bm{x}), where hi^(x)\hat{h_{i}}(\bm{x}) is the unnormalized log probability of the input sample under the ithi^{th} density, parameterized by the model.

Learning a Classifier to Visually Ground Task Planning

We estimate the density ratio p(xi,1∣wi,g)p(xi,1∣g)\frac{p(x_{i,1}|w_{i},g)}{p(x_{i,1}|g)} with a multi-class classifier fϕ(xi,1,{wj},g)f_{\phi}(x_{i,1},\{w_{j}\},g) trained to distinguish samples amongst the conditional distributions p(xi,1∣wi,g),...,p(xi,1∣wM,g)p(x_{i,1}|w_{i},g),...,p(x_{i,1}|w_{M},g) and the marginal distribution p(xi,1∣g)p(x_{i,1}|g). Upon convergence, the classifier learns to assign high scores to (xi,1,wi,g)(x_{i,1},w_{i},g) if wiw_{i} is the subgoal corresponding to the observation xi,1x_{i,1} and task gg and low scores otherwise.

Architecture

We parameterize fϕf_{\phi} as a 4-layer multi-layer perceptron (MLP) on top of an ImageNet-pretrained vision encoder (ResNet-18 ) and a frozen pretrained language encoder (Flan-T5-Base ). The vision encoder encodes the observation xi,1x_{i,1}, and the text encoder encodes the subgoals wjw_{j} and the goal gg. The encoded observation, the encoded subgoals, and the encoded goal are concatenated, and passed through a MLP with 33 hidden layers of sizes 512512, 256256, and 128128. The output dimension for MLP (i.e., number of classes for multi-classification) MM is 66 for paint-block environment, 55 for object-arrange environment and 44 for kitchen-tasks environment.

Choice of Large Language Model

We use GPT3.5-turbo as our large language model.

B.2 Visual Planning

Diffusion Probabilistic Models learn the data distribution h(x)h(\bm{x}) from a dataset D≔{xi}{\mathcal{D}}\coloneqq\{\bm{x}^{i}\}. The data-generating procedure involves a predefined forward noising process q(xk+1∣xk)q(\bm{x}_{k+1}|\bm{x}_{k}) and a trainable reverse process pϕ(xk−1∣xk)p_{\phi}(\bm{x}_{k-1}|\bm{x}_{k}), both parameterized as conditional Gaussian distributions. Here, x0≔x\bm{x}_{0}\coloneqq\bm{x} is a sample, x1,x2,...,xK−1\bm{x}_{1},\bm{x}_{2},...,\bm{x}_{K-1} are the latents, and xK∼N(0,I)\bm{x}_{K}\sim\mathcal{N}(\bm{0},\bm{I}) for a sufficiently large KK. Starting with Gaussian noise, samples are then iteratively generated through a series of “denoising” steps. Although a tractable variational lower-bound on log⁡pϕ\log p_{\phi} can be optimized to train diffusion models, Ho et al. propose a simplified surrogate loss:

The predicted noise ϵθ(xk,k)\epsilon_{\theta}(\bm{x}_{k},k), parameterized with a deep neural network, estimates the noise ϵ∼N(0,I)\epsilon\sim{\mathcal{N}}(0,I) added to the dataset sample x0\bm{x}_{0} to produce noisy xk\bm{x}_{k}.

Guiding Diffusion Models with Text

Diffusion models are most notable for synthesizing high-quality images and videos from text descriptions. Modeling the conditional data distribution q(x∣y)q(\bm{x}|\bm{y}) makes it possible to generate samples satisfying the text description y\bm{y}. To enable conditional data generation with diffusion, Ho and Salimans modified the original training setup to learn both a conditional ϵϕ(xk,y,k)\epsilon_{\phi}(\bm{x}_{k},\bm{y},k) and an unconditional ϵϕ(xk,k)\epsilon_{\phi}(\bm{x}_{k},k) model for the noise. The unconditional noise is represented, in practice, as the conditional noise ϵϕ(xk,∅,k)\epsilon_{\phi}(\bm{x}_{k},\emptyset,k), where a dummy value ∅\emptyset takes the place of y\bm{y}. The perturbed noise \epsilon_{\phi}(\bm{x}_{k},\emptyset,k)+\omega\bigl{(}\epsilon_{\phi}(\bm{x}_{k},\bm{y},k)-\epsilon_{\phi}(\bm{x}_{k},\emptyset,k)\bigr{)} (i.e. classifier-free guidance) is used to later generate samples.

Video Diffusion in Latent Space

As diffusion models generally perform denoising in the input space , the optimization and inference become computationally demanding when dealing with high-dimensional data, such as videos. Inspired by recent works , we first use an autoencoder vencv_{\text{enc}} to learn a latent space for our video data. It projects an observation trajectory τx\tau_{x} (i.e., video) into a 2D tri-plane representation τz=[τzT,τzH,τzW]\tau_{z}=[\tau_{z}^{T},\tau_{z}^{H},\tau_{z}^{W}] where τzT\tau_{z}^{T}, τzH\tau_{z}^{H}, τzW\tau_{z}^{W} capture variations in the video across time, height, and width respectively. We then diffuse over this learned latent space .

Latent Space Video Diffusion for Visual Planning

Our video diffusion model pϕ(τxi∣wi,xi,1)p_{\phi}(\tau_{x}^{i}|w_{i},x_{i,1}) generates video τxi\tau_{x}^{i} given a language subgoal wiw_{i} and the current observation xi,1x_{i,1}. It is parameterized through its noise model ϵϕ((τzi)k,wi,xi,1,k)≔ϵϕ((τzi)k,lenc(wi),venc(xi,1),k)\epsilon_{\phi}((\tau_{z}^{i})_{k},w_{i},x_{i,1},k)\coloneqq\epsilon_{\phi}((\tau_{z}^{i})_{k},l_{\text{enc}}(w_{i}),v_{\text{enc}}(x_{i,1}),k) where τzi≔venc(τxi)\tau_{z}^{i}\coloneqq v_{\text{enc}}(\tau_{x}^{i}) is the latent representation of video τxi\tau_{x}^{i} over which we diffuse. We condition the noise model ϵϕ\epsilon_{\phi} on subgoal wiw_{i} using a pretrained language encoder lencl_{\text{enc}} and on current observation xi,1x_{i,1} using video encoder vencv_{\text{enc}}. To use vencv_{\text{enc}} with a single observation xi,1x_{i,1}, we first tile the observation along the temporal dimension to create a video.

Architecture

We now detail the architectures of different components:

Language Encoder We use Flan-T5-Base as the pretrained frozen language encoder lencl_{\text{enc}}.

Noise Model We borrow PVDM-L architecture which uses 2D UNet architecture, similar to the one in Latent Diffusion Model (LDM) , to represent p(τz∣τz′)p(\tau_{z}|\tau_{z}^{\prime}). In our case, τz=venc(τxi)\tau_{z}=v_{\text{enc}}(\tau_{x}^{i}) and τz′=venc(xi,1)\tau_{z}^{\prime}=v_{\text{enc}}(x_{i,1}). To further condition noise model ϵϕ\epsilon_{\phi} on lenc(wi)l_{\text{enc}}(w_{i}), we augment the 2D UNet Model with cross-attention mechanism borrowed by LDM .

For implementing these architectures, we used the codebase https://github.com/sihyun-yu/PVDM which contains the code for PVDM and LDM.

Classifier for Consistency between Visual Planning and Action Planning

To ensure consistency between visual planning and action planning, we want to sample observation trajectories that maximizes both conditional observation trajectory likelihood from diffusion and the likelihood of sampled actions given the observation trajectory (see equation 4). To approximate likelihood calculation of action trajectory, we learn a binary classifier gψg_{\psi} that models if the observation trajectory leads to a high likelihood action trajectories. Since diffusion happens in latent space and we use gradients from gψg_{\psi} to bias the denoising of the video diffusion, gψ(τzi)g_{\psi}(\tau_{z}^{i}) takes the observation trajectory in latent space. The binary classifier gψg_{\psi} is trained to distinguish between observation trajectories in latent space sampled from video dataset τzi=venc(τxi),τxi∼Dvideo\tau_{z}^{i}=v_{\text{enc}}(\tau_{x}^{i}),\tau_{x}^{i}\sim{\mathcal{D}}_{\text{video}} (i.e. label of 11) and observation trajectories in latent space sampled from video dataset whose frames where randomly shuffled (τzi)′=venc(σ(τxi)),τxi∼Dvideo(\tau_{z}^{i})^{\prime}=v_{\text{enc}}(\sigma(\tau_{x}^{i})),\tau_{x}^{i}\sim{\mathcal{D}}_{\text{video}} (i.e. label of ). Here, σ\sigma denotes the random shuffling of frames. To randomly shuffle frames in an observation trajectory (of length 5050), we first randomly select 55 frames in the observation trajectory. For each of the selected frame, we randomly permute it with its neighboring frame (i.e. either with the frame before it or with the frame after it). Once gψg_{\psi} is trained, we use it to bias the denoising of the video diffusion

Here, ϵ^\hat{\epsilon} is the noise used in denoising of the video diffusion and ω,ω′\omega,\omega^{\prime} are guidance hyperparameters.

Classifier Architecture

The classifier gψ(τz=[τzT,τzH,τzW])g_{\psi}(\tau_{z}=[\tau_{z}^{T},\tau_{z}^{H},\tau_{z}^{W}]) has a ResNet-9 encoder that converts τzT\tau_{z}^{T}, τzH\tau_{z}^{H}, and τzW\tau_{z}^{W} to latent vectors, then concatenate those latent vectors and passes the concatenated vector through an MLP with 22 hidden layers of sizes 256256 and 128128 and an output layer of size 11.

B.3 Action Planning

To do action planning, we learn an inverse dynamics model to pψ(ai,t∣xi,t,xi,t+1)p_{\psi}(a_{i,t}|x_{i,t},x_{i,t+1}) predicts 77-dimensional robot states si,t=pψ(xi,t)s_{i,t}=p_{\psi}(x_{i,t}) and si,t+1=pψ(xi,t+1)s_{i,t+1}=p_{\psi}(x_{i,t+1}). The first 66 dimensions of the robot state represent joint angles and the last dimension of the robot state represents the gripper state (i.e., whether it’s open or closed). The first 66 action dimension is represented as joint angle difference ai,t[:6]=si,t+1[:6]−si,t[:6]a_{i,t}[:6]=s_{i,t+1}[:6]-s_{i,t}[:6] while the last action dimension is gripper state of next timestep ai,t=si,t+1a_{i,t}=s_{i,t+1}.

Appendix C Training and Evaluation

We use a softmax cross-entropy loss to train the multi-class classifier fϕ(xi,1,{wj}j=1M,g)f_{\phi}(x_{i,1},\{w_{j}\}_{j=1}^{M},g) to classify an observation xi,1x_{i,1} into one of the MM given subgoal. We train it using the classification dataset Dclassify≔{xi,1,g,{wj}j=1M,i}{\mathcal{D}}_{\text{classify}}\coloneqq\{x_{i,1},g,\{w_{j}\}_{j=1}^{M},i\} consisting of observation xi,1x_{i,1}, goal gg, candidate subgoals {wj}j=1M\{w_{j}\}_{j=1}^{M} and the correct subgoal label ii. The classification dataset for paint-block, object-arrange, and kitchen-tasks consists of 58k, 82k and 50k datapoints respectively.

Vision-Language Model (VLM) as a Classifier

We use a frozen pretrained Vision-Language Model (VLM) (MiniGPT4 ) as a classifier. We first sample a list of all possible subgoals W={wi}i=1MW=\{w_{i}\}_{i=1}^{M} from the LLM given the language goal gg. We then use the VLM to eliminate subgoals from WW that have been completed. For each subgoal, we question the VLM whether that subgoal has been completed. For example, consider the subgoal "Place white block in yellow bowl". To see if the subgoal has been completed, we ask the VLM "Is there a block in yellow bowl?". Consider the subgoal "Place green block in brown box" as another example. To see if the subgoal has been completed, we ask the VLM "Is there a green block in brown box?". Furthermore, if the VLM says "yes" and the subgoal has been completed, we also remove other subgoals from WW that should have been completed, such as "Place white block in green bowl". Once we have eliminated the completed subgoals, we use the domain knowledge to determine which subgoal to execute out of all the remaining subgoals. As an example, if the goal is to "Stack green block on top of blue block in brown box" and we have a green block in green bowl and a blue block in blue bowl, we should execute the subgoal "Place blue block in brown box" before the subgoal "Place green block on blue block". While this process of asking questions from VLM to determine the remaining subgoals and then sequencing the remaining subgoals doesn’t require any training data, it heavily relies on the task’s domain knowledge.

Failure Modes of VLM

We observe two common failure modes of the VLM approach in object-arrange environment and visualize them in Figure 8. First, because the model is not trained on any in-domain data, it often fails to recognize uncommon objects, such as computer hard drives, in the observations. Second, it occasionally hallucinates the presence of objects at certain locations and thus leads to incorrect visual reasoning.

VLM as a Subgoal Predictor

We also tried to prompt the VLM with 55 examples of goal gg and subgoal candidates {wi}i=1M\{w_{i}\}_{i=1}^{M} and then directly use it to generate the next subgoal wiw_{i} given the observation xi,1x_{i,1} and the goal gg. However, it completely failed. We hypothesize that the VLM fails to directly generate the next subgoal due to its inability to perform in-context learning.

Evaluation

We evaluate the trained classifier fϕf_{\phi} and the frozen VLM for subgoal prediction accuracy on 5k unseen datapoints, consisting of observation, goal, candidate subgoals and correct subgoal, generated from test tasks Ttest{\mathcal{T}}_{\text{test}}. We average over 44 seeds and show the results in Figure 7.

C.2 Visual Planning

We pre-train on canonical clips of the Ego4D dataset which are text-annotated short clips made from longer videos. We further divide each canonical clip into 10sec segments from which we derive 5050 frames. We resize each frame to 48×6448\times 64. We create a pretraining Ego4D dataset of (approximately) 344k short clips, each consisting of 5050 frames and a text annotation. We use the loader from R3M codebase (https://github.com/facebookresearch/r3m) to load our pretraining Ego4D dataset.

Training Objective and Dataset

Classifier Training Objective and Dataset

We use a binary cross-entropy loss to train the binary classifier gψ(τz)g_{\psi}(\tau_{z}) that predicts if the observation trajectory in latent space τz=venc(τx)\tau_{z}=v_{\text{enc}}(\tau_{x}) leads to high-likelihood action trajectory. It is trained using trajectories from video dataset τx∼Dvideo\tau_{x}\sim{\mathcal{D}}_{\text{video}}.

C.3 Action Planning

We train inverse dynamics pψp_{\psi} on a dataset Dinv{\mathcal{D}}_{\text{inv}}. Since actions are differences between robotic joint states, we train pψp_{\psi} to directly predict robotic state si,t=pψ(xi,t)s_{i,t}=p_{\psi}(x_{i,t}) by minimizing the mean squared error between the predicted robotic state and ground truth robotic state. Hence, Dinv≔{τxi,τsi}{\mathcal{D}}_{\text{inv}}\coloneqq\{\tau_{x}^{i},\tau_{s}^{i}\} consists of 1k paired observation and robotic state trajectories, each having a length of T=50T=50, in paint-block and object-arrange domains. In kitchen-tasks domain, it consists of 3.5k paired observation and robotic state trajectories, each having a length of T=50T=50.

Evaluation

We evaluate the trained pψp_{\psi} (i.e., VC-1 initialized model and other related models in Figure 6) on 100 unseen paired observation and robotic state trajectories generated from test tasks Ttest{\mathcal{T}}_{\text{test}}. We use mean squared error to evaluate our inverse dynamics models. We use 44 seeds to calculate the standard error, represented by the shaded area in Figure 6.

C.4 Hyperparameters

We train fϕf_{\phi} for 5050 epochs using AdamW optimizer , a batch size of 256256, a learning rate of 1e−31e-3 and a weight decay of 1e−61e-6. We used one V100 Nvidia GPU for training the multi-class classifier.

Visual Planning

We borrow our hyperparameters for training video diffusion from the PVDM paper . We use AdamW optimizer , a batch size of 2424 and a learning rate of 1e−41e-4 for training the autoencoder. We use AdamW optimizer, a batch size of 6464, and a learning rate of 1e−41e-4 for training the noise model. During the pretraining phase with the Ego4D dataset, we train the autoencoder for 55 epochs and then the noise model for 55 epochs. During the finetuning phase with Dvideo{\mathcal{D}}_{\text{video}}, we train the autoencoder for 1010 epochs and then the noise model for 4040 epochs. We used two A6000 Nvidia GPUs for training these diffusion models. We train gψg_{\psi} for 1010 epochs using AdamW optimizer, a batch size of 256256 and a learning rate of 1e−41e-4. We used one V100 Nvidia GPU for training the binary classifier. During classifier-free guidance, we use ω=4\omega=4 and ω′=1\omega^{\prime}=1.

Action Planning

We train VC-1 initialized inverse dynamics model for 2020 epochs with AdamW optimizer , a batch size of 256256 and a learning rate of 3e−53e-5. We trained other randomly initialized ViT-B inverse dynamics models and randomly initialized ResNet-18 inverse dynamics models for 2020 epochs with AdamW optimizer, a batch size of 256256, and a learning rate of 1e−41e-4. We used one V100 Nvidia GPU for training these inverse dynamics models.

Appendix D Implementation Details for Gato and SayCan

In training the visual and action planning for HiP, we use 100k robot videos for visual planner and the inverse dynamics, when trained from scratch, utilizes 10k state-action trajectory pairs. In order to ensure fair comparison, we use 110k datapoints for training Gato and SayCan .

We borrow the Gato architecture from Vima Codebase and use it for training a language conditioned policy with imitation learning. We use 110k (langauge, observation trajectory, action trajectory) datapoints in each of the three domains for training Gato. Furthermore, we provide oracle subgoals to Gato.

SayCan

We borrow the SayCan algorithm from the SayCan codebase and adapt it to our settings. Following the recommendations of SayCan codebase, we use CLIPort policies as primitives. CLIPort policies take in top-down RGBD view and outputs pick and place pixel coordinates. Then, an underlying motion planner picks the object from the specified pick-coordinate and places the object at the specified place-coordinate. We train CLIPort policies on 110k (language, observation, action) datapoints in paint-block and object-arrange domain. The SayCan paper uses value function as an affordance function to select the correct subgoal given current observation and high level goal. However, CLIPort policies don’t have a value function. The SayCan codebase uses a hardcoded scoring function which doesn’t apply to object-arrange domain. To overcome these issues, we use the LLM grounding strategy from huang2023grounded. It uses unnormalized logits over the pixel space given by CLIPort policies as affordance and uses it to ground LLM to current observation and thus predict the subgoal. We then compare SayCan with HiP and other baselines on paint-block and object-arrange domain in Table 1. While SayCan outpeforms other baselines, HiP still outperforms it both on seen and unseen tasks of paint-block and object-arrange domain. We couldn’t run SayCan on kitchen-tasks domain as there’s no clear-cut primitive in that domain. This points to a limitation of SayCan which requires tasks to be expressed in terms of primitives with each primitive paired with an affordance function.

Appendix E Additional Ablation Studies

To make task planning consistent with visual planning, we need to select subgoal wi∗w_{i}^{*} which maximizes the joint likelihood (see equation 2) of LLM pLLM(wi∣g)p_{\text{LLM}}(w_{i}|g) and video diffusion pϕ(τxi∣wi,xi,1)p_{\phi}(\tau_{x}^{i}|w_{i},x_{i,1}). While generating videos for different subgoal candidates wiw_{i} and calculating the likelihood of the generated video is computationally expensive, we would still like to evaluate its performance in subgoal prediction given it is theoretically grounded. To this end, we first sample MM subgoals W={wj}j=1MW=\{w_{j}\}_{j=1}^{M} from the LLM. Then, we calculate wi∗=arg max⁡w∈Wlog⁡pϕ(τxi∣w,xi,1)w_{i}^{*}=\operatorname*{arg\,max}_{w\in W}\log p_{\phi}(\tau_{x}^{i}|w,x_{i,1}) and use wi∗w_{i}^{*} as our predicted subgoal. Since log⁡pϕ(τxi∣w,xi,1)\log p_{\phi}(\tau_{x}^{i}|w,x_{i,1}) is intractable, we estimate its variational lower-bound as an approximation. We use this approach for subgoal prediction in paint-block environment and compare its performance to that of the learned classifier. It achieves a subgoal prediction accuracy of 54.3±7.2%54.3\pm 7.2\% whereas the learned classifier achieves a subgoal prediction accuracy of 98.2±1.5%98.2\pm 1.5\% in paint-block environment. Both approaches outperform the approach of randomly selecting a subgoal from WW (i.e., no task plan refinement), which yields a subgoal prediction accuracy of 16.67%16.67\% given M=6M=6. The poor performance of the described approach could result from the fact that the diffusion model only coarsely approximates the true distribution p(τxi∣wi,xi,1)p(\tau_{x}^{i}|w_{i},x_{i,1}), which results in loose variational lower-bound and thus uncalibrated likelihoods from the diffusion model. A larger diffusion model could better approximate p(τxi∣wi,xi,1)p(\tau_{x}^{i}|w_{i},x_{i,1}), resulting in tighter variational lower-bound and better-calibrated likelihoods.

E.2 Consistency between visual planning and action planning

To make visual planning consistent with action planning, we need to select observation trajectory (τxi)∗(\tau_{x}^{i})^{*} which maximizes joint likelihood (see equation 4) of conditional video diffusion pϕ(τxi∣wi,xi,1)p_{\phi}(\tau_{x}^{i}|w_{i},x_{i,1}) and inverse model ∏t=1T−1pψ(ai,t∣xi,t,xi,t+1)\prod_{t=1}^{T-1}p_{\psi}(a_{i,t}|x_{i,t},x_{i,t+1}). While sampling action trajectories and calculating their likelihoods during every step of the denoising process is computationally inefficient, we would still like to evaluate its effectiveness in visual plan refinements. However, we perform video diffusion in latent space while our inverse model is in observation space. Hence, for purpose of this experiment, we learn another inverse model p‾ψ‾(τai∣τzi)\overline{p}_{\overline{\psi}}(\tau_{a}^{i}|\tau_{z}^{i}) that uses a sequence model (i.e. a transformer) to produce an action trajectory τai\tau_{a}^{i} given an observation trajectory in latent space τzi\tau_{z}^{i}. We train p‾ψ‾\overline{p}_{\overline{\psi}} for 2020 epochs on 10k paired observation and action trajectories, each having a length of T=50T=50. We use AdamW optimizer, a batch size of 256256 and a learning rate of 1e−41e-4 during training. To generate an observation trajectory that maximizes the joint likelihood, we first sample 3030 observation trajectories from video diffusion pϕ(τxi∣wi,xi,1)p_{\phi}(\tau_{x}^{i}|w_{i},x_{i,1}) conditioned on subgoal wiw_{i} and observation xi,1x_{i,1}. For each generated observation trajectory τxi\tau_{x}^{i}, we sample a corresponding action trajectory τai\tau_{a}^{i} and calculate its corresponding log-likelihood log⁡p‾ψ‾(τai∣venc(τxi))\log\overline{p}_{\overline{\psi}}(\tau_{a}^{i}|v_{\text{enc}}(\tau_{x}^{i})). We select the observation trajectory τxi\tau_{x}^{i} with highest log-likelihood. Note that we only use p‾ψ‾\overline{p}_{\overline{\psi}} for visual plan refinement and use pψp_{\psi} for action execution to ensure fair comparison. If we use this approach for visual plan refinement with HiP, we obtain a success rate of 72.5±1.972.5\pm 1.9 on unseen tasks in paint-block environment. This is comparable to the performance of HiP with visual plan refinements from learned classifier gψg_{\psi} which obtains a success rate of 72.8±1.772.8\pm 1.7 on unseen tasks in paint-block environment. In contrast, HiP without any visual plan refinement obtains a success rate of 71.1±1.371.1\pm 1.3 on unseen tasks in paint-block environment. These results show that gψg_{\psi} serves as a good approximation for estimating whether an observation trajectory leads to a high-likelihood action trajectory, while still being computationally efficient.

We use a transformer model to represent p‾ψ‾(τa∣τz=[τzT,τzH,τzW])\overline{p}_{\overline{\psi}}(\tau_{a}|\tau_{z}=[\tau_{z}^{T},\tau_{z}^{H},\tau_{z}^{W}]). We first use a ResNet-9 encoder to convert τzT\tau_{z}^{T}, τzH\tau_{z}^{H}, and τzW\tau_{z}^{W} to latent vectors. We then concatenate those latent vectors and project the resulting vector to a hidden space of 6464 dimension using a linear layer. We then pass the 6464 dimensional vector to a trajectory transformer model which generates an action trajectory τa\tau_{a} of length 5050. The trajectory transformer uses a transformer architecture with 44 layers and 44 self-attention heads.

E.3 How granularity of subgoals affects performance of HiP ?

We conduct a study in paint-block environment to analyze how granuality of subgoals affect HiP. In our current setup, a subgoal in paint-block domain is of form "Place block in/on/to " and involves a pick and a place operation. We refer to our current setup as HiP (standard). We introduce two additional level of subgoal granuality:

Only one pick or place operation: The subgoal will be of form "Pick block in/on " or "Place block in/on/to ". It will involve either one pick or one place operation. We refer to the model trained in this setup as HiP (more granular).

Two pick and place operations: The subgoal will be of form "Place <1st block color> block in/on/to and Place <2nd block color> block in/on/to ". It will involve two pick and place operations. We refer to the model trained in this setup as HiP (less granular).

Note that UniPi has the least granuality in terms of subgoals as it tries to imagine the entire trajectory from goal description. Table 2 in the rebuttal document compares HiP (standard), HiP (more granular), HiP (less granular) and UniPi on seen and unseen tasks in paint-block environment. We observe that HiP (standard) and HiP (more granular) have similar success rates where HiP (less granular) has a lower success rate. UniPi has the lowest success rate amongst these variants. We hypothesize that success rate of HiP remains intact when we decrease the subgoal granuality as long as the performance of visual planner doesn’t degrade. Hence, HiP (standard) and HiP (more granular) have similar success rates. However, when the performance of visual planner degrades as we further decrease the subgoal granuality, we see a decline in success rate as well. That’s why HiP (less granular) sees a decline in success rate and UniPi has the lowest success rate amongst all variants.

E.4 Ablation on Visual Planning Model

To show the benefits of video diffusion model, we perform an ablation where we use (text-conditioned) recurrent state space model (RSSM), taken from DreamerV3 [hafner2023mastering], as visual model for HiP. We borrow the RSSM code from dreamerv3-torch codebase. To adapt RSSM to our setting, we condition RSSM on subgoal (i.e. subgoal encoded into a latent representation by Flan-T5-Base) instead of actions. Hence, sequence model of RSSM becomes ht=f(ht−1,zt−1,w)h_{t}=f(h_{t-1},z_{t-1},w) where w is latent representation of subgoal. Furthermore, we don’t predict any reward since we aren’t in a reinforcement learning setting and don’t predict continue vector since we decode for a fixed number of steps. Hence, we remove reward prediction and continue prediction from the prediction loss. To make the comparisons fair, we pretrain RSSM with Ego4D data as well. We report the results in Table 3. We see that HiP with video diffusion model outperforms HiP with RSSM in all the three domains. While the performance gap between HiP(RSSM) and HiP (i.e. using video diffusion) is small in paint-block environment, it widens in object-arrange and kitchen-tasks domains as the domains become more visually complex.

E.5 Analyzing sensitivity of iterative refinement to hyperparameters

The subgoal classifier doesn’t introduce any test time hyperparameters and we use standard hyperparameters (1e−31e-3 learning rate, 1e−61e-6 weight decay, 256256 batch size, 5050 epochs, Adam optimizer) for its training which remains fixed across all domains. We observed that the performance changes minimally across different hyperparameters, given a learning rate decay over training. However, the observation trajectory classifier gψg_{\psi} introduces an additional test time hyperparameter ω′\omega^{\prime} which appropriately weights the gradient from observation trajectory classifier. Table 4 in the rebuttal document varies ω′\omega^{\prime} between 0.5 and 2 in intervals of 0.25 and shows success rate of HiP. We see that HiP gives the best performance when ω′∈{1,1.25}\omega^{\prime}\in\{1,1.25\} but it’s performance degrades for higher values of ω′\omega^{\prime}.

Appendix F Analyzing Runtime of HiP

We provide average runtime of HiP for a single episode in all the three domains in Table 5 of the rebuttal document. We average across 10001000 seen tasks in each domain. We break the average runtime by different components: task planning (subgoal candidate generation and subgoal classification), visual planning, action planning and action execution. We execute the action plan for a subgoal in open-loop and then get observation from the environment for deciding the next subgoal. From Table 5, we see that majority of the planning time is taken by visual planning. Recent works have proposed techniques to reduce sampling time in diffusion models, which can be incorporated into our framework for improving visual planning speed in the future.