3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Tsung-Wei Ke, Nikolaos Gkanatsios, Katerina Fragkiadaki

I Introduction

Many robot manipulation tasks are inherently multimodal: at any point during task execution, there may be multiple actions which yield task-optimal behavior. Indeed, human demonstrations often contain diverse ways that a task can be accomplished. A natural choice is then to treat policy learning as a distribution learning problem: instead of representing a policy as a deterministic map πθ(x)\pi_{\theta}(x), learn the entire distribution of actions conditioned on the current robot state p(y∣x)p(y|x) . Recent works use diffusion models for learning such state-conditioned action distributions for robot manipulation policies from demonstrations and show they outperform deterministic or other alternatives, such as variational autoencoders , mixture of Gaussians , combination of classification and regression objectives , or energy-based objectives . Specifically, they exhibit better action distribution coverage and higher fidelity (less mode hallucination) than alternative formulations . They have so far been used with low-dimensional engineered state representations or 2D image encodings .

Lifting features from perspective views to a bird’s eye view (BEV) or 3D robot workspace map has shown strong results in robot learning from demonstrations . Policies that use such 2D-to-BEV or 2D-to-3D scene encodings generalize better than their 2D counterparts across camera viewpoints and can handle novel camera viewpoints at test time . Our conjecture is that this improved performance comes from the fact that the 3D visual scene content and the robot’s end-effector poses live in a common 3D space. BEV or 3D policy formulations e.g., Transporter Networks , C2F-ARM , PerAct , Act3D and Robot View Transformer , discretize the robot’s workspace for localizing the robot’s end-effector, such as the 2D BEV map for pick and place actions , the 3D workspace , or multiview re-projected images . Such 3D policies have not been combined yet with diffusion objectives.

In this paper, we propose 3D Diffuser Actor, a model that marries diffusion policies for handling action multimodality and 3D scene encodings for effective spatial reasoning. 3D Diffuser Actor trains a denoising neural network that takes as input a 3D scene visual encoding, the current estimate of the end-effector’s future trajectory (3D location and orientation), as well as the diffusion iteration index, and predicts the error in 3D translation and rotation. Our model achieves translation equivariance in prediction by representing the current estimate of the robot’s end-effector trajectory as 3D scene tokens and featurizing them jointly with the visual tokens using relative-position 3D attentions , as shown in Figure 2.

We test 3D Diffuser Actor in learning robot manipulation policies from demonstrations on the simulation benchmarks of RLBench and CALVIN , as well as in the real world. Our model sets a new state-of-the-art on RLBench, outperforming existing 3D policies and 2D diffusion policies with a 18.1% absolute gain on multi-view setups and 13.1% on single-view setups (Figure 1). On CALVIN, our model outperforms the current SOTA in the setting of zero-shot unseen scene generalization by a 7% relative gain (Figure 1). We additionally show that 3D Diffuser Actor outperforms all existing policy formulations that either do not use 3D scene representations or do not use action diffusion, as well as ablative versions of our model that do not use 3D relative attentions. We further show 3D Diffuser Actor can learn multi-task manipulation in the real world across 12 tasks from a handful of real world demonstrations. Simulation and real-world robot execution videos, as well as our code and trained checkpoints are publicly available on our website: 3d-diffuser-actor.github.io.

II Related Work

Learning robot manipulation policies from demonstrations Though state-to-action mappings are typically multimodal, earlier works on learning from demonstrations train deterministic policies with behavior cloning . To better handle action multimodality, other approaches discretize action dimensions and use cross entropy losses . However, the number of bins needed to approximate a continuous action space grows exponentially with increasing dimensionality. Generative adversarial networks , variational autoencoders and combined Categorical and Gaussian distributions have been used to learn from multimodal demonstrations. Nevertheless, these models tend to be sensitive to hyperparameters, such as the number of clusters used . Implicit behaviour cloning represents distributions over actions by using Energy-Based Models (EBMs) . Optimization in EBMs amounts to searching the energy landscape for minimal-energy action, given a state. EBMs are inherently multimodal, since the learned landscape can have more than one minima. EBMs are also composable, which makes them suitable for combining action distributions with additional constraints during inference . Diffusion models are a powerful class of generative models related to EBMs in that they model the score of the distribution, else, the gradient of the energy, as opposed to the energy itself . The key idea behind diffusion models is to iteratively transform a simple prior distribution into a target distribution by applying a sequential denoising process. They have been used for modeling state-conditioned action distributions in imitation learning from low-dimensional input, as well as from visual sensory input, and show both better mode coverage and higher fidelity in action prediction than alternatives. They have not been yet combined with 3D scene representations.

Diffusion models in robotics Beyond policy representations in imitation learning, diffusion models have been used to model cross-object and object-part arrangements and visual image subgoals . They have also been used successfully in offline reinforcement learning , where they model the state-conditioned action trajectory distribution or state-action trajectory distribution . ChainedDiffuser proposes to replace motion planners, commonly used for keypose to keypose linking, with a trajectory diffusion model that conditions on the 3D scene feature cloud and the predicted target 3D keypose to denoise a trajectory from the current to the target keypose. 3D Diffuser Actor instead predicts the next 3D keypose for the robot’s end-effector alongside the linking trajectory, which is a much harder task than linking two given keyposes. Lastly, image diffusion models have been used for augmenting the conditioning images input to robot policies to help the latter generalize better .

2D and 3D scene representations for robot manipulation End-to-end image-to-action policy models, such as RT-1 , RT-2, GATO , BC-Z , RT-X , Octo and InstructRL leverage transformer architectures for the direct prediction of 6-DoF end-effector poses from 2D video input. However, this approach comes at the cost of requiring thousands of demonstrations to implicitly model 3D geometry and adapt to variations in the training domains. Another line of research is centered around Transporter networks , demonstrating remarkable few-shot generalization by framing end-effector pose prediction as pixel classification. Nevertheless, these models are usually confined to top-down 2D planar environments with simple pick-and-place primitives. Direct extensions to 3D, exemplified by C2F-ARM and PerAct , involve voxelizing the robot’s workspace and learning to identify the 3D voxel containing the next end-effector keypose. However, this becomes computationally expensive as resolution requirements increase. Consequently, related approaches resort to either coarse-to-fine voxelization or efficient attention operations to mitigate computational costs. Act3D foregoes 3D scene voxelization altogether; it instead computes a 3D action map of variable spatial resolution by sampling 3D points in the empty workspace and featurizing them using cross-attentions to the 3D physical points. Robotic View Transformer (RVT) re-projects the input RGB-D image to alternative image views, featurizes those and lifts the predictions to 3D to infer 3D locations for the robot’s end-effector. Both Act3D and RVT show currently the highest performance on RLBench . Our model outperforms them by a big margin, as we show in the experimental section.

Instruction-conditioned long-horizon policies To generate actions based on language instructions, early works employ a text encoder to map language instructions to latent features, then they deploy 2D policies that condition on these language latents to predict actions. HULC++ alternates between a free-space policy that reaches subgoals (next object to interact with) and a local policy for low-level interaction. However, these models struggle to infer 3D actions precisely from 2D observations and language instructions and do not generalize to new environments with different texture. Recent works explore the use of large-scale pre-training to boost their performance. SuSIE proposes to synthesize visual subgoals using InstructPix2Pix , an instruction-conditioned image generative model pre-trained on a subset of 5 billion images . Actions are then predicted by a low-level diffusion policy conditioned on both current observations and synthesized visual goals. RoboFlamingo fine-tunes existing vision-language models, pre-trained for solving vision and language tasks, to predict the robot’s actions. GR-1 pre-trains a GPT-style transformer for video generation on a massive video corpus, then jointly trains the model for predicting robot actions and future observations. We show that 3D Diffuser Actor can exceed the performance of these policies thanks to effective use of 3D representations.

III Method

The architecture of 3D Diffuser Actor is shown in Figure 3. It is a conditional diffusion model that takes as input visual observations, a language instruction, a short history of the robot’s end-effectors and the current estimate for the robot’s future action trajectory, and predicts the error in the end-effector’s 3D translations and 3D orientations for each predicted timestep. We review Denoising Diffusion Probabilistic models in Section III-A and describe the architecture and training details of our model in Section III-B.

A diffusion model learns to model a probability distribution p(x)p(x) by inverting a process that gradually adds noise to a sample xx. For us, xx represents a sequence of 3D translations and 3D rotations for the robot’s end-effector. The diffusion process is associated with a variance schedule {βt∈(0,1)}t=1T\{\beta_{t}\in(0,1)\}_{t=1}^{T}, which defines how much noise is added at each time step. The noisy version of sample xx at time tt can then be written as xt=αˉtx+1−αˉtϵx_{t}=\sqrt{\bar{\alpha}_{t}}x+\sqrt{1-\bar{\alpha}_{t}}\epsilon where ϵ∼N(0,1),\epsilon\sim\mathcal{N}(\mathbf{0},\mathbf{1}), is a sample from a Gaussian distribution (with the same dimensionality as xx), αt=1−βt\alpha_{t}=1-\beta_{t}, and αˉt=∏i=1tαi\bar{\alpha}_{t}=\prod_{i=1}^{t}\alpha_{i}. The denoising process is modeled by a neural network ϵ^=ϵθ(xt;t)\hat{\epsilon}=\epsilon_{\theta}(x_{t};t) that takes as input the noisy sample xtx_{t} and the noise level tt and tries to predict the noise component ϵ\epsilon.

Diffusion models can be easily extended to draw samples from a distribution p(x∣c)p(x|\textbf{c}) conditioned on input c, which is added as input to the network ϵθ\epsilon_{\theta}. For us c is the visual scene captured by one or more calibrated RGB-D images, a language instruction, as well as a short history of the robot’s end-effector’s poses. Given a collection of D={(xi,ci)}i=1N\mathcal{D}=\{(x^{i},\textbf{c}^{i})\}_{i=1}^{N} of end-effector trajectories xix^{i} paired with observation and robot history context ci\textbf{c}^{i}, the denoising objective becomes:

This loss corresponds to a reweighted form of the variational lower bound for log⁡p(x∣c)\log p(x|\textbf{c}) .

In order to draw a sample from the learned distribution pθ(x∣c)p_{\theta}(x|\textbf{c}), we start by drawing a sample xT∼N(0,1)x_{T}\sim\mathcal{N}(\mathbf{0},\mathbf{1}). Then, we progressively denoise the sample by iterated application of ϵθ\epsilon_{\theta} TT times according to a specified sampling schedule , which terminates with x0x_{0} sampled from pθ(x)p_{\theta}(x):

where z∼N(0,1)\mathbf{z}\sim\mathcal{N}(\mathbf{0},\mathbf{1}).

III-B 3D Diffuser Actor

We next describe the architecture of our model that predicts the end-effector’s trajectory error given the current noisy estimate, a visual scene encoding, a language instruction, the denoising iteration index and the end-effector’s pose history.

Keyposes 3D Diffuser Actor inherits the temporal abstraction of demonstrations into end-effector keyposes from previous works . Keyposes are important intermediate end-effector poses that summarize a demonstration and can be extracted using simple heuristics, such as a change in the open/close end-effector state or local extrema of velocity/acceleration. We describe the heuristics we use for each dataset in the Appendix (Section A-A). During inference, 3D Diffuser Actor can either predict and execute the full trajectory of actions up to the next keypose (including the keypose), or just predict the next keypose and use a sampling-based motion planner to reach it, similar to previous works . In the rest of this section we use the term “trajectory" to refer to 3D Diffuser Actor’s output. Keypose prediction is a special case where the trajectory contains only one future pose.

Scene and language encoder We use a scene and language encoder similar to . We describe it here to make the paper self-contained. The input to our scene encoder is a set of posed RGB-D images. We first extract multi-scale visual tokens for each camera view using a pre-trained 2D feature extractor and a feature pyramid network . Next, we associate every 2D feature grid location in the 2D feature maps with a depth value, by averaging the depth values of the image pixels that correspond to it. We use camera intrinsics and the pinhole camera equation to map a pixel location and depth value (x,y,d)(x,y,d) to a 3D location (X,Y,Z)(X,Y,Z), and “lift” the 2D feature tokens to 3D, to obtain a 3D feature cloud. The language encoder maps language task descriptions or instructions into language feature tokens. We use the pre-trained CLIP ResNet50 2D image encoder to encode each RGB image into a 2D feature map and the pre-trained CLIP language encoder to featurize the language task instruction.

We contextualize all tokens, namely visual oo, language ll, proprioception cc and action tokens (current estimate) atpos,atrot\mathbf{a}_{t}^{pos},\mathbf{a}_{t}^{rot} using 3D relative position attention layers. Inspired by recent work in visual correspondence and 3D manipulation , we use rotary positional embeddings with the property that the dot product of two positionally encoded features xi,xj\mathbf{x}_{i},\mathbf{x}_{j} is:

We use a modified version of Equation 2 to update the current estimate of each element of the end-effector’s trajectory:

where zpos,zrot∼N(0,1)\mathbf{z}^{\small pos},\mathbf{z}^{\small rot}\sim\mathcal{N}(\mathbf{0},\mathbf{1}) variables of appropriate dimension.

We use the following two noise schedulers:

a scaled-linear noise scheduler βt=(βmax⁡−βmin⁡)t+βmin⁡\beta_{t}=(\beta_{\max}-\beta_{\min})t+\beta_{\min}, where βmax⁡,βmin⁡\beta_{\max},\beta_{\min} are hyperparameters, set to 0.020.02 and 0.00010.0001 in our experiments,

a squared cosine noise scheduler \beta_{t}=\frac{1-\cos\big{(}\frac{(t+1)/T+0.008}{1.008}*\frac{\pi}{2}\big{)}^{2}}{\cos\big{(}\frac{t/T+0.008}{1.008}*\frac{\pi}{2}\big{)}^{2}}.

We found using a scale-linear noise schedule for denoising end-effector’s 3D positions and a squared cosine noise schedule for denoising the end-effector’s 3D orientations to converge much faster than using squared cosine noise for both.

where w1,w2w_{1},w_{2} are hyperparameters estimated using cross-validation.

Implementation details We lift 2D image features to 3D by calculating xyz-coordinates of each image pixel, using the sensed depth and camera parameters. We augment RGB-D observations with random rescaling and cropping. Nearest neighbor interpolation is used for rescaling RGB-D observations. To reduce the memory footprint in our 3D Relative Transformer, we use Farthest Point Sampling to sample a subset of the points in the input 3D feature cloud. We use FiLM to inject conditional input, including the diffusion step and proprioception history, to every attention layer in the model. We include a detailed architecture diagram of our model and a table of hyper-parameters used in our experiments in the Appendix (Section A-E and Section A-F).

IV Experiments

We test 3D Diffuser Actor in learning multi-task manipulation policies on RLBench and CALVIN , two established learning from demonstrations benchmarks. We compare against the current state-of-the-art on single-view and multi-view setups, as well as against ablative versions of our model that do not consider 3D scene representations or relative attention. We also test 3D Diffuser Actor’s ability to learn multi-task policies from real-world demonstrations.

Setup RLBench is built atop the CoppelaSim simulator, where a Franka Panda Robot is used to manipulate the scene. Our 3D Diffuser Actor is trained to predict the next end-effector keypose and we employ the low-level motion planner BiRRT , native to RLBench, to reach the predicted pose, following previous works . Our model does not perform collision checking for our experiments. We train and evaluate 3D Diffuser Actor on two experimental setups of multi-task manipulation:

PerAct setup introduced in : This uses a suite of 18 manipulation tasks, each task has 2-60 variations, which concern scene variability across object poses, appearance and semantics. The tasks are specified by language descriptions. There are four cameras available (front, wrist, left shoulder, right shoulder). There are 100 training demonstrations available per task, evenly split across task variations and 25 unseen test episodes for each task. Due to the randomness of the sampling-based motion planner, we test our model across 5 random seeds. During evaluation, models are allowed to predict a maximum of 25 keyposes, unless they receive earlier task-completion indicators from the simulation environment.

GNFactor setup introduced in : This uses a suite of 10 manipulation tasks (a subset of PerAct’s task set). Only one RGB-D camera view is available (front camera). There are 20 training demonstrations per task, evenly split across variations. The evaluation setup mirrors that of PerAct, with the exception that we test the final checkpoint using 3, rather than 5, random seeds on the test set.

Baselines We consider the following baselines:

C2F-ARM-BC , a 3D policy that iteratively voxelizes RGB-D images and predicts actions in a coarse-to-fine manner. Q-values are estimated within each voxel and the translation action is determined by the centroid of the voxel with the maximal Q-values.

PerAct , a 3D policy that voxelizes the workspace and detects the next voxel action through global self-attention.

Hiveformer , a 3D policy that enables attention between features of different history time steps.

PolarNet , a 3D policy that computes dense point representations for the robot workspace using a PointNext backbone .

RVT , a 3D policy that deploys a multi-view transformer to predict actions and fuses those across views by back-projecting to 3D.

Act3D , a 3D policy that featurizes the robot’s 3D workspace using coarse-to-fine sampling and featurization.

GNFactor , a 3D policy that co-optimizes a neural field for reconstructing the 3D voxels of the input scene and a PerAct module for predicting actions based on voxel representations.

We report results for RVT, PolarNet and GNFactor based on their respective papers. Results for CF2-ARM-BC and PerAct are presented as reported in RVT . Results for Hiveformer are copied from PolarNet .

We observed that Act3D does not follow the same setup as PerAct on RLBench. Specifically, Act3D uses different 1) 3D object models, 2) success conditions, 3) training/test episodes and 4) maximum numbers of keyposes during evaluation. For fair comparison, we retrain and test Act3D on the same setup.

We also compare to the following ablative versions of our model:

2D Diffuser Actor, our implementation of . We remove the 3D scene encoding from 3D Diffuser Actor and instead use per-image 2D representations by average-pooling features within each view. We add learnable embeddings to distinguish between different views. We use standard attention layers for joint encoding the action estimate and 2D image features.

3D Diffuser Actor w/o Rel. Attn., an ablative version of our model that uses standard (non-relative) attentions to featurize the current rotation and translation estimate with the 3D scene feature cloud. This version of our model is not translation-equivariant.

Evaluation metrics Following previous work , we evaluate policies by task completion success rate, which is the proportion of execution trajectories that achieve the goal conditions specified in the language instructions.

Results on the PerAct setup We show quantitative results in Table I. Our 3D Diffuser Actor outperforms all baselines on most tasks by a large margin. It achieves an average 81.3%81.3\% success rate among all 18 tasks, an absolute improvement of +18.1%+18.1\% over Act3D, the previous state-of-the-art. In particular, 3D Diffuser Actor achieves big leaps on tasks with multiple modes, such as stack blocks, stack cups and place cups, which most baselines fail to complete. We obtain substantial improvements of +39.5%+39.5\%, +18.4%+18.4\%, +41.6%+41.6\% and +20.8%+20.8\% on stack blocks, put in cupboard, insert peg and stack cups respectively.

Results on the GNFactor setup We show quantitative results in Table II. We train Act3D on this setup using its publicly available code. 3D Diffuser Actor outperforms both GNFactor and Act3D by a significant margin, achieving absolute performance gains of +46.4%+46.4\% and +13.1%+13.1\% respectively. Notably, 3D Diffuser Actor and Act3D utilize similar 3D scene representations—sparse 3D feature tokens—while GNFactor featurizes a scene with 3D voxels and learns to reconstruct them. Even with a single camera view, both Act3D and 3D Diffuser Actor outperform GNFactor significantly. This suggests that the choice of 3D scene representation is a more crucial factor than the completion of 3D scenes in developing efficient 3D manipulation policies.

Ablations We ablate the use of 3D scene representations and the use of relative attentions in the PerAct experimental setup in Table III. We draw the following conclusions: 1. 3D Diffuser Actor largely outperforms its 2D counterpart, 2D Diffuser Actor, underlining the importance of 3D scene representations. 2. 3D Diffuser Actor with absolute 3D attentions (3D Diffuser Actor w/o Rel. Attn.) is much worse than 3D Diffuser Actor with relative 3D attentions. This shows that translation equivariance through relative attentions is very important for generalization. Notably, this baseline already outperforms all prior arts in Table I, proving the effectiveness of marrying 3D representations and diffusion policies.

IV-B Evaluation on CALVIN

The CALVIN benchmark is build on top of the PyBullet simulator and involves a Franka Panda Robot arm that manipulates the scene. CALVIN consists of 34 tasks and 4 different environments (A, B, C and D). All environments are equipped with a desk, a sliding door, a drawer, a button that turns on/off an LED, a switch that controls a lightbulb and three different colored blocks (red, blue and pink). These environments differ from each other in the texture of the desk and positions of the objects. CALVIN provides 24 hours of tele-operated unstructured play data, 35% of which are annotated with language descriptions.

We train 3D Diffuser Actor on the subset of play data annotated with language descriptions on the environments A, B and C and evaluate it on 1000 unique instruction chains on environment D, following prior works . Each instruction chain includes five language instructions that need to be executed sequentially. We devise an algorithm (Appendix, Section A-A) to extract keyposes on CALVIN, since prior works do not use keyposes on this benchmark. We train our 3D Diffuser Actor to predict both the end-effector pose for each keypose and the corresponding trajectory to reach the predicted pose, instead of using a motion planner. Prior works predict a maximum of 360 actions to complete each instructional task, while, on average, it takes only 60 actions to complete each task using ground-truth trajectories. Our model predicts both keyposes and corresponding trajectories, and, on average, it takes 10 keyposes of each demonstration to complete an instructional task. We thus allow our model to predict a maximum of 60 keyposes for each task. For reference, we also allow our model to predict 360 times and show the influence of this hyperparameter in the Appendix (Section A-B).

Baselines. We consider the following baselines:

MCIL , a multi-modal goal-conditioned 2D policy that maps three types of goals–goal images, language instructions and task labels–to a shared latent feature space, and conditions on such latent goals to predict actions.

HULC , a 2D policy that uses a variational autoencoder to sample a latent plan based on the current observation and task description, then conditions on this latent to predict actions.

RT-1 , a 2D transformer-based policy that encodes the image and language into a sequence of tokens and employs a Transformer-based architecture that contextualizes these tokens and predicts the arm movement or terminates the episode.

RoboFlamingo , a 2D policy that adapts existing vision-language models, which are pre-trained for solving vision and language tasks, to robot control. It uses frozen vision and language foundational models and learns a cross-attention between language and visual features, as well as a recurrent policy that predicts the low-level actions conditioned on the language latents.

SuSIE , a 2D policy that deploys an large-scale pre-trained image generative model to synthesize visual subgoals based on the current observation and language instruction. Actions are then predicted by a low-level goal-conditioned 2D diffusion policy that models inverse dynamics between the current observation and the predicted subgoal image.

GR-1 , a 2D policy that first pre-trains an autoregressive Transformer on next frame prediction, using a large-scale video corpus without action annotations. Each video frame is encoded into an 1d vector by average-pooling its visual features. Then, the same architecture is fine-tuned in-domain to predict both actions and future observations.

We report results for HULC, RoboFlamingo, SuSIE and GR-1 from the respective papers. Results from MCIL are borrowed from . Results from RT-1 are copied from .

Evaluation metrics Following previous works , we report the success rate and the average number of completed sequential tasks.

Results We present our results in Table IV. 3D Diffuser Actor achieves competitive results compared to the state-of-the-art models. In comparison to GR-1, our model achieves absolute performance gains of +6.8%+6.8\%, +7.5%+7.5\%, +4.3%,+1.5%+4.3\%,+1.5\% and +1.1\forcompletingthefivetasksintheinstructionchain.Onaverage,3DDiffuserActorcompletefor completing the five tasks in the instruction chain. On average, 3D Diffuser Actor complete3.27$ tasks in a row, setting a new state of the art. Note that GR-1’s contribution is large-scale pre-training on videos, which is a scheme orthogonal to our contribution. Lastly, allowing our model to predict more times increases the performance significantly. For this performance, we found stronger conditioning on the language instructions to be important, as we explain in our appendix (Figure 6b).

IV-C Evaluation in the real world

We validate 3D Diffuser Actor in learning manipulation tasks from real-world demonstrations. We use a Franka Emika robot and capture visual observations with a Azure Kinect RGB-D sensor at a front view. Images are originally captured at 1280×7201280\times 720 resolution and downsampled to a resolution of 256×256256\times 256. Camera extrinsics are calibrated w.r.t the robot base. We use 12 tasks: 1) close a box, 2) put ducks in bowls, 3) insert a peg vertically into the hole, 4) insert a peg horizontally into the torus, 5) put a computer mouse on the pad, 6) open the pen, 7) press the stapler, 8) put grapes in the bowl, 9) sort the rectangle, 10) stack blocks with the same shape, 11) stack cups and 12) put block in a triangle on the plate. Details on the task definitions and success criteria are included in the Appendix (Section A-C).

We collect 15 demonstrations per task where we record the keyposes, most of which naturally contain noise and multiple modes of human behavior. For example, we pick one of two ducks to put in the bowl, we insert the peg into one of two holes and we put one of three grapes in the bowl, as shown in Figure 4. 3D Diffuser Actor conditions on language descriptions and is trained to predict the next end-effector keypose. During inference, we utilize the BiRRT planner provided by the MoveIt! ROS package to reach the predicted poses. We evaluate 10 episodes for each task are report the success rate.

We show quantitative results in Table V and video results on our project webpage.

IV-D Run time

We measure the latency of our 3D Diffuser Actor on CALVIN in simulation, using an NVIDIA GeForce 2080 Ti graphic card. The wall-clock time of 3D Diffuser Actor is 600 ms. Notably, our model predicts end-effector keyposes sparsely, resulting in better efficiency than methods that predict actions densely at each time step. On CALVIN, on average, it takes 10 keyposes / 60 actions to complete each task using ground-truth trajectories. Our 3D Diffuser Actor can thus perform manipulation efficiently.

IV-E Limitations

Our framework currently has the following limitations: 1. Our model conditions on 3D scene representations, which require camera calibration and depth information. 2. All tasks in RLBench and CALVIN are quasi-static. Extending our method to dynamic tasks and velocity control is a direct avenue of future work.

V Conclusion

We present 3D Diffuser Actor, a 3D robot manipulation policy with action diffusion. Our method sets a new state-of-the-art on RLBench and CALVIN by a large margin over existing 3D policies and 2D diffusion policies. We introduce important architectural innovations, such as 3D token representations of the robot’s pose estimate and 3D relative position denoising transformers that endow 3D Diffuser Actor with translation equivariance, and empirically verify their contribution to performance. We further test our model in the real world and show it can learn manipulation tasks from a handful of demonstrations. Our future work will attempt to train 3D Diffuser Actor policies in domain-randomized simulation environments at a large scale, to help them transfer from simulation to the real world.

VI Acknowledgements

This work is supported by Sony AI, NSF award No 1849287, DARPA Machine Common Sense, an Amazon faculty award, and an NSF CAREER award. The authors would like to thank Moritz Reuss for useful training tips on CALVIN; Zhou Xian for help with the real-robot experiments; Brian Yang for discussions, comments and efforts in the early development of this paper.

References

Appendix A Appendix

We explain our algorithms for keypose discovery on all benchmarks in Section A-A. We show the effect of maximum number of predicted keyposes on CALVIN in Section A-B. We detail the real-world and RLBench tasks and success conditions in Section A-C and Section A-D. We present a detailed version of our architecture in Section A-E. We list the hyper-parameters for each experiment in Section A-F. Lastly, we visualize the importance of a square cos variance scheduler for denoising the rotation estimate in Section A-G.

For RLBench we use the heuristics from : a pose is a keypose if (1) the end-effector state changes (grasp or release) or (2) the velocity’s magnitude approaches zero (often at pre-grasp poses or a new phase of a task). For our real-world experiments we maintain the above heuristics and record pre-grasp poses as well as the poses at the beginning of each phase of a task, e.g., when the end-effector is right above an object of interest. We report the number of keyposes per real-world task in this Appendix (Section A-C). Lastly, for CALVIN we adapt the above heuristics to devise a more robust algorithm to discover keyposes. Specifically, we track end-effector state changes and significant changes of motion, i.e. both velocity and acceleration. For reference, we include our Python code for discovering keyposes in CALVIN here:

A-B Evaluation on CALVIN with varying number of maximum keyposes allowed at test time

We show zero-shot long-horizon performance on CALVIN with varying number of keyposes allowed at test time in Figure 5. The performance saturates with more than 300 keyposes.

A-C Real-world tasks

We explain the the real-world tasks and their success conditions in more detail. All tasks take place in a cluttered scene with distractors (random objects that do not participate in the task) which are not mentioned in the descriptions below.

Close a box: The end-effector needs to move and hit the lid of an open box so that it closes. The agent is successful if the box closes. The task involves two keyposes.

Put a duck in a bowl: There are two toy ducks and two bowls. One of the ducks have to be placed in one of the bowls. The task involves four keyposes.

Insert a peg vertically into the hole: The agent needs to detect and grasp a peg, then insert it into a hole that is placed on the ground. The task involves four keyposes.

Insert a peg horizontally into the torus: The agent needs to detect and grasp a peg, then insert it into a torus that is placed vertically to the ground. The task involves four keyposes.

Put a computer mouse on the pad: There two computer mice and one mousepad. The agent needs to pick one mouse and place it on the pad. The task involves four keyposes.

Open the pen: The agent needs to detect a pen that is attached vertically to the table, grasp its lid and pull it to open the pen. The task involves three keyposes.

Press the stapler: The agent needs to reach and press a stapler. The task involves two keyposes.

Put grapes in the bowl: The scene contains three vines of grapes of different color and one bowl. The agent needs to pick one vine and place it in the bowl. The task involves four keyposes.

Sort the rectangle: Between two rectangle cubes there is space for one more. The task comprises detecting the rectangle to be moved and placing it between the others. It involves four keyposes.

Stack blocks with the same shape: The scene contains of several blacks, some of which have rectangular and some cylindrical shape. The task is to pick the same-shape blocks and stack them on top of the first one. It involves eight keyposes.

Stack cups: The scene contains three cups of different colors. The agent needs to successfully stack them in any order. The task involves eight keyposes.

Put block in a triangle on the plate: The agent needs to detect three blocks of the same color and place them inside a plate to form an equilateral triangle. The task involves 12 keyposes.

The above tasks examine different generalization capabilities of 3D Diffuser Actor, for example multimodality in the solution space (5, 8), order of execution (10, 11, 12), precision (3, 4, 6) and high noise/variance in keyposes (1).

A-D RLBench tasks under PerAct’s setup

We provide an explanation of the RLBench tasks and their success conditions under the PerAct setup for self-completeness. All tasks vary the object pose, appearance and semantics, which are not described in the descriptions below. For more details, please refer to the PerAct paper .

Open a drawer: The cabinet has three drawers (top, middle and bottom). The agent is successful if the target drawer is opened. The task on average involves three keyposes.

Slide a block to a colored zone: There is one block and four zones with different colors (red, blue, pink, and yellow). The end-effector must push the block to the zone with the specified color. On average, the task involves approximately 4.7 keyposes

Sweep the dust into a dustpan: There are two dustpans of different sizes (short and tall). The agent needs to sweep the dirt into the specified dustpan. The task on average involves 4.6 keyposes.

Take the meat off the grill frame: There is chicken leg or steck. The agent needs to take the meat off the grill frame and put it on the side. The task involves 5 keyposes.

Turn on the water tap: The water tap has two sides of handle. The agent needs to rotate the specified handle 90∘90^{\circ}. The task involves 2 keyposes.

Put a block in the drawer: The cabinet has three drawers (top, middle and bottom). There is a block on the cabinet. The agent needs to open and put the block in the target drawer. The task on average involves 12 keyposes.

Close a jar: There are two colored jars. The jar colors are sampled from a set of 20 colors. The agent needs to pick up the lid and screw it in the jar with the specified color. The task involves six keyposes.

Drag a block with the stick: There is a block, a stick and four colored zones. The zone colors are sampled from a set of 20 colors. The agent is successful if the block is dragged to the specified colored zone with the stick. The task involves six keyposes.

Stack blocks: There are 8 colored blocks and 1 green platform. Each four of the 8 blocks share the same color, while differ from the other. The block colors are sampled from a set of 20 colors. The agent needs to stack N blocks of the specified color on the platform. The task involves 14.6 keyposes.

Screw a light bulb: There are 2 light bulbs, 2 holders, and 1 lamp stand. The holder colors are sampled from a set of 20 colors. The agent needs to pick up and screw the light bulb in the specified holder. The task involves 7 keyposes.

Put the cash in a safe: There is a stack of cash and a safe. The safe has three layers (top, middle and bottom). The agent needs to pick up the cast and put it in the specified layer of the safe. The task involves 5 keyposes.

Place a wine bottle on the rack: There is a bottle of wine and a wooden rack. The rack has three slots (left, middle and right). The agent needs to pick up and place the wine at the specified location of the wooden rack. The task involves 5 keyposes.

Put groceries in the cardboard: There are 9 YCB objects and a cupboard. The agent needs to grab the specified object and place it in the cupboard. The task involves 5 keyposes.

Put a block in the shape sorter: There are 5 blocks of different shapes and a sorter with the corresponding slots. The agent needs to pick up the block with the specified shape and insert it into the slot with the same shape. The task involves 5 keyposes.

Push a button: There are 3 buttons, whose colors are sampled from a set of 20 colors. The agent needs to push the colored buttons in the specified sequence. The task involves 3.8 keyposes.

Insert a peg: There is 1 square, and 1 spoke platform with three colored spoke. The spoke colors are sampled from a set of 20 colors. The agent needs to pick up the square and put it onto the spoke with the specified color. The task involves 5 keyposes.

Stack cups: There are 3 cups. The cup colors are sampled from a set of 20 colors. The agent needs to stack all the other cups on the specified one. The task involves 10 keyposes.

Hang cups on the rack: There are 3 mugs and a mug rack. The agent needs to pick up N mugs and place them onto the rack. The task involves 11.5 keyposes.

A-E Detailed Model Diagram

We present a more detailed architecture diagram of our 3D Diffuser Actor in Figure 6a. We also show a variant of 3D Diffuser Actor with enhanced language conditioning in Figure 6b, which achieves SOTA results on CALVIN.

The inputs to our network are i) a stream of RGB-D views; ii) a language instruction; iii) proprioception in the form of end-effector’s history poses; iv) the current noisy estimates of position and rotation; v) the denoising step tt. The images are encoded into visual tokens using a pretrained 2D backbone. The depth values are used to “lift" the multi-view tokens into a 3D feature cloud. The language is encoded into feature tokens using a language backbone. The proprioception is represented as learnable tokens with known 3D locations in the scene. The noisy estimates are fed to linear layers that map them to high-dimensional vectors. The denoising step is fed to an MLP.

The visual tokens cross-attend to the language tokens and get residually updated. The proprioception tokens attend to the visual tokens to contextualize with the scene information. We subsample a number of visual tokens using Farthest Point Sampling (FPS) in order to decrease the computational requirements. The sampled visual tokens, proprioception tokens and noisy position/rotation tokens attend to each other. We modulate the attention using adaptive layer normalization and FiLM . Lastly, the contextualized noisy estimates are fed to MLP to predict the error terms as well as the end-effector’s state (open/close).

A-F Hyper-parameters for experiments

A-G The importance of noise scheduler

We visualize the clean/noised 6D rotation representations as two three-dimensional unit-length vectors in Figure 7. We plot each vector as a point in the 3D space. We can observe that noised rotation vectors generated by the squared linear scheduler cover the space more completely than those by the scaled linear scheduler.