GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields
Yanjie Ze, Ge Yan, Yueh-Hua Wu, Annabella Macaluso, Yuying Ge, Jianglong Ye, Nicklas Hansen, Li Erran Li, Xiaolong Wang
Introduction
One major goal of introducing learning into robotic manipulation is to enable the robot to effectively handle unseen objects and successfully tackle various tasks in new environments. In this paper, we focus on using imitation learning with a few demonstrations for multi-task manipulation. Using imitation learning helps avoid complex reward design and training can be directly conducted on the real robot without creating its digital twin in simulation . This enables policy learning on diverse tasks in complex environments, based on users’ instructions (see Figure 1). However, working with a limited number of demonstrations presents great challenges in terms of generalization. Most of these challenges arise from the need to comprehend the 3D structure of the scene, understand the semantics and functionality of objects, and effectively follow task instructions based on visual cues. Therefore, a comprehensive and informative visual representation of the robot’s observations serves as a crucial foundation for generalization.
The development of visual representation for robot learning has mainly focused on learning within a 2D plane. Self-supervised objectives are leveraged to pre-train the representation from the 2D image observation or jointly optimized with the policy gradients . While these approaches improve sample efficiency and lead to more robust policies, they are mostly applied to relatively simple manipulation tasks. To tackle more complex tasks requiring geometric understanding (e.g., object shape and pose) and with occlusions, 3D visual representation learning has been recently adopted with robot learning . For example, Driess et al. 2022 train the 3D scene representation by using NeRF and view synthesis to provide supervision. While it shows effectiveness over tasks requiring geometric reasoning such as hanging a cup, it only handles the simple scene structure with heavy masking in a single-task setting. More importantly, without a semantic understanding of the scene, it would be very challenging for the robot to follow the user’s language instructions.
In this paper, we introduce learning a language-conditioned policy using a novel representation leveraging both 3D and semantic information for multi-task manipulation. We train Generalizable Neural Feature Fields (GNF) which distills pre-trained semantic features from 2D foundation models into the Neural Radiance Fields (NeRFs). We conduct policy learning upon this representation, leading to our model GNFactor. It is important to note that GNFactor learns an encoder to extract scene features in a feed-forward manner, instead of performing per-scene optimization in NeRF. Given a single RGB-D image observation, our model encodes it into a 3D semantic volumetric feature, which is then processed by a Perceiver Transformer architecture for action prediction. To conduct multi-task learning, the Perceiver Transformer takes in language instructions to get task embedding, and reason the relations between the language and visual semantics for manipulation.
There are two branches of training in our framework (see Figure 3): (i) GNF training. Given the collected demonstrations, we train the Generalizable Neural Feature Fields using view synthesis with volumetric rendering. Besides rendering the RGB pixels, we also render the features of the foundation models in 2D space. The GNF learns from both pixel and feature reconstruction at the same time. To provide supervision for feature reconstruction, we apply a vision foundation model (e.g., pre-trained Stable Diffusion model ) to extract the 2D feature from the input view as the ground truth. In this way, we can distill the semantic features into the 3D space in GNF. (ii) GNFactor joint training. Building on the 3D volumetric feature jointly optimized by the learning objectives of GNF, we conduct behavior cloning to train the whole model end-to-end.
For evaluation, we conduct real-robot experiments on three distinct tasks across two different kitchens (see Figure 1). We successfully train a single policy that effectively addresses these tasks in different scenes, yielding significant improvements over the baseline method PerAct . We also conduct comprehensive evaluations using 10 RLBench simulated tasks and 6 designed generalization tasks. We observe that GNFactor outperforms PerAct with an average improvement of x and x, consistent with the significant margin observed in the real-robot experiments.
Related Work
Multi-Task Robotic Manipulation. Recent works in multi-task robotic manipulation have led to significant progress in the execution of complex tasks and the ability to generalize to new scenarios . Notable methods often involve the use of extensive interaction data to train multi-task models . For example, RT-1 underscores the benefits of task-agnostic training, demonstrating superior performance in real-world robotic tasks across a variety of datasets. To reduce the need for extensive demonstrations, methods that utilize keyframes – which encode the initiation of movement – have proven to be effective . PerAct employs the Perceiver Transformer to encode language goals and voxel observations and shows its effectiveness in real robot experiments. In this work, we utilize the same action prediction framework as PerAct while we focus on improving the generalization ability of this framework by learning a generalizable volumetric representation under limited data.
3D Representations for Reinforcement/Imitation Learning (RL/IL). To improve manipulation policies by leveraging visual information, numerous studies have concentrated on enhancing 2D visual representations , while for addressing more complex tasks, the utilization of 3D representations becomes crucial. Ze et al. 2023 incorporates a deep voxel-based 3D autoencoder in motor control, demonstrating improved sample efficiency compared to 2D representation learning methods. Driess et al. 2022 proposes to first learn a state representation by NeRF and then use the frozen state for downstream RL tasks. While this work shows the initial success of utilizing NeRF in RL, its applicability in real-world scenarios is constrained due to various limitations: e.g., the requirement of object masks, the absence of a robot arm, and the lack of scene structure. The work closest to ours is SNeRL , which also utilizes a vision foundation model in NeRF. However, similar to NeRF-RL , SNeRL masks the scene structure to ensure functionality and the requirement for object masks persists, posing challenges for its application in real robot scenarios. Our proposed GNFactor, instead, handles challenging muti-task real-world scenarios, demonstrating the potential for real robot applications.
Neural Radiance Fields (NeRFs). Neural fields have achieved great success in novel view synthesis and scene representation learning these years , and recent works also start to incorporate neural fields into robotics . NeRF stands out for achieving photorealistic view synthesis by learning an implicit function of the scene, while it requires per-scene optimization and is thus hard to generalize. Many following methods propose more generalizable NeRFs. PixelNeRF and CodeNeRF encode 2D images as the input of NeRFs, while TransINR leverages a vision transformer to directly infer NeRF parameters. A line of recent works utilize pre-trained vision foundation models such as DINO and CLIP as supervision besides the RGB image, which thus enables the NeRF to learn generalizable features. In this work, we incorporate generalizable NeRF to reconstruct different views in RGB and embeddings from a pretrained Stable Diffusion model .
Method
In this section, we detail the proposed GNFactor, a multi-task agent with a 3D volumetric representation for real-world robotic manipulation. GNFactor is composed of a volumetric rendering module and a 3D policy module, sharing the same deep volumetric representation. The volumetric rendering module learns a Generalizable Neural Feature Field (GNF), to reconstruct the RGB image from cameras and the embedding from a vision-language foundation model, e.g., Stable Diffusion . The task-agnostic nature of the vision-language embedding enables the volumetric representation to learn generalizable features via neural rendering and thus helps the 3D policy module better handle multi-task robotic manipulation. The task description is encoded with CLIP to obtain the task embedding . An overview of GNFactor is shown in Figure 3.
Due to the inefficiency of continuous action prediction and the extensive data requirements that come with it, we reformulate the behavior cloning problem as a keyframe-prediction problem . We first extract keyframes from expert demonstrations using the following metric: a frame in the trajectory is a keyframe when joint velocities approach zero and the gripper’s open state remains constant. The model is then trained to predict the subsequent keyframe based on current observations. This formulation effectively transforms the continuous control problem into a discretized keyframe-prediction problem, delegating the internal procedures to the RRT-connect motion planner in simulation and Linear motion planner in real-world xArm7 robot.
2 Learning Volumetric Representations with Generalizable Neural Feature Fields
where is the ground truth color, is the ground truth vision-language embedding generated by Stable Diffusion, is the set of rays generated from camera poses, and is the weight for the embedding reconstruction loss. For efficiency, we sample rays given one target view, instead of reconstructing the entire image. To help the GNF training, we use a coarse-to-fine hierarchical structure as the original NeRF and apply depth-guided sampling in the “fine” network.
3 Action Prediction with Volumetric Representations
The volumetric representation is optimized not only to achieve reconstruction of the GNF module, but also to predict the desired action for accomplishing manipulation tasks within the 3D policy. As such, we jointly train the representation to satisfy the objectives of both the GNF and the 3D policy module. In this section, we elaborate the training objective and the architecture of the 3D policy.
We employ a Perceiver Transformer to handle the high-dimensional multi-modal input, i.e., the 3D volume, the robot’s proprioception, and the language feature. We first condense the shared volumetric representation into a volume of size using a 3D convolution layer with a kernel size and stride of 5, followed by a ReLU function, and flatten the 3D volume into a sequence of small cubes of size . The robot’s proprioception is projected into a 128-dimensional space and concatenated with the volume sequence for each cube, resulting in a sequence of size . We then project the language token features from CLIP into the same dimensions () and concatenate these features with a combination of the 3D volume, the robot’s proprioception state, and the CLIP token embedding. The result is a sequence with dimensions of .
This sequence is combined with a learnable positional encoding and passed through the Perceiver Transformer, which outputs a sequence of the same size. We remove the last features for the ease of voxelization and reshape the sequence back to a voxel of size . This voxel is then upscaled to with trilinear interpolation and referred to as . is shared across three action prediction heads (, , , in Figure 3) to determine the final robot actions at the same scale as the observation space. To retain the learned features from GNF training, we create a skip connection between our volumetric representation and . The combined volume feature is used to predict a 3D Q-function for translation, as well as Q-functions for other robot operations like gripper openness (), rotation (), and collision avoidance (). The -function here represents the action values of one timestep, differing from the traditional -function in RL that is for multiple timesteps. For example, in each timestep, the 3D -value would be equal to for the most possible next voxel and for other voxels. The model then optimizes the cross-entropy loss like a classifier,
where for and is the ground truth one-hot encoding. The overall learning objective for GNFactor is as follows:
where is the weight for the reconstruction loss to balance the scale of different objectives. To train the GNFactor, we employ a joint training approach in which the GNF and 3D policy module are optimized jointly, without any pre-training. From our empirical observation, this approach allows for better fusion of information from the two modules when learning the shared features.
Experiments
In this section, we conduct experiments to answer the following questions: (i) Can GNFactor surpass the baseline model in simulated environments? (ii) Can GNFactor generalize to novel scenes in simulation? (iii) Does GNFactor learn a superior policy that handles real robot tasks in two different kitchens with noisy and limited real-world data? (iv) What are the crucial factors in GNFactor to ensure the functionality of the entire system? Our concluded results are given in Figure 4.
For the sake of reproducibility and benchmarking, we conduct our primary experiments in RLBench simulated tasks. Furthermore, to show the potential of GNFactor in the real world, we design a set of real robot experiments across two kitchens. We compare our GNFactor with the strong language-conditioned multi-task agent PerAct in both simulation and the real world, emphasizing the universal functionality of GNFactor. Both GNFactor and PerAct use the single RGB-D image from the front camera as input to construct the voxel grid. In the multi-task simulation experiments, we also create a stronger version of PerAct by adding more camera views as input to fully cover the scene (visualized in Figure 10). Figure 2 shows our simulation tasks and the real robot setup. We briefly describe the tasks and details are left in Appendix B and Appendix C.
Simulation. We select challenging language-conditioned manipulation tasks from the RLBench tasksuites . Each task has at least two variations, totaling variations. These variations encompass several types, such as variations in shape and color. Therefore, to achieve high success rates with very limited demonstrations, the agent needs to learn generalizable knowledge about manipulation instead of merely overfitting to the given demonstrations. We use the RGB-D image of size from the single front camera as the observation. To train the GNF, we also add additional camera views to provide RGB images as supervision.
Real robot. We use the xArm7 robot with a parallel gripper in real robot experiments. We set up two toy kitchen environments to make the agent generalize manipulation skills across the scenes and designed three manipulation tasks, including open the microwave door, turn the faucet, and relocate the teapot, as shown in Figure 1. We set up three RealSense cameras around the robot. Among the three cameras, the front one captures the RGB-D observations for the policy training and the left/right one provides the RGB supervision for the GNF training.
Expert Demonstrations. We collect demonstrations for each RLBench task with the motion planner. The task variation is uniformly sampled. We collect demonstrations for each real robot task using a VR controller. Details for collection remain in Appendix D.
Generalization tasks. To further show the generalization ability of GNFactor, we design additional simulated tasks and real robot tasks based on the original training tasks and add task distractors.
Training details. One agent is trained with two NVIDIA RTX3090 GPU for days (k iterations) with a batch size of . The shared voxel encoder of GNFactor is implemented as a lightweight 3D UNet with only M parameters. The Perceiver Transformer keeps the same number of parameters as PerAct (M parameters), making our comparison with PerAct fair.
2 Simulation Results
We report the success rates for multi-task tests on RLBench in Table 1 and for generalization to new environments in Table 2. We conclude our observations as follows:
Dominance of GNFactor over PerAct for multi-task learning. As shown by Table 1 and Figure 4, GNFactor achieves higher success rates across various tasks compared to PerAct, particularly excelling in challenging long-horizon tasks. For example, in sweep to dustpan task, the robot needs to first pick up the broom and use the broom to sweep the dust into the dustpan. We find that GNFactor achieves a success rate of , while PerAct could not succeed at all. In simpler tasks like open drawer where the robot only pulls the drawer out, both GNFactor and PerAct perform reasonably well, with success rates of and respectively. Furthermore, we observe that enhancing PerAct with extra camera views does not result in significant improvements. This underscores the importance of efficiently utilizing the available camera views.
Generalization ability of GNFactor to new tasks. In Table 2, we observe that the change made on the environments such as distractors impacts all the agents negatively, while GNFactor shows better capability of generalization on 5 out of 6 tasks compared to PerAct. We also find that for some challenging variations such as the smaller block in the task slide (S), both GNFactor and PerAct struggle to handle. This further emphasizes the importance of robust generalization skills.
Ablations. We summarize the key components in GNFactor that contribute to the success of the volumetric representation in Table 4. From the ablation study, we gained several insights:
(i) Our GNF reconstruction module plays a crucial role in multi-task robot learning. Moreover, the RGB loss is essential for learning a consistent 3D feature in addition to the feature loss, especially since the features derived from foundation models are not inherently 3D consistent.
(ii) The volumetric representation benefits from Diffusion features and depth-guided sampling, where the depth prior is utilized to enhance the sampling quality in neural rendering. An intuitive explanation is that GNF, when combined with DGS, becomes more adept at learning depth and 3D structure information. This enhanced understanding allows the 3D representation to better concentrate on the surfaces of objects rather than the entire volume. Moreover, replacing Stable Diffusion with DINO or CLIP would not result in similar improvements easily, indicating the importance of our vision-language feature.
(iii) While the use of skip connection is not a new story and we merely followed the structure of PerAct, the result of removing the skip connection suggests that our voxel representation, which distills features from the foundation model, plays a critical role in predicting the final action.
(iv) Striking a careful balance between the neural rendering loss and the action prediction loss is critical for optimal performance and utilizing information from multiple views by our GNF module proves to be beneficial for the single-view decision module.
Furthermore, we provide the view synthesis in the real world, generated by GNFactor in Figure 5 and Figure 6. We also give the quantitative evaluation measured by PSNR . We observe that the rendered views are somewhat blurred since the volumetric presentation learned by GNFactor is optimized to minimize both the neural rendering loss as well as the action prediction loss, and the rendering quality is largely improved when the behavior cloning loss is removed and only the GNF is trained. Notably, for the view synthesis in the real world, we do not have access to ground-truth point clouds for either training or testing. Instead, the point clouds are sourced from RealSense cameras and are therefore imperfect. Despite the limitations in achieving accurate pixel-level reconstruction results, we focus on learning semantic understanding of the whole scene from distilling Diffusion features, which is more important for policy learning.
3 Real Robot Experiments
We summarize the results of our real robot experiment in Table 3. From the experiments, GNFactor outperforms the PerAct baseline on almost all tasks. Notably, in the teapot task where the agent is required to accurately determine the grasp location and handle the teapot from a correct angle, PerAct fails to accomplish the task and obtains a zero success rate across two kitchens. We observed that it is indeed challenging to learn a delicate policy from only demonstrations. However, by incorporating the representation from the embedding of a vision-language model, GNFactor gains an understanding of objects. As such, GNFactor does not simply overfit to the given demonstrations.
The second kitchen (Figure 1) presents more challenges due to its smaller size compared to the first kitchen. This requires higher accuracy to manipulate the objects effectively. The performance gap between GNFactor and the baseline PerAct becomes more significant in the second kitchen. Importantly, our method does not suffer the same performance drop transitioning from the first kitchen to the second, unlike the baseline.
We also visualize our 3D policy module by Grad-CAM , as shown in Figure 7. We use the gradients and the 3D feature map from the 3D convolution layer after the Perceiver Transformer to compute Grad-CAM. We observe that the target objects are clearly attended by our policy, though the training signal is only the Q-value for a single voxel.
Conclusion and Limitations
In this work, we propose GNFactor, a visual behavior cloning agent for real-world multi-task robotic manipulation. GNFactor utilizes a Generalizable Neural Feature Field (GNF) to learn a 3D volumetric representation, which is also used by the action prediction module. We employ the vision-language feature from the foundation model Stable Diffusion besides the RGB feature to supervise the GNF training and observe that the volumetric representation enhanced by the GNF is helpful for decision-making. GNFactor achieves strong results in both simulation and the real world, across RLBench tasks and real robot tasks, showcasing the potential of GNFactor in real-world scenarios.
One major limitation of GNFactor is the requirement of multiple views for the GNF training, which can be challenging to scale up in the real world. Currently, we use three fixed cameras for GNFactor, but it would be interesting to explore using a cell phone to randomly collect camera views, where the estimation of the camera poses would be a challenge.
Acknowledgment. This work was supported, in part, by the Amazon Research Award, Cisco Faculty Award and gifts from Qualcomm.
References
Appendix A Visualizations
Appendix B Task Descriptions
Simulated tasks. We select language-conditioned tasks from RLBench , all of which involve at least variations. See Table 5 for an overview. Our task variations include randomly sampled colors, sizes, counts, placements, and categories of objects, totaling different variations. The set of colors have 20 instances: red, maroon, lime, green, blue, navy, yellow, cyan, magenta, silver, gray, orange, olive, purple, teal, azure, violet, rose, black, and white. The set of sizes includes 2 types: short and tall. The set of counts has 3 instances: 1, 2, 3. The placements and object categories are specific to each task. For example, open drawer has 3 placement locations: top, middle and bottom. In addition to these semantic variations, objects are placed on the tabletop at random poses within a limited range.
Generalization tasks in simulation. We design additional tasks where the scene is changed based on the original training environment, to test the generalization ability of GNFactor. Table 6 gives an overview of these tasks. Videos are also available on yanjieze.com/GNFactor.
Real robot tasks. In the experiments, we perform three tasks along with three additional tasks where distracting objects are present. The door task requires the agent to open the door on an mircowave, a task which poses challenges due to the precise coordination required. The faucet task requires the agent to rotate the faucet back to center position, which involves intricate motor control. Lastly, the teapot task requires the agent to locate the randomly placed teapot in the kitchen and move it on top of the stove with the correct pose. Among the three, the teapot task is considered the most challenging due to the random placement and the need for accurate location and rotation of the gripper. All tasks are set up in two different kitchens, as visualized in Figure 8. The keyframes used in real robot tasks are given in Figure 9.
Appendix C Implementation Details
Voxel encoder. We use a lightweight 3D UNet (only M parameters) to encode the input voxel (RGB features, coordinates, indices, and occupancy) into our deep 3D volumetric representation of size . Due to the cluttered output from directly printing the network, we provide the PyTorch-Style pseudo-code for the forward process as follows. For each block, we use a cascading of one Convolutional Layer, one BatchNorm Layer, and one LeakyReLU layer, which is common practice in the vision community.
Generalizable Neural Field (GNF). The overall network architecture of our GNF is close to the original NeRF implementation. We use the same positional encoding as NeRF and the encoding function is formally
Percevier Transformer. Our usage of Percevier Transformer is close to PerAct . We use attention blocks to process the sequence from multi-modalities (3D volume, language token, and robot proprioception) and output a sequence also. The usage of Perceiver Transformer enables us to process the long sequence with computational efficiency, by only utilizing a small set of latents to attend the input. The output sequence is then reshaped back to a voxel to predict the robot action. The Q-function for translation is predicted by a 3D convolutional layer, and for the prediction of openness, collision avoidance, and rotation, we use global max pooling and spatial softmax operation to aggregate 3D volume features and project the resulting feature to the output dimension with a multi-layer perception. We could clarify that the design for the policy module is not our contribution; for more details please refer to PerAct and its official implementation on https://github.com/peract/peract.
Appendix D Demonstration Collection for Real Robot Tasks
For the collection of real robot demonstrations, we utilize the HTC VIVE controller and basestation to track the 6-DOF poses of human hand movements. We then use triad-openvr package (https://github.com/TriadSemi/triad_openvr) to employ SteamVR and accurately map human operations onto the xArm robot, enabling it to interact with objects in the real kitchen. We record the real-time pose of xArm and RGB-D observations with the pyrealsense2 (https://pypi.org/project/pyrealsense2/). Though the image size is different from our simulation setup, we use the same shape of the input voxel, thus ensuring the same algorithm is used across the simulation and the real world. The downscaled images () are used for neural rendering.
Appendix E Detailed Data
Besides reporting the final success rates in our main paper, we give the success rates for the best single checkpoint (i.e., evaluating all saved checkpoints and selecting the one with the highest success rates), as shown in Table 7. Under this setting GNFactor outperforms PerAct with a larger margin. However, we do not use the best checkpoint in the main results for fairness.
We also give the detailed number of success in Table 8 for reference in addition to the success rates computed in Table 2.
Appendix F Stronger Baseline
To make the comparison between our GNFactor and PerAct fairer, we enhance Peract’s input by using 4 camera views, as visualized in Figure 10. These views ensure that the scene is fully covered. It is observed in our experiment results (Table 1) that GNFactor which takes the single view as input still outperforms PerAct with more views.
Appendix G Hyperparameters
We give the hyperparameters used in GNFactor as shown in Table 9. For the GNF training, we use a ray batch size , corresponding to pixels to reconstruct, and use and to maintain major focus on the action prediction. For real-world experiment, we set the weight of the reconstruction loss to 1.0 and the weight of action loss to 0.1. This choice was based on our observation that reducing the weight of the action loss and increasing the weight of the reconstruction loss did not significantly affect convergence but did help prevent overfitting to a limited number of real-world demonstrations. We uniformly sample points along the ray for the “coarse” network and sample points with depth-guided sampling and points with uniform sampling for the “fine” network.