D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement

Yixuan Wang, Mingtong Zhang, Zhuoran Li, Tarik Kelestemur, Katherine Driggs-Campbell, Jiajun Wu, Li Fei-Fei, Yunzhu Li

I INTRODUCTION

The choice of scene representation is critical in robotic systems. An ideal representation should be simultaneously 3D, dynamic, and semantic to meet the needs of various robotic manipulation tasks in our daily lives. However, previous research into scene representations in robotics often does not encompass all three properties. Some representations exist in 3D space , yet they overlook semantic information. Others focus on dynamic modeling , but only consider 2D data. Some other works are limited by only considering semantic information such as object instance and category .

In this work, we aim to satisfy all three criteria by introducing D3Fields, unified descriptor fields that are 3D, dynamic, and semantic. D3Fields take in arbitrary points in the 3D world coordinate frame and output both geometric and semantic information related to these points. This includes the instance mask, dense semantic features, and the signed distance to the object surface. Notably, deriving these descriptor fields requires no training and is conducted in a zero-shot manner using large foundational vision models and vision-language models (VLMs). Specifically, we first use Grounding-DINO , Segment Anything (SAM) , XMem , and DINOv2 to extract information from multi-view 2D RGB images. We then project the 3D points back to each camera, interpolate to compute representations from each view, and fuse these data to derive the descriptors for the associated 3D points, as shown in Fig. 1 (left). By leveraging the dense semantic feature and instance mask of our representation, we can robustly track 3D points of the target object instance and train dynamics models. These learned dynamics models can then be incorporated into a Model-Predictive Control (MPC) framework to plan for manipulation tasks.

Notably, the derived representations allow for goal specification using 2D images sourced from the Internet, phones, or those generated by AI models. Such goal images have been challenging to manage with previous methods, because they contain varied styles, contexts, and object instances different from the robot’s workspace. Our proposed D3Fields can establish dense correspondences between the robot workspace and the target configurations. These correspondences give us the task objective, enabling us to plan the robot’s actions with the learned dynamics model within the MPC framework. This task execution process does not require any further training, offering a flexible and convenient interface for humans to instruct robots.

We evaluate our method across a wide range of household robotic manipulation tasks in a zero-shot manner. These tasks include organizing shoes, collecting debris, and organizing office desks, as shown in Fig. 1 (right). Furthermore, we offer detailed quantitative comparisons between our method and other state-of-the-art dense descriptor techniques. Our results indicate that our approach significantly outperforms in terms of generalizability and manipulation accuracy.

To summarize our contributions: (1) We introduce a novel representation, D3Fields, that is 3D, dynamic, and semantic. (2) We present a novel and flexible goal specification method using 2D images that incorporate a range of styles, contexts, and instances. (3) Our proposed robotic manipulation framework supports zero-shot generalizable manipulation applicable to a broad spectrum of household tasks.

II Related Works

Foundation models generally refer to those trained on broad data, often using self-supervision at scale, which can then be adapted (e.g., fine-tuned) to various downstream tasks. Large Language Models (LLMs) have showcased promising reasoning abilities for language. Robotics researchers have recently released a series of works that leverage LLMs, including SayCan and Inner Monologue , to directly generate robot plans. Some later works have used LLMs as a code generator: Code as Policies uses 2D object detectors as the perception API, whereas VoxPoser creates a 3D value map. Yet, their perception modules fall short in modeling the precise geometry and dynamics of objects. Our D3Fields aim to address this by focusing on detailed 3D geometry and dynamics.

Meanwhile, foundational vision models, such as SAM and DINOv2 , have demonstrated impressive zero-shot generalization capabilities across various vision tasks. However, their focus is primarily on 2D vision tasks. Grounding these models in a dynamic 3D environment remains a challenge. The recent GROOT project showcases how to construct 3D object-centric representations using foundational models and exhibits notable few-shot generalization capabilities . Still, GROOT does not emphasize learning about object dynamics or achieving zero-shot generalizable robotic manipulation.

II-B Representation for Visual Robotic Manipulation

Scene representation has been a pivotal component in robotic manipulation systems. Some early work relies on 2D representations, such as bounding boxes . Many recent methods construct particle representations of the environment and employ learned dynamics to capture the system’s underlying structure . They demonstrate impressive results in unstructured environments and with non-rigid objects. However, they are not semantic, which can hinder their ability to generalize to new tasks and scenarios. Some research opts for a fixed-dimension latent vector derived from high-dimensional sensory inputs as the representation , but such a representation does not scale well to complex manipulation tasks that require high precision and explicit scene structures. Other approaches use 6 DoF object poses as their representation , though focusing primarily on grasping tasks instead of more dynamic ones. In this work, we aim to address these issues by introducing D3Fields, a representation that models dynamic 3D environments at varying semantic levels.

II-C Neural Fields for Robotic Manipulation

Researchers have presented a variety of works using neural fields as a representation for robotic manipulation . Among them, Neural Descriptor Fields are the most relevant to ours . They build neural feature fields that generalize to different instances with several demonstrations; but they focus on learning geometric, not semantic features, which hinders cross-category generalization.

Recently, a series of works distilled neural feature fields using foundation models such as CLIP and DINO for supervision . LeRF distills neural feature fields to handle open-vocabulary 3D queries and develops task-oriented grasping based on it . Shen et al. use a similar distilled feature field for the grasping task. Both methods require dense camera views to train the neural field. GNFactor addresses this by introducing a voxel encoder . However, distilling foundation models to create neural feature fields has drawbacks: (1) They often require dense camera views for a quality field. (2) Distilled neural fields need retraining for new scenes, limiting their generalization and making them ineffective for dynamic scenes. In contrast, our D3Fields do not need extra training for new scenes and can work with sparse views and dynamic settings.

III Method

In this section, we introduce the problem formulation in Section III-A and define camera transformation and projection notations in Section III-B. The construction of D3Fields is detailed in Section III-C. Section III-D discusses tracking keypoints and learning dynamics, while Section IV-C showcases how our representation enables zero-shot generalizable manipulation skills.

Given a 2D goal image I\mathcal{I}, we denote the corresponding scene representation as sgoal\bm{s}_{\text{goal}}. Our goal is to find the action sequence {at}\{a^{t}\} to minimize the task objective:

where c(⋅,⋅)c(\cdot,\cdot) is the cost function measuring the distance between the terminal representation sT\bm{s}^{T} and the goal representation sgoal\bm{s}_{\text{goal}}. Representation extraction function g(⋅)g(\cdot) takes in the current multi-view RGBD observations ot\bm{o}^{t} and outputs the current representation st\bm{s}^{t}. f(⋅,⋅)f(\cdot,\cdot) is the dynamics function that predicts the future representation st+1\bm{s}^{t+1}, conditioned on the current representation st\bm{s}^{t} and action ata^{t}. The optimization aims to find the action sequence {at}\{a_{t}\} that minimizes the cost function c(sT,sgoal)c(\bm{s}^{T},\bm{s}_{\text{goal}}).

III-B Notation: Camera Transformation and Projection

We assume all cameras’ intrinsic parameters K\mathbf{K} and extrinsic parameters T\mathbf{T} are known. The camera ii extrinsic parameters are defined as follows.

where π\pi performs perspective projection, mapping a 3D vector p=[x,y,z]Tp=[x,y,z]^{T} to a 2D vector q=[x/z,y/z]Tq=[x/z,y/z]^{T}.

III-C D3Fields Representation

We fuse observation ot\mathbf{o}^{t} from multiple views to build the implicit 3D descriptor fields Ft(⋅)\mathcal{F}^{t}(\cdot). For simplicity, we will represent ot\mathbf{o}^{t} as o\mathbf{o}, and Ft(⋅)\mathcal{F}^{t}(\cdot) as F(⋅)\mathcal{F}(\cdot) in this subsection. The implicit 3D descriptor field F(⋅)\mathcal{F}(\cdot) is defined as

We fuse descriptors from all KK views as follows:

where HH is the unit step function and δ\delta is a small value to avoid numeric errors. vi=0v_{i}=0 when x\mathbf{x} is not observable in camera ii, because if x\mathbf{x} is occluded in camera ii, it should not contribute to the descriptor of x\mathbf{x}. In addition, we could only have a confident estimation when x\mathbf{x} is close to the surface. Therefore, wiw_{i} will decay as ∣di∣|\mathbf{d}_{i}| increases. For x\mathbf{x} that is far away, f\mathbf{f} and m\mathbf{m} will degrade to 0T0^{T}.

III-D Keypoints Tracking and Dynamics Training

Since F(⋅)\mathcal{F}(\cdot) is differentiable, we could use a gradient-based optimizer. This method could be naturally extended to multiple-instance scenarios. We found that relying solely on features for tracking is unstable. We added rigid constraints and distance regularization for a more stable tracking.

Keypoint tracking enables dynamics model training on real data. We instantiate the dynamics model f(⋅,⋅)f(\cdot,\cdot) as graph neural networks (GNNs). We follow to predict object dynamics. Please refer to for more details on how to train the GNN-based dynamics model. The trained dynamics will be used for trajectory optimization in Section III-E.

III-E Zero-Shot Generalizable Robotic Manipulation

then we have sgoal,j=∑i=1H×Wwijui\bm{s}_{\text{goal},j}=\sum_{i=1}^{H\times W}w_{ij}\mathbf{u}_{i}, where Wgoalf\mathcal{W}^{\mathbf{f}}_{\text{goal}} is the feature volume extracted from Igoal\mathcal{I}_{\text{goal}} using DINOv2. ss is the hyperparameter to determine whether the heatmap wijw_{ij} is more smooth or concentrating. Although Eq. 9 only shows a single instance case, it could be naturally extended to multiple instances by using instance mask information.

However, sgoal\bm{s}_{\text{goal}} is in the image space, while st\bm{s}^{t} is in the 3D space. We bridge this gap by introducing a reference camera with approximate intrinsic and extrinsic parameters K′\mathbf{K^{\prime}} and T′\mathbf{T^{\prime}}. Instead of rendering images in the reference view, we focus on projecting 3D keypoints into 2D images and define the task cost function in image space as follows:

IV Experiments

In this section, we evaluate our representation across various manipulation tasks with varying goal image styles, instances, and contexts. We visualize D3Fields and showcase tracking results in Section IV-B. Then, we highlight our framework’s zero-shot generalizability in both real-world and simulated tasks in Section IV-C. Finally, a quantitative comparison with baselines in Section IV-D underscores our framework’s generalization and manipulation precision.

In the real world, we employ four OAK-PRO D cameras to gather RGBD observations and use the Kinova® Gen3 for action execution. In simulation, we utilize OmniGibson and deploy Fetch for mobile manipulation tasks . Our evaluations span a variety of tasks, including organizing shoes, collecting debris, tidying the office table, arranging utensils, and more.

We implement the baseline methods using Dense Object Nets (DON) and DINO for feature extraction . We quantitatively evaluate these methods on five object classes for single-instance manipulation tasks in the real world. The results and analysis are presented in Section IV-D.

IV-B Descriptor Fields Visualization and Keypoints Tracking

D3Fields provide a good 3D semantic representation, as shown in Fig. 3(a). We first visualize the mask fields by coloring 3D points according to their most likely instance, and our visualization shows a clear 3D instance segmentation. Additionally, we map the semantic features to RGB space using PCA, as with DINOv2 . Visualization of the descriptor fields reveals that D3Fields retain a dense semantic understanding of objects. In the provided shoe example, even though various shoes have distinct appearances and poses, they exhibit similar color patterns: shoe heels are represented in green, and shoe toes in red. We observed similar patterns when evaluating the model on mugs and forks.

As discussed before, D3Fields can also capture scene dynamics. We evaluate it by tracking the object keypoints. We show two examples of 3D keypoint tracking in Fig. 3(b). In the first example, a shoe is pushed and then flipped. Although only a portion of the shoe is visible from the view, our framework tracks it reliably. In another example, a shoe is lifted and then set down. Despite parts of the shoe being out of the camera’s view, we can robustly track it in 3D.

IV-C Zero-Shot Generalizable Manipulation

We conduct a qualitative evaluation of D3Fields in common household robotic manipulation tasks in a zero-shot manner, with partial results displayed in Fig. 1 and Fig. 4. The following capabilities of our framework are observed:

Generalization to AI-Generated Goal Images. In Fig. 1, the goal image, rendered in a Van Gogh style, depicts shoes distinct from those in the workspace. Since D3Fields encode semantic information, capturing shoes with varied appearances under similar descriptors, our framework can manipulate shoes based on AI-generated goal images.

Compositional Goal Images and 3D Manipulation. Using the office desk organization example in Fig. 1, the robot first arranges the mouse and pen according to the goal image. It then repositions the mug from the box to the mug pad, referencing a goal image of the upright mug.

Generalization across Instances and Materials. Granular objects, unlike rigid ones, have more complex dynamics. Our framework effectively handles these materials, as shown in the debris collection in Fig. 1. Fig. 4 further showcases our framework’s instance-level generalization, where the goal image displays instances different from the workspace.

Generalization across Simulation and Real World. We evaluated our framework on household tasks in the simulator, as shown in the utensil organization and mug organization examples in Fig. 4. Given goal images taken from the real world, our framework can also manipulate objects to the goal configurations. Our framework demonstrates generalization capabilities between simulation and the real world.

IV-D Quantitative Comparisons with Baselines

In Fig. 5(a), we measure performance using the IoU between the goal image mask and the final state mask after manipulation, with higher values indicating better alignment. Evaluating across five object classes, our method consistently outperforms the baselines, underscoring its generalization and manipulation accuracy. While DINO struggles with distinguishing object components, leading to imprecise results, it still works better than DON. Although DON performs well on familiar object classes and configurations, it lacks generalization in novel scenarios.

In Fig. 5(b), we present the correspondence results. We manually label corresponding keypoints on both the goal image and the final manipulation result to evaluate the correspondence accuracy. We calculate the fraction of accurately matched points based on a distance threshold. Our method consistently outperforms the baselines, regardless of the threshold. DINO ranks second, while DON lags behind. Consistent with Fig. 5(a), our method excels in generalization and accuracy, DINO is broadly applicable but less precise, and DON struggles with generalization.

V Conclusion

In this work, we introduce D3Fields, which implicitly encode 3D semantic features and 3D instance masks, and model the underlying dynamics. Our emphasis is on zero-shot generalizable robotic manipulation tasks specified by 2D goal images of varying styles, contexts, and instances. Our framework excels in executing a diverse array of household manipulation tasks in both simulated and real-world scenarios. Its performance greatly surpasses baseline methods such as Dense Object Nets and DINO in terms of generalization capabilities and manipulation accuracy.

References